Thesis
Science needs more than words.
LLMs receive a molecule as a string of numbers and letters, or an oil and gas reservoir as a grid of numbers. The scientific object is never properly embedded in an LLM. SciLM, the Scientific Language Model, is a new kind of multimodal language model, trained from scratch to read and write scientific objects in their native form.
LLMs today see science as text
SciLM sees science in its native form
embedded differently
Strings of numbers. A crystal structure file and a subsurface grid file: the model sees digits, not the geometry, the bonds, the barrier or the geology.
Scientific objects. SciLM sees a molecule as atoms and bonds in space, and the subsurface as a volume of geological bodies. It can watch the ion cross the barrier, and the oil flow into a well.
The tools live inside the model.
SciLM reads scientific data itself, so the simulators and specialist models an LLM has to call can live inside one model, and it can answer scientific questions no one wrote a tool for, just as LLMs write text no one has written before. An LLM works only with text. Ask one how oil will flow through a specific field, or how fast a new battery will charge, and it cannot answer even if you hand it the file describing the field or the coordinates of every atom. It calls a tool, a simulator or a specialist AI model, and reads the tool's output back as text, which never holds everything the physics simulation produced. This works reliably only for questions a tool already exists for; an LLM could in theory write a missing tool itself, but no one has shown one reliably writing a new physics simulator in the middle of an answer. And even when every tool works, a text-only LLM starts each loop from whatever a tool hands it, while SciLM, trained on the data itself, starts from a good first guess, the way an experienced scientist does, and reaches the answer with fewer expensive simulations or experiments.
AI needs more data than science has ever measured. We generate it.
The breakthroughs in AI were built on large datasets that already existed: the web for ChatGPT, the Protein Data Bank for AlphaFold. Measurements from different labs rarely combine into one training set, and much of what matters cannot be measured at all: no experiment watches a reaction at the atomic scale or sees flow in the subsurface. Computational science has generated what experiments cannot see for seventy years, and it can do so at larger scales and in greater quantities.
The webLLMs
Trillions
of words
Written by people, for people.
Protein Data BankAlphaFold
240,000
structures in 55 years
Experimentally measured, one at a time.
SimulationSciLM
102.4 M
structures, generated
More than every laboratory in history has measured experimentally. Plus 1,000,000 oil and gas reservoirs.



