Thesis

Science needs more than words.

LLMs receive a molecule as a string of numbers and letters, or an oil and gas reservoir as a grid of numbers. The scientific object is never properly embedded in an LLM. SciLM, the Scientific Language Model, is a new kind of multimodal language model, trained from scratch to read and write scientific objects in their native form.

LLMs today see science as text

SciLM sees science in its native form

cathode.cif
lattice2.8162.81614.0529090120
Li0.00000.00000.0000
Co0.00000.00000.5000
O0.00000.00000.2396
O0.00000.00000.7604
Li0.66670.33330.3333
Co0.66670.33330.8333
O0.66670.33330.5729
O0.66670.33330.0937
Li0.33330.66670.6667
Co0.33330.66670.1667
2,000 more lines
reservoir.grdecl
SPECGRID
200150601F/
COORD
4512.308890.122410.004512.308890.122680.00
4562.308890.122409.404562.308890.122679.40
ZCORN
2410.002410.002410.202410.202410.102410.10
PORO
0.2140.1980.2310.1870.2050.1760.223
PERMX
312.5188.0405.296.7250.1143.8377.4
millions of lines more
Same object,
embedded differently

Strings of numbers. A crystal structure file and a subsurface grid file: the model sees digits, not the geometry, the bonds, the barrier or the geology.

Scientific objects. SciLM sees a molecule as atoms and bonds in space, and the subsurface as a volume of geological bodies. It can watch the ion cross the barrier, and the oil flow into a well.

The tools live inside the model.

SciLM reads scientific data itself, so the simulators and specialist models an LLM has to call can live inside one model, and it can answer scientific questions no one wrote a tool for, just as LLMs write text no one has written before. An LLM works only with text. Ask one how oil will flow through a specific field, or how fast a new battery will charge, and it cannot answer even if you hand it the file describing the field or the coordinates of every atom. It calls a tool, a simulator or a specialist AI model, and reads the tool's output back as text, which never holds everything the physics simulation produced. This works reliably only for questions a tool already exists for; an LLM could in theory write a missing tool itself, but no one has shown one reliably writing a new physics simulator in the middle of an answer. And even when every tool works, a text-only LLM starts each loop from whatever a tool hands it, while SciLM, trained on the data itself, starts from a good first guess, the way an experienced scientist does, and reaches the answer with fewer expensive simulations or experiments.

tool calltool calltool calltool callProtein foldingFlow simulatorMolecular dynamicsa tool it writesitself, untestedLLMreasons and plans withwordssimulates on its ownnothing
An LLM with tools. Every answer comes back as text, and reliably only for questions a tool already exists for.
can still callany tool,if neededSciLMreasons and plans withwordsscientific objectssimulates on its own, the tools in its weightsProtein foldingFlow simulatorMolecular dynamicsanswersquestions no tool exists for
SciLM. The tools live in its weights: protein folding, a flow simulator, molecular dynamics, in one model. It plans with the science in mind, and answers questions no tool exists for. It can still call the same tools if needed, but it can also read their results in their native form and not only text summaries.

AI needs more data than science has ever measured. We generate it.

The breakthroughs in AI were built on large datasets that already existed: the web for ChatGPT, the Protein Data Bank for AlphaFold. Measurements from different labs rarely combine into one training set, and much of what matters cannot be measured at all: no experiment watches a reaction at the atomic scale or sees flow in the subsurface. Computational science has generated what experiments cannot see for seventy years, and it can do so at larger scales and in greater quantities.

The webLLMs

Trillions

of words

Written by people, for people.

Protein Data BankAlphaFold

240,000

structures in 55 years

Experimentally measured, one at a time.

SimulationSciLM

102.4 M

structures, generated

More than every laboratory in history has measured experimentally. Plus 1,000,000 oil and gas reservoirs.

SciLMshape of the physicsSimulationpretrain: every scale, any quantityExperimentfine-tune: correct the small errorsPredictionsfor questions no tool exists for
Pretrain on simulation. Fine-tune on experiment. Nobody can see a reaction at the atomic scale, and nobody can see the subsurface, so scientists and engineers simulate them. An AI model has to learn the same way to gain a deep understanding of the physical world.

What we have released

If the data or model you need does not exist yet, write to us.

We generate simulation data and train AI models to specification: the physics, the systems, the sampling and the format, for the subsurface and for materials alike. Exclusive or non-exclusive licenses.

  • Datasets: MaterialsSaddles and SiliciclasticReservoirs are open; the next are licensed.
  • Models: ResFlow and SaddleFlow, licensed as tools or trained on your data.
  • Data to order: your system, your scale, in the form your models read.