All posts

Why Uncertainty Quantification is the Path to Digitizing Materials Engineering

Why Uncertainty Quantification is the Path to Digitizing Materials Engineering

There is a persistent and growing tension in the AI-for-materials space. On one hand, there are peer reviewed examples that high-accuracy Density Functional Theory (DFT), when used to train atomic scale foundation models, can replace physical experiments [1]. On the other hand, industry scientists struggle to use open-source atomic scale foundation models because they are trained on datasets that do not reliably translate to the real material systems they care about.

The promise of digitizing materials development is often framed as a simple matter of compute: more GPUs plus more data equals better models. However, integrating machine learning interatomic potentials (MLIPs) into industrial workflows to replace physical experiments faces a challenge machine learning models in other fields do not, namely, the inability of a human to determine if the predicted result is accurate. Therefore, while high-accuracy DFT data has accurately parameterized atomic-scale foundation models that simulate materials across the periodic table, they are often only able to model idealized physical experiments accurately leaving a distinct reliability gap when these models meet the "messy" reality of industrial materials.

Industrial systems are rarely pristine. They are defined by features that frequently reside outside the base training distribution of current foundation models such as heterogeneous interfaces, defect motifs, and complex multi-component mixtures. This distribution shift from what a model is trained on to what it sees in production leads to a low reliability rate in reducing experimental measurements in production environments preventing their mass adoption.

To understand the reliability wall industrial scientists are facing, we can look back at the digital revolution of the 1990s. Computational Fluid Dynamics (CFD) was transitioning from a promising idea to an engineering tool capable of displacing physical wind-tunnel testing. As initially skeptical experimentalists tried CFD, they were able to systematically close the gap between experimental results and computational predictions by using the principled error bounds of numerical methods. These error bounds provided clear insight into what systems and geometries could be simulated accurately at given compute budgets. Moreover, the numerical methods gave a clear signal on how to reduce error even further by using finer meshes. When coupled with exponentially increasing computational power, CFD was quickly able to predict flow properties of realistic systems at fractions of the cost and time of wind-tunnel testing. Every valuable prediction created a positive feedback loop that gave engineers more confidence in CFD making it a primary fixture of engineering development.

The landscape of materials modeling today looks remarkably like fluid dynamics 30 years ago. While we have taken a massive leap forward with foundation models that can predict properties across the periodic table at a 10,000x reduction in cost compared to DFT. There is no principled method to determine if simulated results are accurate. Without a way to quantify errors, experimentalists cannot rely on these models to replace high-stakes physical testing. This is true despite training datasets having grown beyond hundreds of millions of structures [2]. This has made the models better, but since the chemical space of materials is so large, hundreds of millions of structures still only captures a tiny percent of the materials properties engineers want to test.

Even today, state of the art pre-trained foundation models often predict materials properties incorrectly [3]. However, just as with CFD, to move from computational scientists developing the tool to experimental scientists using it as a tool requires that they can rely on the models to replace their experiments. They need to know if the predictions are accurate. It is less that they are incorrect 50% of the time, it is that we need to know which 50% are correct.

Physics Inverted Materials was founded to address the challenge of digitizing materials development and eliminating physical testing. Principally this requires training physics based models that can economically and accurately predict materials properties reliably to replace physical experiments. Phin addresses the reliability gap through its uncertainty quantification technology. Uncertainty quantification illuminates the boundary between accurate and inaccurate model predictions to determine which model predictions are trustworthy. Uncertainty aware models that know when they are inaccurate enables them to be used in production environments with greater confidence.

Phin’s pretrained models therefore extend the performance of pre-trained open source models in one critical dimension, their reliability (see figure). This means that instead of computational scientists having to inspect results to determine their validity ad-hoc, models can be used more widely with strict bounds on validity. This is the watershed of error bounds that allowed CFD to be widely used.

To continuously improve materials models and digitize materials development, we need to not only know the boundary between accurate and inaccurate model predictions, but just like refining grids in CFD, there needs to be a principled way to drive the number of inaccurate predictions to zero. Continuous improvement through automated fine-tuning processes like active learning enable systematic reductions to inaccurate predictions. Fine-tuning uses the predictions from uncertainty quantification to identify additional data points and then recalculate those as an extension to the training dataset.

Fine-tuning open-source foundation potentials is still very challenging due to the data efficiency, the quality of the uncertainty quantification, and the sampling methods to identify what structures to calculate. In practice, this means that for specialized fine-tuned models that cover narrow chemical improvements, greater than 500 structures are needed to generate pre-training data with larger design spaces requiring more than ten thousand structures [4, 5]. Fine-tuning open-source models therefore requires tens or hundreds of thousands of dollars in training data cost. This pushes the cost of fine-tuned open source foundation potentials to be more expensive than the physical experiments they may replace (see figure), making it hard to justify when the payoff requires eliminating hundreds or thousands of experiments.

We have bridged the chasm of fine-tuning models by fundamentally changing the unit economics of the field. Instead of heuristics that identify what data to generate, Phin’s uncertainty quantification efficiently identifies structures that reduce the fine-tuning dataset size by 90%. As data generation is the primary cost, this reduces the six-figure fine-tuning bill to manageable costs, allowing us to amortize the cost over much fewer physical experiments. In production this means that we only need to replace ~15 physical experiments before the unit economics of fine tuning amortize the fine-tuning cost and that we can reduce the cost of physical experiments by 50% for pilot studies with our customers.

With these two approaches, pre-trained models that know their reliability boundary and fine-tuned models that continually improve, Phin is driving the adoption of digital materials development. Scientists no longer have to worry about their model accuracy and instead have principled ways of knowing when to use them. This allows them to digitally develop atomic blueprints of their materials, replacing months or even years of lab-scale prototyping experiments. As Phin’s models improve, there will rarely be a need for laboratory scale experiments as ideas can move directly to full scale experiments to validate the atomic blueprints.

If you are interested in deploying Phin’s AI models in your experimental development, schedule a demo today or try using PhisOS with pretrained models today.