← All posts

PHIN-atomic: a foundation potential that reports its own error

Summary

We introduce PHIN-atomic, our first foundation machine learning interatomic potential (MLIP), trained on MatPES-PBE 1. Every energy and force it predicts produces a calibrated error bar, evaluated alongside physical outputs with negligible computational overhead.

PHIN-atomic’s uncertainty predictions, which improve over state-of-the-art, expensive ensembling methods, remove the jagged performance of pre-trained MLIPs and build trustable inference. We establish this on held-out structures, hold it on chemistry absent from training, and prove its value through downstream computed property prediction accuracy.

Background

Foundation potentials have changed the economics of atomistic simulation. Work that would have consumed a year of DFT allocation now runs in an afternoon, and the constraints have moved beyond generating results to deciding which results to trust. That decision typically cannot come from the model’s own outputs as a foundation potential evaluates any structure received, familiar or not, and gives no sign of which case it is. Without the physical constraints that bound an empirical force field, extrapolation error is unbounded, reaching the user as an unphysical result: a relaxation to the wrong basin, imaginary modes in a phonon spectrum, atom loss in a trajectory. A sweep of ten thousand relaxations contains some number of these, and with no per-prediction uncertainty the choice is to verify all of them, forfeiting the speedup, or none, forfeiting the trust.

An established remedy today is to train an ensemble from independent initializations and treat the disagreement as an uncertainty estimate. Language model serving does the same thing, sending a prompt to several models in parallel and having a judge reconcile their answers: OpenRouter built Fusion on exactly this, and its strongest panel outscored every individual model in it on their deep-research evaluation, because it purchases several completions instead of one. A deep ensemble buys reliability the same way, multiplying training, storage, and every force evaluation by the number of members, and its estimate remains overconfident and loosely correlated with the error it predicts. Alternatively, uncertainty estimation via latent-space distance to the training set is cheaper to evaluate, but it measures distance from the training distribution rather than predictive error. Here, we use a single model backbone and small uncertainty layer to predict the calibrated error directly, which enables trust bands around each prediction, set against a tolerance in physical units. A scientist states the error a workflow can absorb, and the model returns the subset that meets it.

Experimental Method

The uncertainty estimate is produced in the same forward pass as the energy and the forces, and calibrated on data withheld during training. The training, validation, and test datasets are segmented to withhold whole materials rather than scattered structures, so no material appears in both training and evaluation. The uncertainty predicts a distribution, which we simplify into three trust bands. Each of the three trust bands pairs a tolerance on a structure’s mean force error with how often predictions may exceed it. We compare this to the best alternative, ensembling, by training 3 further copies from independent initializations and evaluate all four on the 1,064 structures withheld by every one. We evaluate the model and uncertainty performance by calculating emergent properties using MatCalc-Bench 1,2. In addition, we compare our model to published results from five models trained from scratch on MatPES-PBE 3,5,6 and also to MACE-MatPES-PBE-0† 4, an OMAT-pretrained model fine-tuned on MatPES-PBE, benchmarked in-house.

Results

Ranking the errors

An error bar is useful only if its width varies with the error it bounds. A bar of constant width reproduces the correct coverage while indicating nothing about which prediction to distrust, leaving a workflow with no basis for discriminating results or spending budget on additional verification.

Figure 1 tests the ranking of each uncertainty metric. Every strip holds the same withheld structures, reordered by one metric and shaded by the force error measured afterwards against DFT. A metric that carries information drives the dark bands toward its low-confidence end, taking the strip from mixed to graded. A model generating its own confidence ranks its errors more effectively than three independently trained models’ disagreement. Latent-space distance recovers a small fraction of what either achieves. The ranking is what makes selective prediction possible, deciding how much error a fixed verification budget removes.

Four horizontal strips of the same withheld structures, each ordered by a different uncertainty metric and shaded by measured force error
Figure 1: The same withheld structures in every strip, each band an equal-sized group shaded by its measured force error, with darker shading indicating larger error. Structures run from most to least confident under the metric named beside each strip. Each percentage is the drop in mean force error when the least confident fifth is set aside, scaled between discarding a random fifth and discarding the worst fifth, over the structures withheld for all four networks.
Figure 1: The same withheld structures in every strip, each band an equal-sized group shaded by its measured force error, with darker shading indicating larger error. Structures run from most to least confident under the metric named beside each strip. Each percentage is the drop in mean force error when the least confident fifth is set aside, scaled between discarding a random fifth and discarding the worst fifth, over the structures withheld for all four networks.

One network against a deep ensemble

The deep ensemble is the standard a new method has to beat. Figure 2 compares them directly by putting the two on the same axis. Predictions are removed in order of confidence, from the full set at the left edge to the most trusted half at the right, so a steeply falling curve indicates a model that removed the right structures first. The ensemble is the more accurate average predictor but the poorer judge of its own reliability, and PHIN-atomic’s own confidence identifies its errors better than three networks’ disagreement. At matched coverage the ensemble interval must be 17% wider for an equivalent claim. Predictive accuracy and the quality of the uncertainty estimate are separable properties, and it is the second that is critical for simulation workflows that drive real scientific decisions.

Force MAE falling as low-confidence structures are removed, compared for PHIN-atomic and a three-member deep ensemble
Figure 2: Force MAE over the retained set, on the 1,064 structures withheld for every network involved. The left edge is every structure and the right edge the most trusted half, and the shaded band is the error our ranking leaves against a perfect oracle. At matched coverage the ensemble interval runs 0.75 eV/Å against 0.64 eV/Å.
Figure 2: Force MAE over the retained set, on the 1,064 structures withheld for every network involved. The left edge is every structure and the right edge the most trusted half, and the shaded band is the error our ranking leaves against a perfect oracle. At matched coverage the ensemble interval runs 0.75 eV/Å against 0.64 eV/Å.

Moreover, the overhead is small enough to calculate the uncertainty throughout inference. Running molecular dynamics in LAMMPS on a 4,096-atom cell, a step costs 9.7 microseconds per atom with an error bar on every atom against 16.5 for MACE-MatPES-PBE-0 at the same precision, and no error bar at all.

Chemistry absent from training

The ranking in Figure 1 can be produced without generalizing at all, by learning which elements proved difficult in training and widening the bar whenever one appears. Such a model is most confident exactly where a user needs it least, on the novel chemistry a screening campaign exists to explore. Figure 3 rules that out by showing that as the chemistry becomes less familiar, the predicted error responds to the local configuration rather than to element identity. Compare this to latent-space distance, which loses up to half of its performance on exactly the structures the estimate exists to cover. PHIN’s single model uncertainty therefore holds on the most important structures and when it is needed most, when a customer arrives with a material the model has never seen.

Selective-prediction score at 80% retention across four populations of decreasing chemical familiarity
Figure 3: The same score at 80% retention, each system on its own withheld set, over four populations of decreasing familiarity: every withheld structure, those whose combination of elements never occurred in training, those built from the twelve species contributing the fewest atoms to training, and those that are both. The split withholds whole materials, so every withheld structure is already a material the model never saw; composition is the axis that separates further.
Figure 3: The same score at 80% retention, each system on its own withheld set, over four populations of decreasing familiarity: every withheld structure, those whose combination of elements never occurred in training, those built from the twelve species contributing the fewest atoms to training, and those that are both. The split withholds whole materials, so every withheld structure is already a material the model never saw; composition is the axis that separates further.

Difficulty by element

Figure 4 tests whether Figure 1’s ranking is right about specific chemistry, not only on average. Each element gets one tile split along its diagonal, with predicted difficulty on one side and realized error on the other. A tile whose halves match is an element the model predicted correctly, and most of the table matches. The recovered ranking is the one a chemist would expect. What the bar does not track is how much of each element the model was trained on, which is what lets one calibration hold across chemistries instead of a refit for each.

Periodic table with each element tile split along its diagonal, predicted uncertainty against measured force error
Figure 4: Force error by element on withheld structures, normalized by each element's own RMS reference force so that the map reflects model difficulty rather than bond stiffness. Each tile is split on its diagonal, with the predicted spread in the upper-left half and the measured error in the lower-right, both colored by rank within their own half. Elements the withheld set does not cover are blank.
Figure 4: Force error by element on withheld structures, normalized by each element's own RMS reference force so that the map reflects model difficulty rather than bond stiffness. Each tile is split on its diagonal, with the predicted spread in the upper-left half and the measured error in the lower-right, both colored by rank within their own half. Elements the withheld set does not cover are blank.

From a ranking to a tolerance

Everything so far establishes a ranking, not whether any prediction satisfies a stated threshold, which is where most uncertainty methods stop. A ranking cannot be written into a specification. Figure 5 shows the value that quantitative predictions create. Setting a tolerance and discarding the predictions that fail leaves a smaller but more reliable workload. The horizontal axis is the share of predictions kept, the vertical axis the share of those that fall within tolerance. Moving left is a stricter setting, where fewer structures survive and more of the survivors are correct. Each curve is one tolerance swept across the risk levels the model accepts, and the three marked points are the bands we embed. Every band lands within a point of the coverage it promises on a split that never informed the cut, making the bands a concrete contract a workflow can be built against.

Coverage against retention, one curve per force-error tolerance, with the three trust bands marked
Figure 5: Measured coverage against retention, with one curve per tolerance and the three bands we provide marked. Tolerances are 0.10, 0.20 and 0.30 eV/Å on a structure's mean force error, each fitted on a calibration split and scored on a split that never informed it. The strict band delivers 90.5% against a promised 90%.
Figure 5: Measured coverage against retention, with one curve per tolerance and the three bands we provide marked. Tolerances are 0.10, 0.20 and 0.30 eV/Å on a structure's mean force error, each fitted on a calibration split and scored on a split that never informed it. The strict band delivers 90.5% against a promised 90%.

Transfer to computed properties

Every result above is measured on withheld MatPES structures, from the same distribution the model trained on. The MatCalc-Benchmark suite goes beyond this to calculate emergent material properties, requesting elastic constants, phonons, formation energies, and relaxed structures, none of them a training target or an input to the calibration. Applying the same trust bands here tests whether filtering on a force interval also improves downstream computed properties.

All five properties improve, with the elastic constants and the relaxed structure moving furthest. The model selects each subset without access to a reference value, which is what matters for deployment: a band fitted on force error, never having seen a property, identifies the structures whose computed properties are wrong.

MatCalc-Benchmark results for five properties, each with an arrow from the full set to the subset the trust band accepts
Figure 6: MatCalc-Benchmark results for five properties, each on its own scale and oriented so that better is to the right, with one glyph per comparator across all five lanes. The suite is scored twice, once over every structure it requests and once over the subset the band accepts, with the arrow running between them under the same weights and protocol. The retention under each property is that band's.
Figure 6: MatCalc-Benchmark results for five properties, each on its own scale and oriented so that better is to the right, with one glyph per comparator across all five lanes. The suite is scored twice, once over every structure it requests and once over the subset the band accepts, with the arrow running between them under the same weights and protocol. The retention under each property is that band's.

Figure 7 sets all five properties against the four strongest comparators on one shape, where a uniformly strong model traces a wide regular outline rather than a spike. Across the seven models compared, PHIN-atomic on its target band ranks first on average and never falls below third on any individual property. Most importantly, no property collapses, which is what allows a potential to be adopted across a campaign.

Radar chart of five benchmark properties comparing PHIN-atomic with the four strongest comparators
Figure 7: The five properties on one shape against the four strongest comparators by mean rank, with each spoke scaled between the best and the worst value the field records on that property. Ours is the target band, scored on the share of each task it retains, and the comparators are on their full sets.
Figure 7: The five properties on one shape against the four strongest comparators by mean rank, with each spoke scaled between the best and the worst value the field records on that property. Ours is the target band, scored on the share of each task it retains, and the comparators are on their full sets.

Perspective

A per-prediction error bar lets a workflow act on its own reliability without supervision. A screening campaign directs its uncertain fraction to DFT and accepts the remainder, and an active learning loop labels only the structures the model cannot resolve. Both need uncertainty predictions at every step, which negligible overhead makes practical.

Within industrial R&D, this shift makes atomistic simulation accessible without requiring a specialist to audit every output. A scientist simply defines the error a specific decision can absorb, and the system returns only the subset of predictions meeting that standard. The determination of which results to trust thus migrates from manual expert oversight directly into the simulation workflow, enabling digital modeling to function as a reliable peer to physical measurements.

†MACE-MatPES-PBE-0 is OMAT-pretrained and MatPES-finetuned rather than trained on MatPES from scratch, so it is not a peer of the other five comparators; we run it in-house on the same suite at the same version.

References

[1] Kaplan, A. D.; Liu, R.; Qi, J.; Ko, T. W.; Deng, B.; Riebesell, J.; Ceder, G.; Persson, K. A.; Ong, S. P. A Foundational Potential Energy Surface Dataset for Materials (2025). arXiv:2503.04070.

[2] Liu, R.; Liu, E.; Riebesell, J.; Rosen, A.; Qi, J.; Ong, S. P.; Ko, T. W. MatCalc, v0.5.1 (2026). github.com/materialyzeai/matcalc.

[3] Zhou, Y.; Hu, S.; Zhang, X.; Wang, H.; Tan, G.; Jia, W. MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interatomic Potentials. ICLR (2026). arXiv:2603.02002.

[4] ACEsuit. MACE-MATPES-PBE-0 (2025). github.com/ACEsuit/mace-foundations, release mace_matpes_0.

[5] MatPES. Benchmarks (retrieved 2026-08-19). matpes.ai/benchmarks.

[6] Kavanagh, S. R. et al. Fast and Accurate Equivariant Foundation Models for Atomistic Simulation, v2 (2026). arXiv:2607.28461.