PHIN-atomic: a foundation potential that reports its own error
Summary
We introduce PHIN-atomic, our first foundation machine learning interatomic potential (MLIP), trained on MatPES-PBE 1. Every energy and force it predicts produces a calibrated error bar, evaluated alongside physical outputs with negligible computational overhead.
PHIN-atomic’s uncertainty predictions, which improve over state-of-the-art, expensive ensembling methods, remove the jagged performance of pre-trained MLIPs and build trustable inference. We establish this on held-out structures, hold it on chemistry absent from training, and prove its value through downstream computed property prediction accuracy.
Background
Foundation potentials have changed the economics of atomistic simulation. Work that would have consumed a year of DFT allocation now runs in an afternoon, and the constraints have moved beyond generating results to deciding which results to trust. That decision typically cannot come from the model’s own outputs as a foundation potential evaluates any structure received, familiar or not, and gives no sign of which case it is. Without the physical constraints that bound an empirical force field, extrapolation error is unbounded, reaching the user as an unphysical result: a relaxation to the wrong basin, imaginary modes in a phonon spectrum, atom loss in a trajectory. A sweep of ten thousand relaxations contains some number of these, and with no per-prediction uncertainty the choice is to verify all of them, forfeiting the speedup, or none, forfeiting the trust.
An established remedy today is to train an ensemble from independent initializations and treat the disagreement as an uncertainty estimate. Language model serving does the same thing, sending a prompt to several models in parallel and having a judge reconcile their answers: OpenRouter built Fusion on exactly this, and its strongest panel outscored every individual model in it on their deep-research evaluation, because it purchases several completions instead of one. A deep ensemble buys reliability the same way, multiplying training, storage, and every force evaluation by the number of members, and its estimate remains overconfident and loosely correlated with the error it predicts. Alternatively, uncertainty estimation via latent-space distance to the training set is cheaper to evaluate, but it measures distance from the training distribution rather than predictive error. Here, we use a single model backbone and small uncertainty layer to predict the calibrated error directly, which enables trust bands around each prediction, set against a tolerance in physical units. A scientist states the error a workflow can absorb, and the model returns the subset that meets it.
Experimental Method
The uncertainty estimate is produced in the same forward pass as the energy and the forces, and calibrated on data withheld during training. The training, validation, and test datasets are segmented to withhold whole materials rather than scattered structures, so no material appears in both training and evaluation. The uncertainty predicts a distribution, which we simplify into three trust bands. Each of the three trust bands pairs a tolerance on a structure’s mean force error with how often predictions may exceed it. We compare this to the best alternative, ensembling, by training 3 further copies from independent initializations and evaluate all four on the 1,064 structures withheld by every one. We evaluate the model and uncertainty performance by calculating emergent properties using MatCalc-Bench 1,2. In addition, we compare our model to published results from five models trained from scratch on MatPES-PBE 3,5,6 and also to MACE-MatPES-PBE-0† 4, an OMAT-pretrained model fine-tuned on MatPES-PBE, benchmarked in-house.
Results
Ranking the errors
An error bar is useful only if its width varies with the error it bounds. A bar of constant width reproduces the correct coverage while indicating nothing about which prediction to distrust, leaving a workflow with no basis for discriminating results or spending budget on additional verification.
Figure 1 tests the ranking of each uncertainty metric. Every strip holds the same withheld structures, reordered by one metric and shaded by the force error measured afterwards against DFT. A metric that carries information drives the dark bands toward its low-confidence end, taking the strip from mixed to graded. A model generating its own confidence ranks its errors more effectively than three independently trained models’ disagreement. Latent-space distance recovers a small fraction of what either achieves. The ranking is what makes selective prediction possible, deciding how much error a fixed verification budget removes.
One network against a deep ensemble
The deep ensemble is the standard a new method has to beat. Figure 2 compares them directly by putting the two on the same axis. Predictions are removed in order of confidence, from the full set at the left edge to the most trusted half at the right, so a steeply falling curve indicates a model that removed the right structures first. The ensemble is the more accurate average predictor but the poorer judge of its own reliability, and PHIN-atomic’s own confidence identifies its errors better than three networks’ disagreement. At matched coverage the ensemble interval must be 17% wider for an equivalent claim. Predictive accuracy and the quality of the uncertainty estimate are separable properties, and it is the second that is critical for simulation workflows that drive real scientific decisions.
Moreover, the overhead is small enough to calculate the uncertainty throughout inference. Running molecular dynamics in LAMMPS on a 4,096-atom cell, a step costs 9.7 microseconds per atom with an error bar on every atom against 16.5 for MACE-MatPES-PBE-0 at the same precision, and no error bar at all.
Chemistry absent from training
The ranking in Figure 1 can be produced without generalizing at all, by learning which elements proved difficult in training and widening the bar whenever one appears. Such a model is most confident exactly where a user needs it least, on the novel chemistry a screening campaign exists to explore. Figure 3 rules that out by showing that as the chemistry becomes less familiar, the predicted error responds to the local configuration rather than to element identity. Compare this to latent-space distance, which loses up to half of its performance on exactly the structures the estimate exists to cover. PHIN’s single model uncertainty therefore holds on the most important structures and when it is needed most, when a customer arrives with a material the model has never seen.
Difficulty by element
Figure 4 tests whether Figure 1’s ranking is right about specific chemistry, not only on average. Each element gets one tile split along its diagonal, with predicted difficulty on one side and realized error on the other. A tile whose halves match is an element the model predicted correctly, and most of the table matches. The recovered ranking is the one a chemist would expect. What the bar does not track is how much of each element the model was trained on, which is what lets one calibration hold across chemistries instead of a refit for each.
From a ranking to a tolerance
Everything so far establishes a ranking, not whether any prediction satisfies a stated threshold, which is where most uncertainty methods stop. A ranking cannot be written into a specification. Figure 5 shows the value that quantitative predictions create. Setting a tolerance and discarding the predictions that fail leaves a smaller but more reliable workload. The horizontal axis is the share of predictions kept, the vertical axis the share of those that fall within tolerance. Moving left is a stricter setting, where fewer structures survive and more of the survivors are correct. Each curve is one tolerance swept across the risk levels the model accepts, and the three marked points are the bands we embed. Every band lands within a point of the coverage it promises on a split that never informed the cut, making the bands a concrete contract a workflow can be built against.
Transfer to computed properties
Every result above is measured on withheld MatPES structures, from the same distribution the model trained on. The MatCalc-Benchmark suite goes beyond this to calculate emergent material properties, requesting elastic constants, phonons, formation energies, and relaxed structures, none of them a training target or an input to the calibration. Applying the same trust bands here tests whether filtering on a force interval also improves downstream computed properties.
All five properties improve, with the elastic constants and the relaxed structure moving furthest. The model selects each subset without access to a reference value, which is what matters for deployment: a band fitted on force error, never having seen a property, identifies the structures whose computed properties are wrong.
Figure 7 sets all five properties against the four strongest comparators on one shape, where a uniformly strong model traces a wide regular outline rather than a spike. Across the seven models compared, PHIN-atomic on its target band ranks first on average and never falls below third on any individual property. Most importantly, no property collapses, which is what allows a potential to be adopted across a campaign.
Perspective
A per-prediction error bar lets a workflow act on its own reliability without supervision. A screening campaign directs its uncertain fraction to DFT and accepts the remainder, and an active learning loop labels only the structures the model cannot resolve. Both need uncertainty predictions at every step, which negligible overhead makes practical.
Within industrial R&D, this shift makes atomistic simulation accessible without requiring a specialist to audit every output. A scientist simply defines the error a specific decision can absorb, and the system returns only the subset of predictions meeting that standard. The determination of which results to trust thus migrates from manual expert oversight directly into the simulation workflow, enabling digital modeling to function as a reliable peer to physical measurements.
†MACE-MatPES-PBE-0 is OMAT-pretrained and MatPES-finetuned rather than trained on MatPES from scratch, so it is not a peer of the other five comparators; we run it in-house on the same suite at the same version.
References
[1] Kaplan, A. D.; Liu, R.; Qi, J.; Ko, T. W.; Deng, B.; Riebesell, J.; Ceder, G.; Persson, K. A.; Ong, S. P. A Foundational Potential Energy Surface Dataset for Materials (2025). arXiv:2503.04070.
[2] Liu, R.; Liu, E.; Riebesell, J.; Rosen, A.; Qi, J.; Ong, S. P.; Ko, T. W. MatCalc, v0.5.1 (2026). github.com/materialyzeai/matcalc.
[3] Zhou, Y.; Hu, S.; Zhang, X.; Wang, H.; Tan, G.; Jia, W. MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interatomic Potentials. ICLR (2026). arXiv:2603.02002.
[4] ACEsuit. MACE-MATPES-PBE-0 (2025). github.com/ACEsuit/mace-foundations, release mace_matpes_0.
[5] MatPES. Benchmarks (retrieved 2026-08-19). matpes.ai/benchmarks.
[6] Kavanagh, S. R. et al. Fast and Accurate Equivariant Foundation Models for Atomistic Simulation, v2 (2026). arXiv:2607.28461.