Guide · updated 2026-09-11

MLIPs vs DFT for catalyst screening

When a machine-learned interatomic potential is the right tool, when it is not, and why acidic OER is the hardest case for general-purpose models.

Density functional theory is the reference method for catalyst energetics, and it is far too slow to look at a composition space. A machine-learned interatomic potential is fast enough to look at a whole space, and is not a reference method. Knowing which job you are doing settles which tool you want.

The two jobs

Triage. You have forty compositions and time to synthesise three. You need a ranking that is approximately right, quickly, cheaply, across all forty. Being wrong about the exact overpotential of candidate seventeen costs you nothing if the ordering is sound.

Confirmation. You have three candidates and you need defensible numbers for a paper, a specification, or a filing. Here approximation is not acceptable and DFT — converged, documented, with stated functional and pseudopotentials — is the right instrument.

Most computational effort in catalyst work is spent doing the second job on candidates that never deserved it, because doing the first job properly was too expensive.

Why acidic OER is the hardest case for a general model

Foundation interatomic potentials are trained predominantly on bulk crystal data. Their accuracy degrades on surfaces, further on surfaces with adsorbates, and further still when those adsorbates are oxygen species on oxide surfaces.

Public benchmarks make this concrete: error on oxide-rich slabs with oxygen adsorbates runs substantially higher than on the metallic, carbon-and-hydrogen-adsorbate majority of the training distribution. The Open Catalyst 2022 dataset exists specifically because oxides are critical to oxygen-evolution catalysis and were underrepresented.

The regime where general models are weakest is precisely the regime PEM electrolysis lives in.

That is an argument for specialisation, not for abandoning the approach. Fine-tuning a foundation potential onto oxide and adsorbate data recovers a large part of the gap, and doing so for one chemistry is tractable for a small team in a way that training a general model is not.

What is still missing, honestly

Public training data concentrates on low-index surfaces, in vacuum, at sparse coverage. Under-represented: realistic coverages, stepped and defective surfaces, transition states, and — most importantly for electrochemistry — electrolyte environments and applied potential.

Any screening result should be read with that in mind. It is a vacuum-ish, idealised-surface estimate being used to rank candidates that will operate in acid, at potential, under load. It is useful for ranking. It is not a simulation of your cell.

How we use them

Activity and stability screening on verna.science runs on interatomic potentials, by design, and we say so on every result along with the model version. The method page states the validity limits, and runs outside the validated domain are refused rather than answered.

Questions

Can an MLIP replace DFT?

Not for a final number. It replaces DFT for the part of the work where you are deciding what deserves a final number — triage across a composition space, where being approximately right about the ranking is worth far more than being exactly right about one candidate.

Why are oxide surfaces the hard case?

Because most training data is bulk crystals. Public benchmarks show notably higher error on oxide-rich slabs with oxygen adsorbates than on metallic surfaces with carbon and hydrogen adsorbates — and oxide surfaces with oxygen adsorbates is exactly what acidic OER is.

Does fine-tuning actually help?

Measurably. Published results on the Open Catalyst benchmarks show substantial accuracy gains from fine-tuning onto oxide-specific data rather than relying on a bulk-trained general model.

Put a number on it

The LCOH calculator turns any of this into euros per kilogram. Free, no account.

Open the calculator