Skip to content

Benchmark compatibility

Two families of metrics, so that a result here can be read next to results elsewhere.

Crystal-generation quality: the LeMat-GenBench families

LeMat-GenBench (Siron et al., 2025) evaluates crystal generative models in eight families. meidnet score reports the same families, from a folder of CIF files produced by any model, with the definitions below. Where this implementation differs from LeMat-GenBench, the table says so; for leaderboard-comparable numbers, run LeMat-GenBench itself.

family LeMat-GenBench meidnet score
validity charge neutrality, minimum interatomic distance, coordination environment, physical plausibility charge neutrality (an oxidation-state assignment that sums to zero), closest pair ≥ max(0.8 Å, 0.6 × the sum of the atomic radii), density 0.5–25 g/cm³, cell lengths 1–60 Å, angles 10–170°. Coordination environments are not checked.
uniqueness BAWL fingerprints or StructureMatcher within the set pymatgen StructureMatcher (ltol 0.2, stol 0.3, angle_tol 5°, primitive cells)
novelty fraction not in LeMat-Bulk (BAWL or StructureMatcher) against a reference you name (--reference: the Perov-5 split, or any CSV with cif/formula columns): by composition (reduced formula) and by structure (StructureMatcher against the reference entries of the same composition)
diversity Vendi scores and Shannon entropy of elements, space groups, site numbers, physical size Shannon entropy (bits) and counts of the elements, space groups (spglib, symprec 0.1) and site numbers; the density's mean and spread
distribution JSD of categorical properties, MMD of volume and density, Fréchet distance of MLIP embeddings Jensen–Shannon distance (base 2) of the element frequencies and of the site numbers against the reference. MMD and Fréchet distances are not computed.
stability energy above the convex hull from several MLIPs (stable ≤ 0, metastable ≤ 0.1 eV/atom) with --mlip: relaxation with MACE-MP and the formation energy per atom against elemental reference phases (meidnet screen's proxy), stable if ≤ 0.1 eV/atom. This is not the energy above the hull.
hhi production and reserve supply risk not computed
sun stable ∧ unique ∧ novel; MetaSUN stable ∧ unique ∧ novel by structure, with --mlip and a reference with CIFs

Conditional inverse-design quality: the MEIDNet extension

A conditional generator is asked for a property; these metrics say whether it delivered. They need targets.csv with a file column (the CIF file name) and, per property, a point target <p>_target, a window <p>_min / <p>_max, or both. A bound such as "formation energy at most 1.0 eV/atom" is a window with only <p>_max = 1.0, not a point target. An optional <p>_value holds the value the submitter reports for the structure, with a source column saying how it was obtained (dft, experiment, predicted). Without a value, --model <checkpoint> predicts it with a MEIDNet model, and the report labels those values model-predicted.

file,dir_gap_target,dir_gap_min,dir_gap_max,dir_gap_value,heat_all_max,heat_all_value,source
cand-001.cif,1.5,1.2,1.8,1.47,1.0,-0.62,dft
metric definition
target success rate share of structures whose value lies inside the window when one is given, else with |value − target| ≤ tolerance, per property; the tolerance is --tolerance p=… or 5 % of the reference's range
target error mean |value − target|, or the distance outside the window for a property with a window and no point target
multi-property success share with every targeted property within tolerance
constraint success valid structures (the validity family) among the successes
conditional diversity distinct compositions among the successes, over the successes
target coverage distinct target vectors with at least one success, over the distinct targets
interpolation share targets inside the reference's property range (needs property columns in the reference)

Oracle efficiency, the number of candidates evaluated per success, is a property of a run, not of a set of structures: the Perov-5 protocol reports it as the candidate budget and the delivered count.

Running it

pip install "meidnet>=2.3.1"                     # or: pip install "meidnet[stability]" for --mlip
meidnet download-data                            # Perov-5 as a reference (data/perov5/)
meidnet score generated/ --reference data/perov5
meidnet score generated/ --reference data/perov5 --targets targets.csv --tolerance dir_gap=0.3
meidnet score generated/ --reference data/perov5 --mlip --out report.json

generated/ is a folder of CIF files (any depth) or a CSV with a cif column. The report prints as a table; --out writes the JSON with the per-structure checks. A run bundle from MEIDNet Matter contains cifs/ and targets.csv in this layout.

Example: the 26 candidate structures shipped with the paper (examples/perov5/paper_results) scored against the full Perov-5 set, without MLIP, in three seconds:

| validity     | valid (all checks)            | 100.0 %            |
| uniqueness   | unique structures             | 24 of 26 (92.3 %)  |
| novelty      | new compositions vs perov5    |  69.2 %            |
| novelty      | new structures vs perov5      |  69.2 %            |
| diversity    | elements (entropy)            | 25 (3.48 bits)     |
| diversity    | space groups (entropy)        | 1 (0.00 bits)      |
| distribution | element JSD vs reference      | 0.717              |

One space group, because every structure is a cubic ABX₃ prototype: the diversity family shows where a prototype-family generator sits next to a free-form one.

Numbers from the metric families go into a submission under results.generation; the contribute page has the record layout.