Project

Learning smooth, chiral 3D molecular descriptors from atomistic foundation models

Steffen Wedig1,2, Felix Burton1, Rokas Elijošius1,*, Christoph Schran1,*, Lars L. Schaaf1,3,*

  1. 1Cavendish Laboratory, Department of Physics, University of Cambridge
  2. 2Max Planck Institute for Polymer Research, Mainz
  3. 3Department of Materials, Imperial College London

arXiv preprint, 2026

Key contributions

  • Foundation-model features reused as descriptors. A frozen machine-learned interatomic potential already encodes chemically rich local environments. Rem3Di contracts its per-atom features into a single fixed-length molecular descriptor — no simulation, no sampling, and no classical 2D fingerprints.
  • Chirality without hand-crafted rules. Two successive tensor products build pseudoscalar channels: invariant under rotation, but sign-flipping under mirror reflection. This lets a single descriptor distinguish enantiomers, which invariant features provably cannot.
  • Pretraining with no experimental labels. A self-supervised denoising objective reconstructs clean atom features from corrupted ones, so the descriptor can be pretrained on large unlabelled molecular datasets where property labels are scarce.
  • Strong across benchmarks, and beyond organic chemistry. Fine-tuned Rem3Di is the best matched-protocol model on all six TDC/MoleculeNet regression endpoints, and the same descriptor organises transition-metal complexes by metal centre, ligand chemistry and geometry without any predefined bonding rules.
Pipeline diagram: a SMILES string becomes a 3D conformer, a foundation MLIP produces per-atom descriptors, and a learnable global aggregation contracts them into one molecular descriptor used for property prediction, similarity screening and retrieval.
Figure 1. Constructing 3D molecular descriptors with the Rem3Di framework. Atom-centred features from atomistic foundation models are combined into a fixed-dimension molecular descriptor that varies smoothly with 3D structure.

Abstract

Foundation machine-learned interatomic potentials (MLIPs) are trained on large quantum-mechanical datasets and generalise across broad regions of chemical and configurational space. Beyond their usual role in accelerating sampling-based simulations, their internal representations encode chemically rich local atomic environments. Here, we introduce Rem3Di, a representation-learning framework that repurposes latent features from atomistic foundation models as transferable molecular descriptors for property prediction and virtual screening. Rem3Di combines a potential’s per-atom features into a single fixed-length descriptor of the whole molecule that varies smoothly with three-dimensional structure and is invariant to the ordering of the atoms. The descriptor can be used directly or fine-tuned for specific prediction tasks.

To capture molecular handedness, Rem3Di constructs pseudoscalar features, which are unchanged by rotation but reverse sign under mirror reflection. This lets the descriptor distinguish enantiomers, which can differ in activity and toxicity. The transformer is pretrained on large molecular datasets by reconstructing corrupted atom features, so no experimental labels are required. Across public drug-property benchmarks, Rem3Di matches or exceeds published baselines without relying on classical 2D fingerprints. Additionally, the same descriptor yields chemically meaningful differentiation of transition-metal complexes without predefined bonding rules or handcrafted representations. Rem3Di therefore provides a route from simulation-trained atomistic representations to transferable, chirality-aware molecular representations for chemical machine learning.

Results

Denoising pretraining structures the representation

Pretraining with the self-supervised denoising objective — no labels at all — already organises chemical space. Molecules within a cluster share functional groups and differ only in where substituents sit around the ring, and conformers of the same molecule taken from short molecular-dynamics runs land in the same region, so the descriptor is stable against small geometric perturbations. Downstream accuracy keeps improving with the amount of unlabelled data used for pretraining.

Panel a: UMAP of pretrained molecular representations coloured by HOMO-LUMO gap, with insets showing chemically similar molecules clustered together. Panel b: relative minimum validation MSE falling as pretraining dataset size grows from 1e5 to 4e5.
Figure 4. Denoising pretraining. (a) UMAP of the resulting pretrained molecular representations, with insets showing that chemically similar molecules are clustered together. (b) Performance on a HOMO–LUMO gap regression task, on a 10,000-molecule subset of QM9, as a function of pretraining dataset size.

Chirality: predicting the sign of optical rotation

A chiral molecule rotates polarised light, and its two enantiomers rotate it by equal amounts in opposite directions — so the sign of the optical rotation (OR-sign) is a clean, controlled test of whether a descriptor encodes handedness. On a Bemis–Murcko scaffold split, where 2D-fingerprint and descriptor baselines fall close to chance, Rem3Di’s chiral encoder reaches 0.77 on R/S and 0.70 on OR-sign. The confusion matrices show this is genuine discrimination rather than class bias: ECFP predicts one sign preferentially, Rem3Di predicts both at similar rates.

Panel a: bar charts of test accuracy for six methods on R/S and OR-sign targets under random and scaffold splits. Panel b: row-normalised confusion matrices for OR-sign comparing ECFP and Rem3Di. Panel c: schematic contrasting random and scaffold train/test splits.
Figure 5. Chirality property prediction under random and scaffold splits on QM9-OR. (a) Test accuracy of six methods on two targets, R/S (handedness of chiral centres) and OR-sign (sign of the optical rotation at 589.3 nm), under a random split and a Bemis–Murcko scaffold split. The test set is mirror-balanced, so chance accuracy is 0.50 (dashed line). Grey hatched bars are the OHECC descriptor baselines of Zhou et al. on their own random split; coloured bars are the methods retrained on our splits. Error bars are ±1 s.d. over four seeds. (b) Row-normalised confusion matrices for OR-sign on the scaffold split, on single-stereocentre test molecules (n = 1229), for Rem3Di (balanced accuracy ≈ 0.70) and ECFP (≈ 0.56). (c) Schematic of the train/test split: a random split places members of one scaffold cluster in both train and test, leaking information into the test set, while a scaffold split reduces that leakage.

Drug-property benchmarks

Under a matched evaluation protocol — same splits, same regression heads — fine-tuned Rem3Di is the best model on all six regression endpoints and on four of the six TDC classification endpoints, and the single best entry overall on seven of the twelve. The advantage is largest on tasks governed by 3D physics, most clearly the solvation free energy FreeSolv, where Rem3Di improves on every 2D and graph baseline and on its own frozen descriptor (RMSE 1.07 → 0.82).

Classification, AUROC ↑
ModelBBBHIAPgpBioav.Tox-AvgCYP-Avg
D-MPNN / Chemprop0.864.0100.976.0040.889.0050.617.0500.821.0190.819.004
AttentiveFP0.855.0110.974.0070.892.0120.632.0390.842.0100.749.008
DeepMol0.774.0230.880.0120.821.0070.509.0260.735.0150.770.008
FPGNN0.888.0180.958.0120.930.0070.666.0350.860.0170.866.004
TranFoxMol0.868.0190.951.0360.875.0110.619.0190.837.0170.860.006
ECFP4 + LightGBM0.884.0050.905.0480.907.0130.591.0360.829.0130.866
UniMol2 (frozen)0.884.0140.902.0880.898.0170.617.0410.833.0200.867
MuMo (frozen)0.893.0190.980.0120.891.0060.626.0480.812.0120.852
Rem3Di (frozen)0.894.0040.949.0190.898.0120.558.0570.863.0130.862
Rem3Di (fine-tuned)0.898.0070.974.0080.911.0060.561.0180.871.0160.867
Regression, error ↓ — TDC scored by MAE, MoleculeNet by RMSE
ModelLD50Caco-2PPBRLIPOESOLFreeSolv
D-MPNN / Chemprop0.607.0220.388.0778.158.3140.448.0141.050.0082.082.082
AttentiveFP0.678.0120.401.0329.373.3350.572.0070.877.0292.073.183
DeepMol0.589.0060.327.0129.533.1620.660.004
GROVER0.831.1201.544.397
MolCLR1.271.0402.594.249
GraphMVP1.029.033
GEM0.798.0291.877.094
Uni-Mol0.788.0291.480.048
FPGNN0.638.0240.326.0408.4651.7090.544.0110.658.0061.106.195
TranFoxMol0.645.0360.487.0689.055.5230.525.0240.930.2611.225.155
ECFP4 + LightGBM0.659.0050.477.0119.646.1220.627.0141.069.1242.178.400
UniMol2 (frozen)0.682.0120.338.0188.985.2440.602.0170.655.0651.169.144
MuMo (frozen)0.693.0300.353.0228.265.2870.677.0080.647.0391.243.111
Rem3Di (frozen)0.656.0150.368.0228.536.3130.625.0060.655.0711.069.205
Rem3Di (fine-tuned)0.638.0310.326.0168.054.2370.499.0150.594.0570.820.177

Cells are meanstd over seeds (our rows: 3 seeds; external baselines as published). Within each block, external published baselines (above the rule) and matched-protocol evaluations (below) are separated. In each column the best and second-best rankable values are shaded; ties share a rank. “–”: endpoint not reported for that model. : coverage < 90% (shown but excluded from ranking). The complete leaderboard, including the MoleculeNet classification suite, is in Appendix E2 of the paper.

Ablations

Both pretraining and fine-tuning are needed for competitive performance, and the choice of frozen foundation model matters a lot. A parameter-free mean-aggregation baseline already scores 0.63 — matching the pretrained frozen descriptor — which supports the starting hypothesis that MLIPs learn chemically rich features usable outside simulation. The gap from 0.63 to 0.96 is what the learned aggregation contributes.

Line plots of relative score across 12 drug-property tasks. Left: random frozen 0.02, random full fine-tune 0.62, mean aggregation 0.63, pretrained frozen 0.63, pretrained LoRA 0.82, pretrained full fine-tune 0.96. Right: OrbNet 0.16, OFF24 0.56, POLAR 0.63.
Figure 6. Rem3Di architectural and training ablations. Relative performance across 12 drug-property tasks. Left: encoder initialisation (random or denoising-pretrained) crossed with the adaptation strategy (frozen descriptor or full fine-tune), together with a parameter-free mean-aggregation baseline. Right: the same comparison across three frozen foundation models. For each task the native metric is oriented so that higher is always better and then min–max normalised, where 0 is the worst and 1 the best of the conditions for that task. Grey lines show the 12 individual tasks; the thicker blue line is the mean, with the shaded band denoting ±1 s.d. across tasks.

Transition-metal complexes

Conventional graph-based descriptors are brittle for transition-metal complexes, where coordination and organometallic bonding involve delocalised, polycentric metal–ligand interactions and the covalent graph is ambiguous. Rem3Di needs no bonding rules, so it applies unchanged. After label-free pretraining on tmQM, the descriptor space organises itself by metal centre, ligand chemistry and coordination geometry — which makes similarity search, QSAR regression and virtual screening available for metallodrug space through the same descriptor evaluated everywhere else here.

UMAP of transition-metal complex descriptors with insets showing clusters for diimine chelate complexes in square-planar and octahedral geometries, oxygen donor-ligand complexes, titanium and iron centres, sandwich and half-sandwich complexes, phosphine ligands and carborane ligands.
Figure 7. Transition-metal complexes. UMAP of the molecular descriptors after denoising pretraining without ground-truth labels. Insets highlight the clustering of chemically similar compounds, which groups TMCs by geometry, ligand chemistry, and metal-centre identity.

Citation

1
2
3
4
5
6
7
8
@article{wedig2026rem3di,
  title   = {Rem3Di: Learning smooth, chiral 3D molecular descriptors
             from atomistic foundation models},
  author  = {Wedig, Steffen and Burton, Felix and Elijo{\v{s}}ius, Rokas
             and Schran, Christoph and Schaaf, Lars L.},
  journal = {arXiv preprint arXiv:2607.19977},
  year    = {2026}
}