Learning smooth, chiral 3D molecular descriptors from atomistic foundation models
- 1Cavendish Laboratory, Department of Physics, University of Cambridge
- 2Max Planck Institute for Polymer Research, Mainz
- 3Department of Materials, Imperial College London
arXiv preprint, 2026
Key contributions
- Foundation-model features reused as descriptors. A frozen machine-learned interatomic potential already encodes chemically rich local environments. Rem3Di contracts its per-atom features into a single fixed-length molecular descriptor — no simulation, no sampling, and no classical 2D fingerprints.
- Chirality without hand-crafted rules. Two successive tensor products build pseudoscalar channels: invariant under rotation, but sign-flipping under mirror reflection. This lets a single descriptor distinguish enantiomers, which invariant features provably cannot.
- Pretraining with no experimental labels. A self-supervised denoising objective reconstructs clean atom features from corrupted ones, so the descriptor can be pretrained on large unlabelled molecular datasets where property labels are scarce.
- Strong across benchmarks, and beyond organic chemistry. Fine-tuned Rem3Di is the best matched-protocol model on all six TDC/MoleculeNet regression endpoints, and the same descriptor organises transition-metal complexes by metal centre, ligand chemistry and geometry without any predefined bonding rules.

Abstract
Foundation machine-learned interatomic potentials (MLIPs) are trained on large quantum-mechanical datasets and generalise across broad regions of chemical and configurational space. Beyond their usual role in accelerating sampling-based simulations, their internal representations encode chemically rich local atomic environments. Here, we introduce Rem3Di, a representation-learning framework that repurposes latent features from atomistic foundation models as transferable molecular descriptors for property prediction and virtual screening. Rem3Di combines a potential’s per-atom features into a single fixed-length descriptor of the whole molecule that varies smoothly with three-dimensional structure and is invariant to the ordering of the atoms. The descriptor can be used directly or fine-tuned for specific prediction tasks.
To capture molecular handedness, Rem3Di constructs pseudoscalar features, which are unchanged by rotation but reverse sign under mirror reflection. This lets the descriptor distinguish enantiomers, which can differ in activity and toxicity. The transformer is pretrained on large molecular datasets by reconstructing corrupted atom features, so no experimental labels are required. Across public drug-property benchmarks, Rem3Di matches or exceeds published baselines without relying on classical 2D fingerprints. Additionally, the same descriptor yields chemically meaningful differentiation of transition-metal complexes without predefined bonding rules or handcrafted representations. Rem3Di therefore provides a route from simulation-trained atomistic representations to transferable, chirality-aware molecular representations for chemical machine learning.
Results
Denoising pretraining structures the representation
Pretraining with the self-supervised denoising objective — no labels at all — already organises chemical space. Molecules within a cluster share functional groups and differ only in where substituents sit around the ring, and conformers of the same molecule taken from short molecular-dynamics runs land in the same region, so the descriptor is stable against small geometric perturbations. Downstream accuracy keeps improving with the amount of unlabelled data used for pretraining.

Chirality: predicting the sign of optical rotation
A chiral molecule rotates polarised light, and its two enantiomers rotate it by equal amounts in opposite directions — so the sign of the optical rotation (OR-sign) is a clean, controlled test of whether a descriptor encodes handedness. On a Bemis–Murcko scaffold split, where 2D-fingerprint and descriptor baselines fall close to chance, Rem3Di’s chiral encoder reaches 0.77 on R/S and 0.70 on OR-sign. The confusion matrices show this is genuine discrimination rather than class bias: ECFP predicts one sign preferentially, Rem3Di predicts both at similar rates.

Drug-property benchmarks
Under a matched evaluation protocol — same splits, same regression heads — fine-tuned Rem3Di is the best model on all six regression endpoints and on four of the six TDC classification endpoints, and the single best entry overall on seven of the twelve. The advantage is largest on tasks governed by 3D physics, most clearly the solvation free energy FreeSolv, where Rem3Di improves on every 2D and graph baseline and on its own frozen descriptor (RMSE 1.07 → 0.82).
| Model | BBB | HIA | Pgp | Bioav. | Tox-Avg | CYP-Avg |
|---|---|---|---|---|---|---|
| D-MPNN / Chemprop | 0.864.010 | 0.976.004 | 0.889.005 | 0.617.050 | 0.821.019 | 0.819.004 |
| AttentiveFP | 0.855.011 | 0.974.007 | 0.892.012 | 0.632.039 | 0.842.010 | 0.749.008 |
| DeepMol | 0.774.023 | 0.880.012 | 0.821.007 | 0.509.026 | 0.735.015 | 0.770.008 |
| FPGNN | 0.888.018 | 0.958.012 | 0.930.007 | 0.666.035 | 0.860.017 | 0.866.004 |
| TranFoxMol | 0.868.019 | 0.951.036 | 0.875.011 | 0.619.019 | 0.837.017 | 0.860.006 |
| ECFP4 + LightGBM | 0.884.005 | 0.905.048 | 0.907.013 | 0.591.036 | 0.829.013 | 0.866 |
| UniMol2 (frozen) | 0.884.014 | 0.902.088† | 0.898.017 | 0.617.041 | 0.833.020† | 0.867 |
| MuMo (frozen) | 0.893.019 | 0.980.012† | 0.891.006 | 0.626.048 | 0.812.012† | 0.852 |
| Rem3Di (frozen) | 0.894.004 | 0.949.019† | 0.898.012 | 0.558.057 | 0.863.013 | 0.862 |
| Rem3Di (fine-tuned) | 0.898.007 | 0.974.008† | 0.911.006 | 0.561.018 | 0.871.016 | 0.867 |
| Model | LD50 | Caco-2 | PPBR | LIPO | ESOL | FreeSolv |
|---|---|---|---|---|---|---|
| D-MPNN / Chemprop | 0.607.022 | 0.388.077 | 8.158.314 | 0.448.014 | 1.050.008 | 2.082.082 |
| AttentiveFP | 0.678.012 | 0.401.032 | 9.373.335 | 0.572.007 | 0.877.029 | 2.073.183 |
| DeepMol | 0.589.006 | 0.327.012 | 9.533.162 | 0.660.004 | – | – |
| GROVER | – | – | – | – | 0.831.120 | 1.544.397 |
| MolCLR | – | – | – | – | 1.271.040 | 2.594.249 |
| GraphMVP | – | – | – | – | 1.029.033 | – |
| GEM | – | – | – | – | 0.798.029 | 1.877.094 |
| Uni-Mol | – | – | – | – | 0.788.029 | 1.480.048 |
| FPGNN | 0.638.024 | 0.326.040 | 8.4651.709 | 0.544.011 | 0.658.006 | 1.106.195 |
| TranFoxMol | 0.645.036 | 0.487.068 | 9.055.523 | 0.525.024 | 0.930.261 | 1.225.155 |
| ECFP4 + LightGBM | 0.659.005 | 0.477.011 | 9.646.122 | 0.627.014 | 1.069.124 | 2.178.400 |
| UniMol2 (frozen) | 0.682.012 | 0.338.018 | 8.985.244 | 0.602.017 | 0.655.065 | 1.169.144 |
| MuMo (frozen) | 0.693.030 | 0.353.022 | 8.265.287 | 0.677.008 | 0.647.039 | 1.243.111 |
| Rem3Di (frozen) | 0.656.015 | 0.368.022 | 8.536.313 | 0.625.006 | 0.655.071 | 1.069.205 |
| Rem3Di (fine-tuned) | 0.638.031 | 0.326.016 | 8.054.237 | 0.499.015 | 0.594.057 | 0.820.177 |
Cells are meanstd over seeds (our rows: 3 seeds; external baselines as published). Within each block, external published baselines (above the rule) and matched-protocol evaluations (below) are separated. In each column the best and second-best rankable values are shaded; ties share a rank. “–”: endpoint not reported for that model. †: coverage < 90% (shown but excluded from ranking). The complete leaderboard, including the MoleculeNet classification suite, is in Appendix E2 of the paper.
Ablations
Both pretraining and fine-tuning are needed for competitive performance, and the choice of frozen foundation model matters a lot. A parameter-free mean-aggregation baseline already scores 0.63 — matching the pretrained frozen descriptor — which supports the starting hypothesis that MLIPs learn chemically rich features usable outside simulation. The gap from 0.63 to 0.96 is what the learned aggregation contributes.

Transition-metal complexes
Conventional graph-based descriptors are brittle for transition-metal complexes, where coordination and organometallic bonding involve delocalised, polycentric metal–ligand interactions and the covalent graph is ambiguous. Rem3Di needs no bonding rules, so it applies unchanged. After label-free pretraining on tmQM, the descriptor space organises itself by metal centre, ligand chemistry and coordination geometry — which makes similarity search, QSAR regression and virtual screening available for metallodrug space through the same descriptor evaluated everywhere else here.

Citation
1
2
3
4
5
6
7
8
@article{wedig2026rem3di,
title = {Rem3Di: Learning smooth, chiral 3D molecular descriptors
from atomistic foundation models},
author = {Wedig, Steffen and Burton, Felix and Elijo{\v{s}}ius, Rokas
and Schran, Christoph and Schaaf, Lars L.},
journal = {arXiv preprint arXiv:2607.19977},
year = {2026}
}