Theoretical and computational chemistry
Machine learning for property and reaction prediction
Learn how molecular representations, data splits, and uncertainty shape machine-learning predictions of chemical properties and reaction outcomes.
IntuitionLearn patterns, not chemical laws
A model maps a molecular representation to a target such as solubility, energy, or a reaction product. It learns correlations from examples; chemical meaning and reliability depend on what the data represent and whether new molecules resemble them.
SchoolA prediction is a mapping
Definition: Feature and target
A feature vector encodes a molecule or reaction; a target is the measured or computed quantity to predict. The learned function approximates the relation in the training distribution.
The loss measures prediction error and regularizes model complexity. Neither a small training loss nor a large neural network guarantees useful chemistry; validation must reflect the intended prediction task.
Example: Choose the right split
A model predicts properties for new scaffolds, but random molecule-level splitting puts close analogues in both train and test sets. What evaluation better tests scaffold generalization?
Solution
Use a scaffold-based holdout: keep entire structural families out of training. It is a harder but more relevant test for this deployment scenario.
UndergraduateMolecular representations and reaction tasks
| Representation | Encodes | Caution |
|---|---|---|
| Fingerprint | Local substructures | May miss geometry or context |
| Graph | Atoms and bonds | Needs careful handling of 3D and stereochemistry |
Reaction prediction requires representing reactants, conditions, and products. A model may rank candidate products or predict edits to a molecular graph. A high product score is conditional on the reaction data and conditions represented, not a guarantee of experimental yield.
AdvancedGeneralization, calibration, and chemical validity
Benchmark design should prevent leakage, separate near-duplicates where appropriate, and report metrics matched to the use case. Calibration asks whether stated confidence matches empirical correctness. Applicability-domain checks and uncertainty estimates help identify extrapolation, but cannot repair biased or incomplete labels.
ResearchFrontier: reliable chemistry from imperfect data
References
- Machine learning for molecular and materials science · Keith T. Butler, Daniel W. Davies, Hugh Cartwright, Olexandr Isayev, and Aron Walsh, 2018
- Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction · Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A. Hunter, Costas Bekas, and Alpha A. Lee, 2019