Abstract
Generative models now design molecules satisfying almost any computable objective, and the field's benchmarks record that success faithfully. Whether the success means anything is a separate question, answered here from three directions: the internal failure modes of the benchmarks, the prospective exercises that test predictions against experiment, and the clinical record of compounds that reached patients. Goal-directed generators score highly on tasks whose global optimum random sampling can reach, produce molecules that retrosynthesis software solves only 30.2% of the time when synthesizability is not explicitly rewarded, and exploit scoring functions so reliably that molecules ranked highest by an optimisation model score below 0.2 on an independently trained control. Beneath them lie biased datasets: a network trained on DUD-E reaches an AUC of 0.98 with the protein deleted at test time and falls to 0.53 once the decoy artefact is removed. Deep learning docking methods produce physically valid, accurate poses for 12% of a post-cutoff benchmark on which AutoDock Vina manages 58%, and cofolding accuracy falls to between 10% and 20% on ligands unlike the training data. Prospectively, 23 teams competing against the LRRK2 WD40 domain achieved a hit rate of 3.7% with affinities between 18 and 140 micromolar. The clinical record is more equivocal than either advocates or critics allow: AI-derived molecules completed phase 1 in 21 of 24 cases but only 4 of 10 phase 2 trials, the historical industry rate. Rentosertib raised forced vital capacity by 98.4 mL against a placebo decline of 20.3 mL in 71 patients and dosed its first phase 3 patient in September 2026, and no AI-designed small molecule has yet completed phase 3. Generative chemistry has largely solved the design of molecules satisfying stated objectives and has not solved the statement of objectives worth satisfying, which is why validation must gate on generalisation and prospective experiment rather than on optimisation score.