Abstract
Formulation development still consumes a disproportionate share of the time and material spent on a drug candidate, largely because the relationships linking molecular structure, excipient choice, composition and process to product performance are nonlinear and only partly understood. Machine learning (ML) has entered this space along three fronts: numerical representation of molecules and formulations, prediction of developability properties such as solubility and amorphous stability, and, most recently, closed-loop experimentation in which a model chooses the next experiment and a robotic platform performs it. This review traces that progression using only in silico, physicochemical and in vitro evidence, and deliberately excludes studies whose conclusions rest on human or animal data. Across the literature, predictive accuracy for well-posed physicochemical tasks now approaches experimental reproducibility, and dosage-form models trained on a few hundred to several thousand curated records reach useful accuracy for solid dispersions, cyclodextrin complexes, long-acting injectables, printed dosage forms and spectroscopic dissolution testing. The dominant constraint, however, is no longer the algorithm. Formulation data are scarce, heterogeneous, mined from publications that rarely report failure, and routinely evaluated with splits that overstate generalisation. Bayesian optimisation and self-driving laboratories matter chiefly because they change how data are produced: reported campaigns reached target formulations after sampling a few per cent of a design space, or with less than half the experiments of conventional design of experiments. The review argues that the next advance depends on system-level representations of the formulation, standardised and FAIR data including negative results, honest uncertainty estimates, and validation practices aligned with the credibility frameworks now emerging from regulators.