Abstract
Formulation development still consumes scarce drug substance, months of stability storage and a great deal of expert intuition. Machine learning (ML) offers a different economy, one in which experiments are chosen for the information they return rather than for their place in a factorial grid. This review traces that shift across three stages of maturity. The first is descriptor-based prediction, where classical equations such as the General Solubility Equation and ESOL gave way to fingerprint, graph and physics-informed models trained on curated sets such as AqSolDB. The evidence examined here shows that aqueous solubility models have reached a floor set largely by the reproducibility of the measurements themselves, which ranges from about 0.15 log unit for standardised potentiometric methods to roughly 1 log unit across heterogeneous literature. The second stage is stability prediction, covering amorphous solid dispersions, glass-forming ability, co-crystal formation, excipient compatibility and protein formulation, where classifiers reach 80 to 97% accuracy on in vitro endpoints but rarely report applicability domains. The third stage closes the loop: Bayesian optimisation and active learning coupled to robotic platforms, which have located high-solubility curcumin formulations after testing about 3% of a 7,776-member design space and produced in-specification tablets within six hours of material characterisation. Only in vitro and in silico evidence is considered; studies relying on human or animal data were excluded by design. The review closes with a critical appraisal of data quality, validation practice and regulatory credibility, drawing a lesson from the correction of a prominent autonomous-synthesis claim, and sets out what is needed before autonomous formulation laboratories can be trusted with regulated products.