Abstract
The formulation of a drug product depends on a dense web of physicochemical and performance properties, from aqueous solubility and dissolution to the physical stability of an amorphous dispersion and the particle size of a nanosuspension. Measuring each property empirically for every candidate and excipient combination is slow and consumes scarce active material, and this cost is what has drawn quantitative structure-property relationship modelling and, more recently, machine learning into formulation science. This review examines how artificial intelligence is now used to predict formulation-relevant properties, and it argues that the discipline is limited less by the choice of learning algorithm than by three underdeveloped foundations: the representation of the formulation rather than the molecule alone, the scale and quality of the underlying data, and the rigour of validation. The prediction targets are surveyed by property class, spanning solubility, amorphous solid dispersion stability, tablet mechanical behaviour, nanoparticle attributes, and permeability, with the reported models and their measured performance placed side by side. The representation problem is then treated directly, comparing expert descriptors, structural fingerprints, and learned graph or sequence representations, and noting where each fails when the input is a mixture and a process rather than a single structure. The validation section sets out what a trustworthy model must demonstrate, drawing on the established principles for structure-property models: a defined applicability domain, honest external testing, and metrics that reflect real predictivity rather than fitted correlation. The synthesis is that formulation artificial intelligence will become decision-grade only when shared, standardised datasets and applicability-domain-aware, externally validated, interpretable models replace the single-laboratory, internally validated demonstrations that dominate the current literature.