🔑 Key Takeaway
A 2025 Pharmaceutical Research study trained several machine learning classifiers on particle size, shape, surface energy, and silica dry-coating parameters from 410 formulations comprising 89 single components and 321 blends. At the coarsest two-category flowability classification, the best-performing models reached about 85 percent accuracy for individual components and 87 percent for blends. The training labels came from shear cell Flow Function Coefficient values, so the model compresses an existing physical test into a faster screening step rather than replacing it. Accuracy dropped sharply for intermediate flow categories, coated systems, and APIs outside the training set, which marks where shear cell or dynamic powder testing is still needed before a process decision is made.

Formulation teams screening new API-excipient combinations often want an early read on flow behavior before committing bench time to a full shear cell or dynamic powder testing panel. A blend that turns out cohesive after tableting trials are already underway costs more than one flagged earlier, during formulation selection.
A 2025 study in Pharmaceutical Research tested whether that early read could come from a machine learning flowability prediction built on the physical properties of the individual API and excipient particles, rather than from testing the blend itself. The authors trained Random Forest, XGBoost, Support Vector Machine, and other classifiers on a dataset of 410 formulations comprising 89 single components and 321 blends drawn from 9 APIs and 18 excipients with varying silica dry-coating levels. At the coarsest two-category flow classification, the best-performing models reached accuracy in the mid-80 percent range. Where that accuracy held, and where it didn’t, is the more useful part of the result for anyone deciding how far to trust the output.
What the Model Was Actually Trained On
The dataset combined 89 single components and 321 blends built from 9 APIs and 18 excipients treated with different levels of silica dry-coating. Most physical-property inputs were measured at constituent level rather than on the finished blend: particle size metrics (d10, d50, d90, and Sauter mean diameter), particle density, dispersive surface energy, and shape descriptors including aspect ratio, sphericity, and elongation. The blend models also included formulation variables such as API loading and coating-related parameters including silica treatment and surface area coverage. Importantly, the models did not require a measured blend flowability value as an input; that shear cell result served as the target the models were trained to predict.
The flowability labels the model was trained to predict did not come from a survey or a simple visual assessment. They came from shear cell testing on an FT4-type powder rheometer, which measures the Flow Function Coefficient across a range of normal stresses. The study grouped FFC values into flow regimes following a Schulze-type scale, from very cohesive to free-flowing, tested at two, three, four, and five-category resolution. Because the ground truth is itself a physical shear test result, the machine learning model functions as a fast proxy for a test that has already been run on the training set, not as an independent measurement of blend behavior.
Why Constituent Properties Carry Predictive Signal
Dry-coating a particle surface with nanosilica is a common way to reduce interparticle cohesion, since the silica acts as a physical spacer that limits the contact area and strength of the van der Waals attraction between primary particles. That mechanism is established powder science on its own. What the study adds is a ranking: SHAP and feature-importance analysis identified coating-related parameters as the single most influential feature category for classifying single-component flow behavior, ahead of particle size distribution, particle density, and shape descriptors, with elongation standing out among the shape variables.
For blends specifically, the feature hierarchy shifted. API coating percentage and API loading, the fraction of the blend made up by the API, ranked above excipient particle size. In this dataset, that suggests when a coated, potentially more cohesive API dominates blend composition, its own surface treatment does more to set the blend’s flow category than the properties of the excipients around it. That is a finding specific to this dataset and this set of materials, not a general rule for every API-excipient system.
Where Machine Learning Flowability Prediction Holds Up and Where It Breaks Down
At the coarsest split in this study’s scheme, not-well-flowing (FFC below 6) versus well-flowing (FFC of 6 or above), the best-performing models reached about 85 percent accuracy for individual components and up to 87 percent for blends. That is a two-category call, closer to a pass/fail screen than a fine-grained flow assessment, and the accuracy figure applies specifically to that resolution rather than to flowability prediction in general.
Accuracy declined as the number of flow categories increased toward the five-regime scheme, and the models had particular difficulty separating intermediate categories such as cohesive from semi-cohesive. They performed better at distinguishing the extremes, clearly free-flowing from clearly very cohesive material. Accuracy also dropped when dry-coated powders were introduced compared with uncoated-only training data, and it dropped further when the model was tested on an API withheld from training, falling from roughly 81 percent at two-category resolution to around 46 percent at five-category resolution for that held-out material (Pharmaceutical Research, 2025).
What Still Needs Shear Cell or Dynamic Testing
A constituent-property prediction like this one is best treated as a decision input for prioritizing which formulations to test first, not as a determining criterion for process readiness. Because the training labels came from shear cell FFC values measured under a defined set of consolidation stresses, the model inherits that method’s scope. It says nothing directly about aeration-related behavior relevant to hopper discharge or pneumatic transfer, which needs separate dynamic or aerated powder testing, about wall friction against a specific process surface, or about how flow properties shift after storage consolidation covered in powder memory effects.
The accuracy drop for held-out APIs and coated systems matters for how the tool gets used day to day. A favorable prediction for a formulation similar to the training set supports scheduling that blend for confirmatory testing sooner. A favorable prediction for a genuinely new API, a new coating regime, or a mid-range flow category carries less weight, and in this study’s results specifically, warrants the same shear cell or dynamic test that generated the original training labels before any process decision follows from it.
Practical Interpretation Checklist
Treat a favorable machine learning flowability prediction as a reason to schedule a formulation for physical testing sooner, not as a substitute for a flow function measurement.
Weight the prediction differently depending on where it falls in the classification scheme. A call at the extremes, clearly free-flowing or clearly very cohesive, carries more confidence than an intermediate call between adjacent cohesive categories.
Apply extra caution when the formulation includes an API, excipient, or coating level that falls outside the model’s training data, since accuracy in this study degraded well outside its coarsest-category range under those conditions.
Route any blend flagged as cohesive or borderline into shear cell testing, and check aerated or dynamic flow behavior separately if the process involves pneumatic transfer, hopper discharge, or repeated fill and empty cycling, since FFC alone does not capture aeration-driven behavior addressed in powder flow test method selection.
Keep the physical shear cell dataset that trains or validates the model current for the materials in active use, since a constituent-property model is only as reliable as the process-relevant test data it was built from.



