Publications
Selective Cleaning Enhances Machine Learning Accuracy for Drug Repurposing: Multiscale Discovery of MDM2 Inhibitors
The peer-reviewed method reduced RMSE from 0.74 to 0.58, a 21.6% reduction in model error, and reached R² of 0.87 on the MDM2 benchmark.
Conflicting measurements from different assay procedures should not be averaged into a value that describes no real experiment.
Why this matters
Teams usually arrive asking for a model. More often the blocker is upstream — data pooled from different labs, assays and instruments that quietly disagree with each other.
- Structures and units standardised before modelling
- Conflicting measurements grouped by assay procedure
- Duplicates resolved and plausibility rules applied
What Stream does
Selective Cleaning groups measurements by the experimental procedure that produced them, selects the value backed by the largest body of evidence, and records why that choice was made instead of averaging incompatible procedures.

Cleaning is not a preprocessing step you run once. It is a decision you have to be able to defend.
On the published MDM2 benchmark, Selective Cleaning reduced RMSE from 0.74 to 0.58 — a 21.6 percent reduction in model error — and reached R² of 0.87. The method is peer reviewed and the pipeline is public for independent reproduction.
- MDPI Moleculesmethod