Publications

Selective Cleaning Enhances Machine Learning Accuracy for Drug Repurposing: Multiscale Discovery of MDM2 Inhibitors

The peer-reviewed method reduced RMSE from 0.74 to 0.58, a 21.6% reduction in model error, and reached R² of 0.87 on the MDM2 benchmark.

Mohammad Firdaus Akmal1 July 2025 · 1 min read

Conflicting measurements from different assay procedures should not be averaged into a value that describes no real experiment.

Why this matters

Teams usually arrive asking for a model. More often the blocker is upstream — data pooled from different labs, assays and instruments that quietly disagree with each other.

  • Structures and units standardised before modelling
  • Conflicting measurements grouped by assay procedure
  • Duplicates resolved and plausibility rules applied

What Stream does

Selective Cleaning groups measurements by the experimental procedure that produced them, selects the value backed by the largest body of evidence, and records why that choice was made instead of averaging incompatible procedures.

Assay procedures across the pooled measurement set
The selected value remains connected to its assay procedure and selection rationale.
Cleaning is not a preprocessing step you run once. It is a decision you have to be able to defend.

On the published MDM2 benchmark, Selective Cleaning reduced RMSE from 0.74 to 0.58 — a 21.6 percent reduction in model error — and reached R² of 0.87. The method is peer reviewed and the pipeline is public for independent reproduction.

  • MDPI Moleculesmethod

Want to talk to us in person?

Tell us what you are working on and we will find a time.

Bring your challenge.
Let’s find the right next step.

Get in Touch

Leave your details and someone from the team will be in touch.

Interested inPick one or more

We'll only use this to reply about your enquiry.