Insilico Medicine opens its internal drug discovery benchmark to outside AI developers

The Drug Discovery and Development Benchmark as a Service (DDD BaaS) launched on 30 July covers two evaluation suites. Drug Discovery Foundations comprises more than 300 tasks spanning disease biology, molecular property prediction, retrosynthesis, structure-based drug design, and clinical development. Drug Candidate Essentials tests a model's ability to run an end-to-end discovery program, from hit identification through to preclinical candidate nomination.

The rationale is straightforward: Insilico argues that most public AI benchmarks are contaminated, meaning models score well because they have seen variations of the test questions during training rather than because they can actually navigate the ambiguity of real drug programs. Its own datasets come from proprietary validated programs, with explicit decontamination steps intended to prevent memorisation.

For an outside organisation evaluating whether to build AI into drug discovery workflows, that distinction matters. A model that scores 90% on a public chemistry benchmark but struggles with novel synthesis planning against proprietary targets is a different proposition to one that performs consistently against Insilico's closed dataset.

Insilico itself has used generative AI across its pipeline since its founding, nominating 31 preclinical candidates in six years and compressing the typical 2.5-4 year timeline to preclinical nomination down to roughly 12-18 months. Its lead program, Rentosertib (ISM001-055), is in Phase III for idiopathic pulmonary fibrosis. The company is listed on the Hong Kong Stock Exchange under the ticker 3696.

The service is available at dddbench.insilico.com to any organisation developing frontier AI for drug discovery or using foundation models in research workflows. Pricing and tiering details were not disclosed in the announcement.

The broader context is a market where a large number of AI drug discovery companies are competing for pharma partnership deals, and where differentiation on benchmark performance has become a marketing as much as a scientific exercise. An evaluation framework that makes those scores harder to game could reshape how pharma companies assess AI vendors — or, if Insilico's dataset construction holds up to scrutiny, give it a second business line alongside its pipeline.

To stay across the latest in cloud, AI and enterprise tech analysis from Compare the Cloud, subscribe to our weekly newsletter at https://www.comparethecloud.net/newsletter

More News