Drug discovery remains a high-cost, high-risk process, with failure rates exceeding 90% and development timelines averaging 10-15 years. The pharmaceutical industry is increasingly relying on AI to improve success rates and reduce timelines. Paul Belcher, director of protein research strategy at Cytiva, emphasized that AI can help identify and optimize drug candidates more efficiently, reducing the risk of costly failures later in development. "The main cost in drug discovery is still the clinical phase, so trying to reduce risk and increase your success rates there is obviously hugely beneficial," Belcher said. "AI is one approach that drug companies hope will not only save time and compress timelines, but enable better quality candidates to reach the clinic." AI is already showing promise in early-stage applications, such as hit identification, where it can design drug candidates and predict interactions with disease targets before physical testing. However, AI cannot reliably predict the kinetics or developability of new compounds, meaning every candidate still needs lab validation. This shift from empirical screening to predictive design has increased the demand for higher-throughput, information-rich technologies to validate and characterize AI-generated compounds. Source: mittr

AI has also highlighted the need for better, more complete data to train models effectively. Many earlier AI models were trained on publicly available datasets, but these datasets lack the structure, labeling, and diversity needed to keep models accurate and free of bias. Belcher noted that publication bias exacerbates this issue, as most datasets and scientific publications focus exclusively on positive results, leaving out the failed experiments and compounds that don’t bind. "We often joke that there should be a journal of negative data," he said. "It’s often buried in lab notebooks, and it’s never used to inform or guide future research." This lack of negative data creates a fundamental problem: without access to a broad range of data, models can’t be adequately trained to avoid bias. "In all machine learning applications, the model’s performance relies heavily on the quality and scope of the training data," Belcher added. Source: mittr

Belcher also pointed to the growing concern around data integrity, particularly with the rise of generative AI making fabrication easier. He cited research by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images. "Manipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences," Belcher said. Some vendors are addressing this challenge with tools like Cytiva’s Image Integrity Checker, which uses secure hash algorithms to detect tampered images. "We’re starting to see a lot of interest from publishing houses that want to adopt this as standard because it’s a quick way to ensure that what gets published in the literature is genuine," he added. Source: mittr

Source: mittr