University of Bristol researchers propose a framework for testing medical AI systems for reliability, modeled on drug approval standards. The approach, called 'Learning Ensemble,' aims to prevent AI systems from failing in clinical settings by systematically evaluating their limits, reliability across patient groups, and suitability for clinical use.
The framework is inspired by how medicine handles drugs with uncertain effects, where structured information packages detail conditions for safe use. Similarly, the researchers suggest creating a toolkit for medical AI developers that includes three key areas to document and check before deploying systems.
The first area focuses on system limits, such as intended use, hardware requirements, and training data. A 2021 study showed an AI system misinterpreted X-ray images, focusing on incidental details rather than disease signs, leading to failure when deployed at a different clinic.
The second area assesses reliability across patient groups, ensuring systems do not systematically misdiagnose underserved populations. Another 2021 study found AI systems were less likely to detect disease in these groups, highlighting the need for equitable performance.
The third and most critical area evaluates whether the system fits its intended clinical purpose. An AI system that rated asthma patients with pneumonia as low mortality risk was found to be ineffective for triage, as survival rates were influenced by aggressive emergency room treatment.
The authors see their work as a starting point for building reliable medical AI systems, emphasizing the need for trial and error, expert review, and continuous adjustment. Their framework aims to provide a shared language and structure for developers to catch problems early.
Source: thedecoder