Researchers at King’s College London have developed a new mathematical framework to test whether AI systems that appear to explain their decisions are genuinely transparent.
Published in the Journal of Machine Learning Research , the study addresses a growing issue as AI becomes more common in healthcare and other high-stakes settings: how to ensure people can understand, check and even challenge AI-generated decisions.
AI systems can make highly accurate predictions, including identifying signs of disease or predicting patient outcomes. However, many operate as ‘black boxes’, producing an answer without making their reasoning clear to the person using them.
This can make it difficult for clinicians and other experts to identify whether an AI system has made a mistake or intervene in its decision-making. It also presents a challenge as emerging AI regulations increasingly emphasise transparency and meaningful human oversight.
One approach to making AI more understandable is the use of concept-based AI models. These systems are designed to make decisions using clinically meaningful concepts, such as blood pressure, the presence of fever, tumour size or abnormalities visible in a medical image, rather than raw data alone. In principle, this should allow clinicians and other experts to see which concepts influenced a decision, giving them the opportunity to both understand and review the decision with more information. However, concept-based AI models can still suffer from a problem known as ‘information leakage’.
This happens when the concepts used by a model contain additional, unintended information that is not visible to the person reviewing the decision. As a result, the AI system may appear interpretable while still relying on information that a human cannot see or assess. Effectively, it behaves like a ‘black box’ model disguised as an interpretable model.
King’s researchers Dr Chris Banerji and Dr Enrico Parisini, working with colleagues at the Alan Turing Institute and with support from The Turing-Roche Strategic Partnership , developed a new framework to precisely define and measure information leakage in concept-based AI systems. The framework consists of two measures: concepts-task leakage (CTL), which captures hidden information linked to the final prediction, and interconcept leakage (ICL), which captures hidden information shared between concepts.
Testing the framework across several datasets, the researchers found it could reliably detect leakage and predict how models would respond when their concepts were deliberately changed.
The research provides practical guidance for designing concept-based models that minimise leakage and offer more meaningful transparency.
Dr Christopher Banerji , AI+ Senior Fellow (Clinical-Academic) and senior author of the paper, said : “Rather than putting the cart before the horse, while the field of AI is moving quickly to apply models to real-world problems, we have taken a step back to consider what is needed to make these systems safe and reliable. Our work has focused on understanding limitations and addressing them, so that we can move towards deploying these models in clinical practice in a safer way.”
Dr Enrico Parisini , Senior Research Fellow in Machine Learning and first author of the paper, said: “Concept-based AI has the potential to make AI systems more transparent, but our research shows that models can appear interpretable while still relying on information that is hidden from the person using them. By identifying and measuring this hidden information, we can take steps towards developing AI systems that are more transparent.”
The framework gives researchers and developers a clear way to assess whether AI systems are transparent enough to support meaningful human oversight, including in healthcare. The team is now working to apply these approaches to real clinical problems.
The research was supported by the Turing–Roche Strategic Partnership, the King’s College London AI+ Fellowship and PharosAI.
Journal of Machine Learning Research