Research Summary · Artificial Intelligence
What interpretability actually requires
A decade of explanation methods has produced impressive demonstrations and remarkably few guarantees. A reading of the formal literature, and of what regulators are actually asking for.
The phrase "explainable AI" entered institutional vocabulary sometime around 2017 and has since acquired the useful vagueness of all terms that everybody supports. Procurement documents require it. Regulatory guidance recommends it. Very little of that language specifies what would count as satisfying the requirement.
The technical literature is more precise, and considerably less reassuring. A substantial body of work now establishes that the dominant family of methods — post-hoc explanation, in which a second procedure produces an account of a first model's behaviour — cannot guarantee that its account is faithful to the computation actually performed. The explanation may be locally plausible. It may be persuasive to a reviewer. It may also be wrong in ways that are undetectable without access to the very understanding the explanation was supposed to supply.
Where the decision at issue is a film recommendation, this is a curiosity. Where it is a benefits determination, it is a problem of a different order, because the administrative law of most jurisdictions requires not merely that a decision be correct but that reasons be given for it. An unfaithful explanation satisfies the form of that requirement while defeating its function.
The alternative is architectural. A smaller literature has pursued models for which explanation is structural rather than approximated — where the account of the decision is the decision procedure, not a reconstruction of it. The standard objection has been performance: constrained models, it was assumed, must sacrifice substantial accuracy.
The empirical picture is now more favourable than that assumption suggested. Across several consequential-decision domains, the measured cost of constraint has been modest — low single-digit percentages in a number of published comparisons. This is not universal, and the domains where it holds are those with structured, moderate-dimensional inputs rather than raw perceptual data. But it is enough to shift the burden of argument. The claim that interpretability is unaffordable now requires evidence in the specific case rather than assertion in general.
What regulators appear to want. Reading recent guidance across several jurisdictions, a consistent shape emerges. The requirement is rarely for a full account of the model. It is for the capacity to state, in a specific case, the grounds on which that case was decided, in terms a person could contest. This is a narrower demand than the technical literature often assumes, and a harder one than most deployed systems can meet, because it is case-specific and adversarial rather than aggregate and cooperative.
The practical consequence is that the useful question is not whether a system is interpretable but whether it can survive being challenged on a particular decision by someone with an interest in the outcome. Very few procurement processes currently ask that question. Those that have begun to are producing noticeably different results.
Corrections? Contact the desk.
More on artificial intelligence
Scholar Conversation
“Nobody needs to know what the system usually does”
Dr. Aarav Mehta on formal verification, the value of a negative result, and what an industrial partner taught him about his own assumptions.
Artificial Intelligence Sep 7, 2026 8 min read