Healthcare systems increasingly integrate artificial intelligence into routine clinical operations, ranging from triage and diagnostic reasoning to medication management and clinical documentation. As algorithmic decision-support tools touch patient care directly, standard safety reporting and statistical benchmarks face structural limitations. Traditional quality assurance evaluates aggregate model shifts, while hospital incident reporting systems document adverse outcomes without preserving the algorithmic context necessary to diagnose breakdowns.
The Breakdown of Aggregate Monitoring
Aggregate performance tracking reliably flags dataset drift and calibration shifts, yet it obscures the precise mechanics of single-event failures and near-misses. A machine learning model can easily satisfy statistical benchmark thresholds across an entire cohort while failing under specific workflow constraints, interface designs, or user interactions. Standard patient safety reporting captures clinical harm, but it routinely omits crucial technical evidence such as model inputs, outputs, audit logs, active system prompts, user roles, and software versions.
To bridge this operational gap, a research team led by Paulius Mui of X = Primary Care, alongside Dean F. Sittig of UT Health McWilliams School of Biomedical Informatics and Informatics Review LLC, Steve Labkoff of Luminant Consulting, and Sanjay Basu of UCSF and Waymark, proposed AI Morbidity and Mortality (AI M&M). The method adapts the surgical M&M tradition, converting individual adverse events into structured, blameless reviews to identify contributing system factors.
"AI M&M is intended to complement, rather than replace, model monitoring, patient safety reporting, and regulatory oversight by converting individual AI-in-workflow failures into actionable institutional learning."
The authors emphasize that clinical AI introduces unique reproduction challenges compared to traditional surgery. Algorithmic tools are updateable, versioned, vendor-mediated, and sensitive to hidden system instructions, meaning the exact risk profile depends on user expertise, verification workflows, and deployment environments.
The Four-Axis Investigation Architecture
The AI M&M protocol structures case investigations through standardized intake, evidence preservation, investigator-level reconstruction, tool-in-loop attribution, and corrective-action tracking. Rather than treating an error as a simple model miscalculation, the protocol evaluates events across four sequentially linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. This separation distinguishes the conditions that exposed a vulnerability from the underlying process generating risk, the actual consequence in patient care, and the remediation assigned.
To test this taxonomy, the researchers evaluated five illustrative outpatient medication and clinical decision-support cases. Two independent clinician reviewers categorized each event across all four dimensions. The reviewers achieved complete agreement across all 20 axis-level classifications, demonstrating that clinical reviewers can systematically decouple interface, model, and workflow errors.
For HealthTech leaders and clinical executives, AI M&M shifts deployment strategy from blind trust in macro benchmarks to operational forensic readiness. Bridging model logs with clinical liability protocols allows institutions to pinpoint whether a failure stems from a prompt update, UI design flaw, or clinician override before scaling models across clinical operations. While prospective validation across larger health systems remains necessary, establishing rigorous incident-level investigations is becoming an essential prerequisite for managing clinical liability.