Standard enterprise benchmarks still evaluate AI agents on final task accuracy—a fundamentally flawed heuristic for regulated industries. Getting the right business answer is worthless if an agent arrives at it by citing expired waivers, unverified drafts, or phantom authorities. In an independent evaluation, researcher Jesus Salas demonstrated that autonomous workflows routinely reach the expected business outcome while leaning on invalid source context, closing tasks without verified proof, and silently breaking dependency chains.
To bridge this architectural gap, Salas developed Matrix, an external deterministic causal-state layer built around probabilistic models. Rather than relying on the LLM's self-reported success, Matrix tracks explicit lineage across versioned authority, facts, tasks, and cryptographic completion receipts. In controlled evaluations, while unconstrained agents marked tasks finished on pure hallucinated self-assertion, the governed framework required verifiable execution receipts and automatically invalidated downstream artifacts the moment source policies changed.
Regulatory frameworks like Article 12 of the EU AI Act and the NIST AI Risk Management Framework do not accept black-box output transcripts as compliance artifacts. As enterprise architectures mature, deployment gates must pivot from raw accuracy benchmarks to deterministic provenance integrity: proving not just what the model decided, but whether it had the verified authority to decide it at all.