Anthropic quietly operated its frontier models without active biological and chemical threat classifiers for nearly an entire year—from May 2025 through April 2026—according to the company's latest safety disclosure. Guardrails explicitly engineered to prevent models from generating actionable CBRN (chemical, biological, radiological, and nuclear) weapon designs were bypassed entirely for all traffic originating from external human-feedback contractors.

The operational loophole encompassed roughly 50,000 third-party data annotators who ran an estimated 133 million conversational sessions unmonitored by critical safety filters. Anthropic conceded that these contractors were screened solely via third-party vendors whose vetting protocols proved largely inadequate. While the lab claims its post-hoc audit found no evidence of weaponization attempts, the lapse underscores an alarming breakdown in basic infrastructure-level guardrail governance.

The disclosure stands in jarring contrast to the hawkish public rhetoric of CEO Dario Amodei, who routinely positions AI-assisted bioweapons proliferation as a civilizational peril eclipsing cyberwarfare. Preaching catastrophic biorisk in Washington while leaving routing exceptions unpatched in production reveals the persistent chasm between frontier AI marketing and operational engineering discipline.

AI SafetyAnthropicLarge Language ModelsAI RegulationCybersecurity