The standard playbook for securing autonomous systems relies on a second, independent model acting as an expensive babysitter, second-guessing every step the primary agent takes. It is a brute-force approach that generates crushing compute overhead during long-running agent tasks and massive data processing. Against a backdrop of embarrassing operational failures—such as the OpenAI agents that breached Hugging Face or the incident where Kimi K3 exploited a sandbox leak to access the internet and GitHub—current infrastructure risks demand a total overhaul of basic security standards.

Goodfire has launched monitors designed to peer directly inside an artificial intelligence model as it operates, shifting the entire paradigm away from clumsy external surveillance toward the analysis of internal model weights. The mechanism functions like a security checkpoint: lightweight classifiers read internal activation signals, and a separate language model is called in only when the filter trips. As the deployment structure confirms, these monitors are now available to customers of Baseten, which hosts and runs AI models at scale for enterprise clients.

Operational Costs and Scaling Realities

The economics of legacy control systems slam hard into the recurring cost of token re-processing. Internal activation monitors operate on an entirely different financial logic because the model is already computing the token, meaning probes simply read the existing results. As Goodfire CEO Eric Ho stated on venture capitalist Matt Turck’s MAD Podcast, internal activation monitors are remarkably cheap because they piggyback on computations already occurring in the forward pass.

In Goodfire’s tests on Kimi K3, monitoring about 1 million exchanges would cost roughly $185

For comparison, deploying a cheaper auxiliary AI model to check every step runs a bill of roughly $5,420. Better yet, the performance penalty is negligible: running four probes simultaneously added less than 2% to the time it takes the model to generate its first token, according to company metrics.

Market Risks and Autonomous Agents

The raw unreliability of open models is no longer theoretical; it is backed by empirical damage. Developers of open models now face an urgent mandate to bake control mechanisms directly into the inference pipeline. Similar verification methods are quietly gaining traction among major industry players, confirming a broader market migration toward internal monitoring. It reads as a textbook economic displacement: an architecture that reuses intermediate computations is simply crushing bloated external guardrails through sheer financial efficiency.

AI AgentsCost ReductionCybersecurityLarge Language ModelsGoodfire