The Economics of Agent Testing

Back in 2023 when we first started building agents—which we mostly called LLM apps at the time—the primary approach to evals was code-based. Later that year, researchers introduced LLM-as-a-judge, and since then, code-based and LLM-as-a-judge models have been the main ways to evaluate agents. Code-based evaluators check deterministic conditions like tool calls or pattern matches, while LLM-as-a-judge evaluators parse free text to grade traces at a significant cost and slower operational speed. Now, according to the assessment of LangSmith, Jev is available as a specialized system judge for production evaluations, directly integrating into existing workflows.

System One models are a class of AI models built to make fast, structured decisions that software can use directly. Jev isn't a traditional LLM and doesn't generate text; instead, it functions as a System One model that evaluates a state and returns typed answers and probabilities without the computational overhead of parsing unstructured prose.

Parallel State Evaluation in Production

As LangSmith integrated Jev into its evaluation suite, it closed a glaring gap for businesses demanding fast, cheap assessment of open-ended autonomous agent behavior. Jev can answer three types of questions: a noul returns a yes/no probability, a choice picks one option from a set, and a score rates the state on an ordered scale. Each answer is typed directly rather than generated as text that requires parsing.

Jev evaluates every question in a request together, so scoring a trace against multiple feedback keys, such as PII leakage, user intent, and user frustration, occurs in a single pass.

Running evaluations across production traffic no longer requires sacrificing coverage for cost. Teams can score every trace instead of a sample and check multiple criteria simultaneously without inflating operational expenses, shifting project management from guesswork to predictable execution.

Automated evaluation at this speed shifts monitoring from a periodic cost center into a real-time control mechanism. As project leads and technical directors well know, cheap structured classification finally removes the financial barrier that previously restricted rigorous agent testing to well-funded research labs.

Artificial IntelligenceLarge Language ModelsAI AgentsAutomationCost Reduction