General-purpose language models are proving too expensive for the routine steps of autonomous agents. On September 15, TypeSafe released Jev, dubbing it the first System One class model. Within 16 days, five more implementations of the same concept hit the market: Clef and Clef-flash from Cloudflare, pplx-decider-v1-27b from Perplexity, Strands Decider 2B from AWS, Laya from Convai Innovations, and GLiDE from Fastino. The infrastructure dead end of generating tokens at every intermediate agent action has become glaringly obvious to developers.

In many agentic systems, an LLM is invoked at every single step, even though text generation is entirely unnecessary for most of them. An agent chooses a tool, checks whether a task is complete, or decides whether a command can be executed. Full generation on such steps costs more and runs slower than the decision process requires, and the output still needs to be parsed and validated against a specific format. TypeSafe proposed a separate fast loop for these steps, dubbing it System One by analogy with the fast, intuitive thinking that Daniel Kahneman termed System 1. The name Jev references Jevons' paradox: the cheaper a resource becomes, the more of it is consumed. Jev itself is priced at $0.042 per million input tokens, with TypeSafe claiming a 70–500 ms latency and a 32K token context window. LangChain has integrated this type of validation directly into the middleware of its Jev integration.

Architectural shift and inference economics

In a standard LLM, a transformer translates the input into hidden states, and the final layer—the LM head—projects the vector of the final position onto a vocabulary of tens or hundreds of thousands of tokens. Generation consists of two phases: prefill, which reads the entire input in a single pass, and decode, which adds one token at a time per step. A decision model drops everything except the prefill. Instead of projecting onto a vocabulary, it uses an output layer that evaluates valid options, followed by a softmax function. While the system interface remains largely unified, under the hood lie five distinct solutions: a frozen LLM with a trainable output layer, an LLM with a replaced LM head, a pointer mechanism, a 322-million-parameter encoder, and reasoning triggered on demand. Cloudflare states that Clef is fully compatible with the Jev API, while Fastino delivers GLiDE through the exact same /v1/systemone endpoint and has published a migration guide from TypeSafe.

Clef utilizes a frozen Qwen3.8-27B, while Clef-flash uses Qwen3.5-9B. Their weights are open-sourced under the Apache 2.0 license and can be deployed via Transformers, vLLM, or SGLang. On Workers AI, the models cost $0.24 and $0.09 per million input tokens with a 65,536 token context. On BANKING77, Clef outperforms Jev by 14.5 points, though it trails by 8.6 points on When2Call, where the model decides whether to invoke a tool immediately. An arXiv preprint proposes deploying Jev and Laya at four distinct nodes of an autonomous penetration-testing agent. LangChain compared Jev against three LLMs acting as judges across 500 agent response evaluations, demonstrating the practical viability of the approach.

Artificial IntelligenceLarge Language ModelsAI AgentsCost Reduction