Autonomous agent harnesses make dozens of micro-decisions per second—routing queries, picking tools, filtering irrelevant retrieved context, or checking for prompt injections. In standard setups, these routing steps lean on the exact same heavy autoregressive LLMs doing the actual heavy lifting, wasting hundreds of milliseconds and a full generation cycle just to output a single categorical label. Enter System-1 decision models: single-pass classifiers designed to handle these routing choices via raw class probabilities instead of bloated text generation.

Evaluating LAYA and JEV across Agent Decision Points

To test whether these lightweight models can actually shoulder infrastructure workloads without breaking production logic, researcher Jiawei Li put them through a rigorous paired evaluation. The benchmark pits an open-weight model (LAYA) from Laya AI against a hosted System-1 alternative (JEV) from TypeSafe AI, both released in 2026. Built across 18 public sources, the benchmark covers 11 specific decision points spanning 7,283 base cases and 6,640 robustness variants. Testing included byte-identical inputs, paired tests, and cross-hardware reproducibility checks. The complete dataset, raw outputs, and analysis code sit openly in the project's GitHub repository.

Across the 11 tested decision points, JEV decisively outperformed LAYA on nine of them, posting accuracy margins ranging from +10.8 to +46.0 percentage points.

"JEVis significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp)."

Yet neither system proved bulletproof in production conditions. Both models completely bombed zero-shot model routing, failing to beat random chance, and tied squarely on RAG relevance gating. Worse, LAYA demonstrated a brittle sensitivity to prompt formatting: simply reversing the order of the provided options caused the open-weight model to flip 30% of its verdicts. When forced to navigate large or crowded toolsets, LAYA's accuracy collapsed to 31% at 50 nearest-neighbor tools, whereas JEV maintained a solid 98% accuracy on items carrying a unique correct tool.

Pipeline Self-Audit and Economic Reality

Methodological scrutiny of the benchmark itself laid bare how easily teams miscalculate deployment economics. Li identified three major analysis errors and one design confound that heavily distorted initial performance and cost claims. Most glaringly, omitting pre-screen compute overhead initially inflated a modest 4.3% actual cost saving into a misleading 23.9% headline figure.

Similar analytical slips warped perceived model reliability. Initial reports framed raw gate accuracy as end-to-end pipeline quality at 58% versus 98%, while tuning thresholds on in-sample data masked a stubborn 17% miss rate on held-out sets. Meanwhile, a supposed channel effect on injection false positives evaporated entirely once researchers tested against channel-native content.

Replacing heavy autoregressive LLMs with System-1 classifiers introduces brutal, unvarnished trade-offs in agent orchestration layers. While hosted architectures like JEV handle large candidate toolsets with professional competence, open-weight models like LAYA remain dangerously fragile when presented with basic option shuffling. More importantly, when actual end-to-end savings shrivel from double-digit fantasies down to a meager 4.3% once you factor in pre-screening overhead, engineering leads should audit routing reliability on held-out distributions before betting their production infrastructure on single-pass classification.

Artificial IntelligenceLarge Language ModelsAI AgentsCost ReductionMachine Learning