Modern enterprise deployments increasingly rely on multi-step reasoning models and autonomous agents that execute dozens of sequential operations for a single prompt. While frontier model capabilities expand rapidly, data centers have slammed into severe electrical power caps. For enterprise balance sheets, current hardware forces an ugly compromise between interactive latency and raw throughput—leaving operators footing an exorbitant compute bill. OpenAI has now published the first operational benchmarks for its custom inference chip, Jalapeño, designed specifically to dismantle this bottleneck.

According to an engineering update released by OpenAI, Jalapeño delivers higher throughput and reduced time-between-tokens simultaneously within a unified design. In head-to-head benchmarking against leading commercial accelerator setups, the custom ASIC delivered 1.5 to 1.9 times more AI work per kilowatt at peak throughput alongside 1.7 to 3.6 times lower end-to-end latency across diverse open model architectures.

Benchmarking Across Diverse Architectures

To evaluate the silicon under rigorous production constraints, OpenAI benchmarked Jalapeño across high-throughput batching and ultra-low-latency interactive operating points. The evaluation covered dense and mixture-of-experts architectures at varied scales: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Jalapeño placed directly on the Pareto frontier across all three targets, proving that its architectural advantages extend beyond proprietary internal codebases to third-party open architectures.

For interactive reasoning workloads, Jalapeño delivered 2.1 to 4.1 times higher performance across the evaluated models.

"Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two."

This latency profile is decisive for sequential reasoning agents, where delays compound across every intermediate generation step, planning phase, and tool call.

Full-Stack Optimization and Silicon Design

OpenAI engineered the silicon through a closed-loop engineering pipeline. Early internal models assisted engineering teams during the physical chip's design and bring-up, while current reasoning models automate kernel optimization, software compilation, and runtime scheduling. Jalapeño was architected from the ground up so that neural networks can program and manage the hardware dynamically.

This vertical integration ties models, serving frameworks, silicon, networking, and memory into a cohesive stack calibrated by real-world API workloads. Jalapeño is not a speculative prototype; it is functional first-party silicon delivering measurable unit-cost reductions and shielding OpenAI's enterprise margins from Nvidia's pricing premiums.

OpenAI plans to ramp production over the coming quarters to power its API infrastructure and enterprise platform. Custom silicon has officially shifted from an R&D hedge into an existential operational lever for any lab seeking to scale reasoning agents within fixed power budgets.

AI ChipsOpenAINVIDIACost ReductionAI in Business