IBM has quietly rolled out its open-weight Granite 4.2 family in 3B, 8B, and 30B parameter architectures under a truly permissive Apache 2.0 license. Trained from scratch across roughly 15 trillion tokens, the lineup scales up to a 512k-token context window. While proprietary labs push ever-higher API lock-in, IBM is handing enterprise infrastructure teams fully inspectable weights out of the box on Hugging Face, Ollama, and GitHub.
Granite's real leverage lies in its inference economics. IBM baked in toggleable "thinking" and "low-effort" reasoning modes, giving platform engineers direct dial control over token latency and GPU compute per request. Under the hood, the 8B and 30B models incorporate reinforcement learning tailored for agentic behavior—natively handling tool calls, web lookups, and Python execution in sandboxed environments. Because the entire family supports standard OpenAI-format tool calling across serving runtimes like vLLM and SGLang, dropping them into existing private cloud or on-prem pipelines requires virtually zero orchestration overhaul.
The release also pairs with Granite Speech 5.0 Turbo CTC, a lightweight 470M-parameter ASR engine that chews through three hours of audio in a single second. For technical leads balancing soaring monthly API invoices against strict data sovereignty constraints, Granite 4.2 makes self-hosted, agentic enterprise workloads look less like an experimental chore and more like a defensible balance-sheet upgrade.