Nvidia is building Nemotron 4, an open-weight foundation model scaling past one trillion parameters, according to reporting from The Information. The planned architecture doubles the footprint of Nemotron 3 Ultra, with an earliest target release slated for this fall. Backing this internal development drive, Nvidia has committed $28 billion in cloud infrastructure spending through 2031 to train in-house frontier systems.
The push highlights Nvidia racing to close an architectural scale gap opened by Chinese AI labs. Moonshot AI's Kimi K3 already scales to 2.8 trillion parameters, while DeepSeek's V4 Pro operates at 1.6 trillion parameters. While Nemotron 3 Ultra was framed as the leading domestic open-weight release upon its June rollout, it fell short on key benchmarks, trailing Kimi K2.6. On the Artificial Analysis Intelligence Index, Nemotron 3 Ultra registers 38 points against Kimi K3's roughly 60 points.
Beyond technical parity, Nvidia's push into trillion-parameter open weights serves a transparent hardware play: establishing a reference architecture for enterprise self-hosting anchors demand for on-premise GPU clusters. Even as open-model advocacy puts Nvidia into awkward friction with cloud customer OpenAI and geopolitical scrutiny over Chinese competitors, selling the picks and shovels requires proving that open-weight architectures justify multi-rack clusters.
Engineering leadership should audit on-premise cluster capacity and memory footprints now. Serving an unquantized trillion-parameter open model internally demands substantial dedicated compute, making total cost of ownership evaluations against hosted frontier APIs an immediate priority.