Developer akopytko has released an optimized NVFP4 GGUF quantization of Qwen3.8-27B tailored for Nvidia's Blackwell architecture, delivering a 50% jump in prompt prefill throughput over standard 4-bit baselines. Critically for engineering teams, the custom N4_0 format maintains out-of-the-box compatibility with mainline llama.cpp, operating seamlessly across hardware from workstation RTX 50-series cards up to enterprise DGX Spark, B200, and B300 accelerators.

Hardware benchmarks on an RTX 5090 (stock 575W TDP) clock the N4_0 format at 6,258.74 tokens per second during prompt processing (pp2048, ubatch 512)—eclipsing Unsloth's NVFP4 (6,018.71 t/s) and leaving standard Q4_0 (4,139.00 t/s) well behind. akopytko squeezed out an extra 4% to 7% in speed by discarding Nvidia’s global per-tensor scale overhead, while introducing an offline Mean Squared Error (MSE) scale search during quantization to isolate the optimal UE4M3 representation and minimize tensor reconstruction error. When paired with an embedded NVFP4 MTP draft head, the pipeline hit a sustained 152 t/s decoding throughput at a 0.75 draft acceptance rate.

Perplexity metrics confirm that speed does not degrade output quality: Wikitext evaluations show N4_0 registering a 0.066254 Kullback-Leibler Divergence (KLD) mean and 88.525% Top-P agreement against baseline BF16 weights, slightly besting Unsloth’s 0.070824 KLD. For infrastructure leads and technical founders, native 4-bit execution on Blackwell shifts 27B-parameter models from costly multi-GPU clusters into single-card workstation and edge deployments, resetting the unit economics of local enterprise inference.

Large Language ModelsAI ChipsOpen Source AINVIDIACost Reduction