Deploying large language models remains an expensive exercise in hardware acquisition, keeping researchers hard at work on compression tricks. While structural pruning trims the fat without requiring custom serving runtimes, lopping off parameters permanently damages model capacity. If you want that lost accuracy back, traditional fine-tuning demands a massive compute budget, turning model compression into an expensive engineering headache.
Unifying expert construction and routing
To sidestep this performance penalty, a new method called ToMoE converts dense LLMs into Mixture-of-Experts architectures without touching the original weights. Sparse models like DeepseekMoE match dense performance at a fraction of the active parameter count, but getting there usually requires training from scratch.
"MoE inherently exists within dense models and can be uncovered without updating model weights."
Instead of forcing teams to choose between permanent parameter destruction and expensive retraining, ToMoE demonstrates that expert sub-networks are already buried inside dense layers. The method uses dynamic structural pruning to unify expert construction and router training in a single step, keeping conversion costs right down.
Layer-specific dynamic routing
The conversion process targets attention and feedforward layers through tailored mechanisms. For multi-head self-attention layers, the framework applies top-K routing and static pruning for compression. For MLP layers, it transforms them into MoE layers using top-1 expert routing.
Empirical evaluations across Phi-2, LLaMA-2, LLaMA-3, and Qwen-2.5 models show ToMoE outperforming standard pruning and existing MoE techniques, all without a single weight update. The code is already public on GitHub for engineering teams brave enough to test it.
By proving that sparse routing mechanisms can be extracted directly through dynamic pruning, ToMoE eliminates the expensive retraining phase that usually kills sparse architecture projects. For businesses, this bypasses the usual infrastructure upgrade cycles and slashes the server capacity needed for inference. However, open questions remain about how these un-finetuned expert configurations hold up under extreme production concurrency.