The Hidden Cost of Imitation
Supervised fine-tuning on offline agent trajectories serves as the default playbook for shoehorning large language models into complex workflows. By forcing models to mimic reasoning steps and actions token by token, standard fine-tuning jacks up performance on narrow target benchmarks. But there is a catch: the process quietly shreds foundational capabilities like general reasoning and code generation.
To counter this regression, practitioners have traditionally leaned on KL penalties or update limits to curb distributional drift. As researchers Ronghua Li, Zi Liang, Zhishan Li, and Shinan Liu demonstrate in their study, simply throwing a KL penalty at the model or capping update magnitudes fails to stop the bleeding. The foundational skills still erode.
Dynamic Supervision with Information Gain
To break this zero-sum compromise, the researchers introduce Privilege-Guided SFT, a method that uses turn-level information gain across agent trajectories to dial supervision strength up or down. Instead of uniformly anchoring the policy to the base model, the approach dynamically modulates learning pressure where it actually matters.
As the authors show through experiments across four benchmarks, PG-SFT shifts the trade-off curve. It substantially reduces distributional drift and capability degradation, paying for it with only a minor drop in target-task performance.
What this means
The findings prove that keeping a model smart while teaching it new tricks isn't about blank-check anchoring to base behavior. It is about surgical precision on where and how hard supervision departs from the baseline. Engineering teams can finally customize agents for specialized enterprise environments without systematically lobotomizing their core reasoning or code generation skills. While evaluations remain tied to specific benchmark subsets, this approach points toward specialized models that actually survive long-horizon workflows without losing their general competence.