FundamentalsDeep dive9 min
From Experts to Sub-Experts: Smarter Fine-Tuning for MoE Models
NSFT cuts each expert in a Mixture-of-Experts model into channel groups and trains only the ones that matter — beating LoRA and ESFT on accuracy per trained parameter, and forgetting less along the way.