跳到正文
原文
Hugging Face Blog·· 2026-08-10AI 评分51

Multiverse Computing 提出离线 top-K logits 与融合分块 KL 损失,大幅降低 LLM 知识蒸馏显存成本

Making Knowledge Distillation Cheap Enough to Run at Scale

AI 导读

Multiverse Computing 发布论文,用缓存教师模型 top-100 logits 的离线蒸馏和融合分块 KL 损失两项系统改动降低知识蒸馏成本。

来源:Hugging Face Blog · huggingface.co