Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
Jun 5, 2026·,,,,,
Jingchen Sun
Shaobo Han
Deep Patel
Wataru Kohno
Can Jin
Changyou Chen
Abstract
Knowledge distillation establishes a learning paradigm that learns from both data supervision and teacher guidance. However, the optimal balance between learning from data and learning from the teacher is hard to determine, as some samples are data-noisy while others are teacher-uncertain. This raises a pressing need to adaptively balance data and teacher supervision. We propose Beta-weighted Knowledge Distillation (Beta-KD), an uncertainty-aware distillation framework that adaptively modulates how much the student relies on the teacher guidance. Specifically, we formulate teacher–student learning from a unified Bayesian perspective and interpret teacher supervision as a Gibbs prior over student activations. This yields a closed-form, uncertainty-aware weighting mechanism and supports arbitrary distillation objectives and combinations. Extensive experiments are conducted on multimodal VQA benchmarks by distilling a student Vision-Language Model from a large teacher VLM. The results demonstrate that Beta-KD consistently outperforms existing knowledge distillation methods. Code is available at https://github.com/Jingchensun/beta-kd.
Type
Publication
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)