Small Multimodal Models: From Recognition to Reasoning

projects

Many real-world deployments require small models, while scaling trends favor larger ones. How far can we push the capability frontier at small-model scale?

Developing small multimodal models for real-world settings where compute, memory, latency, privacy, or large-scale parallel processing make model capacity a first-class constraint. The broader goal is to advance small multimodal models from open-vocabulary recognition to multimodal understanding and reasoning.

Research progresses from memory-augmented few-shot adaptation that combines explicit knowledge from support-set exemplars with implicit knowledge encoded in model and adapter parameters (CLAP-S, ICASSP 2025), to noise-robust multimodal learning by distilling a pretrained teacher into students trained under clean and noisy input conditions and adaptively fusing their representations through minimum-entropy selection, allowing both interpolation and extrapolation beyond the training conditions (Mix-CLAP, ICASSP 2026), and toward small multimodal foundation models for domain-grounded question answering and reasoning.