Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
Abstract
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first f...
Description / Details
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of on OCRBench, on MMStar Fine-Grained Perception, and on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.
Source: arXiv:2608.09931v1 - http://arxiv.org/abs/2608.09931v1 PDF: https://arxiv.org/pdf/2608.09931v1 Original Link: http://arxiv.org/abs/2608.09931v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Aug 11, 2026
Computer Vision
Computer Vision
0