Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision
Abstract
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model ge...
Description / Details
We introduce ST (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, ST improves VSTAT accuracy by as a single model, with souping, and with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by on VSTAT-YouTube state-tracking questions and on MVBench Action Count.
Source: arXiv:2609.04203v1 - http://arxiv.org/abs/2609.04203v1 PDF: https://arxiv.org/pdf/2609.04203v1 Original Link: http://arxiv.org/abs/2609.04203v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Sep 4, 2026
Computer Vision
Computer Vision
0