ExplorerComputer VisionComputer Vision
Research PaperResearchia:202609.04005

Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision

Shravan Venkatraman

Abstract

We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model ge...

Submitted: September 4, 2026Subjects: Computer Vision; Computer Vision

Description / Details

We introduce S3^3T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, S3^3T improves VSTAT accuracy by +1.74+1.74 as a single model, +2.38+2.38 with souping, and +2.70+2.70 with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by +7.95+7.95 on VSTAT-YouTube state-tracking questions and +4.50+4.50 on MVBench Action Count.


Source: arXiv:2609.04203v1 - http://arxiv.org/abs/2609.04203v1 PDF: https://arxiv.org/pdf/2609.04203v1 Original Link: http://arxiv.org/abs/2609.04203v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 4, 2026
Topic:
Computer Vision
Area:
Computer Vision
Comments:
0
Bookmark