Explorerโ€บBiomedical Engineeringโ€บEngineering
Research PaperResearchia:202609.30038

CASR: Content-Adaptive Neural Super-Resolution Post-Filter for Versatile Video Coding via Low-Rank Overfitting

Khoa Pham-Dinh

Abstract

The use of super-resolution as a post-processing step following a video codec allows videos to be encoded at reduced spatial resolution at the encoder side and to be reconstructed and upsampled to the original resolution at the decoder side. In this way, the required bitrate is reduced and the quality of the reconstructed frames is improved without modifying the core coding architecture. However, generic SR models are typically trained offline and lack sufficient adaptability to the diverse cont...

Submitted: September 30, 2026Subjects: Engineering; Biomedical Engineering

Description / Details

The use of super-resolution as a post-processing step following a video codec allows videos to be encoded at reduced spatial resolution at the encoder side and to be reconstructed and upsampled to the original resolution at the decoder side. In this way, the required bitrate is reduced and the quality of the reconstructed frames is improved without modifying the core coding architecture. However, generic SR models are typically trained offline and lack sufficient adaptability to the diverse content characteristics and compression artifacts produced by video codecs, which limits their effectiveness during test time. To address the limitation, this paper proposes CASR (Content-Adaptive Super-Resolution), a content-adaptive SR post-filter framework for Versatile Video Coding (VVC), based on encoder-side overfitting on each input sequence. In order to limit the bitrate overhead required for signalling the content adaptation signal, i.e. the weight-update, Low-Rank Adaptation (LoRA) is leveraged. The method freezes the convolution kernels of a pretrained SR network and fine-tunes only lightweight rank-r matrices attached to selected convolution layers, using VVC decoded frames and quantization-parameter (QP) maps of test sequences as supervision. The resulting low-rank update is compressed with the MPEG Neural Network Compression and Representation (NNR) standard. Experiments on the JVET common test conditions (CTC) class A1 and A2 sequences indicate that LoRA-based content adaptation provides bitrate savings over a non-adapted SR post-filter at a small signalling cost. Compared with the VVC Test Model (VTM21), the proposed method achieves BD-rate savings of -10.93% (Y), -15.39% (U), -24.41% (V) under random access and -13.43% (Y), -5.75% (U), -22.94% (V) under all-intra. An ablation of the LoRA rank r further shows that r=4 provides the best trade-off between coding gain and signalling cost.


Source: arXiv:2609.37328v1 - http://arxiv.org/abs/2609.37328v1 PDF: https://arxiv.org/pdf/2609.37328v1 Original Link: http://arxiv.org/abs/2609.37328v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 30, 2026
Topic:
Biomedical Engineering
Area:
Engineering
Comments:
0
Bookmark