Explorerโ€บData Scienceโ€บMachine Learning
Research PaperResearchia:202610.05058

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Vladislav Gromadskii

Abstract

Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we p...

Submitted: October 5, 2026Subjects: Machine Learning; Data Science

Description / Details

Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to 32ร—32\times fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.


Source: arXiv:2610.03641v1 - http://arxiv.org/abs/2610.03641v1 PDF: https://arxiv.org/pdf/2610.03641v1 Original Link: http://arxiv.org/abs/2610.03641v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Oct 5, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark
IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models | Researchia