ExplorerArtificial IntelligenceAI
Research PaperResearchia:202603.31006

Stepwise Credit Assignment for GRPO on Flow-Matching Models

Yash Savani

Abstract

Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textures (high-frequency details). Moreover, assigning uniform credit based solely on the final image can inadvertently reward suboptimal intermediate steps, especially when errors are corrected later in th...

Submitted: March 31, 2026Subjects: AI; Artificial Intelligence

Description / Details

Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textures (high-frequency details). Moreover, assigning uniform credit based solely on the final image can inadvertently reward suboptimal intermediate steps, especially when errors are corrected later in the diffusion trajectory. We propose Stepwise-Flow-GRPO, which assigns credit based on each step's reward improvement. By leveraging Tweedie's formula to obtain intermediate reward estimates and introducing gain-based advantages, our method achieves superior sample efficiency and faster convergence. We also introduce a DDIM-inspired SDE that improves reward quality while preserving stochasticity for policy gradients.


Source: arXiv:2603.28718v1 - http://arxiv.org/abs/2603.28718v1 PDF: https://arxiv.org/pdf/2603.28718v1 Original Link: http://arxiv.org/abs/2603.28718v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Mar 31, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark
Stepwise Credit Assignment for GRPO on Flow-Matching Models | Researchia