Explorerโ€บData Scienceโ€บMachine Learning
Research PaperResearchia:202609.18023

dQwen3.5: Hybrid-Attention Diffusion Language Models

Anton Xue

Abstract

Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapti...

Submitted: September 18, 2026Subjects: Machine Learning; Data Science

Description / Details

Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attention DLMs in any-order decoding behavior and performs strongly under parallel decoding.


Source: arXiv:2609.20751v1 - http://arxiv.org/abs/2609.20751v1 PDF: https://arxiv.org/pdf/2609.20751v1 Original Link: http://arxiv.org/abs/2609.20751v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 18, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark