ExplorerComputer VisionComputer Vision
Research PaperResearchia:202608.21009

Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models

Taihang Hu

Abstract

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing sup...

Submitted: August 21, 2026Subjects: Computer Vision; Computer Vision

Description / Details

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.


Source: arXiv:2608.20334v1 - http://arxiv.org/abs/2608.20334v1 PDF: https://arxiv.org/pdf/2608.20334v1 Original Link: http://arxiv.org/abs/2608.20334v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 21, 2026
Topic:
Computer Vision
Area:
Computer Vision
Comments:
0
Bookmark
Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models | Researchia