ExplorerArtificial IntelligenceAI
Research PaperResearchia:202609.09078

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

Yunpeng Xu

Abstract

Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit...

Submitted: September 9, 2026Subjects: AI; Artificial Intelligence

Description / Details

Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band (10%10\%-40%40\%) is best for all five domains, and a calibrated permutation test for quadratic interiority gives P0.010P\approx0.010; the fitted mid-training-only curves, with 8B peaks between 9.9%9.9\% and 35.1%35.1\%, reproduce for curve shape but not peak location. Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean +4.32%+4.32\%) yet bridges 0/2400/240 pairs at a 5%5\% threshold and 30/24030/240 at a 10%10\% ratio, an equal-budget uniform control behaves almost identically, and a permutation null would bridge 13.8±3.313.8\pm3.3 and 77.9±8.577.9\pm8.5 pairs (P<0.001P<0.001). Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift. An exploratory θθ^* allocation attains the largest full-pipeline gain (+4.36%+4.36\% vs. +0.80%+0.80\%/+0.64%+0.64\%,pp) but is marginal under Welch test.


Source: arXiv:2609.09081v1 - http://arxiv.org/abs/2609.09081v1 PDF: https://arxiv.org/pdf/2609.09081v1 Original Link: http://arxiv.org/abs/2609.09081v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 9, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark
Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training | Researchia