The Modality Balance Problem Is the Hidden Constraint on Visual On-Policy Distillation
Published August 7, 2026 · 2060 words · 11 min read
A pattern has been building across the multimodal training literature this cycle, and OPD-V is its sharpest expression yet. For roughly two years, the on-policy distillation research community has been solving one half of the problem for vision-language models: trajectory sampling, divergence selection, and trust-region mechanics have all matured into a settled engineering discipline. The half that hasn't been solved — the one OPD-V goes after directly — sits upstream of all that: which modality is actually doing the learning during distillation, and whether the training procedure even knows it's creating an imbalance.
The timing matters. On-policy distillation has moved from research technique to production default over the past year, and it's now a standard post-training stage at several frontier labs: Alibaba runs it in Qwen3, Xiaomi in MiMo, Zhipu in GLM-5, and NVIDIA in Nemotron-Cascade 2. DeepSeek-V4's post-training pipeline uses on-policy distillation for its unified-model consolidation stage, after domain-expert SFT and GRPO. Once a technique is embedded that widely in production stacks, its failure modes stop being academic footnotes and start being operational risk — and modality imbalance is exactly that kind of failure mode.
The underlying problem has been documented from multiple independent angles this cycle. Modality imbalance describes a model's tendency to lean on text while underusing, or outright ignoring, other modalities such as vision — even when a task explicitly calls for visual evidence. Both data and architecture feed it: text is a denser, easier-to-exploit training signal than images, and the standard MLLM recipe compounds this by bolting a vision encoder trained on a comparatively small dataset onto an LLM backbone pretrained on trillions of text tokens.
One frequently cited number puts a scale on the problem. A CVPR 2025 study on VLM quantization (Li et al., "MBQ: Modality-Balanced Quantization for Large Vision-Language Models") measured the average gradient magnitude of language-token features against vision-token features in large MLLMs and found language tokens carrying more than ten times the gradient weight — meaning a same-sized perturbation to a vision token moves the loss roughly a tenth as much as an equivalent nudge to a language token. It's a narrower, more mechanical finding than the broader "text dominance" research from papers like "When Language Overrules" (arXiv:2508.10552), which looks at attention behavior and output reliance rather than raw gradients, but it points at the same asymmetry from a different angle.
OPD-V's proposal is structurally different from earlier fixes. Instead of treating modality imbalance as something to patch after training — through data curation or a loss-weighting heuristic — it selects which on-policy tokens to train on using a trust region built between two teacher variants: one given a zoomed-in crop of the image, one given a masked version of it. The gap between what those two teachers say becomes the signal the student learns from, so modality balance isn't a correction layered on top of the objective — it's built into what counts as the training signal in the first place. That's a genuine reframe from prior approaches, which mostly used privileged context (crops, higher resolution, ground-truth annotations) to make a single teacher smarter, rather than using the disagreement between two differently-sighted teachers as the signal itself. The reported results back up the reframe: on a Qwen3.5-4B backbone, OPD-V lifts average accuracy by 15.71 percentage points while cutting training-step latency by 31.8% — enough for a compact 4B-parameter model to beat several larger state-of-the-art MLLMs outright.
The space around OPD-V is crowded, which is itself a signal of where attention is concentrating. "Distill What the Student Can See" (Fisher-Projected On-Policy Distillation) keeps only the part of the teacher's correction that falls inside the student's own visual "tangent space," so the student is never asked to chase a target beyond what it's currently capable of representing. "Seeing Before Reasoning" (arXiv:2606.19120, proposing a method called ViGOS) splits supervision across the rollout itself: an image-only teacher grades the model's visual description, while a separate, more privileged teacher grades the reasoning that follows — closing off the shortcut where a model reaches the right answer from the text prompt alone without ever really using the image. And OPOD (arXiv:2607.20918) handles the three-modality case — text, image, and audio — by layering its guidance at three levels: individual tokens, whole modalities, and full trajectories. At the modality layer specifically, a learned controller sets a different guidance strength for each of the three inputs, rather than applying one shared schedule to all of them.
Several independent groups converging on different angles of the same problem within months of each other is usually a leading indicator, not noise. It's roughly the pattern that preceded GRPO and DAPO settling into place as the default recipe for text-only reasoning, and it looks to be repeating here on the vision side.
The infrastructure signal backs this up. verl (github.com/volcengine/verl), the open-source RL training library for LLMs built around ByteDance's HybridFlow design, counts PPO, GRPO, ReMax, and REINFORCE++ among its supported algorithms, with built-in support for vision-language models and multimodal RL already in place. The scaffolding a modality-aware distillation method would need is already sitting in the dominant training framework.
EasyOPD (arXiv:2607.11012), released July 13, 2026, brings more than ten separate OPD variants — including cross-tokenizer, self-distillation, and step-wise approaches — under one framework built on top of verl, letting a team switch between them by changing a single line of YAML rather than swapping codebases.
THUNLP's OPD research group tells a similar story on a shorter timeline: verl absorbed their top-k overlap diagnostics as a built-in metric on May 26, 2026, and the day before that, MiniCPM5-1B's team folded their OPD findings directly into that model's own training recipe. Methods are showing up as default features in production frameworks within weeks of being published rather than the usual multi-year lag — a reasonable proxy for how close modality-aware weighting is to becoming a standard toggle rather than a bespoke patch.
There's also a concrete data point for how much self-distillation can close a modality gap — with an important caveat about what it's actually measuring. "Reading, Not Thinking" (arXiv:2603.09095, March 2026) studies a narrower case than general photo or scene understanding: what happens when the same textual content — a GSM8K math problem, for instance — is shown to a model as a rendered image instead of as text tokens. The baseline gap was large: Qwen3-VL-8B solved the image-rendered version of GSM8K correctly only 30.71% of the time, well below its text-mode accuracy. A small amount of self-distillation — training the model on its own text-mode reasoning traces, paired with the image version of the same problem — closed that gap almost entirely, pushing image-mode accuracy to 92.72% while holding text-mode performance steady or better. That's a genuine, order-of-magnitude closure of a modality gap through self-distillation, and it's encouraging for the broader OPD-V thesis. But it's evidence from a "text-as-pixels" reading task, not from photographic or scene-level visual understanding — so it's suggestive of what's possible, not a direct benchmark of what OPD-V itself achieves on the kind of visual content OPD-V actually targets.
There's a genuine contrarian case against this whole cluster, though. The field is treating modality imbalance almost entirely as a training-time problem to solve with increasingly elaborate distillation architectures — but a growing line of work suggests part of it is fixable much more cheaply, at inference time. "The Cost of Language" (arXiv:2604.14363) found that, across seven models, erasing a model's internal text representations costs roughly four times more accuracy than erasing its visual representations — a striking asymmetry — and showed that contrasting a model's output against a version of itself with those text representations erased recovers up to 16.9 percentage points of accuracy on individual tasks, without touching the training loop at all. If a meaningful share of the modality gap can be closed this cheaply after the fact, the OPD-V class of solutions may be over-engineered for a lot of real deployments. Training-time fixes are where the research attention and the GPU budgets are; inference-time fixes don't generate survey papers, but for a team fine-tuning a VLM for document understanding or visual inspection on an ordinary budget, "which intervention is affordable at our scale" matters more than which one wins on a leaderboard. None of the papers in this cluster, including OPD-V, really engages with that trade-off.
The hiring data points the same direction, with less precision than the research signal but a consistent trend. AI and ML job postings grew roughly 163% year-over-year into 2025, with 500,000-plus open roles and AI/ML engineers facing what's been measured as a 63% talent shortage. Postings aren't yet asking for on-policy self-distillation by name, but they're increasingly asking for the layer just above it — post-training pipelines, multimodal fine-tuning, VLM evaluation — and, per general hiring-market coverage, cross-modal alignment skills specifically.
Our own collectors corroborate the direction: GatiFlow detected OPD-V across arXiv and HackerNews with 9 mentions from 2 sources in a single day — a newly active entry, confidence 0.689 — which is roughly where a technique sits when research absorption has started but production-workflow integration hasn't caught up yet.
If you're building a VLM fine-tuning pipeline right now, the conversation worth having with your team in the next few weeks isn't which distillation loss to adopt — it's whether your evaluation harness can even detect a modality imbalance in the first place. Most teams still benchmark on aggregate accuracy (MMMU or similar), which averages away exactly the text-versus-image gap this whole cluster of papers is about. A modality-aware diagnostic layer — something as simple as splitting your eval by whether the correct answer actually depends on reading the image — has to come before a distillation intervention, not after, or you're tuning blind.
Forward catalysts: ECCV 2026 runs September 8–13 in Malmö, Sweden. The MARS2 workshop on multimodal reasoning and slow thinking, co-located with ECCV and running September 8–9 with its full speaker lineup confirmed, is the closest near-term venue to this cluster's core question of how reasoning models intersect with computer vision — though its flagship competition track leans specifically toward advertisement and marketing-video reasoning, so it's a partial rather than a direct match. ECCV's second workshop on multimodal LLMs for unified comprehension and generation (MUCG) is a second plausible venue for modality-balance work to surface formally. No confirmed release window turned up for verl's next major version or a vision-specific EasyOPD extension in the August 6–20 window; either shipping before ECCV would compress the research-to-tooling timeline further.
The irony is that the modality imbalance was visible in the architecture from day one — bolting a vision encoder onto a text-pretrained LLM backbone was never a secret design choice. It just took on-policy distillation reaching production scale to make ignoring it expensive enough to fix.
Sources:
- GitHub - nick7nlp/Awesome-LLM-On-Policy-Distillation (https://github.com/nick7nlp/Awesome-LLM-On-Policy-Distillation)
- A Survey of On-Policy Distillation for Large Language Models, Song & Zheng, arXiv:2604.00626 (https://arxiv.org/abs/2604.00626)
- When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models, arXiv:2508.10552 (https://arxiv.org/abs/2508.10552)
- MBQ: Modality-Balanced Quantization for Large Vision-Language Models, Li et al., CVPR 2025 (https://openaccess.thecvf.com/content/CVPR2025/papers/Li_MBQ_Modality-Balanced_Quantization_for_Large_Vision-Language_Models_CVPR_2025_paper.pdf)
- OPD-V — Qwen3.5-4B benchmark figures, via alphaXiv (https://www.alphaxiv.org/)
- The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models, arXiv:2604.14363 (https://arxiv.org/abs/2604.14363)
- GitHub - chrisliu298/awesome-on-policy-distillation (https://github.com/chrisliu298/awesome-on-policy-distillation)
- Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation, arXiv:2606.19120 (https://arxiv.org/abs/2606.19120)
- OPOD: On-Policy Omni Distillation, arXiv:2607.20918 (https://arxiv.org/abs/2607.20918)
- GitHub - volcengine/verl (https://github.com/volcengine/verl)
- EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models, arXiv:2607.11012 (https://arxiv.org/abs/2607.11012)
- GitHub - thunlp/OPD: Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe (https://github.com/thunlp/OPD)
- Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs, arXiv:2603.09095 (https://arxiv.org/abs/2603.09095)
- Software Engineering Job Market 2026: Data, Trends and Outlook (https://www.finalroundai.com/blog/software-engineering-job-market-2026)
- Hire LLM Engineers | Onboard in Jun, 2026 (https://www.secondtalent.com/hire-developers/llm/)
- ECCV 2026 - European Conference on Computer Vision: Everything You Need to Know - AI Expert Magazine (https://www.aiexpertmagazine.com/eccv-2026-european-conference-on-computer-vision-everything-you-need-to-know/)
- MARS2 Workshop | ECCV 2026 (https://mars2workshop.github.io/eccv2026/)
- ECCV 2026 Schedule (https://eccv.ecva.net/virtual/2026/calendar)
Disclaimer: This article is generated by GatiFlow Intelligence for informational purposes only. It does not constitute investment advice, recruitment recommendations, or legal guidance. All data is derived from public sources and AI analysis — verify independently before making decisions. Past trends do not guarantee future results.
Where this came from
Every Deep Dive starts from GatiFlow's own pipeline: 13 public developer sources, collected every six hours, with a confidence score and the evidence behind each signal. The same signals, filtered to the topics you follow, are a JSON API.
No credit card required.
Get the next one by email
One article every Saturday morning in your time zone. No account needed, and nothing else is sent to the address.
Double opt-in: you confirm by email first. What we store, and for how long, is in the privacy policy.
Tell me I am wrong
Corrections, the version of this you have lived through, or what you would like covered next. It reaches me directly and is never published. It is kept for two years so it can be read and answered; the privacy policy has the details.
0/2000