DAILY RESEARCH INDEX

Distillation 进展

不是论文列表,而是按研究方向整理的每日增量。

聚合近期 arXiv 更新,保留摘要、分类、发布日期和原文入口,帮助你更快判断今天哪些论文值得继续阅读与验证。

查看全部研究方向

共 1104 篇 · 多个关键词用空格分隔,按发布日期排序。

01 TOPIC

Distillation 进展

cs.AI

Continuous Context Management

作者William Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan

展开完整摘要收起摘要

Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.

ARXIV 2609.35540 ↗
cs.CL

MAS-OPD: On-Policy Distillation for Multi-agent Systems

作者Qiyong Zhong, Mao Zheng, Mingyang Song, Houcheng Jiang, Jiajie Su, Huwei Ji, Li Zhang, Junfeng Fang

展开完整摘要收起摘要

Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.

ARXIV 2609.34234 ↗
cs.CV

SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

作者Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki Karaimer, Mozhgan Nasr Azadani, Jiayun Luo, Minh Ngoc Le, Gholamali Aminian, Hugo Buurmeijer, Yongchao Chen, Leonid Sigal, Igor Gilitschenski, Konstantinos G. Derpanis, Marco Pavone, Babak Taati

展开完整摘要收起摘要

Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.

ARXIV 2609.34044 ↗
cs.AI

BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation

作者Anglin Liu, Yanlin Wu, Ruichao Chen, Yuting Zhang, Qingyuan Zeng, Pengxiang Cai, Ziqi Gong, Muchen Li, Jintai Chen

展开完整摘要收起摘要

Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model's relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM's own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model's preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.

ARXIV 2609.33713 ↗
cs.CV

Position Aware Layer Queries for Test Time Training in Vision Language Models

作者Rajat Modi, Priyank Pathak, Xin Liang, Yogesh Singh Rawat

展开完整摘要收起摘要

Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM, incurring substantial computational cost. We observe that one forward pass with all the intermediate layer outputs already yields far more signal than the final embedding from all augmentations. We introduce Layer Query Network (LQN), a lightweight approach that can adapt a frozen VLM (teacher) in a single forward pass of the VLM via a small model (student). LQN uses Position-Aware Distillation (PAD) to mimic the teacher VLM's intermediate-layer spatial tokens by querying spatial coordinates of intermediate tokens. LQN additionally relies on Location Consistency Regularization (LCR), a self-supervision technique, replacing expensive O(H x W) image augmentation with O(1) coordinate sampling. Integrating these, LQN i) adapts and improves zero-shot CLIP ViT-B/16 by 9.8% Top-1 on OOD ImageNet, ii) outperforms the previous best GS-Bias on fine-grained classification by 3.9% Top-1, iii) achieves faster convergence than TPS for CLIP ResNet-50 (47 mins vs 55 mins), iv) generalizes adaptation to VLMs like SigLIP, EVA-CLIP, and CoCa, and lightweight students like MLP, ResNet, VGG, and v) extends to panoptic, instance, and semantic segmentation.

ARXIV 2609.34021 ↗
cs.CV

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

作者Shuning Wang, Zhiheng Wu, Xun Zhou, Chongyang Cui, Chen Jia, Bowen Liu, Chuanjie Li, Xiang Chen, Yi Yang, Yumeng Zhang, Wenjie Huang

展开完整摘要收起摘要

Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.

ARXIV 2609.33414 ↗
cs.RO

VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation

作者Jianan Wang, Haoquan Zhai, Siyang Zhang, Bin Li, Juan Chen, Jingtao Qi, Zhuo Zhang, Enze Wang, Haoxiang Jin, Chen Qian

展开完整摘要收起摘要

World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action preconditions and subsequent state transitions remains an open challenge. In this paper, we propose VIDEAS, a data distillation framework that transforms continuous physical dynamics from operational videos into explicit action semantics for foundation models. Specifically, it deconstructs visual demonstrations into discrete action trajectories and utilizes advanced vision-language models (VLMs) to extract structured knowledge encapsulating action preconditions and effects. To ensure physical consistency, we introduce a prior-guided trajectory simulation mechanism grounded within a text-based environment to rigorously validate the extracted knowledge. Notably, we incorporate negative trajectories to enrich knowledge completeness and enhance data diversity to mitigate cognitive bias. Furthermore, we present VIDEAS-WM, an 8B/9B-parameter suite of language-based world models trained on 34K high-quality samples derived from AgiBot-World dataset. Extensive experiments demonstrate that VIDEAS-WM establishes state-of-the-art performance in high-level embodied action semantic reasoning, exhibiting profound physical understanding and robust generalization across unseen scenarios.

ARXIV 2609.33464 ↗
cs.AI

SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought

作者Xiaoshu Chen, Sihang Zhou, Ke Liang, Xinwang Liu

展开完整摘要收起摘要

Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.

ARXIV 2609.33181 ↗
cs.CL

Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

作者Yadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo, Ruochen Zhao

展开完整摘要收起摘要

Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.

ARXIV 2609.33290 ↗
cs.CR

GenoTrace: Inheritable Watermarks for Genome Foundation Model Distillation

作者Guang Yang, Fengchen Liu

展开完整摘要收起摘要

Can a genome model retain a detectable record of the synthetic sequences used to train it? We study watermark inheritance through distillation with GenoTrace, a codon-aware extension of green-list watermarking. Two token-level factors modulate the teacher's generation bias using codon position and organism-specific codon usage. The resulting sequences train a smaller student, whose outputs are audited without an active watermark processor. In a three-seed GenomeOcean-500M-to-100M experiment, the joint configuration achieves a mean audit score of 17.88 and 94.5% detection at a fixed threshold. It retains 49.0% detection after key-aware token substitution, compared with 0% for the available single-seed plain-watermark comparator, and 47.0% after combined mechanism-targeted nucleotide edits. Additional experiments establish inherited signal across five organism-conditioned datasets and teacher-student size ratios up to 40. Component ablations and computational sequence-quality assays reveal distinct operating points for detection strength and coding coverage. GenoTrace provides a practical token-level construction and an empirical account of how genomic structure shapes inherited watermark signals. The findings concern shared-tokenizer distillation and the tested editing procedures, with calibration and biological utility treated as separate evaluation requirements.

ARXIV 2609.35881 ↗
cs.RO

Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning

作者Boyuan Zhang, Yingjun Du, Xiantong Zhen, Ling Shao

展开完整摘要收起摘要

Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not capture how prediction errors propagate under recursive rollout. We decompose multi-step rollout error into the errors introduced at individual steps and their propagation through subsequent transitions. We show that state-affine dynamics are precisely the differentiable transitions with state-independent Jacobians, eliminating the nonlinear propagation residual and making the error propagation operators depend only on the action sequence. Guided by this result, we introduce SALT (State-Affine Latent Transition), an action-conditioned state-affine dynamics model in which the action modulates both the state transformation and the additive update. We train SALT through recursive multi-step rollout supervision, feeding each predicted latent state back into the transition so that training matches how the model is used during planning. Across four visual planning environments, SALT exhibits $1.48$--$2.19\times$ higher one-step prediction error than the matched LeWM baseline, yet improves closed-loop success in every environment by $10.0$ percentage points on average. On OGBench-Cube, the fraction of episodes that fail with a sharp rise in model-predicted cost after execution decreases from $23.3%$ to $2.0%$.

ARXIV 2609.33595 ↗
cs.CV

Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation

作者Zhihao Chen, Yiyuan Ge, Ziyang Wang, Pu Cao, Lu Yang

展开完整摘要收起摘要

Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deploy due to their large parameter counts and computational requirements. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence selection explicit and compressible. The teacher introduces a small set of learnable query slots to extract global and local action-sufficient navigable evidence from panoramic observations via a Navigable Query Generator, then progressively grounds these evidence tokens to the instruction with an Instruction-Query Aligner for policy prediction. Second, using this explicit query bottleneck as a distillation interface, we train a compact student by transferring both where to attend and what to do. We distill the teacher's global and local navigable queries with a navigation-aware token-adaptive objective, then further match action distributions during fine-tuning. Experiments on standard VLN benchmarks demonstrate that our student nearly matches the teacher's navigation performance while reducing the number of parameters by 93.65% compared to the teacher.

ARXIV 2609.33097 ↗
cs.LG

Leaky Students: Membership Inference against On-Policy Distillation

作者Zhexi Lu, Mingzhi Zhu, Stacy Patterson, Lei Yu

展开完整摘要收起摘要

On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher during distillation remains poorly understood. To the best of our knowledge, we present the first systematic study of membership inference in this setting. We find that fresh student trajectories expose sparse membership signals that fixed reference-answer losses often miss. These signals are mixed with probability changes caused by training on other records. We introduce Leaky, which samples fresh trajectories from the target model and compares its token log-probabilities with the maximum across matched reference models trained without the candidate records. It applies Leaky ReLU to the resulting gaps, preserving positive gaps and downweighting negative gaps as an approximate correction for incidental positive gaps in non-members. Across fifteen targets spanning mathematics, medical question answering, and code generation, Leaky outperforms all evaluated baselines and achieves mean AUROC 0.875, compared with 0.614 for the strongest baseline on each target in the main evaluation. On the same sampled trajectories, the strongest baseline achieves mean AUROC 0.826. These results show that students trained through OPD can expose the membership of records used for teacher supervision, even when fixed reference-answer losses provide little evidence of membership.

ARXIV 2609.33136 ↗
cs.LG

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

作者Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam

展开完整摘要收起摘要

On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.

ARXIV 2609.33200 ↗
cs.LG

Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

作者Mingju Chen, Can Lv, Jinrong Liu, Huan Zhang, Heng Chang, Shiji Zhou

展开完整摘要收起摘要

Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify Decision--Timestamp Mismatch: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 % and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD

ARXIV 2609.33391 ↗
cs.LG

ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning

作者Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Xuemin Chen, Tianshu Yu

展开完整摘要收起摘要

Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but this raises two questions: how should specialization be organized, and how should specialist guidance be integrated? We introduce ChemOPD, which addresses both. We estimate task affinities from supervised fine-tuning gradients and solve a constrained mixed-integer program(MIP) to construct partially overlapping specialist groups. During distillation, we retain a generalist teacher trained on all tasks so that specialist guidance supplements rather than replaces its supervision. Our anchor-residual objective gradually increases the routed specialist's contribution on student-generated responses. On ChemCoTBench, affinity-guided specialization produces task-dependent gains over the generalist teacher and improves several capabilities beyond semantic task grouping. Yet stronger teacher-side performance does not automatically yield stronger students: with the same specialists and routes, anchor-residual OPD improves most reported metrics over specialist-only distillation and realizes a larger share of the available teacher gains. These results highlight specialization and capability integration as connected but distinct design problems in chemical reasoning.

ARXIV 2609.33838 ↗
cs.LG

Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs

作者Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. Brinton

展开完整摘要收起摘要

Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.

ARXIV 2609.34009 ↗
cs.LG

UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents

作者Wenbo Zhang, Pengcheng Xu, Weizhi Du, Jing Zhang, Hengrui Cai

展开完整摘要收起摘要

On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to $15.8%$ relative to standard OPD.

ARXIV 2609.34036 ↗
cs.CL

TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

作者Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng, Zaiyi Zheng, Ruocheng Guo, Yushun Dong, Jundong Li

展开完整摘要收起摘要

Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.

ARXIV 2609.33426 ↗
cs.CL

Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation

作者Yuhao Sun, Binrui Wu, Zhuoer Xu, Ming Wen, Haoxiang Xu, Bin Chen, Yan Lin, Qianzijing Zhang

展开完整摘要收起摘要

On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future behavior, while consequential guidance may be beyond the current student's reach or fail to persist without privileged input. We formulate long-horizon OPD as hierarchical supervision allocation and argue that productive guidance lies at the intersection of future utility and current learnability. Crucially, this intersection evolves as the student learns. Based on this principle, we propose LENS-OPD, a coarse-to-fine framework that organizes supervision through Locate, Validate, and Refine. Locate adapts trajectory exposure to the student's evolving competence and proposes a candidate decision for intervention. Validate tests whether teacher guidance at that decision improves the same student's subsequent behavior. Refine internalizes the beneficial guided behavior into the deployable policy and concentrates token-level supervision on decisive teacher-student conflicts within the validated turn. These stages are nested: each finer allocation is conditioned on the coarser decision, rather than being optimized as an independent importance score. Experiments across multiple long-horizon agent benchmarks and student-teacher configurations show that LENS-OPD consistently improves task performance over vanilla OPD and strong curriculum- and selection-based baselines. Our results suggest that effective long-horizon distillation requires teaching at the right depth, the right decision, and the right token.

ARXIV 2609.33409 ↗
cs.AI

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

作者Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo

展开完整摘要收起摘要

On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher--student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.

ARXIV 2609.33455 ↗
cs.AI

Dual-Vocabulary Language Model for Cross-Tokenizer Distillation

作者Kedi Chen, Chen Lin, Yutao Sun, Wei Zhang

展开完整摘要收起摘要

On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.

ARXIV 2609.33816 ↗
stat.ML

A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions

作者Luyang Fang, Haoran Lu, Jiazhang Cai, Tao Wang, Huimin Cheng, Wenxuan Zhong, Ping Ma

展开完整摘要收起摘要

Knowledge Distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient ``student'' counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive settings. Although KD is widely used in practice, it is often viewed primarily as an engineering technique, with a unified statistical perspective remaining less developed. This review bridges that gap by presenting a unified Bayesian formulation of KD that formulates teacher predictions as prior information. This provides a principled interpretation of how teacher information is incorporated into student learning and establishes a rigorous connection to uncertainty quantification. We demonstrate how this foundational lens reconciles classical distillation with modern extensions in generative and foundation-model systems, showing that contemporary developments remain rooted in these same statistical principles. By synthesizing theory with emerging methodologies and diverse applications, this review provides a conceptual roadmap and identifies critical open problems for the future of the field.

ARXIV 2609.33727 ↗
cs.LG

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

作者Wenze Lin, Jiyuan Long, Jiale Zhao, Shenzhi Wang, Xitai Jiang, Ce Luo, Rui Lan, Qianli Ma, Fukang Wen, Hui Wu, Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Gao Huang

展开完整摘要收起摘要

Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in this work, we find that KL divergence may not be necessary for OPD. We show that simply preserving the update direction is sufficient for effective OPD. As long as the update direction is toward the teacher, OPD works. More precisely, it is not the direction of every token, but the direction of a small subset of tokens where the teacher and student disagree strongly. We first show that simply assigning a reward of (+1) to tokens where the teacher probability is higher than the student probability and (-1) where it is lower, which merely encourages updates toward the teacher, reproduces almost the same training mode as OPD with reverse KL. We further show that only the direction of a small subset of tokens with large teacher-student disagreement is critical, and training works as long as their update direction is toward the teacher, even if other tokens are pulled away from the teacher. And as an application of these findings, we introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to improve Multi-Teacher On-Policy Distillation (MOPD). Unlike MOPD, which routes each sample to a single teacher and may cause capability conflicts across domains, C-MOPD lets every sample be supervised by all teachers. Experiments show that C-MOPD consistently outperforms MOPD on both math and code benchmarks. Our code is available at https://github.com/LeapLabTHU/KL-Free-OPD.

ARXIV 2609.33791 ↗
cs.CL

Hesitation-Aware On-Policy Distillation for Diffusion Language Models

作者Jianguo Huang, Lipeng Wan, Yanchen Deng, Bo An

展开完整摘要收起摘要

Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD) builds on this process by matching the student to a stronger teacher, yet only at the committed positions. We argue that this discards much of the useful signal, which resides in the uncommitted proposals, where the student has made a prediction but is not yet confident enough to commit it. We call these proposals hesitations. In our pilot study on an SDAR-4B student, hesitations make up only 24% of supervisable state-position pairs but carry 66% of the teacher-student divergence. To exploit this signal, we propose Hesitation-Aware On-Policy Distillation (HOPD), which extends teacher distribution matching to every masked position of each denoising step. Because hesitations are not equally informative, we further allocate supervision using hindsight from the completed trajectory, placing more weight on positions whose proposal was later disagreed with the final token and on blocks where first-step proposals rarely survive. Since both models already produce distributions at all masked positions, HOPD requires no additional forward passes over TOPD. The only extra cost is evaluating the loss at more positions. With SDAR-1.7B and SDAR-4B students distilled from TraDo-8B-Instruct, HOPD achieves the best average score among the evaluated methods on five math and coding benchmarks, under both static and dynamic decoding and at both scales. It also speeds up decoding. On SDAR-4B, the HOPD student hesitates less and commits 11% more tokens per denoising step than TOPD, while reaching higher accuracy.

ARXIV 2609.33301 ↗
cs.LG

DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation

作者Ao Yu, Weibo Gao, Heng Zhou, Linan Yue, Rui Li, Suyi Liu, Yu Yan, Yizhong Zhang, Qi Liu

展开完整摘要收起摘要

On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.

ARXIV 2609.33711 ↗
cs.RO

AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models

作者Rui Wang, Xiangyu Wang, Donglin Yang, Yibo Li, Canyang Chen, Zhongrui Wang, Xiaojuan Qi

展开完整摘要收起摘要

World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.

ARXIV 2609.33748 ↗
cs.CV

RepFlow: Reciprocal Supervision Improves Generation and Representation in Flow Models

作者Weili Zeng, Feng Tian, Shengqi Liu, Yichao Yan

展开完整摘要收起摘要

Generative models learn visual structure through denoising, yet their internal states are entangled with both noise level and network depth, making it difficult to obtain a stable visual representation from the generator itself. We introduce RepFlow, which learns such a representation from the generator's evolving computation and uses it to guide generation. Specifically, a separate timestep-free encoder is trained, through a timestep-conditioned predictor, to recover generator states across depths and noise levels from masked clean images. By excluding the reference noise and masked-out content from the encoder's input, we encourage the encoder to distill visual information that is recoverable from the visible image context and predictive of generator states. The representation learned from the generator's evolving states is then fed back to guide and improve the generator, whose updated states provide supervision for further representation learning, forming a reciprocal learning process. This reciprocal process improves multi-step generation and representation quality, as measured by frozen linear probing on ImageNet with latent-space SiT and pixel-space JiT, without an externally pretrained representation teacher. Across the two unconditional settings, FID decreases by $19.2$--$40.8%$ relative to native training, while linear-probe accuracy improves by 6.80--10.04 percentage points over the best searched raw generator features. The learned representation also serves as a distributional metric for one-step JiT post-training, extending its role from instance-level alignment to distribution-level supervision.

ARXIV 2609.33217 ↗
cs.CV

MM-OPD: Towards One More Bottleneck Between Perception and Reasoning

作者Jintao Tong, Yujing Lou, Zhanming Shen, Jiaqi Gu, Lubin Fan, Ruixuan Li, Yue Wu, Jieping Ye, Yixiong Zou

展开完整摘要收起摘要

Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding the model, question, and decoding fixed, we replace images with their caption or code representations (symbolic views), which seems to be redundant given the clear image structures, but the performance surprisingly improves by 10.2% to 23.6% across model scales and datasets. We term this performance gap as the Symbolic Visual Gap and then take a closer look at it. Through experiments, we find that although the visual evidence can already appear in the reasoning trace for the image-input model, the symbolic-view-input model shows much higher attention to the correct evidence than the image-input model. This suggests that despite good capabilities from current works in perception and reasoning themselves, another bottleneck exists between perception and reasoning in selecting perceived visual information as appropriate evidence for subsequent reasoning. To handle this bottleneck, since the symbolic view steers attention toward correct evidence and is readily obtained at scale, it provides supervision for evidence selection without manually labeled evidence. Building on this, we introduce MM-OPD, a multimodal on-policy self-distillation framework for symbolic-to-visual correction that transfers guidance from symbolic-conditioned behavior to the image-conditioned policy through residual token-level targets, steering the model toward correct visual evidence. Experiments across benchmarks and model scales show that MM-OPD improves a broad range of multimodal abilities, with gains in visual perception, chart and document understanding, mathematical reasoning, and general VQA.

ARXIV 2609.32690 ↗
cs.AI

PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents

作者Junchi Chen, Changtao Miao, Yuxiao Xiang, Zhenchao Jin, Haojie Yuan, Qi Chu, Tao Gong, He Liu, Bo Zhang, Jiansheng Cai, Zhe Li, Nenghai Yu

展开完整摘要收起摘要

Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors assess subtasks in isolation. To address this gap, we introduce PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment. For training and evaluation, we construct a Multi-Step Plan Safety (MSP-Safe) dataset through paired task construction, plan generation using diverse planners, and safety annotation by three judges. Task-oriented SFT on MSP-Safe establishes fundamental plan-safety assessment capabilities, yet a substantial gap remains between compact models suitable for real-time deployment and stronger but costlier large models. Accordingly, we propose Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision along their on-policy trajectories. It combines token-level distribution transfer from a fine-tuned strong teacher with probability-routed sequence-level compensation, retaining student-generated targets when the student favors the reference safety decision and using teacher-reconstructed targets otherwise. Across all test subsets, PlanGuard-2B achieves average 87.15% ACC and 87.21% F1, demonstrating effective whole-plan physical-risk detection at compact model scale. Code and dataset will be publicly released.

ARXIV 2609.32801 ↗