본문으로 건너뛰기

#On-Policy Distillation

31개의 포스트

[논문리뷰] Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

댓글 수 로딩 중

[논문리뷰] ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

댓글 수 로딩 중

[논문리뷰] Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

댓글 수 로딩 중

[논문리뷰] OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

댓글 수 로딩 중

[논문리뷰] DanceOPD: On-Policy Generative Field Distillation

댓글 수 로딩 중

[논문리뷰] OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

댓글 수 로딩 중

[논문리뷰] OPRD: On-Policy Representation Distillation

댓글 수 로딩 중

[논문리뷰] Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation

댓글 수 로딩 중

[논문리뷰] CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation

댓글 수 로딩 중

[논문리뷰] GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

댓글 수 로딩 중

[논문리뷰] AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

댓글 수 로딩 중

[논문리뷰] Healthcare AI GYM for Medical Agents

댓글 수 로딩 중

[논문리뷰] Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

댓글 수 로딩 중

[논문리뷰] KAT-Coder-V2 Technical Report

댓글 수 로딩 중

[논문리뷰] Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

댓글 수 로딩 중

[논문리뷰] Typhoon-S: Minimal Open Post-Training for Sovereign Large Language Models

댓글 수 로딩 중