본문으로 건너뛰기

#On-policy Distillation

19개의 포스트

[논문리뷰] Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

댓글 수 로딩 중

[논문리뷰] DAPD: Dual-Anchored Policy Distillation

댓글 수 로딩 중

[논문리뷰] CAST: Game Solvers as Turn-Level Teachers for LLM Agents

댓글 수 로딩 중

[논문리뷰] OvisOCR2 Technical Report

댓글 수 로딩 중

[논문리뷰] AsyncOPD: How Stale Can On-Policy Distillation Be?

댓글 수 로딩 중

[논문리뷰] Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation

댓글 수 로딩 중

[논문리뷰] Trajectory-Refined Distillation

댓글 수 로딩 중

[논문리뷰] Trust-Region Behavior Blending for On-Policy Distillation

댓글 수 로딩 중

[논문리뷰] HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

댓글 수 로딩 중