[논문리뷰] DAPD: Dual-Anchored Policy Distillation본 논문은 OPSD 과정에서 발생하는 Privilege Illusion 문제를 해결하기 위해 고안되었습니다. 기존 연구들은 privileged information을 직접 활용하거나, 교사 신호를 선별적으로 재가중하는 방식을 취했으나, 이는 근본적인 정보 비대칭성을 해결하지 못하는 한계가 있습니다 .#Review#Policy Distillation#Language Model Post-training#Information Asymmetry#Privilege Illusion#Reasoning Models#On-policy Distillation2026년 8월 3일댓글 수 로딩 중
[논문리뷰] Guava: An Effective and Universal Harness for Embodied Manipulation본 논문은 Embodied Manipulation 환경에서 복잡한 저수준 제어를 직접 학습하는 기존의 End-to-End VLA(Vision-Language-Action) 모델의 데이터 비효율성과 낮은 복구 능력을 해결하기 위해 Guava 프레임워크를 제안합니다.#Review#Embodied Manipulation#Harness Framework#Vision-Language Models#ReAct#Tool Use#Policy Distillation#Sim2Real2026년 6월 17일댓글 수 로딩 중
[논문리뷰] Hybrid Policy Distillation for LLMs본 연구는 LLM 압축 과정에서 발생하는 divergence direction, optimization strategy, data regime 간의 복잡한 상호작용 문제를 해결하고자 합니다.#Review#Knowledge Distillation#Large Language Models#Forward-Reverse KL#Policy Distillation#Logit-level Reweighting#On-policy Sampling2026년 4월 23일댓글 수 로딩 중