[논문리뷰] Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning본 논문은 Long-CoT 추론 학습 시 기존 GRPO와 같은 방식이 응답 내 모든 토큰에 동일한 중요도를 부여하여 발생하는 최적화 비효율 문제를 해결하고자 합니다.#Review#Long-CoT Reasoning#RLVR#GRPO#On-policy Self-distillation#Counterfactual Sensitivity#Credit Reallocation2026년 8월 2일댓글 수 로딩 중
[논문리뷰] β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation본 논문은 기존의 OPSD가 가지는 취약성과 엔지니어링의 어려움을 해결하는 것을 목표로 합니다. 기존 연구들은 학생 모델이 Privileged Teacher를 직접 모방하도록 강제하며, 학습 과정에서 Reference Policy로부터 얼마나 멀어져야 하는지에 대한 명확한 제어 메커니즘을 제공하지 못합니다 .#Review#On-policy Self-distillation#Policy Optimization#KL-regularization#Logit Interpolation#Return-to-go Credit Assignment2026년 7월 30일댓글 수 로딩 중
[논문리뷰] Visual Contrastive Self-Distillation기존의 OPSD는 Teacher와 Student 간의 정보 비대칭성(Asymmetry)을 확보하기 위해 privileged answer, Reasoning trace, 또는 시각적 증거(Visual evidence)와 같은 추가적인 외부 정보에 의존해왔습니다.#Review#On-policy Self-distillation#Vision-Language Models#Visual Grounding#Contrastive Decoding#Knowledge Distillation#Input Conditioning2026년 7월 23일댓글 수 로딩 중
[논문리뷰] Learning from the Self-future: On-policy Self-distillation for dLLMs본 논문은 기존의 OPSD 방법론들이 Autoregressive (AR) 모델에 최적화되어 있어, dLLMs의 고유한 특성인 비자기회귀적 생성 방식과 충돌한다는 문제를 해결하고자 합니다.#Review#On-policy Self-distillation#Diffusion Large Language Models#dLLMs#Step-level Divergence#Self-future#Reasoning Benchmarks2026년 6월 16일댓글 수 로딩 중