[논문리뷰] Negative Self-Distillation: Learning to Reason by Avoiding Flaws본 논문은 기존 On-Policy Self-Distillation (OPSD) 패러다임이 Large Language Models (LLMs)의 복잡한 추론 task 성능을 저하시키는 문제를 해결하고자 Negative Self-Distillation (NSD)를 제안합니다.#Review#Negative Self-Distillation#LLM Reasoning#Unlikelihood Training#Self-Bootstrapped Learning#Mathematical Reasoning#Adaptive Gating#Reflection2026년 9월 10일댓글 수 로딩 중