[논문리뷰] Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals본 논문은 기존 LLM 사후 학습 방식이 탐색(exploration)과 분포 정렬(distribution alignment)을 강하게 결합하여 컴퓨팅 효율성과 확장성을 저해하는 문제를 해결합니다.#Review#Post-training#Proxy Exploration#Update Signal Transfer#LLM Alignment#Modular Training#Weak-to-Strong Generalization2026년 7월 13일댓글 수 로딩 중