[논문리뷰] Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR본 논문은 RLVR이 단일 샘플 정확도(avg@k)는 향상시키지만, 모델의 근본적인 추론 범위(pass@k)를 충분히 확장하지 못하는 문제를 해결하고자 합니다. 기존 연구들은 최적화 과정에 집중할 뿐, 학습 시의 롤아웃(rollout) 구조를 고정된 병렬 방식으로 유지하여 탐색이 제한적이라는 한계가 있습니다.#Review#Reinforcement Learning#Large Reasoning Models#RLVR#Tree-Structured Policy Optimization#Difficulty-Adaptive#Reasoning Coverage2026년 9월 9일댓글 수 로딩 중