본문으로 건너뛰기

[논문리뷰] Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

링크: 논문 PDF로 바로 열기

Some URLs do not have content.

I have browsed the paper. Now I will proceed with the summarization following the outlined steps.

1. Extract Metadata:

  • Authors: Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang
  • Keywords: From the abstract and introduction, I can identify: Reinforcement Learning with Verifiable Rewards (RLVR), Group Relative Policy Optimization (GRPO), Large Language Models (LLMs), Off-Policy Learning, Cross-Model Trajectory Exchange, Complementary Learning.

2. Identify Key Terms & Definitions:

  • RLVR (Reinforcement Learning with Verifiable Rewards): LLM의 추론 능력을 향상시키기 위한 post-training 패러다임으로, 검증 가능한 보상을 사용하여 정책을 최적화한다.
  • GRPO (Group Relative Policy Optimization): 샘플링된 응답들 간의 상대적인 보상을 사용하여 정책을 최적화하는 RLVR의 한 변형으로, 높은 보상을 받는 trajectory의 가능성을 높이고 낮은 보상을 받는 trajectory를 억제한다.
  • Off-Policy-Aware: 행동 정책(behavior policy)과 업데이트 정책(target policy)이 다를 때 발생하는 mismatch를 인지하고 제어하는 접근 방식이다.
  • Trajectory Exchange: 서로 다른 모델(peers) 간에 생성된 응답 시퀀스(trajectory), 로그 확률, 보상 정보를 공유하여 학습에 활용하는 기법이다.
  • All-Fail Group: RLVR에서 모든 sampled response가 실패하여 reward-based policy-gradient signal이 0이 되는 경우를 지칭한다.

3. Understand Motivation & Problem Statement: The paper identifies a critical limitation in RLVR methods like GRPO: they rely on successful self-generated trajectories. However, finite rollout budgets often lead to "all-fail groups" where no successful trajectory is found, resulting in a zero reward-based policy-gradient signal. Existing solutions like increasing rollout budgets or reshaping rewards are costly or still depend on the learner's own exploration. The motivation stems from the observation that heterogeneous LLMs often succeed on complementary prompts, suggesting an opportunity for mutual learning without a designated "stronger teacher". Standard single-model RLVR fails to utilize these complementary successes. The core problem is how to effectively leverage these peer successes, specifically which prompts warrant peer supervision and how to learn from potentially mismatched peer responses under cross-model policy mismatch.

4. Analyze Method & Key Results: The proposed method is GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework for cross-model trajectory sharing in RLVR. GRAFT addresses two main design questions:

  • Which trajectories to transfer: GRAFT selects prompts where the receiver model fails entirely (an "all-fail group") but the peer model produces both successful and unsuccessful responses. It also implements Balanced Exchange to prevent directional imbalance in transferred data.
  • How to learn from peer responses: GRAFT replaces the receiver's all-fail groups with the corresponding peer groups, retaining their source-computed advantages. It controls peer influence through a Compatibility Gate (sequence-level compatibility weighting based on average token log-likelihoods) and Token-level Importance Ratio Clipping. Furthermore, it uses Peer-last Updates, processing peer-containing minibatches after on-policy minibatches to ensure the receiver's own data takes priority and clipping can moderate peer influence. The objective function for the receiver model J(θB) incorporates w(oj) (compatibility weight) and a^j (source-computed advantage).

Key Results:

  • Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO (with n=8 rollouts) by 2.1 points on average and up to 4.5 points in model-level average performance [cite: 1, Table 1].
  • In two of the three pairs, both models using GRAFT matched or exceeded GRPO trained with four times the rollout budget (n=32) [cite: 1, Table 1].
  • GRAFT outperforms HACPO and SGT by 4.0 and 1.5 points on average, respectively.
  • The gains largely persist when reusing stored peer trajectories (from independent GRPO runs), improving over GRPO by 1.8 points on average without simultaneous co-training, demonstrating compute-efficient gains [cite: 1, Figure 3b, Table 6].
  • Ablation studies show that the compatibility gate and peer-last updates are critical components, with their removal or alteration leading to significant performance degradation [cite: 1, Table 2a].

5. Formulate Conclusion & Impact: The paper concludes that GRAFT effectively addresses the limitations of single-model exploration in RLVR by enabling heterogeneous LLMs to learn from each other's complementary successes. By strategically replacing all-fail groups with informative peer trajectories and employing off-policy-aware mechanisms like compatibility weighting and token-level clipping, GRAFT significantly boosts the performance of both participating models. This research impacts the field by demonstrating that mutual learning through trajectory exchange offers a compute-efficient alternative to simply increasing rollout budgets or relying on a single strong teacher [cite: 1, Figure 3a]. It suggests a promising direction for improving LLM reasoning capabilities, especially in scenarios where computational resources for extensive rollouts are limited or where diverse model strengths can be leveraged. The finding that stored trajectories can also yield most of the benefits opens avenues for asynchronous or offline cross-model learning, reducing the need for synchronous co-training.

6. Select Figures: I will select Figure 1 (Complementary successes), Figure 2 (Overall Framework of GRAFT), and Figure 3 (Average score of Pair 1 vs. total GPU-hours & Stored Trajectories gains).

  • Figure 1: 2609.37868v1/fig_concept.png
  • Figure 2: 2609.37868v1/fig_graft_overview.png (Based on the text: "Figure 2: Overall Framework of GRAFT.") - Correction: The image in the HTML is 2609.37868v1/fig_graft_overview.png, not just fig_concept.png. I need to ensure the correct image URL is extracted.
  • Figure 3: This figure has two sub-figures (a) and (b). I'll use the combined figure if available or refer to its concept. The HTML shows 2609.37868v1/fig_compute.png for Figure 3.

Let me verify image URLs. Figure 1: ![Refer to caption](2609.37868v1/fig_concept.png) -> https://arxiv.org/html/2609.37868v1/fig_concept.png Figure 2: "Figure 2: Overall Framework of GRAFT." The image tag for this is ![Refer to caption](2609.37868v1/fig_graft_overview.png) -> https://arxiv.org/html/2609.37868v1/fig_graft_overview.png Figure 3: "Figure 3: (a) Average score of Pair 1 vs. total GPU-hours. (b) Replacing the co-trained partner with its stored trajectories..." The image tag for this is ![Refer to caption](2609.37868v1/fig_compute.png) -> https://arxiv.org/html/2609.37868v1/fig_compute.png

Looks good.저자: Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang, et al.

1. Key Terms & Definitions (핵심 용어 및 정의)

  • RLVR (Reinforcement Learning with Verifiable Rewards): Large Language Models (LLMs)의 추론 능력을 향상시키기 위한 post-training 패러다임으로, 검증 가능한 보상(verifiable rewards)을 사용하여 정책을 최적화하는 방법론입니다.
  • GRPO (Group Relative Policy Optimization): RLVR의 한 변형으로, 모델이 샘플링한 응답들(responses) 간의 상대적인 보상(relative rewards)을 사용하여 정책을 최적화하며, 높은 보상을 받는 trajectory의 확률을 높이고 낮은 보상을 받는 trajectory를 억제합니다.
  • All-Fail Group: RLVR 학습 과정에서 특정 프롬프트(prompt)에 대해 모든 sampled response가 실패하여 reward-based policy-gradient signal이 0이 되는 상황을 지칭합니다.
  • Off-Policy-Aware: 행동 정책(behavior policy)과 업데이트 정책(target policy)이 다를 때 발생할 수 있는 policy mismatch를 인지하고 이를 제어하여 학습 안정성 및 효율성을 높이는 접근 방식입니다.
  • Cross-Model Trajectory Exchange: 서로 다른 LLM 모델(peers) 간에 생성된 응답 시퀀스(trajectory), 생성 로그 확률(generation log-probabilities), 그리고 보상(rewards) 정보를 공유하여 각 모델의 학습에 활용하는 기법입니다.

2. Motivation & Problem Statement (연구 배경 및 문제 정의)

본 논문은 Reinforcement Learning with Verifiable Rewards (RLVR) 기법, 특히 Group Relative Policy Optimization (GRPO)의 핵심적인 한계를 해결하고자 합니다. 기존 GRPO 방식은 성공적인 self-generated trajectory에 크게 의존하지만, 제한된 rollout budget으로 인해 모든 sampled response가 실패하는 All-Fail Group이 빈번하게 발생할 수 있으며, 이는 reward-based policy-gradient signal을 0으로 만들어 학습을 저해합니다. Rollout budget을 늘리거나 reward reshaping을 통해 학습 신호를 복구하려는 시도는 추가적인 비용이 발생하거나 여전히 모델 자체의 탐색(exploration)에 의존한다는 한계가 있습니다.

저자들은 이 문제를 해결하기 위해, 서로 다른 Heterogeneous LLMs가 종종 상호 보완적인 프롬프트(complementary prompts)에서 성공하는 경향이 있음을 발견했습니다 [cite: 1, Figure 1]. 즉, 한 모델의 rollout에서 실패한 성공적인 응답이 다른 모델에서는 이미 발견되었을 수 있으며, 이는 지정된 "강한 교사(stronger teacher)" 없이도 상호 학습(mutual learning)의 기회를 제공합니다. 그러나 기존의 single-model RLVR은 이러한 상호 보완적인 성공 사례를 활용하지 못하고 있습니다. 따라서 본 연구의 핵심 문제는 어떤 프롬프트에 대해 peer supervision이 필요한지(which)와 교차 모델(cross-model) 간의 policy mismatch 상황에서 peer response를 어떻게 학습할 것인지(how)를 효과적으로 결정하여 이러한 상호 보완성을 학습 이득으로 전환하는 것입니다.

3. Method & Key Results (제안 방법론 및 핵심 결과)

본 논문은 GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories)를 제안합니다. GRAFT는 RLVR에서 cross-model trajectory sharing을 위한 off-policy-aware 프레임워크로, All-Fail Group 문제를 해결하고 heterogeneous models 간의 상호 보완적인 성공을 활용합니다 [cite: 1, Figure 2].

제안하는 방법론은 두 가지 핵심 질문에 답합니다. 첫째, 어떤(which) peer trajectory를 전송할 것인가? GRAFT는 receiver model이 모든 rollout에서 완전히 실패하고, peer model은 성공 및 실패 응답을 모두 생성한 프롬프트를 선택합니다. 또한, 양방향 전송 볼륨의 불균형을 줄이기 위해 Balanced Exchange 메커니즘을 적용합니다. 둘째, peer response로부터 어떻게(how) 학습할 것인가? GRAFT는 선택된 receiver group을 해당 peer group으로 대체하고, peer가 계산한 advantages를 그대로 사용합니다. Peer의 영향력은 Compatibility Gate를 통해 sequence-level 호환성 가중치(compatibility weighting)로 제어되며, Token-level Importance Ratio Clipping을 통해 cross-model mismatch를 제어합니다. 또한, receiver의 자체 데이터가 우선적으로 처리되도록 Peer-last Updates 전략을 사용하여, peer가 포함된 minibatch를 on-policy minibatch 이후에 처리합니다.

주요 실험 결과는 다음과 같습니다:

  • GRAFT는 3개의 heterogeneous model pairs와 5개의 mathematical reasoning benchmarks에서 모든 모델 블록에 걸쳐 표준 GRPO (n=8 rollouts) 대비 평균 2.1점, 최대 4.5점의 model-level 평균 성능 향상을 보였습니다 [cite: 1, Table 1].
  • 두 쌍의 모델에서, GRAFT를 사용한 모델들은 4배 더 많은 rollout budget (n=32)으로 훈련된 GRPO와 유사하거나 이를 능가하는 성능을 달성했습니다 [cite: 1, Table 1].
  • GRAFT는 기존 cross-model 학습 baseline인 HACPO 대비 평균 4.0점, SGT 대비 평균 1.5점 더 높은 성능을 기록했습니다.
  • 특히, 동시 co-training 없이 기존의 저장된(stored) peer trajectory를 재사용했을 때도, GRAFT는 GRPO 대비 평균 1.8점의 성능 향상을 유지하며, compute-efficient한 이득을 입증했습니다 [cite: 1, Figure 3b, Table 6].

4. Conclusion & Impact (결론 및 시사점)

본 논문은 RLVR 환경에서 cross-model trajectory exchange를 위한 off-policy-aware 프레임워크인 GRAFT를 성공적으로 도입했습니다. GRAFT는 heterogeneous model 간의 상호 보완적인 성공 사례를 활용하여, receiver model의 all-fail rollout group을 정보가 풍부한 peer group으로 대체함으로써 기존 GRPO의 한계를 극복합니다. 이 과정에서 Compatibility Gate, token-level clipping 및 peer-last updates를 통해 cross-model mismatch를 효과적으로 제어합니다.

이 연구는 해당 분야에 다음과 같은 중요한 시사점을 제공합니다. 첫째, GRAFT는 동일한 rollout budget 내에서 표준 GRPO보다 일관되게 우수한 성능을 보여, 모델의 추론 능력 향상에 효과적인 접근 방식임을 입증했습니다. 둘째, GRAFT는 단순히 rollout budget을 늘리는 것보다 더 나은 Compute-Efficient Gains를 제공하며, 이는 제한된 컴퓨팅 자원을 가진 환경에서 특히 중요합니다 [cite: 1, Figure 3a]. 셋째, 저장된 peer trajectory를 활용하여 대부분의 성능 이득을 유지할 수 있다는 발견은 synchronous co-training이 필수가 아님을 시사하며, 이는 향후 비동기 또는 오프라인 cross-model learning 연구의 가능성을 열어줍니다. GRAFT는 LLM post-training 패러다임에서 상호 협력적 학습의 중요성을 강조하며, 다양한 모델의 강점을 결합하여 전반적인 성능을 향상시키는 새로운 방향을 제시합니다.

Figure 1: 모델 간 상호 보완적 성공 예시

Figure 1 — 모델 간 상호 보완적 성공 예시

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글