본문으로 건너뛰기

[논문리뷰] QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

링크: 논문 PDF로 바로 열기

저자: Weiqi Wang, Yuxin Zhou, Mouxiang Chen, et al.

1. Key Terms & Definitions (핵심 용어 및 정의)

  • xLong-horizon tasks: 단일 실행이 수 시간, 수백 번의 모델-환경 상호작용, 그리고 롤아웃당 거의 1M 토큰에 달하는 Large Language Model (LLM) 에이전트 태스크를 의미한다.
  • Elastic Scheduler: 라이브 execution을 중단하지 않고 rollout과 training 사이에 GPU를 탄력적으로 재할당하여 GPU 유휴 시간을 최소화하는 QwenGyre의 핵심 구성 요소이다.
  • Trajectory Processor: 복잡한 non-linear branching을 가진 execution에서 trajectory의 중복을 제거하고 partial progress를 평가하여 training efficiency를 향상시키는 QwenGyre의 구성 요소이다.
  • Colocate: 기존 online Reinforcement Learning (RL) 프레임워크 중 하나로, GPU pool을 rollout과 training 사이에 교대로 사용한다.
  • Async: 기존 online RL 프레임워크 중 하나로, GPU를 rollout과 training을 위한 전용 pool로 정적으로 분할하여 사용한다.
  • TITO (Token-in, Token-out): model call의 정확한 input 및 output token ID, behavior log-probabilities, 그리고 conditioning context를 기록하는 메커니즘이다.
  • Trajectory Tree: execution 내의 model call 기록을 shared prefixes와 함께 조직하여 branching history를 효과적으로 표현하는 데이터 구조이다.

2. Motivation & Problem Statement (연구 배경 및 문제 정의)

본 논문은 xLong-horizon tasks를 수행하는 LLM agents의 online Reinforcement Learning (RL)에 있어 GPU underutilization 및 trajectory redundancy 문제를 해결하고자 한다. LLM agents가 repository generation 및 codebase-wide refactoring과 같은 xLong-horizon tasks를 해결함에 따라, 단일 rollout execution이 수 시간 동안 수백 번의 model-environment interaction과 약 1M 토큰을 처리하게 된다. 이러한 xLong-horizon 환경에서 기존 online RL 프레임워크들은 두 가지 주요 challenge에 직면한다.

첫째, xLong-horizon에서의 extreme rollout variance와 prolonged rollout delays는 massive GPU idling을 초래한다 [Figure 1a]. Colocate 방식은 모든 rollout이 종료될 때까지 training이 지연되어 GPU가 유휴 상태에 놓이며, Async 방식은 GPU를 정적으로 분할하여 training GPU가 데이터 대기 상태에 놓이거나 rollout GPU의 유휴 용량을 활용하지 못한다. 둘째, black-box harnesses가 saturated histories를 compact하고 sub-agents에 sub-tasks를 위임하며 failed execution paths를 retry하는 등 complex non-linear branching을 생성하여 massive trajectory redundancy가 발생하고 training efficiency를 저해한다 [Figure 1b]. 이로 인해 training pipeline이 엄청난 양의 paths와 tokens에 압도될 위험이 있다.

3. Method & Key Results (제안 방법론 및 핵심 결과)

본 논문은 xLong-horizon online RL을 위한 end-to-end framework인 QwenGyre를 제안하며, 이는 elastic scheduler와 trajectory processor라는 두 가지 핵심 구성 요소로 이루어진다 [Figure 3]. QwenGyre의 elastic scheduler는 live execution을 중단하지 않고 rollout과 training 사이에 GPU를 동적으로 재할당한다. 이는 topology-aware resource organization을 통해 elastic cells을 관리하고, rollout waterlevel에 따라 idle capacity를 training에 할당한다. 또한, request rerouting 및 KV-cache migration을 포함하는 harness-preserving role transitions를 통해 GPU role changes 중에도 harness state를 유지하며 inference service를 보장한다. centralized dynamic data parallelism을 갖춘 streaming training은 새로운 GPU node들이 진행 중인 batch에 합류할 수 있도록 하여 idle gaps를 최소화한다.

Trajectory processor는 기록된 model call을 shared prefixes를 가진 trajectory trees로 조직화하여 execution의 branching history를 보존한다 [Figure 2]. TITO 메커니즘을 통해 각 output의 original conditioning context를 정확히 기록하며, partial scoring 기능을 통해 timeout 이후에도 assessable partial progress에 대한 유효한 score를 부여한다. Trajectory sampling은 role priority에 따라 execution당 최대 Jmax (5)개의 trajectory를 선택하고, shared target은 한 번만 contribute하도록 masking하여 training cost를 효과적으로 제어한다. loss는 execution 내의 selected trainable tokens에 대해 평균화하여 execution의 aggregate weight가 trajectory 수에 따라 증가하지 않도록 한다.

실험 결과, QwenGyre는 NL2RepoBench에서 Qwen 3.8 2.4T 모델을 사용하여 48 training steps 동안 52.5%에서 58.5%로 absolute gain 6.0%를 달성했다 [Figure 8f]. 또한, 다양한 training datasets(NL2RepoBench, DeepSWE, TerminalBench) 및 모델(Qwen 3.6 122B, Qwen 3.8 2.4T)에 걸쳐 Colocate 대비 최대 1.85x, Async 대비 최대 1.78x의 end-to-end speedups를 달성했다 [Figure 7d, Figure 8d, Table 2]. 이는 GPU budget 및 scheduling staleness가 동일한 조건에서 달성된 결과이다. 특히 NL2RepoBench와 Qwen 3.8 2.4T 조합에서는 Async 대비 1.78x, Colocate 대비 1.21x의 speedup을 보였다 [Figure 8d]. Ablation study를 통해 streaming training과 fine-grained elastic allocation이 end-to-end time을 효과적으로 단축하며 GPU utilization을 크게 향상시킴을 입증했다 [Table 1].

4. Conclusion & Impact (결론 및 시사점)

본 연구는 unmodified black-box agent harnesses를 통해 xLong-horizon online RL을 가능하게 하는 elastic scheduling과 trajectory processing을 결합한 QwenGyre 프레임워크를 제시한다. QwenGyre는 live execution을 보존하면서 GPU를 유연하게 재할당하고, original contexts 및 balanced execution-level contributions를 가진 bounded training samples를 구성함으로써 GPU underutilization과 trajectory redundancy라는 xLong-horizon RL의 주요 challenge를 효과적으로 해결한다.

이러한 QwenGyre의 기능은 flagship model인 Qwen 3.8 2.4T의 xLong RL training을 통해 입증되었다. 결과적으로 QwenGyre는 세 가지 workload에 걸쳐 Async 대비 최대 1.78x, Colocate 대비 최대 1.85x의 speedup을 달성했으며, 동시에 training scores는 comparable 수준을 유지했다. 이 연구는 LLM agents가 복잡하고 장기적인 software engineering tasks를 수행하는 데 필요한 RL infrastructure의 scalability와 efficiency를 크게 향상시키는 중요한 implication을 갖는다. 이는 fully autonomous end-to-end software development와 같은 frontier agentic capabilities의 발전을 가속화하는 데 기여할 것으로 기대된다.

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글