[논문리뷰] Rufus-Air: An Open LLM Post-Training Recipe
링크: 논문 PDF로 바로 열기
I have browsed the paper. Now I will proceed with summarizing it according to the specified format and extracting figure information.
Authors: Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He, Hongye Jin, Linwei Li, Shiyang Li, Fenglin Liu, Xin Liu, Priyanka Nigam, Haoyang Wen, Zhenghao Xu, Zhuocheng Xu, Bing Yin, Qingyu Yin, Chao Zhang, Rongzhi Zhang, Zhihan Zhang, Zixuan Zhang, Zixuan Zhang, Tuo Zhao (Amazon) I need to format the author list as "제1저자, 제2저자, et al." The paper indicates authors are listed alphabetically by surname. So I'll just pick a few and add et al.
For Keywords, I'll look for terms frequently used and highlighted in the abstract and introduction.
Rufus-Air, Post-Training Recipe, LLM, GLM-4.5-Air-Base, SFT, Reinforcement Learning, Agentic Capabilities, Reproducibility, Open-Source Components. I need 5-8, academic terms.
Figures to extract:
- Figure 1: Recipe overview (pipeline architecture). This is critical.
src="2609.29421v1/pipeline.png"->https://arxiv.org/html/2609.29421v1/pipeline.png
- Figure 2: SFT composition (data distribution). This explains data.
src="2609.29421v1/sft_data.png"->https://arxiv.org/html/2609.29421v1/sft_data.png
- Table 1: Performance comparison. This shows key results. I'll pick this as a crucial Table, but the instruction states "Figure/Table" and the JSON format expects
image_url. I need to re-check if tables can be image. If not, I'll stick to image-based figures.- Looking at the HTML, Table 1 is rendered as an HTML table, not an
<img>. So I cannot select it for the JSON part. - I need another Figure. Let's check Figure 3, SFT training dynamics. This shows training progress.
src="2609.29421v1/sft_loss_metrics.png"->https://arxiv.org/html/2609.29421v1/sft_loss_metrics.png
- Looking at the HTML, Table 1 is rendered as an HTML table, not an
- Table 2: Stagewise progression. This is also a key result table. Again, it's an HTML table.
So, I will select Figure 1, Figure 2, and Figure 3.
Let's start drafting the summary.
Part 1: 요약 본문
메타데이터
저자: Chia-Yuan Chang, Renyuan Cheng, et al.
## 1. Key Terms & Definitions (핵심 용어 및 정의)
- Post-Training Recipe: 사전 학습된 LLM(Large Language Model)의 성능과 특정 능력을 향상시키기 위해 Supervised Fine-Tuning (SFT)과 다양한 Reinforcement Learning (RL) 단계를 순차적으로 적용하는 일련의 훈련 과정.
- GLM-4.5-Air-Base (106B-A12B): 본 연구의 기반 모델로 사용된 Mixture-of-Experts (MoE) 아키텍처의 오픈-웨이트 Large Language Model. 총 106B 파라미터 중 12B 파라미터가 활성화된다.
- Supervised Fine-Tuning (SFT): 레이블링된 고품질 데이터셋을 사용하여 모델을 지도 학습 방식으로 미세 조정하는 초기 단계로, 모델의 기초적인 능력과 일관된 출력 형식을 구축하는 데 중점을 둔다.
- Reinforcement Learning with Verifiable Rewards (RLVR): 코드 실행 결과나 수학적 검증 등 외부 도구를 통해 정확성을 객관적으로 검증할 수 있는
Hard Reward를 활용하여 모델의 Reasoning 및 Coding 능력을 강화하는 Reinforcement Learning 기법. - Reward Hacking: 모델이 실제 목표 달성보다 Reward Mechanism의 허점을 찾아내어 높은 Reward를 얻도록 학습되는 현상. 본 논문에서는
Reward Reliability에 따라 훈련 단계를 배치하여 이를 완화한다.
## 2. Motivation & Problem Statement (연구 배경 및 문제 정의)
본 논문은 오픈-웨이트 기반 LLM 체크포인트가 Post-Training 연구의 진입 장벽을 낮추었음에도 불구하고, 실제 Reproducible한 Post-Training Recipe가 부족하다는 문제를 제기한다. 기존 Post-Training에 대한 보고서는 종종 시스템 카드처럼 기능 설명 위주로 작성되어, 다른 연구팀이 재현하는 데 필요한 상세한 Data, Reward Design, Infrastructure, Stage Order 등의 구체적인 정보가 불투명하다는 한계점이 있다. 이에 저자들은 GLM-4.5-Air-Base 모델을 기반으로 한 Rufus-Air라는 오픈 Post-Training Recipe를 제안하며, 이는 Open-Source Components와 Public Data만을 사용하여 재현 가능성을 극대화하는 것을 목표로 한다. 특히, Compute Footprint를 최소화하여 대규모 Frontier Labs 외의 연구팀도 활용할 수 있도록 한다.
## 3. Method & Key Results (제안 방법론 및 핵심 결과)
저자들은 GLM-4.5-Air-Base 체크포인트에서 시작하여 SFT → Reasoning RL → Coding RL → Instruction-Following RL → General Agent → Coding Agent → Search Agent → RLHF의 총 8단계로 구성된 순차적인 Post-Training Pipeline인 Rufus-Air를 제안한다 [Figure 1, cite: 1]. 각 단계는 이전 단계에서 생성된 체크포인트를 기반으로 훈련되며, Reward Reliability 원칙에 따라 Hard, Verifiable Rewards를 사용하는 단계가 먼저, Softer Judge-based Signals를 사용하는 단계가 나중에 배치된다.
Supervised Fine-Tuning (SFT) 단계에서는 General Agent, General Chat, STEM, Math, Code, Coding Agent 등 6가지 Capability Category에 걸쳐 9.01M개의 샘플과 27.0B개의 학습 토큰으로 모델을 훈련하여 광범위한 Capabilities와 일관된 출력 형식을 구축한다 [Figure 2, cite: 1]. 이 SFT 체크포인트는 RL 단계의 강력한 기반을 마련하며, GLM-4.5-Air 공개 릴리스 대비 IFEval에서 +5.33 포인트, IFBench에서 +24.15 포인트, AIME 25에서 +6.63 포인트 향상된 성능을 보인다.
이어지는 Reinforcement Learning 단계들에서는 Difficulty Filtering을 통해 모델이 Productive Learning Band 내의 프롬프트에 집중하도록 하여 Learnability를 최적화한다. 예를 들어, Reasoning RL 단계에서는 GPQA 벤치마크 점수를 +5.32 포인트 상승시키며, Coding RL 단계에서는 LiveCodeBench v6 Pass@1 점수를 +7.3 포인트 향상시킨다 [Figure 4, cite: 1]. 특히 Instruction-Following RL 단계는 Multi-challenge에서 +24.7 포인트, IFBench에서 +14.0 포인트의 상당한 개선을 가져왔다 [Figure 5, cite: 1].
최종 Rufus-Air 모델은 GLM-4.5-Air 공식 Post-Training 릴리스를 Arena-Hard v2 Creative Writing 벤치마크를 제외한 모든 벤치마크에서 능가하며, INTELLECT-3 및 Nemotron-3-Super와 같은 유사 규모의 오픈 모델들과도 경쟁력 있는 성능을 보여준다 [Table 1, cite: 1]. 예를 들어, Arena-Hard v2 (HP)에서 89.1%, IFEval에서 95.4%의 Pass@1을 달성했다. 이러한 결과는 다양한 High-Quality SFT 데이터가 강력한 Capability Floor를 설정하고, Difficulty Filtering이 Productive Learning Range 내에서 RL 프롬프트를 유지하며, Reward Reliability가 효과적인 Stage Order 원칙을 제공한다는 저자들의 주요 발견을 뒷받침한다. Infrastructure 및 Engineering Choices 또한 Recipe의 중요한 부분으로, Distributed Rollout 및 Token-in/Token-out과 같은 기술적 선택이 Agentic RL의 안정성과 효율성에 기여한다 [Figure 9, cite: 1].
## 4. Conclusion & Impact (결론 및 시사점)
본 연구는 GLM-4.5-Air-Base 체크포인트를 활용한 오픈 Post-Training Recipe인 Rufus-Air를 제시하며, 총 8단계의 순차적인 훈련 파이프라인과 그 구현 세부 사항을 투명하게 공개한다. Rufus-Air는 Supervised Fine-Tuning과 Reinforcement Learning의 조합을 통해 Reasoning, Coding, Instruction-Following, Agentic Capabilities 등 다양한 영역에서 GLM-4.5-Air의 공식 릴리스를 능가하는 성능을 달성했다.
이 연구는 Reproducible하고 Reusable한 Post-Training Recipe를 제공함으로써, 대규모 컴퓨팅 자원이 부족한 연구팀도 Frontier LLM 개발에 참여할 수 있도록 LLM Post-Training 연구의 Democratization에 기여한다. 특히 Open-Source Components와 Public Data의 활용은 LLM Community의 투명성과 협력적 발전을 촉진하며, 미래 LLM 개발의 Baseline으로서 중요한 시사점을 제공한다. 저자들은 이 Recipe가 다른 연구팀에 의해 검토되고 개선되기를 기대하고 있다.

Figure 1 — Rufus-Air 파이프라인 개요
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
- [논문리뷰] VideoGen-Agent: Reinforcing Video Generation Agents
- [논문리뷰] Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
- [논문리뷰] Iris: Climbing to the Search Frontier
- [논문리뷰] Post-Training Language Models for Gold-Medal Performance in Coding Competitions
Review 의 다른글
- 이전글 [논문리뷰] Rate-distortion optimization for full-reference image quality metrics via stochastic Hessian estimates
- 현재글 : [논문리뷰] Rufus-Air: An Open LLM Post-Training Recipe
- 다음글 [논문리뷰] Training Object Permanence in World Models
댓글