[논문리뷰] CompoWorld: Compositional Environment Scaling for General Agents
링크: 논문 PDF로 바로 열기
I have browsed the paper. Now I need to extract the information for the summary and figure part.
Part 1: Summary
- Authors: Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan (AllSpark Team)
- Keywords: From the abstract: "Compositional Environment Scaling", "General Agents", "Reusable Services", "World Model", "Completion-Focused Rubric Reward", "Reinforcement Learning", "Supervised Fine-Tuning", "Cross-Environment Tasks". I'll select 5-8 of these.
Compositional Environment Scaling,General Agents,Reusable Services,Cross-Environment Tasks,Supervised Fine-tuning (SFT),Reinforcement Learning (RL),Completion-Focused Rubric Reward
- Key Terms & Definitions:
- Compositional Environment Scaling (CompoWorld): A framework that constructs cross-environment tasks by composing a finite library of reusable services, expanding the task space for training general agents.
- Typed Environment Modeling: Standardizes service states and interaction interfaces using typed Python schemas, enabling agents to infer service entities and construct Pydantic models for environment states.
- Completion-Focused Rubric Reward: A reward mechanism that assigns higher weights to criteria with lower pass rates within each rollout group during reinforcement learning, encouraging full task completion.
- Service-level Dependency Graph: A graph connecting independently executable services, where directed edges indicate information or state dependencies required for task completion.
- Agentic Synthesis and Verification: A process where coding agents generate state models, tool implementations, and test suites, iteratively repairing implementations and conducting independent validation passes using adversarial test cases.
- Motivation & Problem Statement:
- LLM agents need interactive environments for training, but existing approaches primarily generate tasks within a single environment.
- Real-world workflows often require agents to connect information and actions across multiple services, a capability not adequately addressed by current environment scaling methods.
- CompoWorld aims to tackle three challenges: ensuring reliably executable and compatible automatically generated services, synthesizing meaningful and verifiable cross-service dependent tasks, and providing useful supervision for both SFT and RL, especially for partial workflow completion.
- Method & Key Results:
- Methodology: CompoWorld proposes a three-part framework [Figure 2]. First, it builds individual services through Typed Environment Modeling using MCP specifications and Pydantic models, with Agentic Synthesis and Verification in an isolated sandbox. For tools that cannot be reliably implemented, a Selective World-Model Simulation using LLMs is employed. Second, it generates cross-environment tasks by sampling services and connecting them into a service-level dependency graph via a random-walk procedure. An agent then performs Agentic Task Generation to instantiate states, goals, and constraints, followed by verification. Third, agents are trained using verified successful trajectories for SFT and task rubrics for RL, specifically introducing the Completion-Focused Rubric Reward to prioritize less frequently satisfied criteria, combined with Group Relative Policy Optimization (GRPO).
- Key Results: CompoWorld, trained on Qwen3.6-35B-A3B using 448 services (10,130 tools), 3K SFT trajectories, and 1K RL tasks, achieves an average gain of 9.17 points across eight challenging agent benchmarks compared to its backbone model [Table 1]. On AutomationBench, it significantly improves task success rate from 10.33% to 32.33% (+22.00 points), surpassing frontier models like Claude Opus 4.6 (25.50%) and GPT-5.4 (27.67%) and leading all compared agent-specialized 35B-A3B models [Table 1]. Furthermore, CompoWorld shows improvements across various domains in AutomationBench 1.0.6 with an average score increase from 41.94% to 72.68%, notably in HR (+49.40 points), Marketing (+32.51), and Sales (+31.45) [Table 2]. Training with composed environments consistently yields substantial performance improvements over training on single environments, demonstrating more effective training signals [Figure 5 (a)].
- Conclusion & Impact:
- CompoWorld successfully demonstrates a framework for compositional environment scaling, enabling agents to learn by connecting reusable services through task-specific causal dependencies.
- The framework's automated construction and verification pipeline, coupled with the Completion-Focused Rubric Reward, effectively generates diverse and complex cross-environment tasks and encourages full task completion.
- This research highlights service composition as a critical dimension for environment scaling, significantly improving LLM agents' ability to coordinate familiar capabilities in novel, complex workflows, thereby advancing the development of more general and capable AI agents in both academia and industry.
Part 2: Important Figure Information
I need to find the <img> tags and their src attributes.
- Figure 1:
2609.33665v1/intro-performance.png->https://arxiv.org/html/2609.33665v1/intro-performance.png - Figure 2:
2609.33665v1/overview.png(I searched for "Figure 2" and found it) ->https://arxiv.org/html/2609.33665v1/overview.png - Figure 3:
2609.33665v1/domains.png(I searched for "Figure 3" and found it) ->https://arxiv.org/html/2609.33665v1/domains.png - Figure 4:
2609.33665v1/sft-scaling.png(I searched for "Figure 4" and found it) ->https://arxiv.org/html/2609.33665v1/sft-scaling.png - Figure 5:
2609.33665v1/composed-vs-single.png(I searched for "Figure 5" and found it) ->https://arxiv.org/html/2609.33665v1/composed-vs-single.png - Figure 6:
2609.33665v1/ablation-rw.png(I searched for "Figure 6" and found it) ->https://arxiv.org/html/2609.33665v1/ablation-rw.png
I will choose Figure 2 (CompoWorld overview), Figure 1 (performance comparison), and Table 1/Figure 5(a) as a graphical representation of key results or method. Figure 1 shows performance. Figure 2 shows the architecture. Figure 5(a) shows the comparison of composed vs single environment scaling. I will choose Figure 1, Figure 2, and Figure 5.
Let's refine the captions for the chosen figures.
- Figure 1: Performance of CompoWorld and five foundation models on four challenging agent benchmarks. -> "모델 성능 비교"
- Figure 2: CompoWorld overview. -> "CompoWorld 개요"
- Figure 5 (a): Comparison between single-environment scaling and composed-environment scaling; -> "환경 스케일링 비교"
Now I will construct the final output.저자: Xiao-Wen Yang, Weiyi Xu, Wen Da, et al.
1. Key Terms & Definitions (핵심 용어 및 정의)
- Compositional Environment Scaling (CompoWorld): 재사용 가능한 서비스들의 제한된 라이브러리를 구성하여 에이전트 훈련을 위한 작업 공간을 확장하는 프레임워크입니다.
- Typed Environment Modeling: 기계가 읽을 수 있는 Model Context Protocol (MCP) 명세를 활용하여 서비스 상태와 상호작용 인터페이스를 Python 스키마로 표준화하고, Pydantic 모델을 통해 환경 상태를 정의하는 과정입니다.
- Completion-Focused Rubric Reward: RL 훈련 중 각 롤아웃 그룹 내에서 통과율이 낮은 기준에 더 높은 가중치를 부여하여, 에이전트가 작업의 완전한 완료를 지향하도록 유도하는 보상 메커니즘입니다.
- Service-level Dependency Graph: 독립적으로 실행 가능한 서비스들을 연결하는 그래프로, 간선은 작업 완료에 필요한 정보 또는 상태 의존성을 나타냅니다.
- Agentic Synthesis and Verification: 코딩 에이전트가 상태 모델, 도구 구현, 테스트 스위트를 생성하고, 반복적으로 구현을 수정하며, 적대적 테스트 케이스를 사용하여 독립적인 유효성 검사를 수행하는 프로세스입니다.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
본 논문은 대규모 언어 모델(LLM) 에이전트 훈련을 위한 환경 스케일링의 핵심적인 문제를 해결하고자 합니다. 기존의 접근 방식은 주로 단일 환경 내에서 작업을 생성하여 에이전트에게 상호작용 데이터를 제공했지만, 실제 워크플로우는 에이전트가 여러 서비스에 걸쳐 정보와 행동을 연결하는 복합적인 능력을 요구합니다. 이러한 기존 연구의 한계점은 에이전트가 다양한 시스템 간의 의존성을 처리하고 정보를 효과적으로 전달하는 데 어려움을 겪게 합니다.
저자들은 CompoWorld 프레임워크를 통해 세 가지 주요 도전을 제시합니다 [Figure 2]: 첫째, 자동으로 생성된 서비스들이 자체적으로 안정적으로 실행되고 다른 서비스들과 호환 가능해야 합니다. 둘째, 합성된 작업들이 의미 있는 교차 서비스 의존성을 포함하며 해결 및 검증 가능해야 합니다. 셋째, 이러한 작업들이 SFT와 에이전트 RL 모두에 유용한 감독 신호를 제공해야 하며, 특히 에이전트가 워크플로우의 일부만 완료했을 때도 유용해야 합니다.
3. Method & Key Results (제안 방법론 및 핵심 결과)
본 논문은 일반 에이전트 훈련을 위한 Compositional Environment Scaling (CompoWorld) 프레임워크를 제안합니다 [Figure 2]. 이 프레임워크는 세 가지 주요 구성 요소로 이루어져 있습니다. 첫째, 개별 서비스는 Typed Environment Modeling을 통해 MCP 명세와 Pydantic 모델을 사용하여 구축되며, Agentic Synthesis and Verification 과정을 거쳐 신뢰성을 확보합니다. 신뢰성 있는 구현이 어려운 도구의 경우, LLM 기반의 Selective World-Model Simulation을 통해 동적 응답을 예측합니다. 둘째, 교차 환경 작업은 서비스 풀에서 서비스를 샘플링하고 Service-level Dependency Graph를 통해 연결함으로써 생성됩니다. Agentic Task Generation 단계에서 에이전트는 초기 상태, 목표, 제약 조건을 설정하고 검증을 수행하여 작업의 유효성을 확인합니다. 셋째, 에이전트 훈련은 검증된 성공적인 궤적을 이용한 SFT와 작업 루브릭을 이용한 RL로 진행됩니다. 특히, Completion-Focused Rubric Reward를 도입하여 롤아웃 그룹 내에서 통과율이 낮은 기준에 더 높은 가중치를 부여함으로써 완전한 작업 완료를 장려하며, 이는 Group Relative Policy Optimization (GRPO)과 함께 사용됩니다.
CompoWorld는 448개의 재사용 가능한 서비스(총 10,130개의 도구)를 구축하고, Qwen3.6-35B-A3B 모델을 3K SFT 궤적과 1K RL 작업으로 훈련시켰습니다. 실험 결과, CompoWorld는 8개의 도전적인 에이전트 벤치마크에서 백본 모델 대비 평균 9.17점의 성능 향상을 달성했습니다 [Table 1]. 특히, AutomationBench에서 작업 성공률을 10.33%에서 32.33%로 (22.00점 증가) 대폭 향상시켜, Claude Opus 4.6 (25.50%) 및 GPT-5.4 (27.67%)와 같은 최첨단 모델들을 능가하고, 비교된 모든 에이전트 특화 35B-A3B 모델들 중 선두를 차지했습니다 [Table 1]. 또한, AutomationBench 1.0.6의 도메인별 분석에서 CompoWorld는 모든 도메인에서 개선을 보였으며, 평균 점수가 41.94%에서 72.68%로 상승했습니다 [Table 2]. 특히 HR (+49.40점), Marketing (+32.51), Sales (+31.45)에서 큰 폭의 개선을 보였습니다 [Table 2]. 구성된 환경에서의 훈련은 단일 환경 훈련 대비 모든 비교에서 일관되게 상당한 성능 향상을 가져왔으며, 이는 더 풍부한 상호작용 구조와 작업 복잡성, 데이터 다양성을 제공하여 더 강력한 일반화를 촉진함을 시사합니다 [Figure 5 (a)].
4. Conclusion & Impact (결론 및 시사점)
본 논문은 재사용 가능한 실행 가능한 서비스를 구축하고 이를 작업별 인과적 의존성을 통해 연결하는 Compositional Environment Scaling (CompoWorld) 프레임워크를 성공적으로 제시합니다. 이 프레임워크의 자동화된 구축 및 검증 파이프라인은 SFT 및 RL을 위한 교차 환경 작업을 생성하며, Completion-Focused Rubric Reward는 미완성 요구 사항에 중점을 두어 완전한 작업 완료를 유도합니다.
8개 에이전트 벤치마크에 걸친 실험은 이러한 구성된 워크플로우에서의 훈련이 가져오는 이점을 명확하게 보여주었습니다. 특히 AutomationBench에서 CompoWorld는 Claude Opus 4.6과 같은 최첨단 모델을 능가하고 비교된 모든 에이전트 특화 35B-A3B 모델 중 선두를 차지했습니다. 이러한 결과는 서비스 구성이 환경 스케일링의 유망한 차원임을 강조하며, 훈련 데이터의 복잡성과 다양성을 증가시키고 LLM 에이전트가 새로운 워크플로우에서 익숙한 기능을 조정하는 데 도움을 줄 수 있음을 시사합니다. 이는 학계와 산업계 모두에서 더 일반적이고 유능한 AI 에이전트 개발에 중요한 시사점을 제공합니다.

Figure 1 — 모델 성능 비교
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
- [논문리뷰] X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
- [논문리뷰] Monet: Reasoning in Latent Visual Space Beyond Images and Language
- [논문리뷰] Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- [논문리뷰] ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning
Review 의 다른글
- 이전글 [논문리뷰] CoWindow Attention: Full Causal Coverage Is a Collective Property
- 현재글 : [논문리뷰] CompoWorld: Compositional Environment Scaling for General Agents
- 다음글 [논문리뷰] DepthBench: Measuring How Residual Connections Enable More Computational Depth
댓글