[논문리뷰] RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
링크: 논문 PDF로 바로 열기
The paper is about "RRSI: Regularized Recursive Self-Improvement of Agent Harnesses".
I have the content now. I will proceed with generating Part 1 and Part 2.
**Part 1: Summary**
* **Authors**: Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
* **Keywords**: Need to identify 5-8 relevant keywords from the abstract and introduction. Potential candidates: `LLM Agents`, `Harness Evolution`, `Recursive Self-Improvement (RSI)`, `Regularization`, `Overfitting`, `Generalization`, `Agentic Workspace`, `Policy Tokens`.
* **Key Terms & Definitions**:
* **Agent Harness**: Prompts, control flow, tooling, memory, context management surrounding a frozen backbone LLM.
* **Recursive Self-Improvement (RSI)**: Iteratively proposing and selecting component-wise edits of an agent harness using task feedback to improve its behavior.
* **Overfitting**: Adaptive process where harness evolution memorizes training tasks, leading to in-distribution gains that do not transfer to out-of-distribution benchmarks.
* **Policy Tokens**: The number of tokens consumed by the agent's trajectory, used as a proxy for inference cost.
* **Motivation & Problem Statement**:
* Problem: LLM agents' capabilities are significantly influenced by their "harness," which is typically engineered manually. Automating this via iterative harness evolution (RSI) suffers from adaptive overfitting, where improvements on evolve sets do not generalize to unseen tasks.
* Limitations of existing methods: Prior methods often show large in-distribution gains that diminish or vanish on out-of-distribution (OOD) benchmarks, sometimes even increasing test-time computation without creating reusable mechanisms. This overfitting can arise from encoding benchmark-specific patterns, promoting candidates favored by evaluation noise, or accumulating unnecessary complexity.
* **Method & Key Results**:
* Methodology: RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) proposes regularizing both the *proposal* and *selection* phases of the harness evolution loop.
* **Proposal-side Regularization**:
* **L0-Style Annealed Update Sparsity**: Limits the number of independently attributable edits in a single candidate proposal, with a decreasing budget over evolution rounds (from `b_max` to `b_min`) to encourage simpler, more attributable changes.
* **Evidence-Aware Credit Assignment**: Tracks the history of evaluated candidates, their modified components, and outcomes (accepted/rejected) to guide future proposals away from previously falsified hypotheses.
* **Structured Exploration**: During "stalls" (when progress is within noise), a portion of the budget is reserved for unexplored components to encourage diversity.
* **Selection-side Regularization**:
* **Leakage Screening**: A critic rejects edits explicitly encoding task-specific content before full evaluation to prevent benchmark-specific overfitting.
* **Stability-Aware Acceptance**: Requires a candidate's score to exceed a `noise-adjusted floor` (`S_hat(H') >= S* - delta`) to prevent accepting improvements due to stochastic variation.
* **Ridge/L2-Style Complexity-Aware Acceptance**: Constraints `Delta C <= beta_0 + beta_1 * Delta S` where `Delta C` is relative cost change and `Delta S` is score change, ensuring additional inference cost is justified by performance improvement. This discourages unchecked growth in the harness's resource footprint.
* **Lasso/L1-Style Structural Pruning**: Tracks components that fail to produce positive measured gain over a pruning window and instructs the proposer to remove them, leading to a sparser retained harness.
* Key Results:
* RRSI achieves substantial gains: up to **14.1 points** on the evolve split and up to **4.7 points** on five out-of-distribution benchmarks across eight tasks in coding, agentic workspace, and engineering design domains. These gains generalize beyond the evolution environment [Figure 1 (b-d)].
* RRSI consistently outperforms baselines on all held-out datasets. For agentic workspace tasks, RRSI achieved an OOD average of **43.6**, significantly higher than Meta-Harness (**40.6**), AHE (**37.2**), TTHE (**31.7**), and HarnessX (**36.3**) [cite: 1, Table 1].
* RRSI produces a more efficient harness, consuming **30% fewer policy tokens** than unregularized evolution. In ablation studies, removing either proposal-side or acceptance-side regularizers reduces transfer performance and increases token cost [cite: 1, Table 2]. For example, the full RRSI had **2.42 million tokens/trial** on OOD average for agentic workspace, whereas unregularized evolution consumed **3.80 million tokens/trial** for lower performance [Table 2].
* The benefits of RRSI are **policy-agnostic**, showing similar improvements when evolved with different backbone models like **Claude Opus 4.8** and **Gemini 3.5 Flash** [cite: 1, Table 3]. The evolved harness also transfers effectively to an unseen, weaker backbone model (**Gemini 3.1 Flash Lite**) [cite: 1, Table 4].
* **Conclusion & Impact**:
* Conclusion: Iterative harness evolution, a form of recursive self-improvement, necessitates regularization to prevent overfitting to finite evolve sets, evaluation noise, and unnecessary complexity. RRSI successfully addresses this by regularizing both the proposal and selection mechanisms, leading to more generalizable and efficient agent harnesses.
* Impact: This research demonstrates that controlling the *how* of feedback conversion into persistent changes is crucial for making agent systems more capable through RSI. It suggests a pathway for developing more robust and cost-effective LLM agents that can adapt to new tasks and environments without significant manual re-engineering or model re-training, impacting both academic research in agent design and industrial application of LLM agents.
**Part 2: Important Figure Information**
I will look for "Figure 1", "Figure 2", "Figure 3", "Table 1", "Table 2", "Table 3", "Table 4" and "Figure 4" based on the text.
The paper mentions Figure 1(a-d), Figure 2, Figure 3, Figure 4(a-b). Tables 1, 2, 3, 4.
* **Figure 1**: "Evolution overfits on the split it is scored on whereas RRSI generalizes the improvements. (a) Gains on the evolve split against gains out of distribution for the agentic workspace benchmark. Prior methods retain little of their evolve-set gain and several end below H0, the initial harness. (b-d) Out-of-distribution held-out score for H0, average of the four baseline methods and our RRSI on SWE-bench Verified, the mean of JobBench, GDPval and APEX-Agents, and Frontier-Eng." This seems to be a key result figure.
* Image URL: `https://arxiv.org/html/2609.24972v1/x1.png`
* Caption: RRSI의 일반화된 성능 향상 비교
* **Figure 2**: "Overview of RRSI. RRSI regularizes the search trajectory, not restricting the potential harness edit space: proposal-side constraints control how search capacity is used, while selection-side constraints control which measured improvements are allowed to become a permanent state." This is the architecture figure.
* Image URL: `https://arxiv.org/html/2609.24972v1/x2.png`
* Caption: RRSI의 전체 프레임워크 개요
* **Figure 3**: "Main results in all three domains." This is a key result figure.
* Image URL: `https://arxiv.org/html/2609.24972v1/x3.png`
* Caption: 세 가지 도메인에서의 주요 결과
I will pick Figure 1 (overall performance, generalization), Figure 2 (architecture), and Figure 3 (main results across domains).
The output format requires a maximum of 3 figures.
Let's verify the image URLs.
For Figure 1: `src="2609.24972v1/x1.png"` -> `https://arxiv.org/html/2609.24972v1/x1.png`
For Figure 2: `src="2609.24972v1/x2.png"` -> `https://arxiv.org/html/2609.24972v1/x2.png`
For Figure 3: `src="2609.24972v1/x3.png"` -> `https://arxiv.org/html/2609.24972v1/x3.png`
All URLs are correct.**저자:** Peng Xia, Rujun Han, Zifeng Wang, et al.
## 1. Key Terms & Definitions (핵심 용어 및 정의)
- **Agent Harness**: LLM 백본 모델을 둘러싸는 프롬프트, 제어 흐름, 툴링, 메모리 및 컨텍스트 관리 등 시스템의 총체를 의미합니다. 이는 에이전트의 실제 수행 능력에 결정적인 영향을 미칩니다.
- **Recursive Self-Improvement (RSI)**: 에이전트 시스템 수준에서 하네스 컴포넌트의 편집을 반복적으로 제안하고 선택하여 에이전트의 동작을 개선하는 프로세스입니다. 현재 시스템의 피드백을 활용하여 다음 행동을 형성하는 하네스를 개선합니다.
- **Overfitting**: 하네스 진화 과정에서 학습 태스크를 암기하여, 학습(in-distribution) 벤치마크에서는 높은 성능 향상을 보이지만, 보지 못한(out-of-distribution) 태스크에서는 성능 저하가 발생하거나 개선이 사라지는 현상을 지칭합니다.
- **Policy Tokens**: 에이전트의 trajectory(경로) 실행에 사용된 토큰의 수를 나타내며, 추론(inference) 비용을 측정하는 프록시(proxy) 지표로 사용됩니다.
## 2. Motivation & Problem Statement (연구 배경 및 문제 정의)
LLM 에이전트의 역량은 고정된 백본 모델 주변의 프롬프트, 제어 흐름, 툴 인터페이스, 메모리 및 컨텍스트 관리 등을 포함하는 **Harness**에 의해 크게 좌우됩니다. 이러한 하네스 엔지니어링은 주로 수작업에 의존하여 진행되어 왔으나, 최근에는 LLM을 활용하여 하네스 컴포넌트를 태스크 피드백으로부터 자동 최적화하는 **반복적 하네스 진화(Iterative Harness Evolution)** 방식, 즉 에이전트 시스템 수준의 **Recursive Self-Improvement (RSI)** 형태가 등장했습니다. 그러나 이러한 재귀적 진화는 유한한 **Evolve Set**에서 반복적으로 편집을 제안하고 선택함으로써 **Adaptive Overfitting** 위험을 내포합니다 [Figure 1 (a)]. 이는 진화 과정에서 학습된 하네스가 특정 벤치마크 패턴을 암호화하거나, 평가 노이즈에 의해 선호되는 후보를 촉진하거나, 불필요한 복잡성을 축적하여 **In-Distribution** 성능 향상이 **Out-of-Distribution (OOD)** 벤치마크로의 일반화로 이어지지 않는 문제를 야기합니다. 기존 연구들은 진화 성능과 일반화 성능 간의 상당한 격차를 보고하며, 이러한 겉보기 개선이 재사용 가능한 메커니즘보다는 태스크별 적합화 또는 테스트 시간 컴퓨팅 증가에서 비롯될 수 있음을 지적했습니다. 따라서 본 연구는 하네스 기반 RSI의 핵심 과제인 **Overfitting** 문제를 해결하고 일반화 능력을 향상시키는 새로운 접근 방식의 필요성을 제기합니다.
## 3. Method & Key Results (제안 방법론 및 핵심 결과)
본 논문은 에이전트 하네스의 **Recursive Self-Improvement (RSI)** 과정에서 발생하는 Overfitting 문제를 해결하기 위해 **Regularized Recursive Self-Improvement of Agent Harnesses (RRSI)** 프레임워크를 제안합니다. RRSI는 하네스 편집 공간을 완전히 개방한 상태로 유지하면서도, 진화 루프의 **제안(Proposal)** 및 **선택(Selection)** 양측에 정규화(regularization) 원칙을 통합하여 검색 궤적을 제약합니다 [Figure 2].
**제안 측(Proposal-side) 정규화**는 다음과 같은 방법으로 검색 용량이 사용되는 방식과 위치를 제어합니다:
- **L0-Style Annealed Update Sparsity**: 단일 라운드에서 제안될 수 있는 독립적으로 귀속 가능한 편집의 수를 **b_max**에서 **b_min**으로 점진적으로 감소하는 예산으로 제한합니다. 이는 초기에는 넓은 탐색을 허용하고 후기에는 희소하며 귀속 가능한 변화를 유도합니다.
- **Evidence-Aware Credit Assignment**: 이전에 평가된 모든 후보의 컴포넌트, 가설, 결과 등을 포함한 편집 이력을 기록하여, 제안자가 이력에 조건부로 반응하게 함으로써 이미 기각된 가설에 대한 반복적인 탐색을 방지합니다.
- **Structured Exploration**: 검색이 정체되었을 때 (이전 `w` 라운드 동안 진행이 경험적 노이즈 밴드 내에 머무를 때), 탐색 예산의 일부를 아직 탐색되지 않은 컴포넌트에 할당하여 다양성을 촉진합니다.
**선택 측(Selection-side) 정규화**는 측정된 개선 사항 중 어떤 것이 영구적인 상태로 허용되어야 하는지를 제어합니다:
- **Leakage Screening**: 평가 전에 비평가(critic)가 후보 편집이 벤치마크별 태스크 이름, 엔티티, 답 등 특정 정보를 명시적으로 인코딩하는 경우를 거부하여 **Benchmark-Specific Fitting**을 방지합니다.
- **Stability-Aware Acceptance**: 경험적으로 추정된 <strong>Noise Band (δ)</strong>를 사용하여, 후보의 측정된 점수가 이전 최고 점수(`S*`)에서 `δ`를 뺀 값보다 크거나 같아야만 수락되도록 합니다 (`S_hat(H') >= S* - δ`). 이는 확률적 변동으로 인한 개선을 영구적인 상태로 전환하는 것을 방지합니다.
- **Ridge/L2-Style Complexity-Aware Acceptance**: 성능 향상에 상응하는 추론 비용 증가(`ΔC <= β0 + β1ΔS`)를 요구합니다. 이는 하네스의 전체 리소스 사용량의 무제한적인 증가를 억제하여, 측정 가능한 성능 개선 없이는 추가 비용을 허용하지 않습니다.
- **Lasso/L1-Style Structural Pruning**: 최근 `n_prune` 기간 동안 긍정적인 측정 이득을 생성하지 못한 컴포넌트들을 비생산적인 것으로 보고하여, 제안자가 해당 컴포넌트를 삭제하도록 유도함으로써 하네스의 구조적 희소성을 높입니다.
RRSI는 3가지 도메인(코딩, Agentic Workspace, 엔지니어링 설계)의 8개 벤치마크에서 평가되었습니다. 주요 결과는 다음과 같습니다: RRSI는 **Evolve Split**에서 최대 **14.1점**을, 5개의 **Out-of-Distribution (OOD)** 벤치마크에서 최대 **4.7점**의 성능 향상을 달성했습니다. 특히, Agentic Workspace 태스크에서 RRSI는 모든 **Held-Out Dataset**에서 기준선(Baseline) 방법들을 일관되게 능가했습니다 [Table 1]. 예를 들어, OOD 평균 점수에서 RRSI는 **43.6점**을 기록하여 기존의 가장 강력한 Baseline인 Meta-Harness의 **40.6점**을 넘어섰습니다 [cite: 1, Table 1]. RRSI는 또한 **Unregularized Evolution**에 비해 **30% 더 적은 Policy Tokens**를 소비하며 효율적인 하네스를 생성했습니다. Ablation Study 결과, 제안 측(Proposal-side) 또는 선택 측(Acceptance-side) 정규화 중 어느 하나라도 제거하면 일반화 성능이 저하되고 토큰 비용이 증가하는 것으로 나타났습니다 [cite: 1, Table 2]. 예를 들어, 정규화되지 않은 진화(Unregularized Evolution)는 OOD 평균에서 **3.80 million tokens/trial**을 소비한 반면, RRSI는 **2.42 million tokens/trial**로 더 적은 비용으로 더 높은 성능을 달성했습니다 [Table 2]. 또한, RRSI의 이점은 **Claude Opus 4.8** 및 **Gemini 3.5 Flash**와 같은 다른 백본 모델을 사용하더라도 일관되게 유지되었으며 [Table 3], 심지어 진화 과정에 사용되지 않은 약한 모델(**Gemini 3.1 Flash Lite**)에서도 **3.4점**의 전이 이득을 보여 **Policy-Agnostic**하며 재사용 가능한 메커니즘을 학습했음을 시사합니다 [Table 4].
## 4. Conclusion & Impact (결론 및 시사점)
본 연구는 에이전트 시스템 수준의 재귀적 자체 개선(RSI)으로서 반복적인 하네스 진화가 자체적으로 정규화를 필요로 함을 입증했습니다. 유한한 **Evolve Set**이 적응적으로 재사용되기 때문에, 겉보기의 자체 개선은 벤치마크 특정 적합화, 평가 노이즈, 또는 불필요한 복잡성을 반영할 수 있으며, 이는 전이 가능한 진전이 아닐 수 있습니다. **RRSI**는 이러한 문제를 해결하기 위해 하네스 편집 공간을 개방하면서도 **제안(Proposal)** 및 **선택(Selection)** 과정 모두에 정규화 원칙을 적용합니다 [Figure 2].
코딩, Agentic Workspace 및 엔지니어링 설계 태스크 전반에 걸쳐, RRSI를 통해 진화된 하네스는 **Held-Out** 및 **Cross-Benchmark** 성능을 향상시키면서도, 정규화되지 않은 진화보다 더 적은 **Inference Cost**를 사용했습니다 [cite: 1, Figure 3, Figure 4]. 이러한 결과는 재귀적 자체 개선을 통해 에이전트 시스템을 더욱 강력하게 만들려면 무엇을 변경할 수 있는지 뿐만 아니라, 반복적인 피드백이 영구적인 변화로 어떻게 전환되는지를 제어하는 것이 필수적임을 시사합니다. RRSI는 학계에서 에이전트 설계 연구의 새로운 방향을 제시하고, 산업계에서 LLM 에이전트의 견고하고 비용 효율적인 적용을 위한 기반을 마련하는 데 기여할 수 있습니다.
> ⚠️ **알림:** 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
- [논문리뷰] ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
- [논문리뷰] MemEvolve: Meta-Evolution of Agent Memory Systems
- [논문리뷰] Frequency-Adaptive Sharpness Regularization for Improving 3D Gaussian Splatting Generalization
- [논문리뷰] RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
Review 의 다른글
- 이전글 [논문리뷰] Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
- 현재글 : [논문리뷰] RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- 다음글 [논문리뷰] Realtime-Venus: A full-duplex interaction system with asynchronous delegation
댓글