본문으로 건너뛰기

[논문리뷰] The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

링크: 논문 PDF로 바로 열기

The paper "The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks" by Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia (City University of Hong Kong, Independent Researcher, Microsoft) introduces the concept of "taste" in LLM agents and a new benchmark called Taste-Bench.

I need to parse the content to extract information for the summary.

Metadata:

  • Authors: Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia
  • Keywords: I'll look for 5-8 academic terms that are central to the paper. LLM agents, long-horizon tasks, decision fork, taste, Taste-Bench, distillation, software engineering, machine learning research.

Part 1: Summary Body

1. Key Terms & Definitions:

  • Taste: An agent's ability to make good long-horizon decisions, where the influence of a decision is not limited to the current step and its cost often appears much later.
  • Long-Horizon Tasks: Tasks requiring agents to make many sequential decisions whose full impact is not immediately visible.
  • Decision Fork: A point in an agent's trajectory where multiple directions are available, and one leads to a better outcome, identifiable by hindsight.
  • Taste-Bench: A benchmark of 502 taste questions automatically constructed from agent trajectories in software engineering and machine learning research tasks, designed to measure an agent's ability to choose the better direction at a decision fork.
  • Distillation: A training method where the reasoning of a "teacher" model (which has seen the outcome) is transferred to a "student" model (which only sees the question) to improve its judgment.

2. Motivation & Problem Statement: LLM agents increasingly engage in long-horizon tasks, where intermediate decisions significantly impact the final outcome. However, existing benchmarks primarily measure end-to-end success without assessing the quality of decisions made along the way. The core problem is the difficulty in directly measuring these long-horizon decisions because their consequences are not immediately apparent, and human annotation for quality judgment is expensive and hard to scale. This gap means current evaluations do not adequately capture an agent's taste, defined as its ability to make good long-term choices. The paper addresses this by proposing an automatic method to measure taste using hindsight evidence from existing agent trajectories.

3. Method & Key Results: 저자들은 에이전트의 taste를 측정하기 위해 Taste-Bench라는 새로운 벤치마크를 제안한다. 이 벤치마크는 에이전트가 소프트웨어 엔지니어링 및 머신러닝 연구 태스크에서 생성한 기존 궤적(trajectories)에서 decision fork를 자동으로 마이닝하여 구축된다. 각 질문은 에이전트가 여러 방향 중 더 나은 결과를 가져올 방향을 선택해야 하는 결정 분기점(decision fork)을 제시하며, 이때 에이전트는 분기점 이후의 결과를 볼 수 없다. Taste-Bench는 크게 두 가지 방식으로 fork를 구축한다: 동일한 태스크에 대한 여러 번의 시도에서 갈라지는 parallel trajectories와 단일 궤적 내에서 에이전트가 스스로를 수정하는 detour trajectories이다. Figure 2는 Taste-Bench의 구성 및 필터링 파이프라인을 시각화한다.

실험 결과, 최신 frontier models는 Taste-Bench에서 제한적인 taste 능력을 보여주었다. 가장 뛰어난 모델인 GPT-5.6 Sol의 Average accuracy는 59.7%에 불과했으며, GPT-5.5가 59.5%로 그 뒤를 이었다 [cite: 1, Figure 3]. 특히, 결정적 증거가 궤적의 후반부에 나타나는 fork일수록 모든 모델에게 더 어려웠으며, time horizon이 증가함에 따라 평균 정확도가 62.3%에서 21.0%로 크게 하락했다 [cite: 1, Figure 4]. 또한, 더 큰 reasoning budget을 할당하더라도 정확도 향상에는 유의미한 영향을 미치지 않았다. 이는 모델들이 어려운 fork를 인식하고 더 오래 추론하지만, 결정적 증거가 나중에 나타나기 때문에 성능이 개선되지 않음을 시사한다.

이 연구는 distillation을 통해 taste를 훈련할 수 있음을 보여주었다. 정답을 본 teacher model의 추론을 student model에 distill한 결과, student model은 보지 못한 태스크에서 더 나은 판단을 내렸고. held-out SWE-bench Pro 태스크에서 end-to-end success rate를 14.6%에서 33.7%로 향상시키는 결과를 보였다 [cite: 1, Figure 7]. 이는 correct advice가 제공하는 상한선인 39.0%에 근접하는 수치이다.

4. Conclusion & Impact: 본 연구는 에이전트가 결과가 보이지 않는 상태에서 더 나은 방향을 선택하는 능력인 taste를 정의하고 측정하는 방법을 제시한다. 기존 궤적의 후행 증거(hindsight evidence)를 활용하여 Taste-Bench라는 502개의 질문으로 구성된 벤치마크를 성공적으로 구축했다. 현재 frontier models가 Taste-Bench에서 아직 taste 능력이 부족함을 확인했지만, teacher model의 추론을 distill하는 방식으로 taste를 훈련할 수 있음을 입증했다. 이 연구는 에이전트의 long-horizon judgment 능력을 정량적으로 평가하고 개선할 수 있는 실용적인 기반을 제공한다. 이는 미래의 LLM agents가 소프트웨어 개발, 과학 연구 등 복잡한 태스크에서 보다 효율적이고 성공적인 결정을 내릴 수 있도록 돕는 중요한 기여를 할 것으로 기대된다.

Part 2: Important Figure Information (JSON)

I need to find the src attributes for <img> tags for Figure 1, Figure 2, Figure 3, Figure 4, Figure 5, Figure 6, Figure 7.

  • Figure 1: ![Refer to caption](2609.25804v1/decision_fork.png)
    • Full URL: https://arxiv.org/html/2609.25804v1/decision_fork.png
  • Figure 2: ![Refer to caption](2609.25804v1/tastebench_construction.png)
    • Full URL: https://arxiv.org/html/2609.25804v1/tastebench_construction.png
  • Figure 3: ![Refer to caption](2609.25804v1/model_results_system.png)
    • Full URL: https://arxiv.org/html/2609.25804v1/model_results_system.png
  • Figure 4: ![Refer to caption](2609.25804v1/accuracy_by_horizon_and_budget.png)
    • Full URL: https://arxiv.org/html/2609.25804v1/accuracy_by_horizon_and_budget.png
  • Figure 5: ![Refer to caption](2609.25804v1/tastebench_vs_swebench.png)
    • Full URL: https://arxiv.org/html/2609.25804v1/tastebench_vs_swebench.png
  • Figure 6: ![Refer to caption](2609.25804v1/distillation_advisor.png)
    • Full URL: https://arxiv.org/html/2609.25804v1/distillation_advisor.png
  • Figure 7: ![Refer to caption](2609.25804v1/distillation_results.png)
    • Full URL: https://arxiv.org/html/2609.25804v1/distillation_results.png

I need to select up to 3 most important figures.

  1. Figure 2 (tastebench_construction.png): This shows the core methodology of how the benchmark is constructed. Essential for understanding the paper's contribution.
  2. Figure 3 (model_results_system.png): This presents the main experimental results on model performance, showing how current models fare. Crucial quantitative results.
  3. Figure 7 (distillation_results.png): This illustrates the effectiveness of the proposed distillation method, demonstrating that taste is trainable and improves end-to-end performance. Shows the impact and solution.

I will use Figure 2, Figure 3, and Figure 7. Captions:

  • Figure 2: Taste-Bench 구성 및 필터링
  • Figure 3: 모델별 Taste-Bench 성능
  • Figure 7: Taste Distillation 결과

Review the whole output for all constraints: technical terminology, natural Korean, metadata, four sections, direct answer rule, bolding, citations, Figure/Table references, JSON format for figures. The metadata should be at the very top of the markdown. Keywords: I will choose more specific ones. LLM Agents, Long-Horizon Tasks, Decision Forks, Taste-Bench, Automated Evaluation, Knowledge Distillation, Software Engineering Tasks, Machine Learning Research.

One more check for direct answer rule:

  1. Key Terms & Definitions: Yes, starts with definition.
  2. Motivation & Problem Statement: Yes, starts with problem.
  3. Method & Key Results: Yes, starts with proposed method.
  4. Conclusion & Impact: Yes, starts with conclusion.

Figure/Table reference rule: I have referenced Figure 2, Figure 3, and Figure 7 in the "Method & Key Results" section. Figure 1, Figure 4, Figure 5, Figure 6 are not selected for Part 2, so I should not refer to them as [Figure N]. I can say "Figure 1은..." if I want to refer to them, but not as [Figure N]. I'll make sure the references are only for selected figures. I used [Figure 1] for example, I need to fix that. Ah, the constraint is "Part 2에서 선별한 Figure/Table만 [Figure N]으로 참조할 수 있습니다. 선별하지 않은 것은 참조하지 마세요." I need to remove [Figure 1], [Figure 4], [Figure 5], [Figure 6] references.

I referred to Figure 1 in "Our key observation is that the later part of a trajectory provides hindsight evidence for its earlier decisions. ... Figure(#S1.F1 "Figure 1 ‣ Introduction") shows one such fork from a machine-learning trajectory." - I'll remove [1] and adjust. I referred to Figure 4 in "The left panel of Figure(#S4.F4 "Figure 4 ‣ Effect of the time horizon ‣ Experimental Results") plots accuracy over the four levels, and accuracy falls on average across models as the time horizon increases." - I'll modify this to "Figure 4의 좌측 패널은..." I referred to Figure 5 in "Figure(#S4.F5 "Figure 5 ‣ Comparison with end-to-end benchmarks ‣ Experimental Results") shows the comparison, and the two benchmarks are only partly correlated." - I'll modify this to "Figure 5는..." I referred to Figure 6 in "In this section, we distill this judgment into the weights of a model and test whether it transfers to unseen tasks and improves end-to-end task success (Figure(#S5.F6 "Figure 6 ‣ Generalizing Taste through Distillation"))." - I'll modify this to "Figure 6는..."

Okay, I'll update the summary to reflect these changes regarding figure references.저자: Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia et al.

Part 1: 요약 본문

1. Key Terms & Definitions (핵심 용어 및 정의)

  • Taste: 에이전트가 장기적인 관점에서 좋은 의사결정을 내릴 수 있는 능력으로, 결정의 영향이 즉각적인 단계에 국한되지 않고 비용이 훨씬 나중에 나타나는 경우를 포함한다.
  • Long-Horizon Tasks: 에이전트가 여러 순차적인 의사결정을 내려야 하며, 각 결정의 완전한 영향이 즉시 나타나지 않는 복잡한 태스크를 의미한다.
  • Decision Fork: 에이전트의 궤적(trajectory) 내에서 여러 가능한 방향이 존재하며, 그 중 하나가 더 나은 결과를 가져오는 지점을 말하며, 이는 사후(hindsight) 분석을 통해 식별 가능하다.
  • Taste-Bench: 소프트웨어 엔지니어링 및 머신러닝 연구 태스크에서 에이전트가 생성한 궤적에서 자동으로 구성된 502개의 taste 질문으로 이루어진 벤치마크로, 에이전트가 decision fork에서 더 나은 방향을 선택하는 능력을 측정하기 위해 설계되었다.
  • Distillation: "teacher" 모델(결과를 본 모델)의 추론을 "student" 모델(질문만 본 모델)로 전이시켜 판단 능력을 향상시키는 훈련 방법론이다.

2. Motivation & Problem Statement (연구 배경 및 문제 정의)

LLM 에이전트(agents)는 점차 long-horizon tasks에 투입되고 있으며, 중간에 내리는 결정들이 전체 실행 결과에 큰 영향을 미친다. 그러나 기존 벤치마크들은 에이전트의 end-to-end success만을 측정할 뿐, 과정에서 이루어지는 의사결정의 품질은 평가하지 못하는 한계가 있다. 이러한 long-horizon decisions의 품질을 직접 측정하는 것은 그 결과가 즉시 나타나지 않고, 전문가의 수동 주석(annotation)이 비용이 많이 들고 확장하기 어렵기 때문에 매우 도전적이다. 본 연구는 에이전트의 "taste"(장기적인 관점에서 좋은 결정을 내리는 능력)를 측정하지 못하는 이 문제를 해결하기 위해, 기존 궤적(trajectories)의 사후(hindsight) 증거를 활용하여 에이전트의 taste를 자동으로 측정하는 새로운 방법을 제안한다.

3. Method & Key Results (제안 방법론 및 핵심 결과)

저자들은 에이전트의 taste를 측정하기 위해 Taste-Bench라는 새로운 벤치마크를 제안한다. 이 벤치마크는 에이전트가 소프트웨어 엔지니어링 및 머신러닝 연구 태스크에서 생성한 기존 궤적(trajectories)에서 decision fork를 자동으로 마이닝하여 구축된다. 각 질문은 에이전트가 여러 방향 중 더 나은 결과를 가져올 방향을 선택해야 하는 결정 분기점(decision fork)을 제시하며, 이때 에이전트는 분기점 이후의 결과를 볼 수 없다. Taste-Bench는 크게 두 가지 방식으로 fork를 구축한다: 동일한 태스크에 대한 여러 번의 시도에서 갈라지는 parallel trajectories와 단일 궤적 내에서 에이전트가 스스로를 수정하는 detour trajectories이다. [Figure 2]는 Taste-Bench의 구성 및 필터링 파이프라인을 시각화한다.

Figure 2: Taste-Bench 구성 및 필터링

Figure 2 — Taste-Bench 구성 및 필터링

실험 결과, 최신 frontier models는 Taste-Bench에서 제한적인 taste 능력을 보여주었다. 가장 뛰어난 모델인 GPT-5.6 Sol의 Average accuracy는 59.7%에 불과했으며, GPT-5.5가 59.5%로 그 뒤를 이었다 [cite: 1, Figure 3]. 특히, 결정적 증거가 궤적의 후반부에 나타나는 fork일수록 모든 모델에게 더 어려웠으며, time horizon이 증가함에 따라 평균 정확도가 62.3%(in-prefix)에서 21.0%(more-work)로 크게 하락했다. 또한, 더 큰 reasoning budget을 할당하더라도 정확도 향상에는 유의미한 영향을 미치지 않았다. 이는 모델들이 어려운 fork를 인식하고 더 오래 추론하지만, 결정적 증거가 나중에 나타나기 때문에 성능이 개선되지 않음을 시사한다.

이 연구는 distillation을 통해 taste를 훈련할 수 있음을 보여주었다. 정답을 본 teacher model의 추론을 student model에 distill한 결과, student model은 보지 못한 태스크에서 더 나은 판단을 내렸고, held-out SWE-bench Pro 태스크에서 end-to-end success rate를 14.6%에서 33.7%로 향상시키는 결과를 보였다 [cite: 1, Figure 7]. 이는 correct advice가 제공하는 상한선인 39.0%에 근접하는 수치이다.

4. Conclusion & Impact (결론 및 시사점)

본 연구는 에이전트가 결과가 보이지 않는 상태에서 더 나은 방향을 선택하는 능력인 taste를 정의하고 측정하는 방법을 제시한다. 기존 궤적의 사후 증거(hindsight evidence)를 활용하여 Taste-Bench라는 502개의 질문으로 구성된 벤치마크를 성공적으로 구축했다. 현재 frontier models가 Taste-Bench에서 아직 taste 능력이 부족함을 확인했지만, teacher model의 추론을 distill하는 방식으로 taste를 훈련할 수 있음을 입증했다. 이 연구는 에이전트의 long-horizon judgment 능력을 정량적으로 평가하고 개선할 수 있는 실용적인 기반을 제공한다. 이는 미래의 LLM agents가 소프트웨어 개발, 과학 연구 등 복잡한 태스크에서 보다 효율적이고 성공적인 결정을 내릴 수 있도록 돕는 중요한 기여를 할 것으로 기대된다.

Figure 3: 모델별 Taste-Bench 성능

Figure 3 — 모델별 Taste-Bench 성능

Figure 7: Taste Distillation 결과

Figure 7 — Taste Distillation 결과

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글