[논문리뷰] TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation기존의 multi-reference image generation 벤치마크는 task-oriented 방식으로 구성되어 combinatorial setting에 부적합하며, 이는 incomplete coverage, no failure diagnosis, 그리고 uncontrolled complexity라는 세 가지 주요 한계를 야기한다.#Review#Multi-Reference Image Generation#Capability-Oriented Benchmark#Diagnostic Tree Analysis#Atomic Operators#Evaluation Protocol#Failure Localization#Compositional Formulas#Attribute Disentanglement2026년 8월 17일댓글 수 로딩 중
[논문리뷰] UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks본 논문은 현대의 Proactive Agents를 평가하기 위한 기존 벤치마크들의 구조적 한계를 해결하기 위해 UniClawBench를 제안한다. 기존 연구들은 샌드박스화된 고립 환경과 단일 턴(Single-turn) 평가 방식에 의존하여, 실제 환경의 복잡성과 반복적인 사용자 피드백 루프를 반영하지 못한다 .#Review#Proactive Agents#Capability-Oriented Benchmark#Closed-loop Evaluation#Real-World Tasks#Multimodal Understanding#Tool Usage#Docker-based Environment2026년 7월 9일댓글 수 로딩 중