#Benchmark Agent

1개의 포스트

[논문리뷰] Benchmark Everything Everywhere All at Once

본 논문은 기존의 수동적인 벤치마크 구축 방식이 가진 한계인 노동 집약성, 재사용 불가능성, 그리고 모델 성능 향상에 따른 빠른 벤치마크 포화(Saturation) 문제를 해결하고자 합니다.

#Review #Benchmark Agent #Autonomous Evaluation #Benchmark Construction #MLLM-as-a-Judge #Agentic Workflow #Performance Saturation

2026년 6월 4일