본문으로 건너뛰기

#Evaluation Protocol

8개의 포스트

[논문리뷰] Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

댓글 수 로딩 중

[논문리뷰] SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

댓글 수 로딩 중

[논문리뷰] TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

댓글 수 로딩 중

[논문리뷰] From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

댓글 수 로딩 중

[논문리뷰] From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents

댓글 수 로딩 중