[논문리뷰] ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step본 논문은 현재의 대규모 언어 모델(LLM) 기반 에이전트들이 도구 사용 능력을 평가할 때, 실제적인 행동 추론보다는 API 이름과 같은 의미론적 우선순위(semantic priors)에 과도하게 의존하고 있다는 문제를 제기합니다 .#Review#Autonomous Agents#Tool-Use#Benchmark#Behavioral Reasoning#Semantic Obfuscation#Mapping Drift#Stochastic Failure#Persistent Memory2026년 8월 3일댓글 수 로딩 중