[논문리뷰] MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations본 논문은 기존의 LLM Agent 벤치마크들이 단기적인 작업 완수나 즉각적인 피드백 기반의 태스크에 편향되어 있다는 문제를 해결하고자 한다. 실제 환경에서는 상점 운영과 같이 장기적인 전략 수립이 필요하며, 과거의 의사결정이 미래의 선택을 제약하고 피드백이 불균일하게 발생하는 복합적인 상황이 빈번하다.#Review#LLM Agents#Long-Term Coherence#E-Commerce#Order-Level Simulation#Partially Observable Markov Decision Process#Benchmark2026년 8월 4일댓글 수 로딩 중