[논문리뷰] SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information본 논문은 LLM이 모바일 비서로 활용될 때, 여러 애플리케이션에 분산된 개인 정보를 종합하여 사용자의 복잡한 지시를 해결해야 하는 도전 과제를 해결하고자 합니다.#Review#Large Language Models#Mobile Assistants#Personal Information#Benchmark#Information Retrieval#Tool Use#Cognitive Capabilities2026년 8월 11일댓글 수 로딩 중
[논문리뷰] AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games본 논문은 협소하고 정적인 기존 AI 벤치마크의 한계를 극복하고, 인간과 유사한 일반 지능(AGI)을 평가하기 위한 확장 가능하며 개방형의 새로운 접근 방식을 제안합니다. 특히, AI 시스템이 인간이 고안한 모든 게임 을 얼마나 잘 플레이하고 학습하는지를 통해 AGI 역량을 측정하고자 합니다.#Review#Artificial General Intelligence (AGI)#Evaluation Benchmark#General Game Playing#Large Language Models (LLMs)#Human-in-the-loop#Cognitive Capabilities#Vision-Language Models (VLMs)#Game Generation2026년 2월 26일댓글 수 로딩 중