[논문리뷰] ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation본 논문은 기존의 text-to-image 모델들이 복잡한 사실 관계나 외부 지식이 필요한 open-world 작업에서 심각한 정보 결핍과 사실 왜곡 문제를 겪고 있다는 점을 지적합니다.#Review#Agentic Image Generation#Unified Multimodal Model#Reinforcement Learning#Tool-Use#Post-training#Reason-Act-Draw2026년 8월 5일댓글 수 로딩 중
[논문리뷰] ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step본 논문은 현재의 대규모 언어 모델(LLM) 기반 에이전트들이 도구 사용 능력을 평가할 때, 실제적인 행동 추론보다는 API 이름과 같은 의미론적 우선순위(semantic priors)에 과도하게 의존하고 있다는 문제를 제기합니다 .#Review#Autonomous Agents#Tool-Use#Benchmark#Behavioral Reasoning#Semantic Obfuscation#Mapping Drift#Stochastic Failure#Persistent Memory2026년 8월 3일댓글 수 로딩 중
[논문리뷰] S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence본 논문은 기존 VLM들이 정적인 단일 프레임 관찰에 의존하여 연속적이고 진화하는 3D 환경에서의 공간 추론에 한계를 보인다는 점을 해결하고자 합니다 . 기존 모델들은 파편화된 2D 시각 정보에 의존하기 때문에 공간적 일관성(spatial consistency) 유지와 고도화된 3D 기하학적 이해가 어렵습니다.#Review#Spatial Intelligence#Vision-Language Models (VLM)#Agentic Paradigm#Spatio-Temporal Reasoning#Tool-Use#Spatial Evidence Accumulation2026년 6월 18일댓글 수 로딩 중
[논문리뷰] PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions본 연구는 기존 모바일 에이전트 평가가 지나치게 GUI 제어 중심의 task 수행에만 집중되어 있어 실제 사용자 워크플로우를 반영하지 못한다는 한계를 해결하고자 합니다.#Review#Phone Agents#Mixed-Action Space#GUI Control#CLI#Tool-Use#Verifiable Execution#Safety Evaluation2026년 6월 15일댓글 수 로딩 중
[논문리뷰] PREPING: Building Agent Memory without TasksLLM 에이전트는 새로운 환경에 배치될 때 작업별 경험이 부족한 상태에서 발생하는 Cold-Start 문제에 직면합니다. 기존의 메모리 구축 방식은 사전에 수집된 사람의 시연(offline)이나 배포 후 사용자와의 상호작용(online)에 의존하는데, 이는 배포 초기 단계의 실패를 야기하거나 구축 비용을 증가시킵니다 .#Review#Agent Memory#Procedural Memory#Synthetic Practice#Cold-Start#Agentic Context Engineering#Tool-Use#Pre-task Construction2026년 5월 14일댓글 수 로딩 중
[논문리뷰] OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding기존 옴니모달 대규모 언어 모델(OmniLLMs) 이 겪는 미세한 크로스모달 이해(fine-grained cross-modal understanding) 및 멀티모달 정렬(multimodal alignment) 의 한계를 해결하는 것을 목표로 합니다.#Review#Omnimodal Understanding#Audio-Guided Perception#Active Learning Agents#Cross-Modal Alignment#Tool-Use#Video Understanding#Multimodal LLMs2025년 12월 29일댓글 수 로딩 중
[논문리뷰] UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist본 논문은 전문화된 비디오 AI 모델과 실제 비디오 워크플로우 간의 격차를 해소하여 차세대 비디오 일반 인공지능을 구현하는 것을 목표로 합니다.#Review#Video Agents#Multi-modal AI#Plan-Act Architecture#Tool-Use#Long-horizon Reasoning#Open-source#Video Generation#Video Understanding2025년 11월 13일댓글 수 로딩 중
[논문리뷰] SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent이 논문은 기존 3D 장면 합성 방법론들이 고정된 카테고리, 부족한 객체 디테일, 물리적 불일치, 복잡한 사용자 지시와의 낮은 정합성 등의 한계를 가지는 문제를 해결하고자 합니다.#Review#3D Scene Synthesis#Agentic Framework#LLMs#Self-Reflection#Tool-Use#Physical Plausibility#Iterative Refinement#Embodied AI2025년 9월 26일댓글 수 로딩 중