[논문리뷰] Agent Explorative Policy Optimization for Multimodal Agentic Reasoning본 논문은 vision-language models(VLMs)의 agentic reasoning 과정에서 발생하는 '도구 사용의 비효율성' 문제를 해결하고자 합니다.#Review#Multimodal Agentic Reasoning#Reinforcement Learning#GRPO#AXPO#Tool-call Resampling#Thinking-Acting Gap#Vision-Language Models2026년 5월 27일댓글 수 로딩 중
[논문리뷰] AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems본 논문은 LLM 기반의 다중 에이전트 시스템에서 발생하는 조율 불투명성과 고정된 파이프라인의 경직성 문제를 해결하고자 합니다.#Review#Multi-Agent Systems#Online Policy Learning#Coordination Substrate#Large Language Models#Task Signatures#Relative Trajectory Evaluation2026년 5월 27일댓글 수 로딩 중
[논문리뷰] Advancing Creative Physical Intelligence in Large Multimodal Models본 연구는 대규모 다중모달 모델(LMM)이 인식 및 추론 능력은 크게 발전했음에도 불구하고, 비일상적인 상황에서 사물을 창의적으로 재사용하는 물리적 지능이 여전히 부족하다는 문제의식에서 출발합니다.#Review#Multimodal AI#Creative Tool Repurposing#Physical Affordance#Visual Grounding#Direct Preference Optimization (DPO)#Interactive Benchmark2026년 5월 27일댓글 수 로딩 중
[논문리뷰] AI Research Agents Narrow Scientific Exploration본 연구는 AI 연구 에이전트가 과학적 발견의 범위를 실질적으로 확장하는지, 아니면 기존 연구의 주변부에 머무르는지를 규명하는 것을 목적으로 합니다.#Review#AI Research Agents#Scientific Discovery#Ideation#Citation Analysis#Research Breadth#Bibliographic Coupling2026년 5월 27일댓글 수 로딩 중
[논문리뷰] The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence본 논문은 large language model (LLM)이 long-horizon agentic workflow로 전환됨에 따라 발생하는 efficiency 및 cost bottleneck 문제와 intrinsically complex, high-stakes task 해결의 어려움을 다룹니다.#Review#Mixture-of-Experts (MoE)#Mini Activations#Agentic AI#Self-Evolution#Reinforcement Learning (RL)#Multi-Token Prediction (MTP)2026년 5월 26일댓글 수 로딩 중
[논문리뷰] SpatialBench: Is Your Spatial Foundation Model an All-Round Player?본 논문은 현재 Spatial Foundation Models (SFMs)이 standard dataset에서 인상적인 성능을 보여주지만, 다양한 downstream task, 임의의 viewpoint, 변화하는 scene domain, 다양한 input density, 그리고 특정 hardware constraint에 걸쳐 robust하게 generalizing할 수 있는 all-round player인지에 대한 근본적인…#Review#Spatial Foundation Models#3D Reconstruction#Benchmark#Domain Generalization#Input Density#Embodied AI2026년 5월 26일댓글 수 로딩 중
[논문리뷰] Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration본 논문은 Long-horizon Video-to-Video Generation의 핵심 과제인 Long Cinematic Video Remaking 문제를 해결하고자 합니다.#Review#Long-Video Remaking#Multi-Agent System#Dual-Bridge Consistency#Character Identity#Narrative Fidelity#Video-to-Video Generation2026년 5월 26일댓글 수 로딩 중
[논문리뷰] Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling기존 병렬 Test-Time Scaling (TTS) 방법론은 Information Isolation Bottleneck이라는 중요한 한계점을 가지고 있습니다.#Review#Test-Time Scaling#Collaborative Parallel Thinking#Large Language Models#Information Sharing#Redundant Exploration#Accuracy-Latency Pareto Frontier#Mathematical Reasoning2026년 5월 26일댓글 수 로딩 중
[논문리뷰] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research모바일 GUI Agent 연구는 빠른 발전을 보였지만, 현재 평가 및 훈련 환경은 근본적인 Trade-off 문제에 직면해 있다.#Review#Mobile GUI Agent#Simulation Environment#Reinforcement Learning#Verifiable Outcome Signals#Interaction Fidelity#MobileGym-Bench#Sim-to-Real Transfer2026년 5월 26일댓글 수 로딩 중
[논문리뷰] LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV본 논문은 기존 Audio-Visual Generation 벤치마크가 Minute-Scale Content의 평가 요구사항을 충족하지 못하는 문제를 해결하고자 한다.#Review#Audio-Visual Generation#Long Video Generation#Evaluation#Benchmark#T2AV#I2AV#V2AV#MLLM-assisted assessment2026년 5월 26일댓글 수 로딩 중
[논문리뷰] LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box DecodingThe End of the content of the urls browsed.#Review2026년 5월 26일댓글 수 로딩 중
[논문리뷰] Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction본 논문은 Degraded Input Condition 하에서 Multi-view 3D Reconstruction의 Robustness를 향상시키기 위해 Geometry-Aware Representation Denoising (GARD) 프레임워크를 제안한다.#Review#Multi-view 3D Reconstruction#Image Restoration#Representation Denoising#Diffusion Models#Geometry-Aware Features#Feed-Forward Models#Camera Pose Estimation2026년 5월 26일댓글 수 로딩 중
[논문리뷰] EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation본 연구는 generative video foundation models의 빠른 발전으로 professional-grade cinematic synthesis에 대한 수요가 증가함에 따라, Reinforcement Learning (RL) 및 agentic workflows로의 전환에 필요한 신뢰할 수 있는 평가의 bottleneck 문제를 해결하고자 한다.#Review#Video Generation#Benchmarking#Cinematic Quality#VLM#Chain-of-Thought#Human-Machine Alignment#Evaluation Framework#Reinforcement Learning2026년 5월 26일댓글 수 로딩 중
[논문리뷰] D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing본 논문은 D-LLM의 안전성 monitoring 연구가 미흡하며, D-LLM의 오용 가능성이 증대함에 따라 효과적인 방어 메커니즘이 필요하다고 주장합니다.#Review#Diffusion LLMs#Safety Monitoring#Hesitation-Aware Routing#Probe-based Monitors#Multi-step Trajectory#Sample Difficulty#Efficiency-effectiveness Tradeoff#Adversarial Inputs2026년 5월 26일댓글 수 로딩 중
[논문리뷰] Your Embedding Model is SMARTer Than You Think본 논문은 single-vector multimodal retriever가 rich하고 sequential한 token sequence를 단일 global representation으로 압축하면서 발생하는 근본적인 information bottleneck 문제를 해결하고자 합니다.#Review#Multimodal Retrieval#Single-Vector Embeddings#Multi-Vector Embeddings#Late Interaction#Information Bottleneck#Hidden States#Contrastive Learning#Plug-and-Play2026년 5월 25일댓글 수 로딩 중
[논문리뷰] WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation최근 Interactive World Models의 발전에도 불구하고, 기존의 평가 방식은 단편적이며 체계적인 평가를 위한 통합된 표준이 부재하다.#Review#Interactive World Models#Video Generation#Benchmark#Multi-turn Interaction#Evaluation Metrics2026년 5월 25일댓글 수 로딩 중
[논문리뷰] TriSplat: Simulation-Ready Feed-Forward 3D Scene ReconstructionI was unable to access the content of the provided URL: https://arxiv.org/html/2605.26115.#Review2026년 5월 25일댓글 수 로딩 중
[논문리뷰] Toward Native Multimodal Modeling: A Roadmap본 논문은 기존 Large Language Models (LLMs)이 텍스트 전용 인터페이스에 근본적으로 제한되어 실제 세계의 풍부한 센서리 신호(sensory signals)를 통한 그라운딩(grounding)이 부족하다는 문제의식에서 출발합니다.#Review#Native Multimodal Modeling#Cross-modal Fusion#Transformer Architectures#Multimodal LLMs#M2M Symmetric Modeling#Mid-Fusion#Early-Fusion2026년 5월 25일댓글 수 로딩 중
[논문리뷰] ThriftAttention: Selective Mixed Precision for Long-Context FP4 AttentionI am unable to access the content of the provided URL: https://arxiv.org/html/2605.23081. The browsing tool encountered an error while trying to fetch the page.#Review2026년 5월 25일댓글 수 로딩 중
[논문리뷰] SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills본 논문은 LLM Agents가 실제 작업을 해결하면서 축적하는 풍부한 Episodic Experience가 재사용 가능한 Procedural Skills로 증류될 수 있는지 여부가 불분명하다는 핵심 문제를 제기한다.#Review#LLM Agents#Procedural Skills#Skill Formation#Episodic Experience#Benchmarking#Skill Evolution#Abstraction Bottleneck#Deployment Transfer2026년 5월 25일댓글 수 로딩 중