[논문리뷰] AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report본 논문은 현대 게임 개발의 노동 집약적이고 시간 소모적인 파이프라인을 대체할 수 있는 실시간 상호작용형 가상 세계 생성 모델을 제안한다. 기존의 영상 생성 모델들은 긴 시간 동안 일관성(Consistency)을 유지하거나 사용자의 카메라 조작 및 행동 입력에 실시간으로 반응하는 데 한계를 가지고 있다.#Review#World Modeling#Interactive Generation#Video Diffusion Transformer#Long-Horizon#Autoregressive#Spatiotemporal Memory2026년 7월 21일댓글 수 로딩 중
[논문리뷰] AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents본 논문은 LLM Agent에서 발생하는 오류의 가시성 부족과 근본 원인 분석의 어려움을 해결하기 위해 AgentDebugX를 제안합니다. 복잡한 Agent 환경에서는 최종 오류가 발생한 시점과 실제 원인이 된 행동의 시점이 일치하지 않는 경우가 많아, 단순한 Trace replay만으로는 디버깅에 한계가 있습니다.#Review#LLM Agents#Debugging#Observability#Root-Cause Attribution#Error Recovery#Agentic Workflow#Closed-Loop System2026년 7월 21일댓글 수 로딩 중
[논문리뷰] ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU본 논문은 실시간 인터랙티브 환경을 구현하기 위한 비디오 월드 모델의 제어 가능성(Controllability), 일관성(Consistency), 효율성(Efficiency) 간의 결합 문제를 해결하는 데 집중한다.#Review#World Model#Action-Conditioned#Real-Time Inference#Video Generation#Streaming#Long-Horizon#System Co-Design2026년 7월 21일댓글 수 로딩 중
[논문리뷰] WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting본 논문은 LLM의 예측 능력을 실제 시나리오에서 검증하기 위해, 미래의 결과를 예측하고 이를 실제 결과와 대조하는 동적 벤치마크인 WorldCupArena를 제안한다.#Review#Sports Forecasting#Large Language Models#Deep-Research Agents#Football#Benchmark#Temporal Prediction2026년 7월 20일댓글 수 로딩 중
[논문리뷰] Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift본 논문은 LLM이 학습 시 의존한 패턴으로 인해 생성 단계에서 발생하는 환각(hallucination) 문제를 해결하고자 합니다.#Review#Token-Level Off-Policy Learning#Faithful Generation#Distribution Shift#LoRA#Conditional Steering#LLM Alignment#Hallucination2026년 7월 20일댓글 수 로딩 중
[논문리뷰] TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs본 논문은 기존 MLLM이 영상의 내용을 설명(Description)하는 데는 능숙하나, 근거가 되는 구체적인 시간대(Temporal Evidence)를 식별하는 데는 한계가 있다는 점을 해결하고자 합니다.#Review#Multimodal LLM#Temporal Grounding#Video Understanding#Temporal Wasserstein#Long-context#Supervised Grounding2026년 7월 20일댓글 수 로딩 중
[논문리뷰] The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture본 논문은 Transformer의 복잡한 이산적 연산 구조를 연속적인 기하학적 프레임워크로 재해석하여, 모델의 안정성과 성능 한계를 예측 가능한 보편적 언어로 설명하고자 합니다.#Review#Transformer Architecture#Differential Geometry#Semantic Fiber Bundle#Stochastic Calculus#Non-Equilibrium Thermodynamics#Lipschitz Scaling2026년 7월 20일댓글 수 로딩 중
[논문리뷰] ShotPlan: Cinematic Video Generation with Learnable Planning Token본 논문은 기존 비디오 생성 모델이 단일 샷(Single-shot) 생성에는 탁월하지만, 영화나 드라마와 같은 복잡한 내러티브를 위한 멀티 샷(Multi-shot) 생성에는 한계가 있다는 문제를 해결하고자 합니다.#Review#Video Generation#Diffusion Transformer#Learnable Planning Token#Cinematic Video#FRoPE#Multi-shot Generation#Temporal Control2026년 7월 20일댓글 수 로딩 중
[논문리뷰] Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?본 논문은 Self-hosted AI Agent가 자신의 내부 상태(Memory, Instruction, Configuration)를 오염시키는 Self-state Attacks에 취약하다는 점을 해결하고자 합니다 .#Review#Self-hosted AI Agents#Self-state Attacks#OS Security#Anomaly Detection#Workload-conditioned Analysis2026년 7월 20일댓글 수 로딩 중
[논문리뷰] SWE-Pruner Pro: The Coder LLM Already Knows What to Prune본 논문은 coding agents가 다중 턴 환경에서 축적하는 방대한 tool outputs로 인해 발생하는 연산 비용 및 컨텍스트 품질 저하 문제를 해결하고자 한다.#Review#Coding Agents#Context Pruning#Large Language Models#Model Efficiency#Token Compression#Hidden-State Analysis#Software Engineering2026년 7월 20일댓글 수 로딩 중
[논문리뷰] RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model본 논문은 범용적인 멀티모달 모델이 실제 물리적 환경에서 로봇 제어로 전이될 때 발생하는 Embodied Foundation Model의 한계점을 해결하고자 합니다.#Review#Embodied AI#Foundation Model#Vision-Language-Action (VLA)#Multi-embodiment#3D Grounding#Contact-point Prediction#Scaling Laws2026년 7월 20일댓글 수 로딩 중
[논문리뷰] ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams본 논문은 기존의 Long-term memory 모델들이 비디오의 길이를 제한하거나(bounded), 단일 규모의 프레임 단위 저장에 의존하여 장기적인 Entity tracking 및 맥락 유지가 어렵다는 점을 해결하고자 합니다 .#Review#Multimodal Memory#Entity-Oriented#Long-term Memory#Open-Ended Video Streams#Lifelong Perception#Episodic Memory#Semantic Memory2026년 7월 20일댓글 수 로딩 중
[논문리뷰] ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video본 논문은 단일 monocular egocentric video로부터 인간의 동작과 주변 환경을 일관된 4D로 동시에 복원하는 통합 프레임워크의 부재 문제를 해결합니다.#Review#Egocentric Vision#4D Reconstruction#Masked Generative Modeling#Human Motion Estimation#Unified Transformer#Multimodal Learning2026년 7월 20일댓글 수 로딩 중
[논문리뷰] OpenLongTail: Generative Scaling of Long-Tail Driving Data본 논문은 자율주행 VLA 모델의 long-tail 시나리오 대응 능력을 향상시키기 위한 데이터 확보 문제를 해결하고자 한다.#Review#Autonomous Driving#Video Diffusion Model#VLA Model#Generative Data Engine#Extrapolative View Synthesis2026년 7월 20일댓글 수 로딩 중
[논문리뷰] Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning본 논문은 Embodied Intelligence 연구에 필요한 대규모 고품질 조작 데이터의 파편화 문제를 해결하고자 합니다. 기존 데이터셋은 고가의 장비나 특정 환경에 의존하여 확장성이 낮으며, 수집된 원시 비디오를 로봇 학습에 바로 사용할 수 있도록 가공하는 체계적인 인프라가 부족합니다.#Review#Embodied Learning#Egocentric Dataset#Manipulation#Toolchain#MANO#Vision-Language-Action#Robot Learning2026년 7월 20일댓글 수 로딩 중
[논문리뷰] LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks본 논문은 오픈 엔디드(open-ended) 작업에서 Reinforcement Learning (RL) 기반의 보상 최적화가 갖는 정보 병목(information bottleneck) 문제를 해결하고자 합니다.#Review#Experiential Learning#LLM-as-a-Coach#On-policy Context Distillation#Feedback Bandwidth#Non-Verifiable Tasks#Reward Hacking2026년 7월 20일댓글 수 로딩 중
[논문리뷰] JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models본 논문은 VLA 모델의 후속 학습 과정에서 발생하는 자원 비효율성 및 인프라 관리의 복잡성 문제를 해결하기 위해 JoyNexus를 제안합니다.#Review#VLA#Multi-Tenant#Post-Training#Reinforcement Learning#Service-Oriented Architecture#Group Batching#Resource Utilization2026년 7월 20일댓글 수 로딩 중
[논문리뷰] HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis본 연구는 HOI 생성 과정에서 시각적 일관성과 3D 움직임의 물리적 타당성을 동시에 확보하는 것이 어렵다는 문제를 해결합니다.#Review#Hand-Object Interaction#Motion Synthesis#Multi-view Synthesis#Generative Modeling#Human-Object Interaction#3D Perception2026년 7월 20일댓글 수 로딩 중
[논문리뷰] HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement본 논문은 기존의 HOCVP 접근 방식이 갖는 인터-서브젝트(Inter-subject) 및 인트라-서브젝트(Intra-subject) personalization의 한계를 해결하기 위해 제안되었습니다.#Review#HOCVP#Video Personalization#Multimodal Large Language Models#Diffusion Transformers#Global Multimodal Guidance#Modality-Reference Embedding2026년 7월 20일댓글 수 로딩 중
[논문리뷰] Group Entropy-Controlled Policy Optimization본 논문은 현대 LLM의 Reinforcement Learning 과정에서 발생하는 다중 작업(Multi-task) 간의 최적화 불균형 문제를 해결합니다.#Review#Reinforcement Learning#Large Language Models#Group Entropy#Policy Optimization#GRPO#Alignment#Exploration-Exploitation2026년 7월 20일댓글 수 로딩 중