[논문리뷰] WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes본 논문은 최신 3D generative world model들이 생성한 장면의 정적인 특성과 낮은 편집 가능성 문제를 해결하기 위해 WorldAct를 제안합니다.#Review#3D Gaussian Splatting#Scene Decomposition#Agent-Driven Interaction#3D World Modeling#Embodied Simulation#Interactive Content Creation2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Unlocking Dense Metric Depth Estimation in VLMs본 논문은 기존 VLMs가 2D 과업에는 뛰어나지만 3D 이해 능력은 여전히 제한적이라는 핵심 문제에서 출발합니다 . 기존 연구들은 외부의 3D 전문 모델로부터 지식을 증류하거나, 텍스트 기반으로만 학습하여 정밀한 기하학적 정보가 부족하고 오류가 누적되는 한계를 보입니다.#Review#Vision-Language Models#Dense Metric Depth Estimation#3D Geometry#Unified Supervision#Spatial Reasoning2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Steered LLM Activations are Non-Surjective본 연구는 Activation Steering이 유도하는 모델의 내부 행동 변화가 실제 텍스트 프롬프트를 통해서도 동일하게 구현 가능한지라는 근본적인 의문을 해결하고자 합니다.#Review#Activation Steering#Surjectivity#LLM Interpretability#Prompt-Reachability#White-box Intervention#AI Safety2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models본 논문은 CLIP과 같은 대규모 vision-language 모델을 하위 태스크(downstream task)에 맞게 fine-tuning할 때 발생하는 OOD(Out-of-Distribution) 성능 저하 문제를 해결하고자 한다.#Review#CLIP#Sparse Autoencoders#Robust Fine-tuning#Interpretability#Representational Drift#Computer Vision2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution본 논문은 기존 LLM 기반 경쟁 프로그래밍 에이전트들이 가진 상태 비저장(stateless) 구조의 한계를 해결하고자 합니다. 대다수의 최신 프레임워크는 문제 해결 시마다 처음부터 시작하며, 과거의 디버깅 경험이나 실패 기록을 재사용하지 못하는 고립된 구조를 띱니다 .#Review#Large Language Models#Competitive Programming#Agentic Evolution#Reinforcement Learning#Knowledge Network#Code Generation#Multi-Agent System2026년 5월 17일댓글 수 로딩 중
[논문리뷰] ReactiveGWM: Steering NPC in Reactive Game World Models본 논문은 기존의 Game World Models가 NPC를 단순한 배경 요소로 취급하여 상호작용이 결여된 정적인 비디오 렌더러에 머물러 있는 문제를 해결하고자 합니다.#Review#Game World Models#NPC#Controllable Video Generation#Diffusion Models#Strategy Transfer#Cross-Attention#Interaction Logic2026년 5월 17일댓글 수 로딩 중
[논문리뷰] PhysBrain 1.0 Technical Report본 논문은 기존 VLA 시스템이 의존하는 플랫폼 종속적인 로봇 궤적(Trajectory) 데이터 수집의 한계를 극복하고, 물리적 환경에 대한 근본적인 이해(Physical Commonsense)를 확보하는 것을 목표로 합니다.#Review#Vision-Language-Action Models#Embodied Intelligence#Physical Commonsense#Egocentric Video#Data Engine#VLA Adaptation2026년 5월 17일댓글 수 로딩 중
[논문리뷰] PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control본 연구는 기존 GUI 에이전트들이 주로 의존하는 'region-tolerant' 패러다임이 정밀한 기하학적 구성 작업에서 실패하는 근본적인 문제를 해결하고자 한다.#Review#GUI Agents#Geometric Reasoning#Precision-Sensitive#Dependency-Structured Planning#Pixel-Grounded Supervised Tuning#Reinforcement Learning#Semantic-Execution Gap2026년 5월 17일댓글 수 로딩 중
[논문리뷰] OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation본 연구는 로봇 학습을 위한 고품질 데이터 수집의 높은 비용과 확장성 문제를 해결하기 위해, 다양한 humanoid embodiment 간의 cross-embodiment video generation을 수행하고자 합니다.#Review#Cross-embodiment Video Generation#Diffusion Transformer#Embodiment-specific Adaptation#Streaming Inference#Paired-free Learning2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR본 논문은 RLVR 환경에서 고질적인 문제인 탐색의 병목 현상을 해결하고자 합니다. 기존 방식은 탐색 효율을 높이기 위해 샘플링 횟수(Rollout)를 무작정 늘리는 방식을 취하지만, 이는 계산 비용이 극심하고 long-tail에 위치한 희귀한 정답 추론 경로를 발견하는 데 한계가 있습니다 .#Review#RLVR#Reinforcement Learning#Exploration#LLM Reasoning#Strategy Nudging#Inter-Intra Group Advantage#Distillation2026년 5월 17일댓글 수 로딩 중
[논문리뷰] MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware본 논문은 대규모 VLA 모델 학습에 필수적인 장기 시점(long horizon)의 egocentric 데이터를 수집하기 위한 개방형 인프라를 구축하는 데 목적이 있습니다. 기존 데이터셋은 에피소드 길이가 짧고 고가의 하드웨어 장비에 의존해야 하는 등 확장성에 한계를 보입니다.#Review#Egocentric Data#Vision Language Action (VLA)#Long-horizon#SLAM#STERA#Smartphone-based Capture2026년 5월 17일댓글 수 로딩 중
[논문리뷰] MMSkills: Towards Multimodal Skills for General Visual Agents본 논문은 시각적 에이전트가 복잡한 환경에서 성공적인 결정을 내리기 위해 필요한 Multimodal Procedural Knowledge의 부재 문제를 해결하고자 합니다.#Review#Multimodal Agents#Procedural Knowledge#Visual Grounding#Branch Loading#GUI Agents#Skill Representation2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Look Before You Leap: Autonomous Exploration for LLM Agents본 논문은 현대의 LLM 기반 에이전트가 새로운 환경에서 적응하지 못하고 조기 착취(Premature Exploitation) 문제에 빠지는 현상을 해결하고자 합니다.#Review#LLM Agents#Autonomous Exploration#RLVR#GRPO#Exploration Checkpoint Coverage#Explore-then-Act2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation본 논문은 대규모 언어 모델(LLM)의 post-training에서 OPD가 RL보다 높은 효율성을 보이는 근본적인 파라미터 업데이트 메커니즘을 규명하고자 합니다.#Review#On-Policy Distillation#Large Language Models#Parameter Dynamics#Training Efficiency#EffOPD#Subspace Evolution2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards본 논문은 기존 RLVR 패러다임이 가진 sparse binary reward와 weak credit assignment 문제를 해결하여 모델의 추론 능력을 극대화하는 것을 목적으로 합니다.#Review#Reinforcement Learning#Large Language Models#Verifiable Rewards#Policy Optimization#Error Correction#Reasoning Capability2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Learning POMDP World Models from Observations with Language-Model Priors본 연구는 잠재 상태에 대한 정보(Ground-truth state)가 주어지지 않는 완전한 부분 관측 환경(Strict POMDP setting)에서 에이전트가 어떻게 효과적으로 세계 모델(World Model)을 학습할 수 있는지 탐구합니다.#Review#POMDP#World Model#Large Language Models#Program Induction#Sample Efficiency#Partial Observability#Belief-based Filtering2026년 5월 17일댓글 수 로딩 중
[논문리뷰] InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation본 논문은 Autoregressive 모델 기반의 이미지 생성에서 텍스트와 얼굴의 품질이 저하되는 문제를 해결하고자 합니다.#Review#Discrete Tokenization#Autoregressive Image Generation#Perceptual Loss#Text Fidelity#Face Fidelity#Content-Aware Supervision2026년 5월 17일댓글 수 로딩 중
[논문리뷰] Hölder Policy Optimisation본 논문은 LLM의 long-horizon 추론 과제에서 GRPO와 같은 기존 그룹 기반 RL 알고리즘이 사용하는 고정된 aggregation mechanism의 한계를 지적한다.#Review#Reinforcement Learning#Large Language Models#Hölder Mean#Gradient Concentration#Policy Optimisation#Group Relative Policy Optimisation (GRPO)2026년 5월 17일댓글 수 로딩 중
[논문리뷰] HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts본 논문은 기존의 MoE 압축 방식들이 전문가 간의 결합 가능성을 평가할 때 사용하는 pairwise 점수의 구조적 한계를 해결하고자 합니다.#Review#Sparse Mixture-of-Experts#Simplicial Complex#Hodge Decomposition#Harmonic Kernel#Model Compression#Topological Deep Learning2026년 5월 17일댓글 수 로딩 중
[논문리뷰] GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding본 연구는 MLA가 특정 하드웨어(예: NVIDIA H100)의 연산-대역폭 비율에 지나치게 종속되어 있다는 문제를 해결합니다.#Review#Large Language Model#KV-cache#Multi-head Latent Attention#GQLA#Hardware-Adaptive#Roofline Model#Tensor Parallelism2026년 5월 17일댓글 수 로딩 중