[논문리뷰] Omni-Streaming Thinking본 논문은 스트리밍 오디오-비주얼 모델이 불완전하고 순차적으로 관찰되는 데이터에서 적시에 응답해야 하는 근본적인 문제에 직면하고 있음을 지적합니다.#Review#Streaming Omni-modal Models#Premature Cross-Modal Commitment#Claim Verification#Causal Retraction#Audio-Visual Reasoning#OST-DiagBench#Adaptive Response Timing2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training본 논문은 Online RL 기반의 MLLM post-training 과정에서 모든 prompt가 동일한 학습 가치를 지니지 않는다는 비효율성을 해결하고자 합니다.#Review#Multimodal Large Language Models#Reinforcement Learning#Prompt Scaffolding#Exploration Potential Score#Online RL#Data Flywheel#Curriculum Design2026년 9월 14일댓글 수 로딩 중
[논문리뷰] ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs의료용 VLM은 임상 질문에 답변할 때 이미지의 시각적 정보와 함께 제공된 텍스트 리포트(radiology report)에 의존하는 경향이 있는데, 리포트가 이미 답변을 포함하고 있는 경우 모델이 이미지를 실제로 보고 있는지 판별하기 어렵습니다.#Review#Medical VLMs#Image Sensitivity#Radiology Report#Counterfactual Audit#MIMIC-CXR#Attention Knockout2026년 9월 14일댓글 수 로딩 중
[논문리뷰] LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows현재 Video Diffusion Models는 인상적인 Visual Fidelity를 달성했지만, 그 제어는 여전히 어렵습니다.#Review#Multimodal Video Generation#Diffusion Transformer#Agentic Visual Creation#Inference Acceleration#Lightweight VAE#Long Video Generation#MSAVP Benchmark2026년 9월 14일댓글 수 로딩 중
[논문리뷰] LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents기존 autoregressive (AR) 패러다임의 멀티모달 대규모 언어 모델(Large Language Models, LLMs)은 내재적인 순차적(sequential) 특성으로 인해 병렬화(parallelism)가 제한되고 상당한 추론 Latency를 발생시켜, 실시간성이 중요한 GUI Agent 애플리케이션에는 비효율적이다.#Review#Diffusion Language Models#GUI Agents#Vision-Language Models#Block-wise Diffusion#Multimodal Pre-training#Supervised Fine-tuning#Inference Efficiency2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Kaininja: Extending Native 3D Generators to the Part Level기존 Native 3D Generators는 단일 이미지로부터 고품질 3D Geometry를 생성하지만, 출력이 Part-level asset이 아닌 하나의 Fused Mesh 형태여서 편집, 리깅, 시뮬레이션 등 후처리(Downstream) 작업에 적합하지 않다.#Review#Part-level 3D Generation#Native 3D Generators#TRELLIS.2#O-Voxel Representation#Dual-Volume Packing#Image-to-3D Synthesis#Rectified Flow#Agent-authored 3D Assets2026년 9월 14일댓글 수 로딩 중
[논문리뷰] How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus본 논문은 autoregressive (AR) 언어 모델의 디코딩 절차가 본질적으로 순차적이어서 높은 inference 비용과 제한된 하드웨어 활용을 초래하는 문제를 다룹니다.#Review#Lossless Speculative Decoding#Orthrus#Numerical Precision#BF16#FP32#Trajectory Divergence#Inference Acceleration#Perplexity2026년 9월 14일댓글 수 로딩 중
[논문리뷰] HazardAuditor: From Executable Threats to Safer Computer-Use Agents본 논문은 Computer-use Agents의 런타임 행동에서 발생하는 안전성 위험을 효과적으로 감지하기 위한 방법론을 제안한다.#Review#LLM agent safety#agent guard model#computer-use agents#execution-grounded framework#Guard Policy Optimization (GuardPO)2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Grouped Value Attention: Efficient KV Caching via On-Demand Key ReconstructionThe paper is 'Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction' by Vishesh Tripathi, Abhay Kumar, and Ramsha Khan.#Review#KV Cache#Grouped Value Attention#Autoregressive Decoding#Transformer#Memory Efficiency#Decoupled RoPE#Linear Reconstruction2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Expert-Space Exploration in MoE Reinforcement Learning본 논문은 기존의 LLM post-training 연구가 token-level의 sampling에만 집중하고, MoE 모델의 핵심인 expert routing의 다양성 활용에는 소홀했다는 점을 지적하며 ESRL을 제안합니다.#Review#MoE#Reinforcement Learning#Expert Routing#Exploration#Rollout Routing Replay2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Dream-RSI: Recursive Self-Improvement through Evolving Worlds본 논문은 자율 AI 에이전트가 복잡한 도메인에서 고가치 솔루션을 발견하기 위한 recursive self-improvement (RSI) 과정에서 exploration strategy를 관리하고 개선하는 데 따르는 근본적인 병목 현상을 해결하고자 합니다.#Review#Recursive Self-Improvement#Exploration Policy#Discovery History#Replay Simulator#Meta-optimization#Algorithm Engineering#Mathematical Optimization#GPU Kernel Engineering2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Discovery Foundation Models: Toward Open-Ended Discovery Intelligence본 논문은 Foundation Model이 기존 지식에 대한 학습과 추론을 넘어, action과 tool use, outcome feedback을 통한 학습에서 open-ended discovery를 위한 다음 frontier로의 전환, 즉 새로운 문제, 표현, 설명 및 지식 생성 과정에 참여하는 능력인 Discovery Intelligence를 제안합니다.#Review#Discovery Foundation Models#Discovery Intelligence#Open-Ended Discovery#Research State#Discovery Skill#Dry-Lab/Wet-Lab#Scientific AI#Capability Formation2026년 9월 14일댓글 수 로딩 중
[논문리뷰] BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender기존 Video Understanding 벤치마크는 주로 Question Answering (QA) 방식으로 모델을 평가하지만, 이러한 방식은 모델이 Answer Prior나 단일 프레임(Single Frame) 정보에 의존하여 정답을 맞출 수 있어 비디오의 Spatiotemporal한 이해를 온전히 입증하지 못한다.#Review#Agentic Video Understanding#Programmatic Reconstruction#Blender#Benchmark#Multimodal Agents#Dual VQA#Latent Similarity2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning본 연구는 complex, cluttered manipulation scenes에서 3D point-cloud observations의 내재된 모호성 문제를 해결하고자 합니다.#Review#Embodied AI#3D Point Cloud#Diffusion Policy#Object-aware#Attentional Conditioning#Robotic Manipulation#Clutter Robustness2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Atria Dawn: The Dawn of Agentic Superintelligence본 논문은 AI 에이전트가 스스로의 successors 개발에 참여함에 따라 intelligence 생산 방식과 인간 연구자의 역할이 변화하는 문제를 다룬다.#Review#Agentic Superintelligence#Human-AI Collaboration#Recursive Self-Improvement#Foundation Agentic Language Model#Verifiable Experience Pipeline#Scientific Research Automation#Engineering Workflows#Agentic Evaluation2026년 9월 14일댓글 수 로딩 중
[논문리뷰] AlayaVista: Streaming World Modeling from Panoramic States to Perspective VideoInteractive video world models은 camera motion 하에서 broad scene context를 유지하고, high-fidelity observations를 low latency로 생성해야 하는 중요한 과제에 직면해 있습니다.#Review#World Modeling#Panoramic Video#Perspective Video#Streaming Generation#Camera Control#Latent Space#Diffusion Models#Dataset2026년 9월 14일댓글 수 로딩 중
[논문리뷰] Agent as Policy for Robotic Manipulation본 논문은 기존의 로봇 제어 방식이 지닌 유연성 부족 문제를 해결하기 위해, 범용 에이전트가 직접 물리적 로봇을 제어하는 AGP를 제안합니다. 기존의 프로그램 합성 방식은 불확실한 환경 변화에 대응하는 규칙을 미리 예측하기 어렵고, 정책 오케스트레이션 방식은 선택된 정책이 가진 고유한 한계에 묶인다는 단점이 있습니다.#Review#Robotic Manipulation#Multimodal Large Language Models#Agentic Systems#Runtime Programming#Policy Orchestration#Zero-shot Manipulation2026년 9월 14일댓글 수 로딩 중
[논문리뷰] TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents본 논문은 GUI 에이전트의 추론 과정에서 발생하는 높은 latency와 메모리 비용 문제를 해결하기 위해, 재사용 가능한 시각적 상태(visual state) 기반의 효율적인 프루닝 프레임워크를 제안합니다.#Review#GUI Agents#Multimodal Large Language Models#Visual Token Pruning#Lifecycle-aware Pruning#Trajectory-robust Admission#Nested Evidence Ordering2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Studying Without a Syllabus: Task-Agnostic Environment PreprocessingLLM Agent의 Performance는 Prompts, Tools, Context and Memory Management, Corpora 및 External Services의 구성 방식에 크게 의존합니다.#Review#LLM Agent#Task-Agnostic#Environment Preprocessing#Meta-Agent#Artifacts#Study Budget#Downstream Reward#Test-Time Compute2026년 9월 13일댓글 수 로딩 중
[논문리뷰] StepAudio 3 Gen Technical Report기존 오디오 생성 연구는 Text-to-Speech (TTS), Text-to-Audio, Text-to-Music 등 개별 도메인에 특화되어 발전해 왔으며, 이는 다양한 종류의 오디오를 필요로 하는 애플리케이션에서 호환되지 않는 표현, 컨디셔닝 형식, 및 생성 파이프라인 문제를 야기했습니다.#Review#Audio Generation#Discrete Autoregressive Modeling#Residual Vector Quantization (RVQ)#Text-to-Speech (TTS)#Voice Design#Interference-aware Pretraining#RVQ Adaptor#Unified Audio Model2026년 9월 13일댓글 수 로딩 중