[논문리뷰] Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking본 논문은 기존의 VOT 방식들이 task-specific supervised training에 의존하여 unseen 환경에 대한 일반화 능력이 제한적이라는 점을 지적합니다.#Review#Visual Object Tracking#Foundation Models#SAM 2#Nonlinear Motion#Motion Predictor#Error Detection-Recovery2026년 5월 21일댓글 수 로딩 중
[논문리뷰] SceneAligner: 3D-Grounded Floorplan Localization in the Wild본 논문은 대규모 환경 및 상업용 건물의 비정형(in-the-wild) 이미지 컬렉션 내에서 카메라 관측치를 2D floorplan에 로컬라이제이션하는 문제를 다룬다.#Review#Floorplan Localization#3D Foundation Models#Cross-modal Correspondence#Density Map#LoRA#Computer Vision2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws본 논문은 기존의 스케일링 법칙이 최적화기(optimizer)를 고정된 요소로 간주하여, 모델 내부 표현의 구조적 차이를 간과한다는 점을 문제로 지적합니다. 저자들은 동일한 아키텍처와 컴퓨팅 자원을 사용하더라도 최적화기 선택에 따라 FFN 폭이 실제 유효 용량으로 전환되는 효율이 크게 달라질 수 있음을 밝힙니다 .#Review#Spectral Scaling Laws#Optimizer Geometry#Effective Rank#FFN Width#Representation Scaling2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Q-ARVD: Quantizing Autoregressive Video Diffusion Models본 논문은 실시간 인터랙티브 비디오 생성을 위한 ARVDs의 추론 비용 문제를 해결하기 위해 모델 양자화(Model Quantization)를 제안합니다.#Review#Autoregressive Video Diffusion Models#Model Quantization#Frame-wise Sensitivity#Outlier-aware Quantization#Dual-scale Quantization2026년 5월 21일댓글 수 로딩 중
[논문리뷰] PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects기존의 3D 생성 연구들은 주로 시각적인 사실성(photorealism)에만 집중하여 물리 기반 시뮬레이션이나 실제 로봇 제어 환경에서 요구되는 물리적 속성을 결여하고 있습니다. 또한, 기존 방법론들은 rigid, deformable, articulated 등 특정 객체 유형에 국한되어 있어 범용적인 활용이 어렵습니다 .#Review#PhysX-Omni#Simulation-Ready#3D Generation#PhysXVerse#PhysX-Bench#Vision-Language Model2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?본 논문은 MLLM이 인적 자원 관리나 정신 건강 진단 등 인간 중심적인 역할에 배치되면서 핵심적으로 요구되는 성격 인식(personality perception) 능력을 진단하고자 합니다.#Review#Multimodal Large Language Models#Personality Perception#Grounded Personality Reasoning#MM-OCEAN#Prejudice Gap#Holistic-Grounding Rate#Apparent Personality Recognition2026년 5월 21일댓글 수 로딩 중
[논문리뷰] One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems본 논문은 기존의 디지털 단편 드라마 제작 방식이 가진 narrative pacing의 부재, 클립 간 spatial consistency 부족, 그리고 높은 manual review 의존성이라는 세 가지 핵심 문제를 해결하고자 합니다.#Review#Short-Form Drama#Multi-Agent System#3D-Grounded Generation#Narrative Pacing#Spatial Consistency#Production-Level Quality Control2026년 5월 21일댓글 수 로딩 중
[논문리뷰] OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding본 논문은 Omni-modal Large Language Models(MLLMs)의 발전에도 불구하고, 실제 환경에서의 Proactive 스트리밍 이해 능력을 정밀하게 평가할 수 있는 표준화된 벤치마크가 부재하다는 문제점을 해결하고자 합니다 .#Review#Omni-proactive streaming#Video understanding#Benchmark#Multimodal LLMs#Audio-visual perception#Long-horizon evaluation2026년 5월 21일댓글 수 로딩 중
[논문리뷰] More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts본 논문은 정치적 텍스트에서 Schwartz values를 감지할 때, 주변 문맥(Context)과 명시적인 도덕 지식이 모델 성능에 미치는 영향을 체계적으로 분석하고자 한다 . 정치적 발화는 가치가 간접적으로 표현되는 경우가 많아 문장 단위의 분류가 매우 어렵다.#Review#Schwartz Values#Political Text#Retrieval-Augmented Generation (RAG)#DeBERTa#Large Language Models (LLMs)#Context Analysis2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Minimalist Visual Inertial Odometry본 연구는 자원 제약적인 로봇 플랫폼에서 기존 VIO (Visual-Inertial Odometry) 시스템의 높은 전력 소모 및 계산 요구사항이 가지는 한계점을 해결하고자 합니다.#Review#Visual-Inertial Odometry#Minimalist Vision#Planar Odometry#Gabor Masks#Photodiode#Temporal Convolutional Network#Motion Estimation2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles본 논문은 현대 LLM 에이전트가 특정 도메인에 강점을 가진 다양한 전문가 모델과 모듈식 스킬을 효과적으로 활용하지 못하는 Coordination Bottleneck 문제를 해결하고자 합니다.#Review#Reinforcement Learning#Multimodal Agent#Orchestration#Skill Library#Expert Models#Hierarchical Registry2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search본 논문은 LLM이 생성한 Lean 4 증명이 정답은 맞추지만, 지나치게 장황하고 특정 버전의 라이브러리에 취약하다는 점을 해결하고자 합니다 .#Review#Lean 4#Proof Optimization#Agentic Framework#Retrieval-Augmented Generation#Multi-Objective Optimization#Formal Verification2026년 5월 21일댓글 수 로딩 중
[논문리뷰] LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning본 논문은 기존의 Explicit Text CoT 기반 MLLM이 고차원 오디오-비주얼 정보를 텍스트라는 좁은 병목으로 압축함에 따라, 다중 모달 간의 세밀한 시간적 정렬과 의미적 연결을 놓치는 문제를 해결하고자 한다.#Review#Multimodal Large Language Models#Audio-Visual Reasoning#Latent Reasoning#Cross-modal Alignment#Chain-of-Thought#Instruction Tuning2026년 5월 21일댓글 수 로딩 중
[논문리뷰] KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving본 논문은 Disaggregated LLM Serving 환경에서 KV cache 통신이 전체 end-to-end 지연시간의 최대 60%를 차지하는 주요 병목 현상을 해결하고자 한다 .#Review#LLM Serving#KV Cache Compression#Disaggregated Inference#Bayesian Optimization#Service-Aware Control2026년 5월 21일댓글 수 로딩 중
[논문리뷰] GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation본 논문은 오픈 엔드 이미지 생성이 단순한 텍스트 프롬프트 기반의 task를 넘어, 모델의 내부 지식과 외부 리소스를 효과적으로 결합해야 하는 복잡한 에이전트 과정임을 강조합니다.#Review#Image Generation#Agentic Workflow#Self-Evolving#Visual Experience Distillation#Tool-Orchestrated#On-Policy Distillation#Multimodal Agent2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention본 논문은 Linear Attention 기반 모델들에서 메모리 편집의 핵심인 erase(제거)와 write(삽입) 동작이 단일 scalar gate에 의해 묶여 있는 구조적 한계를 해결하고자 합니다.#Review#Linear Attention#Recurrent Neural Networks#Delta Rule#Fast-Weight Memory#Selective State Space#Chunkwise Parallel Training#Long-Context Retrieval2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps본 논문은 Long-context 추론 시 발생하는 full attention의 이차 비용(quadratic cost) 문제를 해결하기 위해 효율적인 스파스(sparse) 구조로의 전환을 제안한다.#Review#Long-context LLM#Sparse Attention#Head Specialization#Dynamic Top-pp Selection#Efficient Inference#Self-distillation2026년 5월 21일댓글 수 로딩 중
[논문리뷰] From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning본 논문은 난도가 높은 추론 문제에 대해 기존의 RLVR 방식이 가지는 효율성 한계를 해결하고자 한다 . 고난도 문제에서는 최종 정답에 도달하는 경로가 매우 희소하여, 모델이 중간 단계에서 올바른 추론을 수행하더라도 이를 학습 신호로 적절히 환원하기 어렵다.#Review#Curriculum Reinforcement Learning#LLM Reasoning#Credit Assignment#Verifiable Rewards#Subproblem Decomposition#RLVR2026년 5월 21일댓글 수 로딩 중
[논문리뷰] FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching본 논문은 비디오 Diffusion 모델의 생성 범위를 학습된 문맥 길이 이상으로 확장하는 과정에서 발생하는 품질 저하와 시간적 일관성 문제를 해결하고자 합니다.#Review#Long Video Generation#Flow Matching#Tweedie Matching#Stochastic Early-Phase Sampling#Inference-time Framework#Diffusion Models2026년 5월 21일댓글 수 로딩 중
[논문리뷰] Diversed Model Discovery via Structured Table Discovery본 논문은 기존의 모델 검색 시스템이 텍스트 중심의 시맨틱 유사도에 과도하게 의존하여 결과의 다양성을 저해하고 비교 가능한 정보를 충분히 제공하지 못하는 문제를 해결하고자 한다.#Review#Model Lake#Model Search#Structured Semantic Search#Table Discovery#Nugget-based Evaluation#Model Cards2026년 5월 21일댓글 수 로딩 중