[논문리뷰] Video Generative Models as Geometry Learner본 논문은 기존의 단일 이미지 기반 기하학적 추정(monocular geometry estimation) 방식이 직면한 데이터 의존성과 정밀도 한계를 해결하고자 한다.#Review#Video Generative Models#Geometry Estimation#Monocular Depth#Surface Normal#Next-frames Prediction#Diffusion Models2026년 8월 30일댓글 수 로딩 중
[논문리뷰] StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing본 논문은 LLM 기반 에이전트의 실질적인 위험(파일 변조, 정보 유출 등)을 사전에 차단하기 위한 Step-level 안전 감독 기술의 부재를 해결하고자 합니다.#Review#LLM Agents#Guardrails#Step-level Supervision#Safety-Utility Balancing#Balance-GRPO#Synthetic Data2026년 8월 30일댓글 수 로딩 중
[논문리뷰] StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments본 논문은 기업 환경(Enterprise Environments)에서 발생하는 LLM 에이전트의 상호작용 불일치와 낮은 태스크 성공률 문제를 해결하고자 한다. 기존 모델 중심의 접근 방식은 복잡한 도메인 관례나 상태 의존성(State Dependencies)이 요구되는 도구 사용 환경에서 최적화에 한계를 보인다.#Review#Harness Evolution#Agentic Workflow#Enterprise Environments#Stratified Sampling#Frozen Model Transfer#Tool Use2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Sliding-window beats linear attention본 논문은 LLM의 quadratic scaling 문제로 인한 추론 비용 및 메모리 문제를 해결하기 위해, 기존의 Linear Attention 기반 변환 기법들이 Sliding Window Attention (SWA) 보다 효율적인지 검증하고자 합니다.#Review#Sliding Window Attention#Linear Attention#KV Cache#Long-context Reasoning#Model Distillation#Inference Efficiency2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Rubric-to-Code Credit Assignment for Reinforcement Learning본 논문은 인터랙티브 웹 애플리케이션 생성 시 발생하는 학습 신호의 불투명성 문제를 해결하고자 합니다.#Review#Reinforcement Learning#Interactive Web Application#Credit Assignment#GRPO#Hierarchical Reward#Rubric-Driven Training2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion본 논문은 autoregressive 비디오 생성 모델이 장기 생성 과정에서 겪는 '객체 영속성(Object Permanence)'의 결여와 장기 기억 용량의 한계를 해결하는 것을 목표로 합니다.#Review#Long Video Generation#Autoregressive Video Diffusion#Long-term Memory#Object Permanence#Ring-Structured Training2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction본 논문은 초장기(ultra-long) 비디오 스트림에서 카메라 모션과 3D 장면을 추론할 때 발생하는 성능 저하와 계산 비용 문제를 해결하고자 합니다.#Review#3D Reconstruction#Streaming Inference#Local Context#Camera Pose Estimation#Transformer#Long-Horizon Stability2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090본 논문은 대규모 언어 모델(LLM)의 pretraining 비용이 지나치게 높아 학계나 소규모 연구실의 접근성이 떨어진다는 문제 의식에서 출발합니다.#Review#Large Language Models#Pretraining#Cost-Efficiency#RTX 5090#FP8#MuonH#Curriculum Model Averaging2026년 8월 30일댓글 수 로딩 중
[논문리뷰] PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control본 논문은 기존 VLA 모델들이 사전 학습된 풍부한 문맥 이해 능력을 로봇 제어의 에피소드 메모리로 활용하지 못하고, 별도의 복잡한 메모리 기법을 추가하는 구조적 비효율성 문제를 해결하고자 합니다 .#Review#Multimodal Large Language Models#Robot Control#Episode Memory#System 2 Cognition#Vision-Language-Action Models#Asynchronous Interface#End-to-End Training2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents본 논문은 multimodal agents가 external tools를 통해 complex digital environments에서 작업을 수행하는 능력이 증가함에도 불구하고, dexterous visual tool use라는 핵심적이지만 잘 탐구되지 않은 역량에 대한 평가가 부족하다는 점을 지적합니다.#Review#Multimodal Agents#Visual Tool Use#Benchmarking#Parameterized Action#Closed-Loop Control#Visual Reconstruction#Trajectory Supervision2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors본 논문은 영어 위주로 개발된 Extractive Prompt Compressor가 비영어권 언어에 적용될 때 심각한 성능 저하와 불공정한 비용 페널티를 유발한다는 문제를 지적한다. 기존의 많은 연구는 영어를 기준으로만 평가되어 왔으나, 비영어권 언어는 이미 더 높은 Token Premium을 지불하고 있다.#Review#Prompt Compression#Cross-Lingual#LLM#Inference Cost#Extractive Compression#Tokenization#Transfer Gap2026년 8월 30일댓글 수 로딩 중
[논문리뷰] LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering본 논문은 Loop Engineering 패러다임 내에서 Large Language Models (LLMs)의 runtime controllers로서의 역량을 효과적으로 평가하기 위한 벤치마크의 필요성을 제기합니다.#Review#Loop Engineering#Benchmarking#Runtime Controllers#Coding Agents#Long-Horizon Tasks#Strict Success Rate#Inference Cost#Loop Contract2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding기존의 MLLM 기반 STVG 모델들은 영상 내 객체의 이동 궤적(Trajectory)을 Autoregressive 방식으로 순차 생성하는데, 이는 Tube 길이가 길어질수록 Decoding Latency가 선형적으로 증가하고, 이전 시점의 위치 정보 오류가 다음 시점으로 전파되는 문제를 야기한다.#Review#Spatio-Temporal Video Grounding#Parallel Tube Decoding#Decoupled Block Attention#Localization-Aware Policy Optimization#Multimodal Large Language Models2026년 8월 30일댓글 수 로딩 중
[논문리뷰] LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation본 논문은 오토레그레시브(autoregressive) 방식의 비디오 생성 모델에서 시간 경과에 따른 개체(subject)나 배경의 일관성 유지 문제를 해결하고자 합니다.#Review#Video Generation#Autoregressive#Diffusion Transformer#Long-Horizon Consistency#Memory Routing#K/V Caching#Cross-Horizon Prediction Matching2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Language Chain in Alignment: Cross-lingual Ranking Preference Optimization대규모 언어 모델(LLM)의 정렬은 주로 영어 중심의 고품질 선호도 데이터에 의존하기 때문에, 다른 언어에서는 최적 이하의 성능을 보이며 응답 정확성에서 일관성이 떨어지는 문제가 발생합니다. 이러한 불균형은 비영어권 질의에 대한 입력-출력 언어 불일치 및 성능 저하와 같은 바람직하지 않은 행동으로 이어집니다.#Review#Cross-lingual Alignment#Preference Optimization#Large Language Models#Learning to Rank#Multilingual LLMs#Instruction Following#LambdaLoss2026년 8월 30일댓글 수 로딩 중
[논문리뷰] LMSM: LLM Security Framework Inspired by Linux Security Modules본 논문은 LLM의 내적 상태를 해석하는 다양한 방법들이 실제 서비스 환경에서 통합된 보안 제어로 작동하지 못하고 파편화되어 있다는 문제를 해결합니다.#Review#LLM Security#Runtime Mediation#Linux Security Modules#Interpretability#Sparse Autoencoders#Adversarial Robustness#Model Serving2026년 8월 30일댓글 수 로딩 중
[논문리뷰] J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data본 논문은 기존의 Zero-data 기반 Self-evolving 모델들이 직면한 평가 성능의 상한선(Performance Ceiling) 문제를 해결합니다. 고정된 Judge를 사용할 경우, Solver가 해당 Judge의 변별력을 넘어서는 순간 더 이상의 학습 신호를 얻지 못하고 성능이 정체되는 현상이 발생합니다.#Review#Large Language Models#Self-evolution#Zero-data#Judge Co-adaptation#Adversarial Co-evolution#Group Relative Policy Optimization#Bradley-Terry Loss2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Generative Semantic Scene Completion본 논문은 기존의 SSC 접근 방식이 가진 데이터 불균형 및 정적인 지도 학습의 한계를 극복하기 위해 생성 모델링 기반의 접근법을 제안합니다 . 기존 모델들은 학습 데이터 내 클래스 빈도에 크게 의존하며, 특정 클래스(예: motorcyclist)는 매우 희귀하여 인식 정확도가 극도로 낮습니다.#Review#Semantic Scene Completion#Discrete Diffusion Models#LiDAR Point Clouds#Synthetic Training Data#Bird’s-Eye View Perception#Autonomous Driving2026년 8월 30일댓글 수 로딩 중
[논문리뷰] GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models본 논문은 최신 Generative VLMs가 특정 인구통계학적 속성에 따라 편향된 결과를 생성하는 문제를 해결하고자 합니다. 기존의 추론 단계 편향 완화 기법들은 주로 CLIP과 같은 정적 임베딩 공간에서 전역적인 편향을 제거하는 방식에 최적화되어 있습니다.#Review#Generative Vision-Language Models#Inference-Time Debiasing#Activation Steering#Spherical Geometry#Geodesic Interpolation#Counterfactual Bias Subspace2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Fast Weight Attention for Continual Learning본 논문은 Transformer의 KV cache가 시퀀스 길이에 따라 O(N²)의 비용을 소모하며 발생하는 비효율성 및 Continual Learning 환경에서의 파괴적 망각(catastrophic interference) 문제를 해결하고자 합니다.#Review#Continual Learning#Fast Weight#Linear Attention#State Space Models#Online Gradient Descent#Ridge Regression2026년 8월 30일댓글 수 로딩 중