[논문리뷰] Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing본 논문은 현재의 확산 모델(Diffusion-based models) 기반 이미지 편집 시스템이 표면적인 지시사항 수행(Surface-level instruction following)에만 치중하여 논리적 일관성이 결여된 결과물을 생성하는 문제를 해결하고자 합니다 .#Review#Image Editing#Reasoning-aware#Benchmark#Diffusion Models#Multi-modal LLMs#Logic Consistency#EditRefine2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction본 논문은 기존의 Video MLLM들이 미래 사건 예측(VEP) 시 텍스트 기반의 Chain-of-Thought(CoT)에 의존함에 따라 발생하는 시각적 정보 손실 문제를 해결하고자 합니다.#Review#Video Event Prediction#Multimodal Large Language Models#Latent Visual Reasoning#Interleaved Reasoning#Reinforcement Learning#Future-L1#LA-DAPO2026년 6월 4일댓글 수 로딩 중
[논문리뷰] ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment본 논문은 자율 연구 에이전트가 기술의 미래 발전 방향을 예측하는 의사결정 영역에서 얼마나 타당한 판단을 내릴 수 있는지에 대한 근본적인 의문을 제기합니다.#Review#LLM Agents#Foresight Evaluation#Scientific Judgment#Temporal Integrity#Benchmark#Research Forecasting2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Flash-WAM: Modality-Aware Distillation for World Action Models본 논문은 WAM이 manipulation 벤치마크에서 강력한 성능을 보임에도 불구하고, 실시간 제어를 저해하는 높은 inference latency 문제를 해결하고자 합니다. 기존 WAM은 video 및 action denoising에 수십 단계의 반복적인 과정을 거쳐야 하므로 실시간 로봇 제어에 부적합합니다.#Review#World-Action Models#Step Distillation#Consistency Models#Robotic Foundation Models#Flow Matching#Modality-Aware Distillation2026년 6월 4일댓글 수 로딩 중
[논문리뷰] EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management기존의 데이터 과학 에이전트는 고정된 작업 워크플로우와 제한적인 Action space에 의존하여, 경험을 체계적으로 축적하거나 재사용하는 능력이 부족합니다.#Review#Data Science Agent#Multi-Agent System#Self-Evolving#Agent Skill#Agentic Reinforcement Learning2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?본 논문은 비디오 생성 모델이 단순히 시각적으로 그럴듯한 영상을 만드는 수준을 넘어, 실제 물리 법칙을 내재화한 'World Model'로서의 기능을 수행하는지 검증하고자 합니다.#Review#Video Generation Models#Robotic Manipulation#Physical Executability#Benchmark#Sim-to-Real#World Models2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning본 논문은 기존 자율주행 시스템이 행동 조건부 동역학(Action-conditioned dynamics)을 명시적으로 모델링하지 못하고, 단순한 Direct State-to-Action Mapping에 의존한다는 근본적인 한계를 해결하고자 한다 .#Review#Autonomous Driving#World Model#Discrete Diffusion#Token Editing#Policy Learning#Counterfactual Reasoning2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Complexity-Balanced Diffusion Splitting본 논문은 표준 확산 모델이 사용하는 단일 모놀리식(monolithic) 구조의 비효율성을 해결하고자 합니다. 기존 방식은 단순한 노이즈부터 복잡한 데이터 구조까지 모든 영역을 하나의 고정된 네트워크가 처리하게 하여, 특정 생성 단계에서 필요한 적정 모델 용량을 적재적소에 할당하지 못하는 한계가 있습니다.#Review#Diffusion Models#Complexity-Balanced Splitting#Temporal Capacity Allocation#De Boor Principle#Dirichlet Energy#Path Acceleration#Generative Flow2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination본 논문은 RLVR의 확장을 가로막는 핵심 병목인 '도전적인 검증 가능(verifiable) 코드 데이터의 희소성' 문제를 해결하고자 합니다.#Review#RLVR#Synthetic Data#Atomic Decomposition#Code Generation#Scaling#Reinforcement Learning2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Benchmark Everything Everywhere All at Once본 논문은 기존의 수동적인 벤치마크 구축 방식이 가진 한계인 노동 집약성, 재사용 불가능성, 그리고 모델 성능 향상에 따른 빠른 벤치마크 포화(Saturation) 문제를 해결하고자 합니다.#Review#Benchmark Agent#Autonomous Evaluation#Benchmark Construction#MLLM-as-a-Judge#Agentic Workflow#Performance Saturation2026년 6월 4일댓글 수 로딩 중
[논문리뷰] ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?본 연구는 기존 RPLA 벤치마크가 캐릭터를 서사 흐름과 무관한 정적인 persona로 간주하여 발생하는 행동 일관성 부족 문제를 해결하고자 합니다.#Review#Role-Playing Language Agents#Character Arc#Narrative Evaluation#Temporal Alignment#Language Model Benchmarking#Persona Grounding2026년 6월 4일댓글 수 로딩 중
[논문리뷰] AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints본 논문은 실세계 복잡한 환경에서 LLM 에이전트가 Progressive Disclosure되는 Dual Constraints 환경 하에서 효과적으로 계획을 수립하고 수정하는 능력이 부족하다는 점을 지적한다.#Review#Large Language Model Agents#Adaptive Planning#Dual Constraints#Progressive Disclosure#Interactive Benchmarking#Constraint-based Planning2026년 6월 4일댓글 수 로딩 중
[논문리뷰] AdaCodec: A Predictive Visual Code for Video MLLMs본 논문은 기존 비디오 MLLMs가 비디오의 시간적 중복성(Temporal Redundancy)을 무시하고 모든 프레임을 독립적인 RGB 이미지로 처리하여 발생하는 비효율성 문제를 해결한다.#Review#Video MLLMs#Predictive Coding#Visual Token#Efficiency#Temporal Redundancy#GOP (Group of Pictures)#Latency2026년 6월 4일댓글 수 로딩 중
[논문리뷰] Absorbing Complexity: An Interaction-Native Knowledge Harness for Financial LLM Agents본 논문은 금융 AI 에이전트가 겪는 '금융 인지 마찰(financial cognition friction)'과 그에 따른 성능 저하 문제를 해결합니다.#Review#Financial LLM Agents#Interaction-Native#Knowledge Harness#Temporal Knowledge Graph#Passive Knowledge Injection#Execution Safety#Cognition Friction2026년 6월 4일댓글 수 로딩 중
[논문리뷰] AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents기존의 LLM 에이전트는 사용자의 Literal query에만 집중하여, 그 이면에 숨겨진 의도(예: '누가 어디에 있는가?'라는 질문 속에 숨겨진 '지금 그 사람이 대화할 여유가 있는가?'라는 의도)를 간과하는 문제가 있다.#Review2026년 6월 4일댓글 수 로딩 중
[sglang] DeepSeek V4의 Prefill 성능을 1.35배 향상시킨 FlashAttention 최적화DeepSeek V4 모델의 Prefill 단계 성능을 획기적으로 개선한 FlashAttention 최적화 분석#AI#LLM#Performance Optimization#FlashAttention#DeepSeek V4#SGLang2026년 6월 3일댓글 수 로딩 중
[feast] Feast 온라인 서빙 성능 튜닝: Sub-2ms 달성을 위한 여정Feast 온라인 피처 서버의 p99 지연 시간을 sub-2ms로 단축하기 위한 성능 튜닝 과정을 상세히 분석합니다.#Feast#성능 최적화#Kubernetes#Redis#Python2026년 6월 3일댓글 수 로딩 중
[vllm] [ROCm CI 최적화] Docker 3단계 빌드 전략으로 빌드 시간 26분 단축하기vLLM 프로젝트의 ROCm CI 빌드 시간을 획기적으로 단축하기 위해 도입된 3단계 Docker 빌드 아키텍처와 Content-addressed 캐싱 기법을 심층 분석합니다.#vLLM#ROCm#Docker#CI/CD#Buildkite#Optimization2026년 6월 3일댓글 수 로딩 중
[transformers] Hugging Face Transformers: Slow Tokenizer 성능 회귀 문제 해결하기PreTrainedTokenizer의 O(T*N*logN) 성능 저하 문제를 O(T)로 복구한 최적화 사례 분석#HuggingFace#Transformers#Python#Optimization#Tokenizer2026년 6월 3일댓글 수 로딩 중
[논문리뷰] ZipSplat: Fewer Gaussians, Better Splats본 논문은 기존의 Feed-forward 3DGS 방식이 3D Gaussian 배치를 입력 이미지의 픽셀 그리드에 고정시킴으로써 발생하는 구조적 비효율성을 해결하고자 합니다.#Review#3D Gaussian Splatting#Feed-forward Reconstruction#Novel View Synthesis#Scene Tokens#Clustering#Pose-free2026년 6월 3일댓글 수 로딩 중