[논문리뷰] GD^2PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization본 논문은 LLM의 다차원적 성능 향상을 위해 사용되는 Multi-Reward RL 환경에서 발생하는 Advantage 상쇄 문제를 해결하고자 한다.#Review#Reinforcement Learning#Multi-Reward Optimization#Policy Optimization#Conflict Mitigation#Dynamic Filtering#Tool Calling#Alignment2026년 6월 15일댓글 수 로딩 중
[논문리뷰] FastContext: Training Efficient Repository Explorer for Coding Agents본 논문은 LLM 기반 코딩 에이전트의 저장소 탐색 단계에서 발생하는 고비용 토큰 소비 및 불필요한 컨텍스트 오염 문제를 해결하기 위해 제안되었다. 기존 에이전트들은 동일한 모델이 탐색과 문제 해결을 모두 수행하여, 탐색 과정에서 누적된 방대한 양의 관련 없는 코드 스니펫이 주 모델의 컨텍스트를 오염시킨다 .#Review#Coding Agents#Repository Exploration#Subagent Architecture#Supervised Fine-Tuning#Reinforcement Learning#Context Efficiency#Token Consumption2026년 6월 15일댓글 수 로딩 중
[논문리뷰] EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video본 논문은 일상적인 상호작용이 담긴 단일 egocentric RGB 영상으로부터 복잡한 변형체(Deformable objects)의 물리적 속성을 파악하여 '디지털 트윈'을 구축하는 난제를 해결하고자 합니다.#Review#Physical Understanding#Real-to-sim#Egocentric Video#Deformable Objects#Digital Twin#Physics-based Simulation2026년 6월 15일댓글 수 로딩 중
[논문리뷰] DreamX-World 1.0: A General-Purpose Interactive World Model본 논문은 다양한 visual domain(photorealistic, game-style, stylized) 전반에서 카메라 탐색 및 이벤트 조작을 실시간으로 지원하는 general-purpose interactive world model 구축을 목표로 합니다 .#Review#Interactive World Model#Camera Control#E-PRoPE#Memory-Conditioned Scene Persistence#Event Instruction Tuning#Autoregressive Distillation#Reinforcement Learning2026년 6월 15일댓글 수 로딩 중
[논문리뷰] Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories본 논문은 데이터 저널리즘에서 발생하는 할루시네이션(Hallucination) 문제와 데이터 투명성 결여를 해결하기 위해 Data2Story를 제안한다.#Review#Data Journalism#Multi-Agent System#Evidence-Grounded#Multimodal Generation#Verifiability#Auditability2026년 6월 15일댓글 수 로딩 중
[논문리뷰] CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?본 논문은 현대의 자율 에이전트가 실제 소프트웨어 엔지니어링이나 데이터 분석 현장에서 겪는 복잡한 데이터 처리 요구사항을 해결하지 못하고 있다는 문제의식에서 출발합니다.#Review#CoDA-Bench#Code Agents#Data-Intensive Tasks#Data Discovery#Autonomous Engineering#Kaggle Ecosystem#Evaluation Benchmark2026년 6월 15일댓글 수 로딩 중
[논문리뷰] BadWorld: Adversarial Attacks on World Models본 논문은 VWMs의 잠재적 취약성을 평가하기 위한 최초의 적대적 공격 프레임워크인 BadWorld를 제안합니다.#Review#Adversarial Attack#Visual World Models#Autoregressive Generation#Flow Matching#Trajectory-Adaptive Optimization#Label-Free2026년 6월 15일댓글 수 로딩 중
[논문리뷰] BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering본 논문은 기존의 Physically-based inverse rendering 모델들이 가지는 물리적 불일치 문제와 Generative 모델들의 제어 불가능성 문제를 동시에 해결하기 위해 BRDFusion 프레임워크를 제안합니다.#Review#Inverse Rendering#3D Gaussian Splatting#Generative Prior#Relighting#Urban Scene#Diffusion Model2026년 6월 15일댓글 수 로딩 중
[논문리뷰] Artificial Intelligence Index Report 2026본 보고서는 AI 기술이 전례 없는 속도로 확산됨에 따라, 기술 발전 속도와 이를 관리하기 위한 거버넌스 및 평가 프레임워크 간의 격차가 심화되는 문제를 제기한다.#Review#Generative AI#AI Sovereignty#Technical Benchmarks#AI Adoption#Responsible AI2026년 6월 15일댓글 수 로딩 중
[논문리뷰] μ_0: A Scalable 3D Interaction-Trace World Model본 논문은 기존 로봇 학습이 직면한 데이터 파라독스, 즉 '액션이 포함된 로봇 데이터의 희소성'과 '비디오 데이터의 높은 가용성' 사이의 간극을 해결하고자 합니다 .#Review#World Model#3D Interaction-Trace#Robot Manipulation#Cross-Embodiment Learning#Semantic Flow Matching#Data Pipeline2026년 6월 14일댓글 수 로딩 중
[논문리뷰] iMaC: Translating Actions into Motion and Contact Images for Embodied World Models본 논문은 Embodied World Model이 로봇 정책(Policy) 평가 시 가지는 행동 조건부(Action-Conditioning) 비디오 생성의 불확실성 문제를 해결하고자 한다.#Review#Embodied World Models#Action-Conditioned Video Generation#Robot Policy Evaluation#Motion Images#Contact Images#URDF/FK#Long-Horizon Manipulation2026년 6월 14일댓글 수 로딩 중
[논문리뷰] World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible본 논문은 기존의 단일 이미지 3D 추정 방식이 가진 '충실도(Faithfulness)'와 '완전성(Completeness)' 사이의 상충 문제를 해결하고자 합니다.#Review#World Tracing#Pixel-Aligned#Geometry Generation#Diffusion Transformer#Flow Matching#Multilayer#3D Vision2026년 6월 14일댓글 수 로딩 중
[논문리뷰] When is Your LLM Steerable?본 연구는 Activation Steering의 성공 여부가 모델, 프롬프트, 개념, 그리고 Steering Strength의 복합적인 요소에 의해 결정되는 취약성 문제를 해결하고자 합니다.#Review#Activation Steering#Steerability Prediction#LLM Inference#Gradient Boosting Decision Trees#ASTEER Dataset#SteerBoost2026년 6월 14일댓글 수 로딩 중
[논문리뷰] WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis본 논문은 3D MRI 합성 시 발생하는 높은 계산 비용과 해부학적 상세 정보 손실 문제를 해결하기 위해 WaveDiT를 제안합니다.#Review#3D MRI Synthesis#Flow Matching#Discrete Wavelet Transform#Heteroscedastic Uncertainty#Generative Models#Brain Age Prediction2026년 6월 14일댓글 수 로딩 중
[논문리뷰] VISTA: View-Consistent Self-Verified Training for GUI Grounding본 논문은 기존의 GRPO를 활용한 GUI Grounding 학습에서 발생하는 보상 퇴화(reward degeneracy) 문제를 해결하는 데 집중합니다.#Review#GUI Grounding#GRPO#Self-Verified Training#View-Consistent#Reinforcement Learning#VLM2026년 6월 14일댓글 수 로딩 중
[논문리뷰] The Hidden Power of Scaling Factor in LoRA Optimization본 논문은 LoRA 학습 시 하이퍼파라미터인 scaling factor $\alpha$의 역할이 체계적으로 연구되지 않았으며, 단순히 learning rate($\eta$)의 보조적 수단으로만 간주되어 온 점을 지적합니다.#Review#LoRA#Scaling Factor#Optimization Dynamics#Signal-Drift Framework#Spectral Suppression#PEFT2026년 6월 14일댓글 수 로딩 중
[논문리뷰] The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment본 논문은 독립적으로는 정렬된(Aligned) 에이전트들이 상호작용하며 발생하는 예측 불가능한 시스템 레벨의 위험을 감지하기 위한 실시간 감사 프레임워크를 제안한다.#Review#Multi-agent Safety#Emergent Misalignment#Alignment Auditing#LLM Agents#AI Control#Budget-constrained Monitoring2026년 6월 14일댓글 수 로딩 중
[논문리뷰] Squeeze-Release: Iterative Pruning with Exact Structural Minimization본 논문은 일반적인 비구조적(Unstructured) Pruning이 파라미터의 중요도에 따라 0으로 만들더라도, 실제 tensor의 물리적 크기를 줄이지 못해 모델 압축 효과가 미비한 문제를 해결하고자 한다. .#Review#Network Pruning#Model Compression#Iterative Pruning#Function-preserving Transformations#Layer Normalization2026년 6월 14일댓글 수 로딩 중
[논문리뷰] Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO본 논문은 GRPO (Group Relative Policy Optimization) 기반 LLM 학습에서 rollout diversity를 향상시키기 위한 새로운 차원을 식별한다.#Review#GRPO#LLMs#Policy-Level Diversity#Token-Level Diversity#S2L-PO#Reinforcement Learning#Mathematical Reasoning#Parameter-Level Compression2026년 6월 14일댓글 수 로딩 중
[논문리뷰] Skip a Layer or Loop It? Learning Program-of-Layers in LLMs본 논문은 모든 입력에 대해 고정된 depth와 순서로 수행되는 기존 LLM의 정적 추론 방식이 비효율적이며, 모델의 잠재적 추론 능력을 충분히 활용하지 못한다는 점을 지적합니다 .#Review#Large Language Models#Dynamic Inference#Program-of-Layers#Test-time Scaling#Layer Skipping#Layer Recurrence#Computational Efficiency2026년 6월 14일댓글 수 로딩 중