[논문리뷰] Trust Region On-Policy Distillation본 논문은 Small Reasoning Models (SRM)을 위한 On-Policy Distillation (OPD)의 학습 불안정성과 비효율성 문제를 해결하고자 합니다.#Review#On-Policy Distillation#Reasoning Models#Trust Region#Policy Gradient#Knowledge Distillation#Language Models2026년 6월 2일댓글 수 로딩 중
[논문리뷰] TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL본 연구는 시각적 추론(visual reasoning)을 위한 RL 학습 시, 정적 데이터셋(static datasets)이 가진 한계를 극복하기 위해 수행되었습니다.#Review#Reinforcement Learning#Visual Reasoning#Online Environment#Multimodal Large Language Models#Rule-Verifiable#Curriculum Learning2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling본 논문은 LLM의 추론 성능을 높이기 위한 Test-Time Scaling이 과도한 연산 비용과 지연 시간(Latency)을 초래한다는 문제를 해결하고자 합니다.#Review#Test-Time Scaling#Adaptive Sampling#Reinforcement Learning#Markov Decision Process#Inference Efficiency#Large Language Models2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Prior Availability in Industrial Visual Sim-to-Real: A Review of CAD-Guided and CAD-Unavailable Regimes본 논문은 산업용 비전 시스템이 직면한 핵심 문제인 '데이터 활용 가능성'과 '실제 배포 환경 간의 도메인 간극'을 체계적으로 재정의한다 . 기존 연구들은 시뮬레이션에서 현실로의 전이를 단순히 합성 이미지에서 실사 이미지로의 변환으로 좁게 해석하는 한계가 있다.#Review#Industrial Visual Sim-to-Real#Prior Availability#CAD-Guided Vision#CAD-Unavailable Inspection#6D Object Pose Estimation#Industrial Anomaly Detection#Domain Gap2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations본 연구는 LLM의 deception detection을 위해 사용되는 Linear Probes가 실전 환경에서 보이는 극심한 성능 저하의 원인을 규명하고자 합니다.#Review#LLM#Deception Detection#Linear Probes#Scaling Laws#Robustness#Geometric Analysis#Activation Engineering2026년 6월 2일댓글 수 로딩 중
[논문리뷰] PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps본 논문은 기존의 Embodied Navigation 연구들이 Vision-Language Navigation (VLN)과 Object Goal Navigation (ObjNav)을 분리된 문제로 다루며, 이들 사이의 연계를 위해 과도한 Cross-modal 학습이나 대규모 VLM 모델에 의존하고 있다는 점을 문제로 지적한다 .#Review#Embodied Navigation#Platonic Representation Hypothesis#Topological Map#Blind Matching#Zero-shot Navigation#Cross-modal Alignment2026년 6월 2일댓글 수 로딩 중
[논문리뷰] PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training본 연구는 고성능 0.9B 파라미터 모델인 PaddleOCR-VL-1.5의 잔여 오류를 해결하여 성능을 극대화하고자 합니다 . 저자들은 단순히 훈련 데이터를 늘리는 것만으로는 긴 꼬리(long-tail) 분포의 문서 레이아웃, 복잡한 테이블, 희귀 스크립트 등에서 발생하는 오류를 근본적으로 해결할 수 없음을 관찰했습니다.#Review#Document Parsing#Vision-Language Model#Under-Optimized Region#Progressive Post-Training#Data Engine#GRPO2026년 6월 2일댓글 수 로딩 중
[논문리뷰] OCC-RAG: Optimal Cognitive Core for Faithful Question Answering본 논문은 범용 LLM이 파라미터 내 방대한 지식에 의존하여 주어진 Context를 무시하거나 할루시네이션(Hallucination)을 생성하는 문제를 해결하고자 합니다.#Review#Small Language Models#Context Question Answering#Multi-hop Reasoning#Faithfulness#Mid-training#Synthetic Data#Abstention2026년 6월 2일댓글 수 로딩 중
[논문리뷰] NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation본 논문은 기존의 Reconstruction-based 자율주행 시뮬레이터가 가진 제약 사항인 데이터 의존성과 새로운 장면(Novel scene)에 대한 일반화 부족 문제를 해결하기 위해 OmniDreams를 제안한다. 기존 방식은 캡처된 데이터 환경 내부에서만 가상 시나리오를 구성할 수 있어 확장성이 매우 제한적이다.#Review#Generative World Model#Autonomous Vehicle Simulation#Closed-Loop#Autoregressive Diffusion#World-Action Model#Vision-Language-Action2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling본 연구는 MLLM이 평가자(Judge)로 활용될 때 발생하는 Perceptual Judgment Bias를 해결하여 평가의 신뢰성을 제고하고자 합니다. 기존 MLLM 평가자들은 시각적으로 잘못된 응답임에도 불구하고 논리적으로 그럴듯한 텍스트가 포함되어 있으면 높은 점수를 부여하는 경향이 있습니다 .#Review#Multimodal LLM-as-a-Judge#Perceptual Judgment Bias#Reward Modeling#Perceptual Perturbation#GRPO#Visual Grounding2026년 6월 2일댓글 수 로딩 중
[논문리뷰] MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection본 논문은 이질적인(Heterogeneous) Mid-training 데이터 혼합물에서 효과적인 데이터 선택이 어렵다는 문제를 해결하고자 합니다.#Review#Mid-training#Data Selection#Rubric Discovery#LLM#Distillation#Source-Aware#Scalability2026년 6월 2일댓글 수 로딩 중
[논문리뷰] MERIT: Learning Disentangled Music Representations for Audio Similarity본 논문은 기존 음악 유사도 모델이 여러 음악적 요소를 하나의 Monolithic 점수로 융합하여 표현함에 따라 발생하는 해석 가능성 및 세밀한 쿼리 제어의 한계를 해결하고자 합니다 .#Review#Music Representation Learning#Disentanglement#Audio Similarity#Representation Learning#Contrastive Learning#Self-Supervised Learning2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories본 논문은 현대의 LLM이 배포 이후 새로운 정보를 지속적으로 학습하지 못하는 '정적(Static)'인 한계와, 업데이트 시 발생하는 Catastrophic Forgetting (CF) 문제를 해결하고자 합니다.#Review#Continual Learning#Language Models#Memory Consolidation#Knowledge Seeding#Self-Improvement#Dreaming#Catastrophic Forgetting2026년 6월 2일댓글 수 로딩 중
[논문리뷰] KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks본 논문은 test-time scaling 환경에서 발생하는 KV-Cache 양자화의 오류 누적 문제를 해결하는 데 집중합니다. 기존의 양자화 방식은 주로 고정된 긴 컨텍스트를 다루는 prefill 설정에서 평가되었으나, 실제 디코딩 과정에서는 토큰 생성마다 오류가 반복적으로 누적되어 추론 품질이 급격히 저하됩니다 .#Review#KV-Cache Quantization#Variance Normalization#Error Accumulation#Reasoning Tasks#Hadamard Rotation#Dual-Scaling2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking본 논문은 기존 휴머노이드 모션 트래킹 연구가 겪고 있는 데이터 및 모델 규모의 한계와 그로 인한 일반화 성능 저하 문제를 해결하고자 합니다. 기존의 연구들은 주로 소규모 MLP 기반 정책에 의존해왔으며, 이는 정교한 모션 추적과 범용적인 일반화 사이의 고질적인 트레이드오프(trade-off)를 유발했습니다 .#Review#Humanoid Motion Tracking#Transformer#Zero-Shot Generalization#Large-scale Motion Data#Harmonic Motion Embedding#DAgger Distillation2026년 6월 2일댓글 수 로딩 중
[논문리뷰] From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain본 논문은 뇌의 시각적 개념 표상을 결정하는 데 있어 기존의 Activation-based 방법론이 갖는 근본적인 한계를 해결하고자 합니다.#Review#fMRI#Causal Representation Discovery#Visual Concept Localization#Generative Models#Counterfactual Stimuli#BrainCause2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning본 논문은 LLM의 도메인 특화 적응(Domain-Specific Adaptation) 과정에서 발생하는 데이터 확보 문제를 해결하고자 한다.#Review#Domain-Specific Data Synthesis#LLMs#Minimal Sufficient Representation Learning#Prompt Tuning#Contrastive Disentanglement#Domain Adaptation2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces본 연구는 답변이 정확한 Long-CoT 데이터라도 그 내부의 추론 궤적에 따라 모델 학습의 유효성이 크게 달라질 수 있다는 점을 문제로 제기합니다. 기존 연구들은 데이터 선택이나 단순한 길이 절삭(truncation)에 의존하여 추론 단계의 품질을 근본적으로 규명하지 못했습니다.#Review#Long-CoT#Supervised Fine-Tuning#Harmful Continuation#Uncertainty–Geometry Mismatch#Reasoning Trace#Boundary Proxy2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation본 논문은 기존 coupled diffusion models가 unified I2I translation 과제에서 겪는 성능 한계를 해결하고자 합니다.#Review#Diffusion Models#Image-to-Image Translation#Domain Harmonization#Data Efficiency#Residual Learning#Manifold Lifting2026년 6월 2일댓글 수 로딩 중
[논문리뷰] Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging본 연구는 대규모 instruction tuning에서 발생하는 Gradient Interference와 시스템 통신 병목이라는 두 가지 핵심 문제를 동시에 해결하고자 한다.#Review#Instruction Tuning#Model Merging#Decentralized Optimization#Gradient Interference#Vision-Language Models#PCA Decomposition2026년 6월 2일댓글 수 로딩 중