[논문리뷰] Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning본 논문은 대규모 언어 모델이 단순히 정보를 암기하는 단계를 넘어 고도의 논리적 추론 능력을 갖추기 위한 핵심 동력으로 Zero RL의 확장성을 주목합니다.#Review#Zero RL#Trillion Parameters#Emergent Reasoning#Reinforcement Learning#Scalability#LLM2026년 7월 15일댓글 수 로딩 중
[논문리뷰] Registers Matter for Pixel-Space Diffusion Transformers본 논문은 Register Tokens가 기존 ViTs에서의 고질적인 문제인 '고 norm 아웃라이어(high-norm patch-token outliers)'를 해결하는 것과 달리, DiTs에서의 구체적인 역할과 효과는 미비하게 탐구되었다는 점에 주목합니다.#Review#Diffusion Transformers#Register Tokens#Pixel-Space#Feature Norms#Attention Sinks#Dual-Stream Architecture2026년 7월 15일댓글 수 로딩 중
[논문리뷰] PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails본 논문은 기존의 이미지 가드레일이 고정된 안전 정책하에서만 작동하며, 실제 산업 현장에서 요구되는 정책적 유연성을 결여하고 있다는 문제를 해결하고자 합니다.#Review#Image Guardrail#Policy-Adaptive#PolicyShiftBench#PolicyShiftGuard#Boundary-Pair Policy Adaptation#Multimodal Safety2026년 7월 15일댓글 수 로딩 중
[논문리뷰] PalmClaw: A Native On-Device Agent Framework for Mobile Phones본 논문은 기존 모바일 에이전트가 주로 의존하는 GUI 기반 조작의 한계를 극복하고, 모바일 기기 환경에서 더 효율적이고 제어 가능한 에이전트 프레임워크를 구축하는 것을 목표로 한다.#Review#Mobile Agent#On-Device#LLM Agent#Device Tools#Execution Boundary#Agent Framework2026년 7월 15일댓글 수 로딩 중
[논문리뷰] OvisOCR2 Technical Report본 논문은 기존의 문서 파싱 방식인 파이프라인(Pipeline) 모델의 복잡한 배포 구조와 단계별 오류 누적 문제를 해결하고자 OvisOCR2를 제안한다. 기존의 파이프라인 방식은 레이아웃 분석, 콘텐츠 인식, 페이지 병합 등 여러 단계가 분리되어 있어 효율성이 낮고, 한 단계의 오류가 후속 단계로 전파되는 한계가 있다.#Review#End-to-End Document Parsing#Markdown Serialization#Multimodal Large Language Model#Reinforcement Learning#On-policy Distillation#OvisOCR22026년 7월 15일댓글 수 로딩 중
[논문리뷰] MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors본 논문은 기존 NVS 방법론들이 겪고 있는 구조적 불일치와 스케일 표류 문제를 해결하고자 합니다. 기존의 명시적 재구성 기반 방식은 국소적인 일관성은 보장하지만, 복잡한 재구성 파이프라인으로 인해 대규모 시점 변화 시 일반화 성능이 제한됩니다 .#Review#Monocular Novel View Synthesis#Diffusion Models#Implicit Geometry Priors#Scale-Awareness#Camera Control#MM-DiT2026년 7월 15일댓글 수 로딩 중
[논문리뷰] KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill본 논문은 기존의 OpenClaw 계열 에이전트가 GUI 환경에서의 복잡한 작업 자동화 시 겪는 구조적 한계를 해결하고자 합니다. 기존 방식은 플랫폼 간의 호환성이 부족하고, 지속적인 학습을 통한 성능 향상 메커니즘이 부재하여 다양한 기기 환경에 적응하기 어렵다는 문제점이 있습니다.#Review#GUI Agents#Personal Assistant#Self-Evolving Memory#Skill Library#Cross-Platform Interaction#POMDP#Task Decomposition2026년 7월 15일댓글 수 로딩 중
[논문리뷰] Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable본 논문은 대규모 Agent Harness의 구조적 복잡성으로 인해 발생하는 Behavior Localization의 어려움을 해결하는 것을 목표로 합니다.#Review#Agent Harness#Behavior Localization#Static Program Analysis#LLM-assisted Behavioral Structuring#Behavior-Guided Progressive Disclosure#Software Engineering2026년 7월 15일댓글 수 로딩 중
[논문리뷰] Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation본 논문은 3D 및 4D 콘텐츠 생성 시 발생하는 공간적·시간적 불일치(hallucination) 문제를 해결하는 것을 목적으로 합니다.#Review#3D Generation#4D Generation#Spatio-temporal Consistency#Multi-Modal Reasoning#Diffusion Models#Hallucination Mitigation2026년 7월 15일댓글 수 로딩 중
[논문리뷰] GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch본 논문은 기존 WAM 방식이 추론 시 명시적인 미래 비디오 생성을 요구하여 발생하는 높은 연산 오버헤드와 실시간 제어의 한계를 해결하는 것을 목표로 합니다.#Review#World Action Models#Robot Control#Mixture-of-Transformers#AutoResearch#Inference Latency#Flow Matching#Visual Dynamics2026년 7월 15일댓글 수 로딩 중
[논문리뷰] From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization본 논문은 장기적(Long-horizon) 에이전트 최적화 시 발생하는 컨텍스트 노이즈 문제를 해결하고자 합니다.#Review#Agent Optimization#Causal Localization#Execution Dependency Graph#Failure Pattern Mining#Structural Trajectory Analysis#Context-Noise Trade-off2026년 7월 15일댓글 수 로딩 중
[논문리뷰] From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World본 논문은 기존의 사이버 보안 벤치마크가 지나치게 제한된 환경(예: Capture-the-Flag)에 국한되어 있어, 실제 환경에서의 복잡한 공격 표면과 전략적 탐색 능력을 평가하지 못하는 한계를 해결하고자 한다 .#Review#AI Pentesting Agents#Vulnerability Discovery#Evaluation Protocol#Ground-Truth Matching#Stochasticity#Agentic Workflow2026년 7월 15일댓글 수 로딩 중
[논문리뷰] Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation본 연구는 기존 오픈소스 생성 모델이 상업적 frontier 모델 대비 복잡한 의도를 해석하는 Understanding 능력이 부족하다는 점을 해결하고자 합니다.#Review#Unified Multimodal#Text-to-Image#Agentic Inference#Data Curation#Diffusion Transformer#Instruction-Driven Generation2026년 7월 15일댓글 수 로딩 중
[논문리뷰] AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities본 논문은 LLM 기반 Agent의 성능을 평가하기 위한 인프라가 극도로 파편화되고 복잡하게 얽혀 있는 문제를 해결하고자 한다. 기존의 평가 방식은 특정 도메인에 고착화되어 있거나, 실행 환경과 평가 프로토콜이 강하게 결합되어 있어 재현성(Reproducibility)을 저해하고 반복적인 엔지니어링 비용을 발생시킨다 .#Review#LLM-based Agents#Evaluation Infrastructure#Benchmarking#Trajectory Analysis#Agentic Capabilities#Reproducibility2026년 7월 15일댓글 수 로딩 중
[논문리뷰] AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow본 논문은 in-the-wild 환경의 감정 분석에서 발생하는 데이터의 내재적 모호성과 표현의 불확실성을 해결하기 위해 AffectFlow-DINO를 제안합니다.#Review#Affective Computing#Conditional Rectified Flow#Multi-Task Learning#Uncertainty-Aware#DINOv3#Facial Affect Estimation#ABAW Challenge2026년 7월 15일댓글 수 로딩 중
[논문리뷰] Towards Autonomous and Auditable Medical Imaging Model Development본 논문은 의료 영상 모델 개발의 자동화 과정에서 발생하는 복잡성과 불투명성 문제를 해결하고자 합니다.#Review#Medical Imaging#Autonomous Agents#Machine Learning Engineering#Model Development#Verification-Guided Optimization#Auditability2026년 7월 14일댓글 수 로딩 중
[논문리뷰] Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation본 논문은 텍스트-이미지 생성(T2I) 모델의 강화학습(RL) 과정에서 효율적이고 신뢰성 높은 보상 모델을 설계하는 것이 어렵다는 점을 해결하고자 합니다 .#Review#SpectraReward#Self-SpectraReward#Text-to-Image Generation#Reinforcement Learning#MLLM#Prompt-Likelihood Reward#Unified Multimodal Models2026년 7월 14일댓글 수 로딩 중
[논문리뷰] Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms본 논문은 딥 강화학습 분야에서 고착화된 평가 패러다임과 그에 내재된 잘못된 가정들을 비판적으로 분석합니다.#Review#Deep Reinforcement Learning#Scaling Laws#Sample Complexity#Evaluation Paradigm#Monotonicity Assumption#Arcade Learning Environment2026년 7월 14일댓글 수 로딩 중
[논문리뷰] MuScriptor: An Open Model for Multi-Instrument Music Transcription기존의 AMT 연구들은 주로 단일 악기(피아노, 기타 등)에 국한되어 있으며, 다중 악기(Multi-instrument) 환경에서의 실질적인 성능은 매우 저조합니다.#Review#Automatic Music Transcription#Multi-Instrument#Transformer#Synthetic Pre-training#Reinforcement Learning#Open-Weight Model2026년 7월 14일댓글 수 로딩 중
[논문리뷰] Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution본 논문은 LLM 기반 coding agent가 repository에 대한 깊이 있는 이해 부족으로 인해 factual errors를 범하고, 결과적으로 복잡한 이슈 해결에 실패하는 문제를 해결하고자 합니다 .#Review#Software Engineering Agents#Knowledge Acquisition#Repository Understanding#Question-Answering (QA)#Automated Issue Resolution#LLM-based Agents2026년 7월 14일댓글 수 로딩 중