[논문리뷰] The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements본 논문은 sparse signal recovery 문제에서 측정 행렬(measurement matrix)의 sparsity가 recovery의 샘플 복잡도(sample complexity)에 미치는 영향을 규명하고자 한다.#Review#Sparse Recovery#Sparse Measurements#Active Sparsification#Information-Theoretic Threshold#Sample Complexity#Maximum-Likelihood Estimator#Chernoff Bound#Phase Transition2026년 9월 9일댓글 수 로딩 중
[논문리뷰] SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators본 논문은 다양한 로봇 환경에서 범용적으로 작동 가능한 Action-conditioned world model의 부재와 환경 변화에 따른 모델의 취약성을 해결하고자 합니다.#Review#World Model#Robotic Manipulation#Visual Calibration#Zero-Shot Simulation#In-Context Adaptation#Action-Conditioned Generation2026년 9월 9일댓글 수 로딩 중
[논문리뷰] StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean본 논문은 기존의 수학 벤치마크들이 경쟁 수학(Competition Math) 문제에만 편중되어 있어, 실제 응용 수학 분야인 확률 과정 문제를 충분히 다루지 못한다는 한계를 해결하고자 합니다.#Review#Stochastic Processes#Lean 4#Automated Theorem Proving#Benchmark#Formalization#Markov Chains#Stochastic Calculus2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Show-Harness: Just a VLM Agent Can Play Robots본 논문은 Foundation VLM의 강력한 지능을 로봇 제어로 전환할 때 발생하는 불투명성과 비효율성 문제를 해결하고자 합니다.#Review#Embodied AI#Vision-Language Models#Robot Manipulation#Semantic Action Interface#Zero-shot Control#Human-Agent Collaboration2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents본 논문은 현대 AI 연구 에이전트가 보고하는 높은 성능 점수가 실제 과학적 발견(Discovery)의 타당성을 보장하지 못한다는 비판적 문제의식에서 출발합니다.#Review#AI Research Agents#Discovery Certification#Benchmarking#Reproducibility#Scientific Discovery#Evaluation Protocol2026년 9월 9일댓글 수 로딩 중
[논문리뷰] SchemeArena: Factorized Stress Testing of Scheming in LLM Agents본 논문은 LLM 에이전트의 위험한 실패 모드인 Scheming이 어떤 조건에서 발생하는지 체계적으로 분석하고자 한다. 기존 연구들은 몇 가지 한정된 시나리오에 의존하여 scheming에 관여하는 개별 요인들의 독립적인 영향력을 입증하는 데 한계가 있었다 .#Review#LLM Agents#Scheming#Safety#Benchmark#Oversight#Instrumental Goals#Scalable Monitoring2026년 9월 9일댓글 수 로딩 중
[논문리뷰] SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents본 논문은 기존의 소프트웨어 엔지니어링 에이전트 평가 표준인 SWE-Bench Pro가 가진 심각한 신뢰성 결여 문제를 해결하고자 한다.#Review#SWE-Bench#LLM Agents#Software Engineering#Reward Hacking#Benchmarking#Task Quality#Evaluation Validity2026년 9월 9일댓글 수 로딩 중
[논문리뷰] SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?본 논문은 Recursive Self-Improvement (RSI) 시스템의 신뢰성과 제어 가능성을 확보하기 위해 필수적인 내부 표현 모니터링 및 감사(auditing)의 부재 문제를 해결하고자 합니다.#Review#AI Agents#Mechanistic Interpretability#Sparse Autoencoders (SAEs)#Recursive Self-Improvement (RSI)#Autonomous Discovery#Contrastive Probes#Causal Steering#Benchmark2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Revisiting Complete Reasoning Traces for Post-Training본 연구는 LLM의 post-training 과정에서 필수적인 것으로 간주되던 긴 reasoning trajectory가 실제로는 많은 중복 정보를 포함하고 있어 효율성이 떨어진다는 문제점을 해결하고자 합니다.#Review#Reasoning Trajectories#Post-Training#Supervised Fine-Tuning (SFT)#Reasoning Redundancy#Endpoint-based SFT (E-SFT)#Chain-of-Thought (CoT)2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States본 연구는 fine-tuning 과정에서 발생하는 LLM의 내재적 편향 변화를 효과적으로 감지하기 위한 reference-based auditing 프레임워크를 제안합니다.#Review#Large Language Models#Bias Detection#Relative Representations#Hidden States#Fine-tuning#Model Auditing#Representational Bias Shift2026년 9월 9일댓글 수 로딩 중
[논문리뷰] RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems본 연구는 기존 LLM 기반의 Emotional Support Conversation (ESC) 시스템들이 주로 일대일 대화에만 치중하고 있어, 실제 다자간 관계에서 발생하는 복잡한 정서적 역학을 반영하지 못한다는 한계를 해결하고자 합니다.#Review#Emotional Support Conversation#Multi-Party Dialogue#Relation-Aware Modeling#RESCUE-Bench#LLM Evaluation#Interpersonal Dynamics2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation본 논문은 기존 speech-driven 제스처 모델들이 공간적 맥락을 고려하지 못하는 한계점을 해결하기 위해 Puppeteer를 제안합니다. 기존 모델들은 주로 audio-gesture 정렬에만 집중하여, 주변 사물(가구 등)과 신체 자세(posture)에 따른 물리적 제약 사항을 반영하지 못했습니다 .#Review#Co-speech Gesture Generation#Diffusion Model#Causal Latent Space#Object-Grounding#Posture-Awareness#SceneGes2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Programmable World Model본 논문은 비디오 월드 모델이 가진 불투명한 세계 상태와 제어의 한계를 해결하기 위해 Programmable World Model을 제안한다. 기존의 월드 모델들은 단순히 시각적 과거를 기반으로 미래를 예측할 뿐, 외부에서 접근하거나 수정 가능한 명시적이고 지속적인 세계 상태를 유지하지 못한다는 문제점이 있다.#Review#Programmable World Model#Generative Renderer#3D Oriented Bounding Boxes#State-Augmented Representation#Control Compilation#Long-Horizon Generation2026년 9월 9일댓글 수 로딩 중
[논문리뷰] PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving본 논문은 자율주행 시스템의 안전성 검증을 위한 Scenario-based Testing 프로세스가 현재 파편화되어 있다는 핵심 문제를 해결하고자 합니다. 기존 방식은 시나리오 생성, 검색, 수정, 실행, 분석 단계가 개별 도구로 분리되어 자동화가 어렵고, 수동으로 처리해야 하는 병목 현상이 존재합니다.#Review#Autonomous Driving Systems#Scenario-based Testing#LLM Agents#Motion Planners#OpenStreetMap#Simulation#Constraint Satisfaction2026년 9월 9일댓글 수 로딩 중
[논문리뷰] OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution본 논문은 재귀적 SR에서 배율이 커짐에 따라 발생하는 Supervision Gap 문제를 해결하고자 합니다.#Review#Recursive Super-Resolution#On-Policy Self-Distillation#Cross-Scale Supervision#Latent Prior#Image Fidelity#CLIPIQA#Hallucination Reduction2026년 9월 9일댓글 수 로딩 중
[논문리뷰] From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data AttributionLarge Language Models (LLMs)의 다양한 Capabilities와 Behaviors는 학습 데이터로부터 형성되지만, 어떤 학습 예제가 이러한 Behaviors를 야기하는지에 대한 이해는 부족한 상태입니다.#Review#Training Data Attribution#Influence Functions#Supervised Fine-tuning#Response Rewriting#Reweighting#Language Model Abstention#Safety Refusal#Behavioral Intervention2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models본 논문은 LLM을 활용한 코드 편집 환경에서 Direct 방식과 Iterative Diff-Based 방식 중 어느 것이 모델 학습 및 추론에 최적인지 검증하고자 한다.#Review#Code Editing#Large Language Models#Iterative Generation#Direct Generation#Flutter/Dart#Task Locality2026년 9월 9일댓글 수 로딩 중
[논문리뷰] Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR본 논문은 RLVR이 단일 샘플 정확도(avg@k)는 향상시키지만, 모델의 근본적인 추론 범위(pass@k)를 충분히 확장하지 못하는 문제를 해결하고자 합니다. 기존 연구들은 최적화 과정에 집중할 뿐, 학습 시의 롤아웃(rollout) 구조를 고정된 병렬 방식으로 유지하여 탐색이 제한적이라는 한계가 있습니다.#Review#Reinforcement Learning#Large Reasoning Models#RLVR#Tree-Structured Policy Optimization#Difficulty-Adaptive#Reasoning Coverage2026년 9월 9일댓글 수 로딩 중
[논문리뷰] DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents본 연구는 파편화된 화학 문헌 데이터와 비표준화된 데이터 형식으로 인해 발생하는 화학 AI 모델 학습의 한계를 극복하고자 합니다. 기존의 데이터셋은 수동 수집의 비용 문제와 데이터 노이즈, 그리고 구조적 정보의 부재로 인해 최신 Generative AI 및 AI Agent 모델의 성능을 온전히 이끌어내기에 부족했습니다.#Review#Organic Reaction#Data Platform#Automated Pipeline#Fine-grained Data#AI Agents#Chemical Intelligence2026년 9월 9일댓글 수 로딩 중
[논문리뷰] DF26: We Cannot Tell Fake From Real Anymore본 논문은 최신 비디오 생성 모델의 발전으로 인해 Deepfake 탐지가 더욱 어려워지고 있음을 지적하며, 기존 벤치마크의 한계를 넘어선 새로운 평가의 필요성을 제기합니다.#Review#Deepfake Detection#AI-generated Video#Text-to-Video#Image-to-Video#Public Speaking Scenarios#Benchmark Dataset#Cross-generator Evaluation#Human Perception2026년 9월 9일댓글 수 로딩 중