[논문리뷰] Realtime-Venus: A full-duplex interaction system with asynchronous delegation본 논문은 실시간 상호작용 환경에서 연속적인 인식(perception)과 즉각적인 응답 사이의 균형을 맞추는 동시에, 복잡한 외부 도구 사용과 대화의 일관성을 유지하는 문제를 해결하고자 합니다.#Review#Full-Duplex#Asynchronous Delegation#Omni-modal Interaction#Streaming Decoding#Conversational Frontend#Speech Generation2026년 9월 21일댓글 수 로딩 중
[논문리뷰] RRSI: Regularized Recursive Self-Improvement of Agent HarnessesLLM 에이전트의 역량은 고정된 백본 모델 주변의 프롬프트, 제어 흐름, 툴 인터페이스, 메모리 및 컨텍스트 관리 등을 포함하는 Harness에 의해 크게 좌우됩니다.#Review#LLM Agents#Harness Evolution#Recursive Self-Improvement (RSI)#Regularization#Overfitting#Generalization#Policy Tokens2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation본 연구는 AI 모델의 성능을 도메인별로 세분화하여 평가할 때 발생하는 높은 비용과 표본 부족 문제를 해결하고자 합니다. 기존의 Direct Estimators(예: HT estimator)는 특정 도메인의 표본 수가 적을 경우 추정치의 분산이 매우 커져 신뢰도가 낮아지는 한계가 있습니다.#Review#Disaggregated Evaluation#Prediction-Powered Inference#Small Area Estimation#Fay-Herriot Model#Cross-Validation#Survey Sampling2026년 9월 21일댓글 수 로딩 중
[논문리뷰] One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents본 논문은 repository-level Software Engineering (SWE) task의 heterogeneous한 특성으로 인해 발생하는 category see-saw 문제를 해결하고자 합니다 .#Review#Software Engineering Agents#Reinforcement Learning#Category-Aware Training#Expert Training#Policy Integration#Multi-Teacher On-Policy Distillation (MOPD)#Refresh-Repair-Expand (RRE)#SWE Labeler2026년 9월 21일댓글 수 로딩 중
[논문리뷰] OmniEdu: Open Foundation Models for Learning and Teaching본 논문은 기존 교육용 언어 모델들이 단순히 정답을 제공하는 것을 넘어, '무엇을 가르칠지', '학습자가 어느 수준인지', '어떻게 반응할지'를 조율하는 총체적인 학습-교수 루프(learning-teaching loop)를 다루지 못하는 한계를 지적한다.#Review#Foundation Models#Educational AI#Instruction Tuning#K-12 Education#Curriculum Grounding#Pedagogical Tutoring#Diagnostic Reasoning#Subject Competence2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene본 논문은 단일 이미지로부터 Compositional 3D Scene을 재구성하는 과정에서 Object Layout을 정확하게 배치하는 문제를 해결하고자 합니다.#Review#Compositional 3D Scene Reconstruction#Pixel-Aligned Layouts#Canonical Coordinate Map (CCM)#Point Cloud Map (PCM)#Multimodal Diffusion Transformer#Geometry-Layout Co-Generation#Robust Geometric Alignment2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles본 논문은 LLM이 생성한 GPU 커널의 정확성 검증을 담당하는 벤치마크 오라클의 약점을 측정하고 개선하기 위한 Mutation Analysis 방법론을 제안합니다.#Review#Mutation Analysis#GPU-Kernel#Benchmark Oracle#Floating-Point Precision#Test Suite Generation#KernelBench-M#Tolerance Vacuity#Validity Ceiling2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents본 논문은 기존 Agentic Memory 시스템이 Autoregressive LLM에 과도하게 의존하여 발생하는 비효율성과 높은 비용 문제를 해결하고자 한다.#Review#Agentic Memory#System-One/System-Two#LLM Agents#Memory Control#Efficiency#Multi-Relational Memory#Adaptive Retrieval2026년 9월 21일댓글 수 로딩 중
[논문리뷰] HuRo: Robotizing Human Videos for Scalable VLA Pretraining본 논문은 Vision-Language-Action (VLA) policies가 pretraining data의 증가로부터 큰 이점을 얻지만, 대규모의 다양한 실제 로봇 상호작용 데이터 수집이 매우 비싸다는 문제에 직면해 있다고 지적한다.#Review#Learning from Human Videos#Robotization#VLA Pretraining#Embodiment Gap#Scalable Data2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Harness-Zero: Harness Distillation via Agent-as-HarnessLLM Agent의 성능은 기반 모델뿐만 아니라 외부 시스템인 Harness의 설계에도 크게 의존하지만, 기존의 Harness 최적화 기법은 그 이점이 배포 시점에 해당 Harness에 종속된다는 핵심적인 한계를 가집니다.#Review#LLM agents#agent harness#harness evolution#harness distillation#agent-as-harness#fine-tuning2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Grounded Action Model: 3D Grounding as a Foundation for Robotics본 논문은 기존 로봇 Foundation Model들이 metric grounding 능력이 부족하여 manipulation generalization에 한계를 가짐을 지적한다.#Review#Robotics#3D Grounding#Action Models#Manipulation#Foundation Models#Generalization#Multi-stream Transformer#Object-Centric Representation2026년 9월 21일댓글 수 로딩 중
[논문리뷰] GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay본 논문은 현대 비디오 게임 환경에서 AI 모델의 복합적인 능력(시각적 이해, 명령어 분해, 목표 계획, 정밀한 액션 제어)을 평가하는 데 있어 기존 데이터셋과 벤치마크의 한계를 해결하고자 한다.#Review#Gameplay AI#Multi-Horizon Instructions#AAA Games#Large-Scale Dataset#Benchmarking#Vision-Language Models#Action Planning#Goal Decomposition2026년 9월 21일댓글 수 로딩 중
[논문리뷰] EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation본 논문은 엔터프라이즈 애플리케이션에서 Tool-Calling LLM Agents의 효과적인 평가 및 최적화를 위한 고품질의 다양하고 현실적인 태스크 데이터셋 확보의 어려움을 핵심 문제로 제기합니다.#Review#Tool-Calling Agents#Edge Cases#Synthetic Data Generation#Finetuning#Harness Optimization#Policy Compliance#Database Grounding#LLM Agents2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown ConversionUday Allu이 arXiv에 게시한 'Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion' 논문에 대한 자세한 리뷰입니다.#Review2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations본 논문은 Large Language Models (LLMs)를 활용한 기존의 페르소나 시뮬레이션 방식이 피상적인 캐릭터 설명에 의존하여 장기적인 상호작용에서 일관된 캐릭터 행동을 유지하지 못하는 핵심 문제를 다룬다.#Review#Role-Playing Agents#Persona Simulation#Large Language Models (LLMs)#Psychologically Grounded Architecture#Evaluation Framework#Clinical Simulations#Dialogue Naturalness#Adversarial Stress Test2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta AttentionLinear RNN은 시퀀스 길이에 선형적으로 확장되는 효율적인 시퀀스 모델링이 가능하지만, 낮은 Rank correction을 포함하는 선형 업데이트로 인해 Expressivity가 제한된다.#Review#Complex KDA#Kimi Delta Attention#Linear RNNs#Expressivity#State Tracking#Rotation#Diagonal-plus-rank-one#Language Modeling2026년 9월 21일댓글 수 로딩 중
[논문리뷰] CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies본 논문은 Vision-Language-Action (VLA) Policies가 로봇 조작에서 강력한 성능을 보임에도 불구하고, 실행이 nominal trajectory에서 벗어날 경우 취약하다는 핵심 문제를 제기합니다.#Review#Robotic Manipulation#Failure Recovery#Vision-Language-Action Policies#Self-Correction#Corrective Data Synthesis#3D Geometric Monitoring#Multi-arm Robotics#FSR-Bench2026년 9월 21일댓글 수 로딩 중
[논문리뷰] A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal본 논문은 LLM이 내부적으로 알고 있는 지식을 외부에 드러내지 않거나, 의도적으로 잘못된 답변을 생성하는 문제를 해결하고자 한다.#Review#Concealed Information Test#Sandbagging#Unlearning Verification#Probe of Internal Recognition (PIR)#Internal Recognition#Activation Probing#Deception Detection2026년 9월 21일댓글 수 로딩 중
[논문리뷰] 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation본 논문은 Sparse OPD에서 발생하는 기울기 추정의 신뢰성 문제를 해결하고자 합니다. 기존의 On-Policy Distillation(OPD) 방법론들은 학생 모델이 생성한 궤적의 모든 토큰에 대해 교사 감독을 적용하거나, 토큰의 '유용성(usefulness)'에 기반하여 중요한 토큰을 선별했습니다.#Review#On-Policy Distillation#Gradient Estimation#Information Geometry#Signal-to-Noise Ratio#Token Selection#Large Language Models#Sparse Supervision2026년 9월 21일댓글 수 로딩 중
[논문리뷰] When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation본 논문은 LLM이 과학적 피어 리뷰에 도입되면서 발생하는 recursive training 기반의 판단 편향 문제를 해결하고자 합니다. 기존 연구들은 생성 데이터에 의존한 반복 학습이 데이터 분포를 왜곡하고 저확률 영역을 소멸시키는 모델 붕괴(model collapse)를 유발함을 지적해 왔습니다.#Review#Large Language Models#Scientific Peer Review#Recursive Training#Model Collapse#Activation Steering#Scientific-Judgment Collapse2026년 9월 20일댓글 수 로딩 중