[논문리뷰] EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses본 논문은 현대의 LLM Agent가 자신의 프롬프트, 도구, 미들웨어 등을 스스로 변경(Self-Evolution)할 때 발생하는 비가역적 상태 변화 문제를 해결하고자 합니다.#Review#LLM Agent#Self-Evolution#Recoverability#Counterfactual Verification#Typed Observational Equivalence#Execution Harnesses2026년 8월 30일댓글 수 로딩 중
[논문리뷰] DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents본 논문은 대규모 언어 모델 기반의 다중 턴 툴 호출 에이전트 학습에서 발생하는 Topological Collapse 문제를 해결하고자 합니다.#Review#Multi-turn Tool-Calling#Self-Distillation#Interaction-State Transition Graph (ISTG)#Critical Topological Breakpoint (CTB)#Localized Supervision#Agent Training#Policy Diversity2026년 8월 30일댓글 수 로딩 중
[논문리뷰] ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL본 논문은 long-horizon agentic task에서 발생하는 과도한 working context 문제를 해결하기 위한 ContextPilot을 제안한다.#Review#Long-horizon Agents#Proactive Context Management#Reinforcement Learning#Credit Assignment#Partial Rollout#Tool-use2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning본 논문은 현대의 Vision-Language Models(VLM)가 물리적 현상을 설명할 수는 있으나, 이를 유발하는 기저의 물리적 메커니즘을 명시적으로 이해하지 못한다는 문제를 해결하고자 합니다 .#Review#Executable World Representations#Physical Reasoning#Agentic Discovery#Vision-Language Models#Physical Intelligence#Sim-to-Real2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge본 연구는 기존 LLM 평가 방법론이 단일 canonical answer만을 정답으로 간주하여, 사실 관계의 복잡성이나 이견(divergence)을 무시하고 있다는 문제에서 출발한다.#Review#LLM#Knowledge Tail#Epistemic Myopia#Benchmark#Parametric Memory#Multi-account QA2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models본 논문은 대규모 로봇 데이터 확보의 물리적 한계로 인해, 기존 VLA 모델들이 단순히 동작 데이터를 모사(Fitting)하는 데 그쳐 범용성을 확보하지 못하는 문제를 해결하고자 합니다.#Review#Vision-Language-Action Models#Continued Pre-training#Representation Learning#Embodied AI#Robotics#Foundation Models2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities본 논문은 기존의 Direct Generation이 복잡한 의존성을 가진 최종 결과물을 생성하는 데 한계가 있음을 지적하며, 이를 극복하기 위한 Agentic Artifact Creation의 개념을 제안합니다 .#Review#Agentic Artifact Creation#Stateful Construction#Generative Foundation Models#Operational Representation#Construction Policy#Runtime Verification#Iterative Refinement2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models본 논문은 Vision-Language-Action (VLA) 모델의 액션 디코더가 단순히 행동을 모방하는 것을 넘어, 해당 행동이 달성하고자 하는 behavior-level intent를 명시적으로 모델링해야 한다고 주장합니다.#Review#Vision-Language-Action Models#Intention Distillation#Behavioral Intent#Robot Manipulation#Semantic Supervision#Flow Matching#Policy Learning#Multimodal AI2026년 8월 30일댓글 수 로딩 중
[논문리뷰] Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization본 논문은 범용적인 로봇 조작(robotic manipulation)을 위해 unseen task에 대한 Zero-shot cross-task generalization을 달성하는 것을 목표로 합니다.#Review#Robotic Manipulation#In-Context Learning#World-Action Modeling#Zero-shot Generalization#Human Video Guidance#HumanGen Dataset2026년 8월 27일댓글 수 로딩 중
[논문리뷰] WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution본 논문은 에이전트의 경험이 최적화 히스토리에 파편화되어 축적되는 문제를 해결하고자 한다. 기존 연구들은 기술을 반복적으로 개발하지만, 이를 뒷받침하는 지식의 조직화가 부족하여 학습한 내용을 시스템적으로 재사용하는 데 한계가 있다.#Review#AI Agents#Skill Evolution#Persistent Knowledge#LLM#Agentic Workflow#Knowledge Consolidation2026년 8월 27일댓글 수 로딩 중
[논문리뷰] What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents본 논문은 LLM 에이전트 학습을 위한 데이터 생성 과정이 직면한 일관성 유지 문제와 데이터의 질적 관리의 모호함을 해결하고자 합니다. 기존 연구들은 도메인별로 파편화되어 있어, 에이전트의 상호작용 데이터 생성 메커니즘을 통합적으로 이해하는 데 어려움을 겪고 있습니다 .#Review#LLM Agents#Agentic Data#Data Generation#Accuracy#Complexity#Diversity#ACE Lens2026년 8월 27일댓글 수 로딩 중
[논문리뷰] What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals본 논문은 기존의 평가 아티팩트(Evaluation Artifact)들이 단순히 수치적인 결과값만을 보고할 뿐, 해당 결과가 어떤 데이터와 해석을 바탕으로 도출되었는지에 대한 '재현 가능성(Claim-replay)'을 보장하지 못한다는 문제를 제기합니다 .#Review#Inspect Evals#Claim-Relative Inference#Evidence Gate#Semantic Multiplicity#Audit Contract#Reproducibility#Finite Frame2026년 8월 27일댓글 수 로딩 중
[논문리뷰] UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City본 논문은 Multimodal Large Language Models (MLLMs) 에이전트가 고립된 장면 인식에는 능숙하지만, 복잡한 실세계 도시 환경에서 이동하며 지속적인 spatial agency를 발휘할 수 있는지에 대해 의문을 제기합니다.#Review#Multimodal Large Language Models#Spatial Agency#UrbanGround#Embodied Navigation#Geospatial Data#Closed-Loop Interaction#3D Sandbox2026년 8월 27일댓글 수 로딩 중
[논문리뷰] Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO본 논문은 LLM reasoning 후학습(post-training)에서 ES가 GRPO와 비교하여 어떠한 최적화 동역학(dynamics)과 장점을 갖는지 체계적으로 분석합니다.#Review#Evolution Strategies#GRPO#LLM Reasoning#Pass@K#Entropy Collapse#Policy Diversity#Catastrophic Forgetting2026년 8월 27일댓글 수 로딩 중
[논문리뷰] Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report본 연구는 라이브 커머스 환경에서 실시간 대응이 필요한 디지털 아바타 에이전트가 고정된 Harness에 과적합되어 Harness 변화 시 성능이 급격히 저하되는 문제를 해결하고자 합니다.#Review#Harness-Aware Training#Digital Avatar#Harness Evolution#Surface-form Overfitting#On-Policy Distillation#Agentic RL2026년 8월 27일댓글 수 로딩 중
[논문리뷰] Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning본 논문은 기존의 영상 편집 모델들이 가진 '무분별한 청킹(blind chunking)' 전략의 한계를 극복하기 위해 MMLVE 과제를 제안한다 .#Review#Multi-Shot Video Editing#Agentic Reasoning#Cross-Shot Editing Consistency#Multi-Instruction Decoupling#Spatiotemporal Structure2026년 8월 27일댓글 수 로딩 중
[논문리뷰] TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback본 논문은 contact-rich manipulation 환경에서 초기 관측치 기반의 VLA 모델이 겪는 실행 시점의 촉각 정보 불일치 문제를 해결하고자 합니다.#Review#Tactile Feedback#Streaming Action Generation#Flow Matching#Vision-Language-Action Models#Robot Manipulation#Execution-Aware Tactile Attention2026년 8월 27일댓글 수 로딩 중
[논문리뷰] TTPO: Test-Time Policy Optimization본 논문은 LLM의 test-time reasoning 능력을 향상시키기 위한 test-time training(TTT) 환경에서 ground-truth label 부재 문제를 해결합니다. 기존의 OPSD나 RLVR 기법은 정답을 필수로 요구하지만, test-time에는 이러한 정보를 얻을 수 없습니다.#Review#Test-Time Policy Optimization#Test-Time Training#On-Policy Self-Distillation#Reinforcement Learning#Chain-of-Thought#Mathematical Reasoning#Asymmetric Objective2026년 8월 27일댓글 수 로딩 중
[논문리뷰] Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher본 논문은 기존의 Flow Matching 모델 정렬(alignment) 방식이 가진 높은 컴퓨팅 비용과 성능 한계를 해결하기 위해 Self-OPD를 제안합니다.#Review#Flow Matching#On-Policy Distillation#Teacher-Free Alignment#Multi-Objective Optimization#Self-Exploration#Reward-Level Fusion2026년 8월 27일댓글 수 로딩 중
[논문리뷰] Procedura: Agentic 3D Modeling with Procedural Control기존의 Native 3D generators는 고품질 메시를 생성할 수 있으나, 정교한 기계적 구조를 표현하는 데 한계가 있으며, 생성된 결과물이 부품 단위로 분해되지 않거나 편집이 불가능하다는 문제점이 있습니다.#Review#3D Modeling#Agentic Framework#Procedural Assembly#CSG#Large Language Models#Typed Mates#Parametric Program2026년 8월 27일댓글 수 로딩 중