[논문리뷰] RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance본 논문은 기존 로봇 학습에서 보상 체계(Reward Supervision)의 불투명성과 낮은 일반화 성능 문제를 해결하기 위해 Temporal Distance 기반의 새로운 가치 평가 프레임워크를 제안합니다 .#Review#Robotic Foundation Model#Temporal Distance#Value Function#Reward Interface#Embodied AI#Reinforcement Learning#Shortcut Suppression2026년 8월 10일댓글 수 로딩 중
[논문리뷰] RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States본 논문은 self-evolving LLM 에이전트의 메모리 시스템이 경험 누적에 따른 utility 상태 공간의 무한 확장과 이로 인한 MRT 문제라는 이중고에 직면해 있음을 지적합니다.#Review#LLM Agents#Memory Reinforcement Learning#Memory-Reward Trap#Reduced-Order Utility States#Feedback Density#Self-Evolving Memory2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution본 논문은 기존 Coding Agent들이 고정된 Harness 하에서만 동작하여 성능 향상에 한계가 있다는 문제점을 해결하고자 한다.#Review#Self-evolving Agents#Coding Agent#Agentic Workflow#Frontier Model#Operational Safety#Recursive Evolution#Version Control2026년 8월 10일댓글 수 로딩 중
[논문리뷰] OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching본 논문은 대규모 언어 모델(LLM)의 long-context 및 agentic workloads가 증가함에 따라 발생하는 HBM의 용량 병목 현상을 해결하고자 한다 .#Review#LLM Inference#KV Cache#Sparse Prefetching#Memory Wall#Speculative Decoding#Lookahead Attention#PD Disaggregation2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Motif 3: Technical Report본 논문은 대규모 모델의 파라미터 확장과 효율적인 컴퓨팅 자원 활용 사이의 간극을 해결하기 위해 Motif 3를 제안합니다. 기존의 일반적인 LLM들은 모델 크기가 커짐에 따라 추론 비용과 메모리 소모가 비효율적으로 증가하며, 특히 학습 과정에서의 expert 쏠림 현상이나 고차원 정보 처리의 안정성 문제가 발생합니다.#Review#Large Language Model#Mixture-of-Experts#Grouped Differential Latent Attention#Multi-token Prediction#Inference Efficiency#Training Stability2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA본 연구는 대규모 언어 모델이 고정된 환경에서 사전 학습된 후 정체되는 한계를 극복하고, 실제 환경에서 경험을 통해 지속적으로 발전하는 Experiential Intelligence 구현을 목표로 합니다.#Review#Experiential Intelligence#Mixture-of-LoRA#Continual Learning#Recursive Self-Improvement#Agent Model2026년 8월 10일댓글 수 로딩 중
[논문리뷰] MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models본 논문은 MLLM이 불완전한 시각적 맥락(Imperfect Context) 하에서 거부와 답변 능력을 적절히 균형 잡지 못하는 신뢰성 문제를 해결하고자 합니다.#Review#Multimodal Large Language Models#Out-of-Context#Benchmark#Refusal#Robust Answering#Shifted In-Context#Model Reliability2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation기존의 LLM 기반 사용자 시뮬레이터는 대화 문맥과 프로필을 기반으로 다음 사용자 턴을 단순히 모방(Imitation)하는 방식에 의존합니다 .#Review#User Simulation#LLM#Reinforcement Learning#Controllable Generation#Interaction Intent#Directive-Conditioned Generation2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval본 논문은 많은 NLP 작업에서 발생하는 입력 데이터와 대상 개념 간의 모호성 문제를 해결하고자 합니다.#Review#Evidence-to-Taxonomy Retrieval#Factorized Hypothesis Search#Retrieval Readiness Gap#Semantic Dimension#Candidate-Level Verifier#Financial Tagging#Clinical Coding2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Evo-Bench: Can Language Models Improve Agent Harness?본 논문은 LLM 기반 에이전트의 자율적 하네스 개선(harness evolution) 역량을 체계적으로 평가하기 위한 새로운 벤치마크인 Evo-Bench를 제안합니다.#Review#Agent Harness#Harness Evolution#LLM#Benchmark#Self-Improvement#Autonomous Agent#Task Sensitivity2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Evidence-RL: Towards Evidence-intensive Visual Reasoning본 논문은 Vision-Language Models (VLMs)가 시각적 증거에 기반하지 않고 언어적 편향(language priors)이나 학습 데이터의 지름길(shortcuts)에 의존하여 답변하는 문제를 해결하고자 한다.#Review#Vision-Language Models#Reinforcement Learning#Counterfactual Evidence Disentanglement#Visual Grounding#Causal Reasoning#Post-training2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Ego-OSCAR: Egocentric Open source Stereo CAptuRe System본 논문은 VLA 모델 및 로봇 정책 학습을 위해 필요한 대규모 egocentric 데이터 수집의 높은 진입장벽을 해결하고자 합니다 .#Review#Egocentric Data Collection#Open-source Hardware#Stereo-inertial Capture#Vision-Language-Action Models#Robotic Pretraining#Hardware-synchronized2026년 8월 10일댓글 수 로딩 중
[논문리뷰] CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems본 논문은 기존 인지 모델의 추상적인 이론적 기초와 상용 게임 엔진의 실시간 성능 중심 구현 사이의 간극을 해소하기 위한 프레임워크를 제안한다.#Review#Intelligent Virtual Agents#Cognitive Architecture#BDI Model#Embodied Cognition#Interactive Computing Systems#Virtual Reality#Multi-agent Systems2026년 8월 10일댓글 수 로딩 중
[논문리뷰] BDH-CQ: In-Context Learning with Recurrent Latent Reasoning본 논문은 In-context learning과 latent reasoning을 결합하여 추론의 효율성과 성능을 극대화하는 것을 목표로 합니다.#Review#In-Context Learning#Recurrent Latent Reasoning#ARC-AGI#Cost Efficiency#Latent Workspace#Generalization#Neural Architecture2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory본 논문은 기존의 메모리 기반 에이전트 시스템이 대규모 모델에서는 성공적이었으나, 상대적으로 추론 및 지시 이행 능력이 낮은 소규모 모델에서는 성능 개선이 미미하다는 한계를 해결하고자 합니다.#Review#Agent Memory Distillation#Small LLM Agents#Hierarchical Memory#Tool-use#Knowledge Distillation#Proactive/Reactive Injection2026년 8월 10일댓글 수 로딩 중
[논문리뷰] A^2E : An End-to-End Agent Auditing Engine본 논문은 기존의 Agent 평가 방식이 최종 성공 여부(Correctness)에만 과도하게 의존하여, 실제 시스템의 성능을 결정짓는 Agent Harness의 고유한 특성을 충분히 포착하지 못한다는 문제를 해결하고자 한다.#Review#Agent Evaluation#LLM Agent#Harness-level Auditing#Agent Task Protocol (ATP)#Observability#Lifecycle-Aligned Evaluation2026년 8월 10일댓글 수 로딩 중
[논문리뷰] Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination본 논문은 기존의 LLM 벤치마크 오염(Contamination) 평가 지표인 G-AP가 모델의 진정한 복원 능력을 왜곡하여 평가한다는 문제를 제기합니다.#Review#Data Contamination#LLM Evaluation#Benchmark Restoration#Mitigation Strategy#Probability Gap#SA-PPG#RailCap2026년 8월 9일댓글 수 로딩 중
[논문리뷰] YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family본 논문은 YOLO 시리즈와 같은 대규모 객체 탐지 모델을 새로운 타겟 도메인에 적용할 때 발생하는 높은 컴퓨팅 자원 및 저장 공간 요구 문제를 해결하고자 한다. 기존의 Full Fine-Tuning 방식은 모델 전체의 가중치를 수정해야 하므로, 배포 시 각 태스크마다 전체 모델을 저장해야 하는 비효율성이 존재한다.#Review#YOLO#Parameter-Efficient Fine-Tuning#PEFT#Transfer Learning#Object Detection#Model Adaptation2026년 8월 9일댓글 수 로딩 중
[논문리뷰] When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents본 논문은 멀티턴 에이전트 학습에서 Privileged On-Policy Distillation이 직면한 상태-참조 불일치 문제를 해결하고자 합니다.#Review#On-Policy Distillation#Multi-Turn Agents#State-Reference Mismatch#Self-Distillation#State-Matched Routing#Contextualized Guidance#Embodied AI2026년 8월 9일댓글 수 로딩 중
[논문리뷰] When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles본 논문은 learned interpretability interfaces, 특히 Activation Oracles (AOs)의 신뢰성 문제를 다룹니다.#Review#Activation Oracles#Interpretability#Large Language Models#Concept-Specific Blind Spots#Anti-reading#Taboo Word Guessing#Readout Suppression2026년 8월 9일댓글 수 로딩 중