[논문리뷰] Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers본 논문은 현대의 agent workflow 프레임워크들이 제공하는 persistence 보장의 불투명성과 그로 인한 신뢰성 문제를 해결하고자 한다.#Review#Conformance Testing#Checkpointing#Crash Recovery#Exactly-Once Semantics#TLA+#Workflow Persistence#Distributed Systems2026년 8월 5일댓글 수 로딩 중
[논문리뷰] Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models본 논문은 서로 다른 아키텍처를 가진 교사 모델들의 보완적인 강점을 단일 소형 모델로 통합하는 Multi-capability Consolidation 문제를 해결한다.#Review#On-Policy Distillation#Flow Matching#Multi-Teacher#Capability-Selectable#Gradient Compatibility#Heterogeneous Models2026년 8월 5일댓글 수 로딩 중
[논문리뷰] OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents본 논문은 Large Language Model (LLM) 기반 에이전트가 long-horizon, cross-environment, 그리고 multimodal 특성을 지닌 일상적인 요청을 처리하는 데 직면하는 한계점을 해결하고자 합니다.#Review#Autonomous Agents#Long-Horizon Tasks#Task Decomposition#Execution Memory#Global Verification and Repair#LLM Harness#AgentIF-OneDay2026년 8월 5일댓글 수 로딩 중
[논문리뷰] OPD-V: Visual On-Policy Self-Distillation with Modality Balance본 논문은 MLLM의 시각적 추론 성능을 향상시키는 On-Policy Self-Distillation (OPSD) 과정에서 발생하는 Modality Imbalance 문제를 해결하고자 합니다.#Review#Multimodal Large Language Models#On-Policy Self-Distillation#Modality Imbalance#Visual Reasoning#Privileged Information#Modality-Balance Trust Region2026년 8월 5일댓글 수 로딩 중
[논문리뷰] NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap본 논문은 LLM이 단순 pattern matching을 넘어 진정한 logical reasoning 능력을 갖추고 있는지에 대한 불확실성과 multilingual evaluation에서 language-specific failure가 단일 performance gap으로 뭉뚱그려지는 문제를 해결하고자 한다.#Review#Puzzle Benchmark#Cross-Lingual Evaluation#Performance Gap#Hangul Jamo#Difficulty Calibration#Procedural Generation#LLM Reasoning#Korean NLP2026년 8월 5일댓글 수 로딩 중
[논문리뷰] Lossless Tensor Compression as Program Synthesis본 논문은 AI 모델 Checkpoint의 급격한 증가하는 규모와 개수로 인해 발생하는 아카이빙, 전송 및 배포 비용 문제를 해결하고자 합니다.#Review#Lossless Compression#Tensor Compression#Program Synthesis#Domain-Specific Language (DSL)#Model Checkpoints#A* Search#Bit-Exact Reconstruction#Hugging Face2026년 8월 5일댓글 수 로딩 중
[논문리뷰] K-EXAONE 2.0 Technical Report본 연구는 글로벌 AI 경쟁 시대에 대응하여 국산 거대언어모델의 독자적인 구축 및 운영 능력을 확보하기 위해 수행되었다. 기존 모델을 폐기하고 처음부터 모델을 학습시키는 것은 방대한 계산 자원과 데이터, 학습 노하우를 낭비하는 결과를 초래한다.#Review#Foundation Model#Mixture-of-Experts#Upcycling#Speculative Decoding#Long-Context#Agentic Capability#Post-training2026년 8월 5일댓글 수 로딩 중
[논문리뷰] HelloWorld: Enabling Socially Interactive Characters in Video World Models본 논문은 기존 video world models이 사용자-캐릭터 간의 Social Interaction 기능을 지원하지 않는다는 핵심적인 문제를 해결하고자 합니다.#Review#Video World Models#Social Interaction#Self-Distillation#Temporal Cross-Attention Mask#HelloWorldBench#Camera-Pose Conditioning#Generative AI2026년 8월 5일댓글 수 로딩 중
[논문리뷰] GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks본 논문은 기존 에이전트 self-evolution 벤치마크가 실질적인 경제적 가치가 높은 도메인을 충분히 다루지 못하며, 학습-테스트 간의 인과 관계가 불분명하여 성과 향상을 신뢰하기 어렵다는 문제점을 해결하고자 합니다.#Review#Agent Self-Evolution#Benchmarking#Enterprise Workflows#Rule Hybridization#Data Contamination#Autonomous Agents2026년 8월 5일댓글 수 로딩 중
[논문리뷰] FocusMem: Factorizing Content, Readout, and Trust in Latent GUI MemoryGUI 에이전트(agents)는 과거 태스크의 유용한 경험과 현재 인터랙션의 미완성된 진행 상황을 효과적으로 기억해야 하지만, 기존 잠재 메모리(latent memory) 방식은 중요한 세부 정보 손실, 다양한 의사결정 단계에 대한 부적합성, 그리고 불필요한 정보로 인한 오도(misleading)와 같은 한계를 가지고 있습니다.#Review#GUI agents#latent memory#episodic memory#working memory#content basis#state-conditioned readout#trust gate#multimodal trajectories2026년 8월 5일댓글 수 로딩 중
[논문리뷰] Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data본 논문은 일반적인 로봇 조작(Manipulation) 정책을 학습하기 위한 대규모 데모 데이터 부족 문제를 해결하고자 한다. 기존의 로봇 데이터 수집 방식은 비용이 많이 들고 노동 집약적이며, 데이터의 다양성 또한 하드웨어 제약으로 인해 제한적이다.#Review#Robot Data Synthesis#Egocentric Data#Generalization Evaluation#Vision-Language-Action#Embodiment#Robot Learning2026년 8월 5일댓글 수 로딩 중
[논문리뷰] Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher GuidanceLLM의 추론 능력 향상을 위한 RLVR(Reinforcement Learning with Verifiable Rewards) 방법 중 GRPO (Group Relative Policy Optimization)는 단순하고 안정적이지만, 희소한 보상 신호(sparse reward signals)와 Negative Zero-Variance Prompts에서 그래디언트 소실(vanishing gradients) 문제를 겪습니다.#Review#Reinforcement Learning#On-Policy Distillation#Large Language Models#GRPO#Knowledge Distillation#Adaptive Teacher Guidance#Zero-Variance Prompts#Supervised Fine-Tuning2026년 8월 5일댓글 수 로딩 중
[논문리뷰] DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack본 연구는 최근 Flow-Matching 기반의 VLA 모델들이 기존의 autoregressive 모델들에 비해 adversarial 공격에 강인하다는 주장이 실제로는 오해일 수 있음을 지적한다.#Review#Vision-Language-Action Models#Flow-Matching#Adversarial Patch Attack#Denoising Velocity Field#Robustness#Gradient Conflict2026년 8월 5일댓글 수 로딩 중
[논문리뷰] Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning본 논문은 차트 이미지, 테이블 데이터, 시각화 코드 간의 cross-representation understanding에서 AI 시스템이 직면하는 근본적인 문제를 해결하고자 합니다.#Review#Cross-Representation Learning#Self-Supervised Learning#Consistency-Driven Co-Evolution#Multimodal Reasoning#Chart Understanding#Code Generation2026년 8월 5일댓글 수 로딩 중
[논문리뷰] BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation기존의 3D Vision-Language-Action(VLA) 모델들은 대규모 데이터셋에 과도하게 의존하며, 작업 수행 중 발생하는 시야 가림(occlusion)이나 과거 상태에 의존해야 하는 작업(memory-dependent task)을 해결하는 데 한계를 보입니다.#Review#Vision-Language-Action Models#3D Manipulation#Spatio-Temporal Memory#Heatmap Prediction#Data-Efficient#Robot Learning2026년 8월 5일댓글 수 로딩 중
[논문리뷰] Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming본 논문은 LLM 기반 에이전트의 취약점을 탐지하기 위한 기존 레드팀 방식의 낮은 효율성과 제한적인 전이 가능성을 해결하고자 합니다. 기존의 RL-based 방식은 특정 타겟 모델에 대한 방대한 학습 데이터와 수만 번의 쿼리가 필요하여 비용이 높고 새로운 모델에 대한 적응력이 떨어집니다.#Review#Prompt Injection#Red Teaming#Agentic System#Hierarchical Memory#Strategy Library#LLM Security2026년 8월 5일댓글 수 로딩 중
[논문리뷰] AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities본 논문은 지시 기반(instruction-based) 비디오 편집 모델의 평가에 있어 중요한 격차를 해결하고자 한다. 특히, 오디오와 시각 신호가 밀접하게 결합된 실제 비디오를 다루는 모델의 능력 평가에 초점을 맞춘다.#Review#Audio-Video Editing#Holistic Evaluation#Benchmark#MLLM-as-Judge#Fidelity Preserving#Instruction Following#Realism#AVE-Agent2026년 8월 5일댓글 수 로딩 중
[논문리뷰] ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment본 논문은 기존의 Long-horizon search agent 학습 방식이 가진 궤적 단위(Trajectory-level) 보상의 불명확성 문제를 해결하고자 합니다.#Review#Long-Horizon Search#Credit Assignment#Reinforcement Learning#Process Supervision#Information Retrieval#Agentic Search2026년 8월 5일댓글 수 로딩 중
[논문리뷰] When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings본 논문은 ALiBi positional encoding이 학습 과정에서 발생하는 Softmax Underflow로 인해 의도치 않은 'Attention Blindness'를 유발한다는 점을 최초로 규명합니다.#Review#ALiBi#Positional Encoding#Softmax Underflow#Attention Mechanism#Floating-Point Precision#Length Extrapolation#Associative Retrieval2026년 8월 4일댓글 수 로딩 중
[논문리뷰] When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills본 논문은 Persona Skills 파이프라인에서 발생하는 구조적 안전성 위협을 규명하고 이를 체계적으로 평가하고자 합니다.#Review#Persona Skills#Agentic Systems#Privacy Leakage#Impersonation Risk#AntiSkillBench#Distillation Strategy#Defense Evaluation2026년 8월 4일댓글 수 로딩 중