[논문리뷰] Uncertainty-Aware World Model for Aerial Image-Goal Navigation본 논문은 aerial navigation 환경에서 발생하는 미래 상태의 불확실성이 기존 세계 모델의 경로 선택(Trajectory scoring) 능력을 저해한다는 문제를 해결하고자 합니다.#Review#Aerial Navigation#World Model#Uncertainty-Aware#OOD Detection#Hierarchical Error Projection#Image-Goal Navigation2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Towards Interpretable Foundation Models for Retinal Fundus Images본 논문은 기존 망막 이미지 분석을 위한 Foundation Model들이 가지는 불투명성(Lack of Interpretability) 문제를 해결하고자 합니다.#Review#Foundation Models#Retinal Fundus Images#Interpretable-by-design#Self-Supervised Learning#BagNet#t-SimCNE2026년 8월 9일댓글 수 로딩 중
[논문리뷰] The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows기존의 프롬프트, 프로그램, ML 워크플로우 최적화 방식은 주로 진화 알고리즘(Evolutionary Search)이나 밴딧(Bandit)과 같은 외부의 명시적인 제어기(Outer-loop controller)에 의존하고 있습니다.#Review#Agentic Search#LLM Optimization#Prompt Engineering#Program Evolution#ML Workflow#Reasoning-driven2026년 8월 9일댓글 수 로딩 중
[논문리뷰] StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding본 논문은 autonomous multimodal agents를 continuous, real-world environments에 배포하는 데 있어 기존 모델과 벤치마크의 한계점을 해결하고자 합니다.#Review#Streaming Video Understanding#Agentic AI#Long-Horizon Memory#Multimodal Perception#Proactive Interaction#Tool Utilization#StreamArena#Two-tier Architecture2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Small Foundation Models of Human Cognition and Behaviour본 연구는 대규모 언어 모델을 활용한 Cognitive Foundation Models의 성능이 단순히 파라미터 규모에 의존하는지, 아니면 실제 과제 구조(task structure)를 학습하는지 규명하고자 합니다.#Review#Foundation Models#Cognitive Science#Behavioral Data#Supervised Fine-tuning (SFT)#Noise Ceiling#Prompt Decomposition#LoRA2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Skaling: Chinchilla's Exponents Meet Kaplan's Coupling본 논문은 기존의 Additive Chinchilla law가 데이터 부족(data-scarce) 및 과잉 학습(overtraining) 극단 영역에서 Systematic prediction bias를 유발한다는 점을 지적한다.#Review#Neural Scaling Laws#Chinchilla#Kaplan#Loss Surface#Compute Allocation#Extrapolation#Sparse Profiling2026년 8월 9일댓글 수 로딩 중
[논문리뷰] SimWAM: A Simple World Action Model for End-to-End Autonomous Driving본 논문은 기존 World-Action Models의 'imagine-then-act' 파이프라인이 갖는 실시간 추론 지연(Latency) 문제를 해결하고자 합니다. 기존 방식은 자율주행 경로를 계획하기 전 미래 프레임을 반드시 생성해야 하므로 컴퓨팅 자원 소모가 크다는 한계가 있습니다.#Review#World-Action Models#End-to-End Autonomous Driving#Flow Matching#Trajectory Planning#Reinforcement Learning#Vision-Language-Action2026년 8월 9일댓글 수 로딩 중
[논문리뷰] SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs본 논문은 LLM의 Multi-task 성능 향상을 위한 두 주류 방법론인 SFT와 RL이 Multi-task 환경에서 보이는 극명한 거동 차이를 규명하고자 합니다. 기존 연구들은 SFT와 RL의 단일 태스크 학습 특성을 비교했으나, Multi-stage 학습 상황에서의 근본적인 차이는 명확히 설명하지 못했습니다.#Review#Large Language Models#Supervised Fine-Tuning#Reinforcement Learning#Multi-Task Learning#Gradient Interference#Parallel-RL2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors본 연구는 오토레그레시브 모델이 long-horizon rollout 과정에서 누적하는 오차(rollout error)를 배포 시점에 측정할 수 없는 문제를 해결합니다 .#Review#Bidirectional Diffusion#Round-Trip Consistency#Self-Supervised Uncertainty Quantification#Rollout Error#Latent Dynamics#Scientific Machine Learning2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression본 논문은 대규모 컨텍스트를 처리하는 LLM의 추론 비용을 줄이기 위한 Hard Prompt Compression 기법이, 정보 조각을 독립적으로 선택함에 따라 필연적으로 발생하는 구조적 결함인 Referential Dangling 문제를 제기합니다 .#Review#Prompt Compression#Hard Prompt Compression#Referential Dangling#Long-Context#LLM Inference#Information Retrieval#Contextual Dependency2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning본 연구는 기존 오디오 추론 모델이 결과 중심의 보상(Outcome-based rewards)이나 고정된 루브릭 보상(Process-based rewards) 체계 하에서 겪는 한계를 해결하고자 합니다.#Review#Audio Reasoning#Reinforcement Learning#Evolving Rubrics#Large Audio Language Models#Process-based Reward#Audio-grounded Supervision2026년 8월 9일댓글 수 로딩 중
[논문리뷰] PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say본 논문은 LLM 기반 에이전트가 외부 툴을 통해 작업 수행 시 과도한 민감 정보를 획득하는 문제를 해결하기 위해 PrivacyPeek을 제안합니다 .#Review#LLM-based Agents#Privacy Leakage#Data Acquisition#Benchmark#Contextual Integrity#Tool-use2026년 8월 9일댓글 수 로딩 중
[논문리뷰] OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction본 논문은 기존 감정 지능 연구가 특정 태스크에만 치중된 Specialist 모델 위주로 구성되어 있어, 감정의 다면적 속성을 포괄하지 못하고 태스크 간 시너지(Synergy)를 활용하지 못하는 문제를 해결하고자 합니다.#Review#Multimodal Large Language Models#Affective Computing#Reinforcement Learning#Reasoning Trajectories#Emotion Perception#Multi-task Learning2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Modular TTT: Rethinking Test-Time Training as Composable Modules본 논문은 기존 TTT 연구들이 각 모델 변형을 개별적으로 hard-coding하여 TTT 디자인 공간을 체계적으로 이해하기 어렵다는 문제에서 출발합니다 .#Review#Test-Time Training#Modular Framework#Sequence Modeling#Directed Acyclic Graph#Inner Learner#Fast Weights2026년 8월 9일댓글 수 로딩 중
[논문리뷰] FATE: Frame-Level Audio-Visual Temporal Embedding본 논문은 오디오-비주얼 모델에서 '무엇(what)'에 해당하는 의미적 정보와 '언제(when)'에 해당하는 시간적 정렬 정보가 효과적으로 결합되지 못하는 문제를 해결합니다.#Review#Audio-Visual Embedding#Temporal Alignment#Contrastive Learning#Event Localization#Cross-Modal Retrieval#Frame-Level Representation2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss본 연구는 자원 제약이 심한 환경에서 대규모 언어 모델(LLM)을 배포하기 위한 효율적인 지식 증류(Knowledge Distillation) 파이프라인 구축을 목표로 합니다.#Review#Knowledge Distillation#Offline Distillation#Fused Chunked KL Loss#Long-context Healing#LLM Compression#Sparse Top-K Logits2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Douyin Multimodal Embedding Model Technical Report본 논문은 대규모 산업용 다중 모달 플랫폼에서 효율적인 검색과 세밀한 의미론적 식별이라는 두 가지 상충하는 요구사항을 동시에 만족시키기 위해 DME를 제안합니다.#Review#Multimodal Embedding#Retrieval#Semantic Sufficiency#Latent Reasoning#Cross-Conditional Reconstruction#Bi-Encoder#Representation Learning2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events본 연구는 LLM 기반 에이전트가 시간이 지남에 따라 경험하는 삶의 사건에 대해 인간과 유사한 심리학적 성격 변화를 보이는지 규명하고자 합니다.#Review#LLM Agents#Personality Evolution#Big Five Personality#Life Events#BFI-Adapt#Psychometric Evaluation2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Characterizing the Quality Profile of AI-Generated C++ in Production본 논문은 대규모 엔터프라이즈 환경에서 AI가 생성한 C++ 코드의 품질, 성능, 유지보수 특성을 실증적으로 분석합니다 . 기존 연구들은 대부분 통제된 환경이나 벤치마크 데이터셋 위주로 수행되어, 실제 프로덕션 환경의 복잡한 코드 리뷰 및 배포 과정을 반영하지 못한다는 한계가 있었습니다.#Review#AI-generated code#Software quality taxonomy#Authoring-time provenance#Static analysis#Code review#Empirical software engineering2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence본 논문은 기존의 VLM들이 embodied agent의 반복적이고 동적인 실행 환경(observation–reasoning–action cycle)을 통합적으로 처리하지 못하고 특정 작업에 편향되어 있다는 점을 핵심 문제로 정의합니다 .#Review#Vision-Language Model#Embodied Intelligence#Capability Specialist#GRPO#TIES#Model Consolidation#Action Guidance2026년 8월 9일댓글 수 로딩 중