[논문리뷰] Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models본 논문은 기존 벤치마크에서 우수한 성능을 보이는 최신 멀티모달 모델들이 인간에게는 사소한 작업에서 여전히 실패하는 문제를 해결하고자 한다 . 대규모 언어 모델과 멀티모달 모델은 이미 많은 표준 벤치마크를 거의 포화 상태로 만들었으나, 이러한 점수가 모델의 실질적인 견고성을 항상 대변하지는 않는다.#Review#Multimodal Models#Benchmarking#Blind Spots#Reasoning Evaluation#Task Taxonomy#AI Evaluation2026년 7월 14일댓글 수 로딩 중
[논문리뷰] Weak-to-Strong Generalization via Direct On-Policy Distillation본 논문은 대규모 언어 모델의 post-training 단계에서 발생하는 RLVR(Reinforcement Learning with Verifiable Rewards)의 높은 컴퓨팅 비용 문제를 해결하고자 합니다.#Review#Weak-to-Strong Generalization#Reinforcement Learning#On-Policy Distillation#Policy Shift#Implicit Reward#Post-Training#Large Language Models2026년 7월 13일댓글 수 로딩 중
[논문리뷰] Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals본 논문은 기존 LLM 사후 학습 방식이 탐색(exploration)과 분포 정렬(distribution alignment)을 강하게 결합하여 컴퓨팅 효율성과 확장성을 저해하는 문제를 해결합니다.#Review#Post-training#Proxy Exploration#Update Signal Transfer#LLM Alignment#Modular Training#Weak-to-Strong Generalization2026년 7월 13일댓글 수 로딩 중
[논문리뷰] NeuroCogMap Reveals Cognitive Organization of Large Language Models본 논문은 LLM이 복잡한 인지적 능력을 발휘함에도 불구하고, 이러한 능력이 내부적으로 어떻게 조직화되어 있는지에 대한 시스템 수준의 설명이 부족하다는 문제의식을 다룹니다.#Review#Large Language Models#NeuroCogMap#Functional Parcellation#Cognitive Hierarchy#Mechanistic Interpretability#Pathology Detection#Cortical Alignment2026년 7월 13일댓글 수 로딩 중
[논문리뷰] Motion4Motion: Motion Transfer Across Subjects at Inference본 논문은 기존 모션 전이 방식이 스켈레톤 구조에 지나치게 의존함으로써 겪는 범용성 부족 문제를 해결하고자 합니다. 대다수의 기존 연구는 인간 중심의 스켈레톤 사전 지식을 강제하여, 동물과 같이 다양한 형태의 캐릭터 간 모션 전이에 적용하기 어렵습니다 .#Review#Motion Transfer#Training-free#Diffusion Transformer#Attention Control#Video Generation#Cross-species#Motion Flow2026년 7월 13일댓글 수 로딩 중
[논문리뷰] Metacognition in LLMs: Foundations, Progress, and Opportunities본 논문은 LLM이 인간의 고유한 지적 능력으로 여겨지는 Metacognition을 어느 수준까지 발휘할 수 있는지, 그리고 이를 어떻게 시스템 수준에서 구현하여 성능과 신뢰성을 높일 수 있는지에 대한 체계적인 분석을 목표로 합니다.#Review#Metacognition#Large Language Models#Confidence Calibration#Self-Correction#Uncertainty Estimation#Artificial Intelligence#Cognitive Psychology2026년 7월 13일댓글 수 로딩 중
[논문리뷰] LightMem-Ego: Your AI Memory for Everyday Life본 논문은 일상생활의 경험을 지속적으로 기록하고 활용해야 하는 개인용 AI 어시스턴트의 메모리 한계 문제를 해결하기 위해 LightMem-Ego를 제안합니다.#Review#Egocentric Perception#Multimodal Memory#Streaming Architecture#Hierarchical Memory#Life Assistant#Experience Retrieval2026년 7월 13일댓글 수 로딩 중
[논문리뷰] Latent-Identity Tuning in Text-to-Image Personalization Models본 논문은 기존의 Text-to-Image personalization 모델이 특정 개인의 정체성을 재현하는 데에는 뛰어나지만, 생성된 정체성을 세밀하게 수정하거나 제어하는 기능이 결여되어 있다는 점을 해결하고자 합니다 .#Review#Text-to-Image#Personalization#Identity Tuning#Latent Space#Q-Former#Fine-grained Editing2026년 7월 13일댓글 수 로딩 중
[논문리뷰] LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow기존의 3D 메시 생성 모델들은 정점의 공간적 위치와 위상적 연결성을 하나의 공유된 latent space에서 동시에 학습하려는 경향이 있어, 통계적으로 이질적인 두 신호를 효율적으로 처리하는 데 한계가 있다.#Review#3D Mesh Generation#Flow Matching#Factorized Representation#Vertex Flow#Topology Flow#Latent Representation2026년 7월 13일댓글 수 로딩 중
[논문리뷰] EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos본 논문은 일반적인 로봇 조작 모델이 실시간 Steerability를 확보하지 못하고, 특정 로봇 환경에 국한되는 한계를 해결하고자 한다.#Review#Steerable Dexterous Manipulation#VLA Models#Egocentric Videos#World Model#Robot Learning#DAgger2026년 7월 13일댓글 수 로딩 중
[논문리뷰] CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation본 논문은 기존 가상 착장(VTO) 시스템이 의류의 스타일, 크기, 공간적 배치와 같은 사용자 수준의 미세한 제어를 지원하지 못한다는 한계를 해결하고자 한다.#Review#Virtual Try-On#Image Editing#Visual-Instance-Prompt Segmentation#Segmentation Masks#Diffusion Transformer#Controllability2026년 7월 13일댓글 수 로딩 중
[논문리뷰] AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification본 논문은 기존의 수학 벤치마크가 고등 수학 및 연구 수준의 증명 능력을 평가하기에는 범위와 입도가 부족하다는 문제를 해결하고자 합니다.#Review#Advanced Mathematics#Proof Generation#Process Verification#LLM-as-Judge#Mathematical Reasoning#Benchmark#Automatic Verification Pipeline2026년 7월 13일댓글 수 로딩 중
[논문리뷰] ABot-N1: Toward a General Visual Language Navigation Foundation Model본 논문은 기존의 단일 통합 정책(Monolithic Policy)이 가진 navigation의 한계점과 확장성 문제를 해결하기 위해 ABot-N1을 제안합니다 .#Review#Visual Language Navigation#Foundation Model#Slow-Fast Architecture#Chain-of-Thought#Pixel Goal#Embodied AI#Cross-Task Generalization2026년 7월 13일댓글 수 로딩 중
[논문리뷰] ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory본 연구는 고수준의 semantic reasoning을 물리적인 다단계 실행(multi-step physical execution)으로 연결하는 과정에서 발생하는 'reasoning-execution gap'을 해결하고자 합니다 .#Review#Embodied Intelligence#Agent Operating System#Multi-modal Memory#Lifelong Self-Evolution#Robot Learning#Hierarchical Reasoning#EmbodiedWorldBench2026년 7월 13일댓글 수 로딩 중
[논문리뷰] 4D Human-Scene Reconstruction from Low-Overlap Captures본 논문은 소수의 low-overlap 카메라만으로도 고품질의 4D 인간-장면 복원(Human-Scene Reconstruction)을 구현하는 문제를 해결합니다.#Review#4D Reconstruction#Gaussian Splatting#Sparse-view#Video Diffusion#Human-Scene Decomposition#Multi-view Pose Estimation2026년 7월 13일댓글 수 로딩 중
[논문리뷰] Video Generation Models are General-Purpose Vision Learners본 논문은 컴퓨터 비전 분야가 여전히 개별 과제에 특화된 모델(Specialized Model) 단계에 머물러 있는 문제를 해결하고자 합니다 .#Review#Video Generation#Foundation Models#Generalist Vision Intelligence#Diffusion Models#Spatiotemporal Priors#Perception Task-Agnostic#Synthetic Data2026년 7월 12일댓글 수 로딩 중
[논문리뷰] VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery본 연구는 고대 그리스 도자기와 같은 문화유산 분야에서 VLM 기반의 디지털 박물관 가이드 시스템이 직면한 신뢰성 부족 문제를 해결하고자 합니다. 기존 모델들은 파편화되거나 불완전한 정보를 바탕으로 과도하게 확신에 찬 답변(Hallucination)을 생성하거나, 검증되지 않은 외부 참조를 인용하는 한계가 있습니다.#Review#Vision-Language Models#Digital Museum#Cultural Heritage#Multimodal Agent#Retrieval-Augmented Generation#Inference-time Reliability Control#GRPO2026년 7월 12일댓글 수 로딩 중
[논문리뷰] Trust Region Policy Distillation본 논문은 기존 On-Policy Distillation (OPD) 방식이 가진 구조적 불안정성과 낮은 샘플 효율성 문제를 해결하기 위해 고안되었습니다.#Review#On-Policy Distillation#Trust Region#Policy Gradient#Proximal Teacher#Gradient Variance#Mathematical Reasoning#Post-training2026년 7월 12일댓글 수 로딩 중
[논문리뷰] Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning본 논문은 LLM이 새로운 지식을 성공적으로 기억함에도 불구하고, 이를 활용한 downstream 추론 작업에서는 낮은 성능을 보이는 문제를 다룬다 . 기존 연구들은 주로 모델의 파라미터 업데이트나 지식 편집에 집중했으나, 지식 저장과 추론 간의 인과적 단절을 메커니즘적으로 설명하는 데에는 한계가 있었다.#Review#LLM Finetuning#Knowledge Generalization#Mechanistic Interpretability#Self-Patching#Knowing-Using Gap#Knowledge-Circuit Misalignment2026년 7월 12일댓글 수 로딩 중
[논문리뷰] Self-Guided Test-Time Training for Long-Context LLMs본 논문은 긴 문맥을 처리하는 LLM의 성능이 문맥의 길이에 따라 저하되는 현상이 단순히 문맥을 모두 담지 못해서가 아니라, 질문에 필요한 핵심 증거를 식별하고 활용하는 능력이 부족하기 때문임을 지적합니다.#Review#Long-Context LLMs#Test-Time Training (TTT)#Evidence Selection#Parameter Adaptation#Context Reasoning#Signal-to-Noise Ratio2026년 7월 12일댓글 수 로딩 중