[논문리뷰] OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation본 논문은 기존의 독립적인 오디오-비디오 VAE 학습 방식이 초래하는 교차 모달 정렬(cross-modal alignment) 부족 문제를 해결하고자 합니다.#Review#OmniVAE#Cross-modal Alignment#Audio-Video Generation#Latent Diffusion#Contrastive Learning#Semantic Distillation2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models본 연구는 물리적 훼손으로 판독이 어려운 역사 기록물을 복원하기 위해 LLM과 외부 지식을 결합하는 새로운 접근 방식을 제안합니다.#Review#Historical Document Restoration#Retrieval-Augmented Generation (RAG)#Large Language Models (LLMs)#Named Entity Recognition#Hanja#Domain-specific Fine-Tuning#Textual Contextualization2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Kimi K3: Open Frontier Intelligence본 논문은 기존 오픈소스 모델들이 1T 파라미터 규모에 정체되어 있는 반면, 상업적 모델들은 더 큰 규모와 복잡한 추론 능력을 갖추며 격차가 벌어지는 문제를 해결하고자 합니다. 특히 long-context와 agentic tasks를 동시에 처리하기 위한 모델의 확장성과 효율성이 부족하다는 점이 핵심 배경입니다.#Review#Mixture-of-Experts#Kimi Delta Attention#Long-context#Reinforcement Learning#Multimodal#Scaling Efficiency#Agentic Intelligence2026년 7월 27일댓글 수 로딩 중
[논문리뷰] JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents본 논문은 기존의 Creative AI 시스템이 고립된 단일 단계 생성에 머물러 있어, 복잡한 다단계 창작 작업에서 필요한 맥락(Context)을 유지하지 못하는 문제를 해결하고자 합니다.#Review#Multimodal Agents#Canvas-Native#Creative AI#Agent Harness#Long-Horizon Creation#Project State2026년 7월 27일댓글 수 로딩 중
[논문리뷰] IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages본 연구는 다국어 및 코드 혼용 대화 시스템 구축에 필수적인 고품질 대규모 말뭉치가 부족한 문제를 해결합니다. 기존의 Indic 언어 데이터셋은 주로 감성 분석이나 개체명 인식과 같은 식별 작업(Discriminative tasks)에 치우쳐 있어, 생성 모델의 instruction tuning에 부적합합니다.#Review#Indic Languages#Code-mixed Conversation#LLM-based Generation#Persona-conditioned Dialogue#Multilingual NLP#Corpus Construction2026년 7월 27일댓글 수 로딩 중
[논문리뷰] GNM Head: A Generative aNthropometric Model of the human head본 논문은 기존의 3DMM들이 가진 해부학적 폐쇄성 문제를 해결하고자 합니다. 기존의 모델들(예: FLAME, BFM)은 인간의 머리를 속이 빈 껍데기 형태로 간주하여 치아나 혀 같은 중요한 내부 기관을 생략해 왔습니다. 이러한 구조적 한계는 생성형 AI나 실시간 렌더링 시 시각적 품질을 저하시키는 원인이 됩니다.#Review#3D Morphable Model (3DMM)#Human Head#Parametric Model#Computer Vision#Neural Rendering#Facial Reconstruction#Anatomy2026년 7월 27일댓글 수 로딩 중
[논문리뷰] From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search본 논문은 Agentic Search 최적화 시 발생하는 Sparse Supervision 문제와 Proprietary-to-Open-Source 지식 증류의 제약사항을 해결하는 것을 목표로 합니다.#Review#Agentic Search#Knowledge Distillation#Multi-Agent System#Reinforcement Learning#JSON Protocol#Policy Alignment#Large Language Models2026년 7월 27일댓글 수 로딩 중
[논문리뷰] FilmBench: A Film-Grade Benchmark for Cinematic Video Generation본 논문은 기존 비디오 생성 벤치마크들이 웹 기반의 단순 텍스트 프롬프트를 사용하여 영상의 'Cinematic Quality'를 평가하는 데 한계가 있다는 점을 지적한다.#Review#Text-to-Video#Reference-to-Video#Cinematic Language#FilmBench#Benchmark#Video Generation#FilmOps2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels본 논문은 Visual Document Understanding 분야에서 발생하는 Attribution Hallucination 문제를 해결하기 위해, 현재의 좌표 기반 인터페이스가 모델의 근거 표현 능력을 제한하고 있는지 분석합니다 .#Review#Vision-Language Models#Evidence Attribution#Attribution Hallucination#Multimodal Retrieval#Reinforcement Learning#Doc-VQA2026년 7월 27일댓글 수 로딩 중
[논문리뷰] DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification본 논문은 실주행 데이터에서 운전자의 개별적인 주행 습관을 식별할 때 발생하는 불명확한 환경적 편향(Confounds) 문제를 해결하고자 한다.#Review#Naturalistic Driving#Driving Style Identification#Behavior Prediction#Driver Re-identification#Multimodal Time Series#Shortcut Learning2026년 7월 27일댓글 수 로딩 중
[논문리뷰] DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes본 논문은 VLM의 사전 학습 데이터 구성을 직관에 의존하는 휴리스틱한 방식에서 재현 가능한 엔지니어링 원칙으로 전환하고자 한다. 기존 연구들은 단순히 고품질 데이터를 필터링하거나 데이터셋을 단순히 쌓는 방식을 사용하며, 데이터셋 추가가 성능에 미치는 인과 관계를 명확히 파악하기 어려운 문제가 있다.#Review#Vision-Language Models#Data Mixture Optimization#Convex Optimization#Dataset-Level Assessment#Scalable VLM Training#Attributable Validation2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Data Pyramid for Embodied Manipulation본 논문은 Embodied AI 모델의 성능을 결정짓는 핵심 요소인 '데이터'가 매우 파편화되어 있다는 문제 의식에서 출발합니다. 기존 연구들은 특정 모델의 학습 레시피에 의존하여 데이터를 수집하거나, 아키텍처 중심의 평가에 치중하여 데이터 소스 간의 관계와 trade-off를 시스템화하지 못했습니다.#Review#Embodied AI#Foundation Models#Data Pyramid#Robotics#Manipulation#Multimodal Learning#Robot Learning2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Codifying the Judge: Scalable Evaluation via Program Distillation본 논문은 기존 LLM-as-a-judge 패러다임이 가진 높은 API 비용, 불투명한 의사결정 과정, 시스템적 편향(Bias), 그리고 유연성 부족 문제를 해결하는 것을 목적으로 한다.#Review#LLM-as-a-judge#Program Distillation#Pajama#Weak Supervision#Model Routing#Pareto Frontier#Reward Modeling2026년 7월 27일댓글 수 로딩 중
[논문리뷰] ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding본 논문은 현재의 의료용 MLLM이 직면한 '시각 중심적 지식 흡수'와 '임상적으로 유효한 평가'라는 두 가지 핵심 문제를 해결하고자 합니다.#Review#Multimodal LLM#Medical Imaging#Vision Encoder#CaSL Fusion#Clinical Report Generation#MedIF-Bench#RoI-Grounded Evaluation2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Characterizing Warp Divergence from Pascal to Blackwell본 논문은 ITS 도입 이후 NVIDIA GPU의 warp divergence 처리 방식이 고정되어 있다는 업계의 통념을 검증하는 것을 목표로 합니다.#Review#GPU#Warp Divergence#Independent Thread Scheduling#Reconvergence#SASS#Microarchitecture#Blackwell2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling본 논문은 기존의 단일 타겟-단일 상태(single-target, single-state) 중심의 바인더 설계 모델이 다중 상태(Multi-state) 단백질의 구조적 변화나 다중 타겟(Multi-target) 결합 요구를 수용하지 못하는 한계를 해결합니다 .#Review#Cross-context Binder Design#In-context Generation#Mixed Sampling#Diffusion Model#Protein Binder Design#Multi-state Design#Multi-target Design2026년 7월 27일댓글 수 로딩 중
[논문리뷰] A Vocabulary for Multi-Agent Automated Research Systems본 논문은 다양한 멀티 에이전트 기반 자동 연구 시스템들을 일관된 기준으로 기술하고 비교할 수 있는 표준화된 Vocabulary의 부재 문제를 해결합니다.#Review#Multi-Agent Systems#Automated Research#System Design#Trajectory#Proxy Optimization#Evaluation2026년 7월 27일댓글 수 로딩 중
[논문리뷰] A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever본 연구는 기존의 거대 언어 모델이 Retraining을 통해 성능을 개선하는 과정에서 발생하는 비용, 불투명성, 비결정론적 출력 문제를 해결하고자 합니다.#Review#Verified Knowledge#Determinism#Memory-augmented LLM#Zero-token Inference#Byte-exact#Merlin Engine2026년 7월 27일댓글 수 로딩 중
[논문리뷰] VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression본 논문은 VLM이 고해상도 이미지를 처리할 때 발생하는 과도한 시각적 토큰으로 인한 inference latency와 KV cache 메모리 오버헤드 문제를 해결하고자 합니다.#Review#Vision-Language Models#Token Compression#Autoencoder#Intrinsic Encoding#Hierarchical Information#Parameter Sharing2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Three-Body Scattering for Generative Modeling본 논문은 현대 generative models이 adversarial critic, prescribed noise-to-data path(예: diffusion models), 또는 autoregressive factorization에 의존하는 한계를 극복하기 위해 새로운 one-step generation 방법을 제안합니다.#Review#Three-Body Scattering Modeling (TBSM)#Generative Modeling#Energy Distance#One-Step Generation#Tracked Scattering#Wasserstein Gradient Flow#FID#ImageNet2026년 7월 26일댓글 수 로딩 중