[논문리뷰] DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation본 논문은 기존의 고성능 VDR 시스템이 요구하는 막대한 컴퓨팅 자원과 인덱싱 비용 문제를 해결하고자 합니다. 현재 최첨단 VDR 모델들은 수십억 개의 파라미터를 사용하며, Multi-vector 기반의 복잡한 연산으로 인해 메모리 사용량과 검색 레이턴시가 지나치게 높습니다.#Review#Visual Document Retrieval#Distillation#Single-Vector#End-to-End#Multimodal#Encoder-only2026년 8월 11일댓글 수 로딩 중
[논문리뷰] Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation본 논문은 기존의 고성능 생성 모델들이 고착화된 모델링 구조로 인해 완전한 End-to-End 학습을 수행하지 못하고 있다는 문제를 제기한다.#Review#Generative Modeling#Explorative Modeling#End-to-End#Mode Forcing#Generative Expressivity#Pretraining Axis2026년 7월 30일댓글 수 로딩 중
[논문리뷰] HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone본 연구는 로봇 학습의 핵심 병목 현상인 '고품질 조작 데이터의 부족' 문제를 해결하고자 합니다. 기존 방식은 로봇 직접 조작(Teleoperation)을 통해 데이터를 수집하지만, 이는 비용이 매우 많이 들고 확장이 어렵다는 근본적인 한계가 있습니다.#Review#Robot Learning#Manipulation Policy#UMI#End-to-End#Data-Production#Deployability2026년 7월 28일댓글 수 로딩 중
[논문리뷰] Xiaomi-GUI-0 Technical Report본 연구는 기존 GUI 에이전트 연구들이 의존하는 정적인 벤치마크나 시뮬레이션 환경이 실제 모바일 기기의 복잡한 상태 분포를 반영하지 못하는 한계를 해결하기 위해 수행되었다.#Review#GUI Agent#VLM#Real-Device#Reinforcement Learning#Data Flywheel#End-to-End#Mobile Automation2026년 6월 30일댓글 수 로딩 중
[논문리뷰] GEAR: Guided End-to-End AutoRegression for Image Synthesis본 논문은 현대의 시각적 생성 모델들이 tokenizer와 generator를 2단계로 분리하여 학습함으로써 발생하는 비효율성을 해결하고자 합니다 .#Review#GEAR#Autoregressive#Tokenizer#End-to-End#Representation Alignment#Vector Quantization#Image Synthesis2026년 6월 30일댓글 수 로딩 중
[논문리뷰] Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models본 논문은 실시간 오디오-비디오 인터랙션의 단절성과 모듈 간의 지연 시간 문제를 해결하기 위해 Wan-Streamer를 제안한다. 기존 연구들은 VAD, ASR, LLM, TTS 등을 결합한 캐스케이드(cascaded) 방식을 사용하여, 모듈 경계에서의 대기 시간과 오차 누적 문제에 직면해 있다 .#Review#End-to-End#Real-time Interaction#Multimodal Foundation Models#Full-duplex#Streaming Inference#Block-causal Attention#Thinker-Performer Pipeline2026년 6월 24일댓글 수 로딩 중
[논문리뷰] Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models본 논문은 실시간 객체 탐지 모델이 가진 NMS 의존성, 불필요한 모델 파라미터 팽창, 학습 효율성 저하, 그리고 소형 객체 탐지 실패 문제를 해결하고자 합니다 .#Review#YOLO26#Real-Time Object Detection#End-to-End#NMS-Free#MuSGD#STAL#Instance Segmentation2026년 6월 2일댓글 수 로딩 중
[논문리뷰] WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training본 논문은 통합적인 End-to-End Spoken Dialogue Model의 의미론적 지능(Intelligence, IQ)과 음성 표현력(Expressiveness, EQ)을 동시에 향상시키는 문제를 해결하고자 한다.#Review#Spoken Dialogue Models#Post-Training#Reinforcement Learning#Preference Optimization#Modality Alignment#End-to-End#Acoustic Expressiveness2026년 4월 22일댓글 수 로딩 중
[논문리뷰] AURA: Always-On Understanding and Real-Time Assistance via Video Streams본 논문은 기존 VideoLLMs 가 대부분 오프라인 분석에 최적화되어 있어, 실시간으로 변화하는 비디오 스트림에 대한 연속적이고 즉각적인 대응에 한계가 있다는 문제점을 해결하고자 합니다.#Review#VideoLLMs#Streaming Video Understanding#End-to-End#Context Management#Proactive Response#Real-Time Inference2026년 4월 6일댓글 수 로딩 중
[논문리뷰] MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction단일 이미지로부터 관절형 3D 객체를 재구성하는 것은 객체의 기하학적 구조, Part 구조 및 motion parameter를 제한된 시각적 증거로부터 함께 추론해야 하므로 여전히 근본적인 도전 과제이다.#Review#Monocular 3D Reconstruction#Articulated Objects#Progressive Structural Reasoning#Kinematic Estimation#PartNet-Mobility#End-to-End2026년 3월 19일댓글 수 로딩 중
[논문리뷰] ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation본 논문은 기존 비디오-오디오 생성 모델이 모노 출력에 국한되어 공간적 몰입감이 부족하며, 기존 바이노럴 접근 방식이 2단계 파이프라인(모노 생성 후 공간화)으로 인한 오류 누적과 시공간 불일치 문제를 겪는 한계를 해결하고자 합니다.#Review#Binaural Audio Generation#Spatial Audio#Video-Driven#End-to-End#Conditional Flow Matching#Multimodal AI#Deep Learning#Audio-Visual Synthesis2025년 12월 2일댓글 수 로딩 중
[논문리뷰] OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion본 논문은 텍스트 전용 번역 LLM이 겪는 지연 시간과 멀티모달 컨텍스트 활용 불가능성, 그리고 MMFM이 가진 다국어 번역 성능 및 커버리지의 한계를 해결하고자 합니다.#Review#Multimodal Translation#Speech Translation#Simultaneous Translation#Large Language Models#Multimodal Foundation Models#Modular Fusion#End-to-End#Gated Fusion#OCR2025년 12월 1일댓글 수 로딩 중