[논문리뷰] Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos본 논문은 기존 Egocentric assistant가 수동적인 Reactive 방식이나 특정 이벤트 발생 시에만 응답하는 Semi-proactive 방식에 머물러 있다는 한계를 지적합니다.#Review#Egocentric Video#Proactive Assistance#Retrieval-Augmented Reasoning#Streaming Memory#Video-LLM#Benchmarking2026년 7월 15일댓글 수 로딩 중
[논문리뷰] Video-Oasis: Rethinking Evaluation of Video Understanding본 논문은 현대 Video-LLM 벤치마크들이 모델의 진정한 시공간적 추론 능력을 측정하지 못하고 있다는 근본적인 문제를 지적한다.#Review#Video-LLM#Diagnostic Suite#Spatio-Temporal Reasoning#Benchmark Audit#Video-Native Challenges2026년 7월 9일댓글 수 로딩 중
[논문리뷰] Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning본 논문은 기존 Autoregressive Video-LLM 기반의 Dense Video Captioning 모델들이 겪는 높은 추론 지연(Latency)과 확장성 문제를 해결하고자 합니다.#Review#Dense video captioning#Parallel decoding#Latent planning#Omni-modal#Video-LLM#Dependency restructuring2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory본 논문은 기존 비디오 에이전트 모델들이 롱폼 비디오를 처리할 때 의존하는 '탐정 스타일'의 반복적 추론(Iterative Reasoning)이 초래하는 과도한 비용과 레이턴시 문제를 해결하고자 합니다 .#Review#Multimodal Long-Term Memory#Agentic Video Understanding#Dual-State Design#Reflexive Response#Retrieval-Augmented Generation#Video-LLM2026년 7월 7일댓글 수 로딩 중
[논문리뷰] When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding본 논문은 기존 Video-LLM의 한계인 불명확한 시간 인코딩, 프레임 수준의 낮은 연속성, 그리고 관심 엔티티에 대한 언어-비전 정렬 불일치를 극복하는 것을 목표로 합니다. 특히 긴 비디오에서 발생하는 이벤트의 정밀한 시간적 위치 파악과 엔티티 수준의 견고한 정렬을 통해 비디오 이해 능력을 향상시키고자 합니다.#Review#Video-LLM#Diffusion Model#Temporal Grounding#Object Segmentation#Long Video Understanding#Multimodal AI#Video Question Answering2025년 8월 22일댓글 수 로딩 중