[논문리뷰] Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding기존의 MLLM 기반 STVG 모델들은 영상 내 객체의 이동 궤적(Trajectory)을 Autoregressive 방식으로 순차 생성하는데, 이는 Tube 길이가 길어질수록 Decoding Latency가 선형적으로 증가하고, 이전 시점의 위치 정보 오류가 다음 시점으로 전파되는 문제를 야기한다.#Review#Spatio-Temporal Video Grounding#Parallel Tube Decoding#Decoupled Block Attention#Localization-Aware Policy Optimization#Multimodal Large Language Models2026년 8월 30일댓글 수 로딩 중
[논문리뷰] AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models본 논문은 기존의 STVG 평가 방식이 일반적인 일상 데이터에만 국한되어 있어, 실제 산업 현장이나 전문 분야에서 요구되는 고차원적 인식 능력을 측정하지 못한다는 문제를 제기합니다 .#Review#Spatio-Temporal Video Grounding#Vision-Language Models#Domain Adaptation#In-Context Learning#Benchmark#Video Understanding2026년 7월 2일댓글 수 로딩 중
[논문리뷰] Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding본 논문은 입력 텍스트 질의를 기반으로 비디오 내에서 대상의 시공간 튜브(spatio-temporal tube)를 찾아내는 시공간 비디오 그라운딩(STVG) 태스크에서, MLLM(Multimodal Large Language Models) 의 잠재력을 활용하여 제로샷(zero-shot) 해결책 을 제시하는 것을 목표로 합니다.#Review#Spatio-Temporal Video Grounding#Multimodal Large Language Models#Zero-Shot Learning#Visual Grounding#Decomposed Spatio-Temporal Highlighting#Logit-Guided Re-attention#Temporal-Augmented Assembling2025년 9월 19일댓글 수 로딩 중