[논문리뷰] VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding본 논문은 기존 Video MLLM이 겪는 일반화 부족, 높은 연산 비용, 그리고 폐쇄적인 연구 생태계라는 세 가지 한계를 해결하는 것을 목표로 한다. 기존 모델들은 짧은 영상에는 강점을 보이지만, 장시간 영상이나 실시간 스트리밍 환경으로의 확장이 어렵고 연산량이 기하급수적으로 증가하는 문제를 안고 있다.#Review#Video MLLM#I3D-ViT#Adaptive Frame Resolution#Video Understanding#Open Source#Streaming Perception2026년 7월 16일댓글 수 로딩 중
[논문리뷰] VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification본 논문은 현재의 Video MLLM 평가 방식이 답변의 정성적 정확도에만 치중하여 실제적인 시공간적 추론 역량을 제대로 측정하지 못한다는 문제를 지적한다. 기존 벤치마크들은 고득점을 기록하지만, 모델이 정답을 도출하기 위해 필요한 핵심적인 시각적 증거를 정확하게 탐색하고 활용하는지 검증하지 못한다 .#Review#Video MLLM#Spatio-Temporal Grounding#Benchmark#Long-Video Understanding#Evidence Verification#Atomic Ability2026년 4월 2일댓글 수 로딩 중
[논문리뷰] ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video UnderstandingVideo MLLM(Multimodal Large Language Models)이 긴 비디오에서 보이는 Semantic Aggregation Hallucination (SAH) 문제를 해결하는 데 목표를 둡니다.#Review#Long Video Understanding#Hallucination#Semantic Aggregation#Video MLLM#Benchmark#DPO#Positional Encoding#VideoQA2025년 9월 3일댓글 수 로딩 중