[논문리뷰] TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining본 논문은 기존의 비디오 SSL(Self-Supervised Learning) 연구들이 구조, 학습 목표, 데이터 노출 등 다양한 요소를 동시에 변화시킴으로써 실제 어떤 요소가 Motion-Centric한 표현력을 만드는지 불분명하게 만든다는 점을 문제로 지적합니다.#Review#Video Pretraining#Self-Supervised Learning#Motion-Centric#Temporal Decoupling#Diff Compression#Representation Learning2026년 9월 28일댓글 수 로딩 중