[논문리뷰] FilmBench: A Film-Grade Benchmark for Cinematic Video Generation본 논문은 기존 비디오 생성 벤치마크들이 웹 기반의 단순 텍스트 프롬프트를 사용하여 영상의 'Cinematic Quality'를 평가하는 데 한계가 있다는 점을 지적한다.#Review#Text-to-Video#Reference-to-Video#Cinematic Language#FilmBench#Benchmark#Video Generation#FilmOps2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels본 논문은 Visual Document Understanding 분야에서 발생하는 Attribution Hallucination 문제를 해결하기 위해, 현재의 좌표 기반 인터페이스가 모델의 근거 표현 능력을 제한하고 있는지 분석합니다 .#Review#Vision-Language Models#Evidence Attribution#Attribution Hallucination#Multimodal Retrieval#Reinforcement Learning#Doc-VQA2026년 7월 27일댓글 수 로딩 중
[논문리뷰] DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification본 논문은 실주행 데이터에서 운전자의 개별적인 주행 습관을 식별할 때 발생하는 불명확한 환경적 편향(Confounds) 문제를 해결하고자 한다.#Review#Naturalistic Driving#Driving Style Identification#Behavior Prediction#Driver Re-identification#Multimodal Time Series#Shortcut Learning2026년 7월 27일댓글 수 로딩 중
[논문리뷰] DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes본 논문은 VLM의 사전 학습 데이터 구성을 직관에 의존하는 휴리스틱한 방식에서 재현 가능한 엔지니어링 원칙으로 전환하고자 한다. 기존 연구들은 단순히 고품질 데이터를 필터링하거나 데이터셋을 단순히 쌓는 방식을 사용하며, 데이터셋 추가가 성능에 미치는 인과 관계를 명확히 파악하기 어려운 문제가 있다.#Review#Vision-Language Models#Data Mixture Optimization#Convex Optimization#Dataset-Level Assessment#Scalable VLM Training#Attributable Validation2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Data Pyramid for Embodied Manipulation본 논문은 Embodied AI 모델의 성능을 결정짓는 핵심 요소인 '데이터'가 매우 파편화되어 있다는 문제 의식에서 출발합니다. 기존 연구들은 특정 모델의 학습 레시피에 의존하여 데이터를 수집하거나, 아키텍처 중심의 평가에 치중하여 데이터 소스 간의 관계와 trade-off를 시스템화하지 못했습니다.#Review#Embodied AI#Foundation Models#Data Pyramid#Robotics#Manipulation#Multimodal Learning#Robot Learning2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Codifying the Judge: Scalable Evaluation via Program Distillation본 논문은 기존 LLM-as-a-judge 패러다임이 가진 높은 API 비용, 불투명한 의사결정 과정, 시스템적 편향(Bias), 그리고 유연성 부족 문제를 해결하는 것을 목적으로 한다.#Review#LLM-as-a-judge#Program Distillation#Pajama#Weak Supervision#Model Routing#Pareto Frontier#Reward Modeling2026년 7월 27일댓글 수 로딩 중
[논문리뷰] ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding본 논문은 현재의 의료용 MLLM이 직면한 '시각 중심적 지식 흡수'와 '임상적으로 유효한 평가'라는 두 가지 핵심 문제를 해결하고자 합니다.#Review#Multimodal LLM#Medical Imaging#Vision Encoder#CaSL Fusion#Clinical Report Generation#MedIF-Bench#RoI-Grounded Evaluation2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Characterizing Warp Divergence from Pascal to Blackwell본 논문은 ITS 도입 이후 NVIDIA GPU의 warp divergence 처리 방식이 고정되어 있다는 업계의 통념을 검증하는 것을 목표로 합니다.#Review#GPU#Warp Divergence#Independent Thread Scheduling#Reconvergence#SASS#Microarchitecture#Blackwell2026년 7월 27일댓글 수 로딩 중
[논문리뷰] Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling본 논문은 기존의 단일 타겟-단일 상태(single-target, single-state) 중심의 바인더 설계 모델이 다중 상태(Multi-state) 단백질의 구조적 변화나 다중 타겟(Multi-target) 결합 요구를 수용하지 못하는 한계를 해결합니다 .#Review#Cross-context Binder Design#In-context Generation#Mixed Sampling#Diffusion Model#Protein Binder Design#Multi-state Design#Multi-target Design2026년 7월 27일댓글 수 로딩 중
[논문리뷰] A Vocabulary for Multi-Agent Automated Research Systems본 논문은 다양한 멀티 에이전트 기반 자동 연구 시스템들을 일관된 기준으로 기술하고 비교할 수 있는 표준화된 Vocabulary의 부재 문제를 해결합니다.#Review#Multi-Agent Systems#Automated Research#System Design#Trajectory#Proxy Optimization#Evaluation2026년 7월 27일댓글 수 로딩 중
[논문리뷰] A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever본 연구는 기존의 거대 언어 모델이 Retraining을 통해 성능을 개선하는 과정에서 발생하는 비용, 불투명성, 비결정론적 출력 문제를 해결하고자 합니다.#Review#Verified Knowledge#Determinism#Memory-augmented LLM#Zero-token Inference#Byte-exact#Merlin Engine2026년 7월 27일댓글 수 로딩 중
[논문리뷰] VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression본 논문은 VLM이 고해상도 이미지를 처리할 때 발생하는 과도한 시각적 토큰으로 인한 inference latency와 KV cache 메모리 오버헤드 문제를 해결하고자 합니다.#Review#Vision-Language Models#Token Compression#Autoencoder#Intrinsic Encoding#Hierarchical Information#Parameter Sharing2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Three-Body Scattering for Generative Modeling본 논문은 현대 generative models이 adversarial critic, prescribed noise-to-data path(예: diffusion models), 또는 autoregressive factorization에 의존하는 한계를 극복하기 위해 새로운 one-step generation 방법을 제안합니다.#Review#Three-Body Scattering Modeling (TBSM)#Generative Modeling#Energy Distance#One-Step Generation#Tracked Scattering#Wasserstein Gradient Flow#FID#ImageNet2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Spectral Prior for Reducing Exposure Bias in Diffusion Models본 논문은 diffusion model의 iterative sampling 과정에서 발생하는 exposure bias로 인한 예측 오류 누적 문제를 해결하고자 한다.#Review#Diffusion Models#Exposure Bias#Spectral Mismatch#Spectral Alignment (SPA)#Generative Models#Guidance#Power Spectrum#Image Generation2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills본 논문은 LLM의 self-evolution 과정에서 발생하는 task diversity와 verification reliability 사이의 trade-off를 해결하고자 합니다.#Review#Large Language Models#Self-Play#Agent Skills#Reinforcement Learning#Curriculum Learning#Task Generation#Co-Evolution2026년 7월 26일댓글 수 로딩 중
[논문리뷰] SceneActBench: Can Agents Act on the 3D Scenes They See?본 논문은 VLM 에이전트가 3D 장면을 단순히 설명하는 단계를 넘어, 실제 3D 환경에서 다중 객체를 대상으로 동작을 수행할 수 있는지 검증하기 위한 SceneActBench를 제안한다.#Review#Vision-Language Model#3D Agent#Benchmarking#Executable 3D Action#Geometric Reasoning2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Scaling Native Multimodal Pre-Training From Scratch본 논문은 전통적인 late-fusion 방식의 멀티모달 모델이 지닌 vision과 language 표현 학습 간의 불균형(asymmetry) 문제를 해결하고자, native multimodal pre-training의 scaling 법칙을 체계적으로 규명합니다.#Review#Native Multimodal Pre-Training#Compute-Optimal Scaling Laws#Pareto Frontier#Transformer#Mixture-of-Experts#Cross-Modal Transfer#In-Context Learning2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Multimodal Speaker Verification as a Threat to Speaker Anonymization본 논문은 기존의 Speaker Anonymization 연구들이 주로 단일 발화(Single utterance)의 음향 특성 제거에만 집중하고 있어, 실제 다중 발화 환경에서 발생할 수 있는 프라이버시 취약점을 간과하고 있음을 지적합니다.#Review#Speaker Anonymization#Automatic Speaker Verification#Multimodal Fusion#Multi-utterance Aggregation#Voice Privacy#Speaker Representation2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making본 논문은 LLM 기반 에이전트가 모든 과제에 대해 대형 모델을 사용하는 비용 효율성 문제와, 적절한 도구 사용 및 모델 전환을 스스로 결정하지 못하는 자기 인식(self-awareness) 부재 문제를 해결하기 위해 Multi-Head Latent Control을 제안한다.#Review#LLM Agent#Latent Control#Model Routing#Agentic Systems#Inference Efficiency#Tool Use#Self-Awareness2026년 7월 26일댓글 수 로딩 중
[논문리뷰] Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning본 논문은 기존의 대규모 강화학습 프레임워크가 가진 복잡성으로 인해 연구자의 반복적인 알고리즘 개선 과정에 큰 오버헤드가 발생하는 문제를 해결하고자 합니다.#Review#Agentic Reinforcement Learning#PyTorch-native#Scalable Training#LLM#Asynchronous Loop#Token-exact#FSDP22026년 7월 26일댓글 수 로딩 중