[논문리뷰] PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages본 연구는 대규모 언어 모델(LLM) 평가 및 학습 데이터가 영어와 중국어 등 고자원 언어에 과도하게 편향되어 있는 문제를 해결하는 것을 목적으로 합니다.#Review#Multilingual Benchmark#Mathematical Reasoning#Large Language Models#Low-resource Languages#Human-in-the-loop2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning본 논문은 기존 Autoregressive Video-LLM 기반의 Dense Video Captioning 모델들이 겪는 높은 추론 지연(Latency)과 확장성 문제를 해결하고자 합니다.#Review#Dense video captioning#Parallel decoding#Latent planning#Omni-modal#Video-LLM#Dependency restructuring2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding본 논문은 기존의 엄격한 순차적 Autoregressive (AR) 디코딩 방식이 가진 낮은 추론 병렬성과 자원 활용도 문제를 해결하기 위해 고안되었습니다.#Review#Language Model#Autoregressive#Diffusion#Self-Speculation#Parallel Decoding#Inference Efficiency#Tri-Mode Decoding2026년 7월 7일댓글 수 로딩 중
[논문리뷰] MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs본 논문은 최신 MLLMs가 일반적인 인식 및 추론 태스크에서는 높은 성능을 보이나, 예술적 창작 의도를 해석하는 전문 영역에서는 여전히 유의미한 한계를 보인다는 문제의식에서 출발합니다.#Review#Multimodal Large Language Models#Audiovisual Arts#Benchmark#Intent-Level Understanding#Video Essay#Interpretation Plurality2026년 7월 7일댓글 수 로딩 중
[논문리뷰] MentalThink: Shaping Thoughts in Mental SVG World본 논문은 기존의 언어 중심 Multimodal CoT가 가진 시각적 접지(Visual Grounding)의 취약성과 할루시네이션(Hallucination) 문제를 해결하고자 합니다.#Review#Multimodal LLMs#Spatial Reasoning#Scalable Vector Graphics#Chain-of-Thought#Reinforcement Learning#Mental Imagery2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory본 논문은 기존 비디오 에이전트 모델들이 롱폼 비디오를 처리할 때 의존하는 '탐정 스타일'의 반복적 추론(Iterative Reasoning)이 초래하는 과도한 비용과 레이턴시 문제를 해결하고자 합니다 .#Review#Multimodal Long-Term Memory#Agentic Video Understanding#Dual-State Design#Reflexive Response#Retrieval-Augmented Generation#Video-LLM2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment본 논문은 Speech 기반 우울증 탐지 모델이 언어적 경계를 넘어 일반화되지 못하는 한계를 해결하고자 합니다.#Review#Cross-lingual Depression Detection#Supervised Contrastive Alignment#WavLM#Speaker-identity Leakage#Layer-wise Analysis#CLeaD2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator본 논문은 Embodied Navigation 학습을 위한 대규모의 고품질 물리 기반 대화형 시뮬레이션 환경이 부족하다는 문제점을 해결하고자 합니다. 기존 연구들은 실제 스캔 데이터와 합성 데이터 사이의 trade-off, 즉 시각적 충실도와 확장성 사이의 한계에 직면해 있습니다 .#Review#Embodied Navigation#Neural Simulator#3D Gaussian Splatting#Pixel Flow#Vision-Language Navigation#Sim-to-Real2026년 7월 7일댓글 수 로딩 중
[논문리뷰] HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better본 논문은 OCR 특화 VLM이 단순한 문서 파싱 도구를 넘어 더 넓은 영역을 커버하고 실제 배포 환경에서 더 빠른 성능을 내야 한다는 필요성에 착안했습니다.#Review#OCR#Vision-Language Model#DFlash#Agentic Data Flow#Speculative Decoding#Document Parsing#Inference Acceleration2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling본 논문은 LLM의 long-context 확장을 저해하는 quadratic computation cost와 length extrapolation 성능 저하 문제를 해결하기 위해, 기존 chunk-wise sparse attention 방식이 갖는 불완전한 chunk 선택 메커니즘을 개선하고자 합니다.#Review#Large Language Models#Long Context Modeling#Sparse Attention#Hierarchical Attention#Chunk-wise Attention#End-to-end Learning2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Gemma 4 Technical Report본 논문은 최신 LLM 생태계에서 요구되는 강력한 multimodal 이해도, 복잡한 추론 능력, 그리고 컴퓨팅 효율성을 동시에 달성하기 위해 Gemma 4 모델 제품군을 제안합니다.#Review#Multimodal#Mixture-of-Experts#Reasoning Trace#Speculative Decoding#Quantization-Aware Training#Long-context#Encoder-free2026년 7월 7일댓글 수 로딩 중
[논문리뷰] From Foundation to Application: Improving VLA Models in Practice본 논문은 기존의 VLA foundation model들이 실험실 환경의 벤치마크에서는 뛰어난 성능을 보이지만, 실제 로봇 환경의 다양한 하드웨어 구성과 복잡한 작업 조건에서는 여전히 한계가 있다는 문제 의식에서 출발합니다.#Review#Vision-Language-Action (VLA)#Mixture-of-Experts (MoE)#Embodiment Generalization#Dual-Query Distillation#Robotic Manipulation#Spatiotemporal Reasoning2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model기존의 비디오 생성 모델은 Bidirectional diffusion과 Autoregressive 모델이라는 두 개의 분리된 패러다임으로 나뉘어 있어, 각각의 장단점이 뚜렷하다는 한계가 있습니다.#Review#Video Diffusion Models#Autoregressive Generation#Bidirectional Generation#Flexible Chunking#Denoising Timesteps#KV Caching#Any-order Editing2026년 7월 7일댓글 수 로딩 중
[논문리뷰] DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation본 논문은 기존 Speculative Decoding 방식이 가진 병렬 생성의 품질 저하와 비효율적인 검증 문제를 해결하기 위해 DSpark를 제안한다. 기존의 Parallel drafter는 토큰 간 의존성을 모델링하지 못해 뒤로 갈수록 수용률이 떨어지는 Suffix Decay 문제를 겪는다.#Review#Speculative Decoding#Semi-Autoregressive Generation#Confidence-Scheduled Verification#Hardware-Aware Scheduler#LLM Inference Acceleration#Throughput Optimization2026년 7월 7일댓글 수 로딩 중
[논문리뷰] CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration본 논문은 복잡한 이미지 생성 및 편집 워크플로우를 수행하는 멀티모달 에이전트의 한계를 해결하기 위해 CanvasAgent를 제안한다.#Review#Multimodal Agents#Image Creation#Tool Orchestration#Reinforcement Learning#Hybrid Reward#Trajectory Optimization2026년 7월 7일댓글 수 로딩 중
[논문리뷰] CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation본 연구는 기존 ego-centric 3D 생성 모델들이 시점 변화에 따른 심각한 Consistency 저하 및 기하학적 왜곡 문제를 겪고 있다는 점을 해결하고자 한다.#Review#3D Scene Generation#Gaussian Splatting#Ego-centric#Consistency#Geometry#Generative Modeling2026년 7월 7일댓글 수 로딩 중
[논문리뷰] Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing본 논문은 현대 학술 연구 과정이 여러 도구로 파편화되어 있어 발생하는 과도한 컨텍스트 전환과 비효율 문제를 해결하고자 한다.#Review#Academic Writing#Agentic Platform#LaTeX#Toolchain Compression#Retrieval-Augmented Generation#Scholarly Infrastructure2026년 7월 7일댓글 수 로딩 중
[논문리뷰] AlayaWorld: Long-Horizon and Playable Video World Generation본 논문은 노동 집약적인 기존 게임 개발 파이프라인의 한계를 극복하고, 확장성과 적응성이 뛰어난 상호작용 가능한 가상 세계를 생성하는 Generative World Models의 기반을 마련하고자 합니다.#Review#Generative World Models#Interactive Video Generation#Long-Horizon Generation#Spatial Memory#Camera Control#Open-ended Action#Prompt-switching2026년 7월 7일댓글 수 로딩 중
[논문리뷰] 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance본 논문은 기존의 Hierarchical VLA 모델들이 직면한 2D 계획과 3D 실행 사이의 표현적 불일치(Representational Misalignment) 문제를 해결합니다.#Review#Vision-Language-Action Models#3D Trajectory Guidance#Hierarchical Robotics#Metric Depth#Point Cloud Policy2026년 7월 7일댓글 수 로딩 중
[vllm] [vLLM 성능 최적화] Kimi-K2.5/K2.6 이미지 전처리 10배 가속화: Numba와 퓨전 기법 활용vLLM에서 Kimi-K2.5/K2.6 모델의 이미지 전처리를 Numba와 룩업 테이블로 최대 10배 최적화한 사례를 분석합니다.#vLLM#성능 최적화#Numba#이미지 전처리#Kimi-K2.5#Python#Deep Learning2026년 7월 6일댓글 수 로딩 중