[논문리뷰] ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes본 논문은 기존 3D tokenizer들이 latent sequence 길이를 줄였을 때 기하학적 재구성 품질이 급격히 저하되는 문제를 해결하고자 합니다. 기존 방식들은 고정된 토큰 예산 하에 정보를 암묵적으로 분산시키므로, 토큰 수를 극단적으로 줄일 경우 전역 구조와 세부 사항 사이의 정보 병목 현상이 발생합니다.#Review#3D tokenization#3D representation learning#Neural fields#Iterative refinement#Nested dropout#Compact prefix2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations본 논문은 기존 공간적 자산 가격 결정 모델이 사전에 주어진 interaction matrix에 의존하여 경제적 연결성을 충분히 설명하지 못하는 한계를 해결합니다.#Review#Spatial Exposure Adjustment#Barycentric Interaction Fields#Wasserstein Barycentric Reconstruction#Language-Model Representations#Spatial Autoregression2026년 9월 2일댓글 수 로딩 중
[논문리뷰] VibeVoice-ASR-Streaming Technical Report본 논문은 기존 Unified End-to-End 모델들이 오프라인 인식에 국한되어 실시간 음성 비서 환경에서 요구되는 저지연(Low-latency) 요구사항을 충족하지 못하는 문제를 해결합니다.#Review#Streaming ASR#Speaker-Attributed ASR#LLM#End-to-End#Multi-talker Recognition#Interleaved Generation#Voice Assistant2026년 9월 2일댓글 수 로딩 중
[논문리뷰] SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models본 논문은 다양한 이기종 데이터 소스와 서로 다른 아키텍처를 가진 Video generator들을 효과적으로 통합하여 일관된 Interactive video world model을 구축하는 핵심 문제에 주목한다.#Review#Interactive Video World Models#Multi-source Data Engine#Backbone-native Adaptation#Long-horizon Inference#Distribution Matching Distillation#Teacher-forced AnyFlow2026년 9월 2일댓글 수 로딩 중
[논문리뷰] SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions본 논문은 모바일 환경에서 일상적으로 발생하는 이미지 및 텍스트 노이즈가 멀티모달 검색 성능에 미치는 영향을 체계적으로 규명하고자 합니다. 기존 벤치마크들은 주로 노이즈가 없는 깨끗한 데이터셋을 사용하거나, 특정 엔티티에 대한 견고성(Robustness)을 독립적으로 평가하지 못한다는 한계가 있습니다 .#Review#Multimodal Retrieval#Mobile Interaction#Robustness#Benchmark#Snap-and-Ask#Modality Fusion2026년 9월 2일댓글 수 로딩 중
[논문리뷰] S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?본 논문은 LLM 기반 에이전트가 환경과의 상호작용 경험을 통해 실제로 스스로를 개선(Self-Improvement)할 수 있는지에 대한 근본적인 의문을 해결하고자 합니다.#Review#LLM Agents#Self-Improvement#Interactive Benchmark#Self-Judging#History ICL#Summary Memory#Parameter Training2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills본 논문은 자율 연구 에이전트가 수행하는 머신러닝(ML) 연구 과정에서 겪는 도메인 특화 전문 지식 부족 문제를 해결하고자 합니다 .#Review#Autonomous Agents#Operational Knowledge#Skill Distillation#GitHub Repositories#Agentic Systems#AI4AI#DisCo2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Post-Training Language Models for Gold-Medal Performance in Coding Competitions본 논문은 competitive programming에서 AI 시스템이 Gold-Medal 이상의 성능을 달성하기 위한 최적의 파이프라인을 구축하는 것을 목표로 합니다.#Review#Competitive Programming#Supervised Fine-Tuning#Reinforcement Learning#Test-Time Compute#GenCorrect#IOI2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations본 논문은 데이터가 부족하거나 고차원적인 패널 환경에서 cross-asset return covariances를 안정적으로 추정하기 어렵다는 금융권의 고질적인 문제를 해결하고자 합니다.#Review#Portfolio Risk#Certified Diversification#Wasserstein Distance#Distributional Fields#Robust Portfolio Choice#Language-Model Representations2026년 9월 2일댓글 수 로딩 중
[논문리뷰] PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation본 논문은 최신 연구 논문을 저장소 수준의 코드로 변환할 때 발생하는 정보 손실 및 구현 불일치 문제를 해결하기 위해 PaperCompiler를 제안합니다.#Review#Paper-to-Code#Repository-level Generation#Specification Compilation#Constraint-guided Generation#Software Engineering Agents2026년 9월 2일댓글 수 로딩 중
[논문리뷰] On the Design Fundamentals of Pixel Text Representation Learning본 논문은 기존 Pixel-based Text Encoder가 고정된 해상도 Pretraining, Visual Shortcut Learning, 취약한 Multimodal Grounding, 그리고 Multilingual Visual Text Understanding 측면에서 겪는 근본적인 문제들을 해결하고자 합니다.#Review#Pixel Text Representation Learning#Vision-Language Models#Multimodal Grounding#Native-Resolution Encoding#Layout Augmentation#Multilingual Text Understanding#Contrastive Learning#Optical Context Compression2026년 9월 2일댓글 수 로딩 중
[논문리뷰] NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference본 논문은 기존 multimodal 모델들이 지나치게 복잡한 아키텍처를 사용하여 비효율적인 추론 및 fine-tuning 구조를 가진 문제를 해결하고자 합니다.#Review#Multimodal Encoder#Bidirectional Transformer#Visual Document Retrieval#Masked Discrete-Diffusion#Late-Interaction#Efficient Inference#Dynamic Resolution2026년 9월 2일댓글 수 로딩 중
[논문리뷰] MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval본 논문은 현대의 Open-ended Query 검색이 특정 단일 관점에 치우쳐(Single-perspective bias) 다양한 측면의 정보를 제공하지 못하는 문제를 해결하고자 합니다.#Review#Information Retrieval#Multi-modal#Benchmark#Perspective-guided Learning#Retrieval Diversification#Multi-vector Retrieval2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Language Models Can Control Their Own AttentionLong-context Language Model (LLM) Inference 과정에서 Transformer 아키텍처는 매 Decoding Step마다 모든 이전 토큰에 대해 Attention을 계산하므로, KV Cache 메모리 접근 Latency가 전체 Decoding Time의 상당 부분을 차지하며 이는 Long-context Regime에서 특히 심화됩니다.#Review2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Kirin: Animal Motion Generation from In-the-Wild Video본 논문은 야생 동물의 동작 데이터를 확보하기 어려운 환경에서, 실제 동물의 다양하고 생동감 있는 동작을 생성하는 데 한계가 있다는 문제점을 해결하고자 한다 . 기존 연구들은 소규모의 인위적인 데이터나 게임용 데이터에 의존하여 실세계 동물 동작의 다양성을 포착하지 못하는 단점이 있다.#Review#Animal Motion Generation#In-the-Wild Video#3D Motion Reconstruction#Diffusion Models#Text-Image Conditioned Generation#AiM3D Dataset2026년 9월 2일댓글 수 로딩 중
[논문리뷰] It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning본 논문은 현대의 검색 시스템에서 검색(Retrieval) 단계의 오류가 하위 랭킹(Ranking) 단계에서 복구되기 어렵다는 문제점을 지적하며, 효과적인 검색을 위한 정밀도와 재현율의 균형 문제를 해결하고자 합니다.#Review#Generative Retrieval#Reinforcement Learning#Co-Evolution#Keyword Generation#Retrieval F1#LLM2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers본 연구는 역사적 신문 자료의 불규칙한 레이아웃과 낮은 OCR 품질로 인해 학계의 계산적 접근이 제한되는 문제를 해결하고자 합니다. 기존의 전통적인 OCR 엔진은 복잡한 다단 구성이나 장식적 헤딩을 파싱하는 데 한계가 있으며, 대규모 데이터 처리에 필요한 비용 효율성을 확보하기 어렵습니다.#Review#Historical Newspapers#OCR#Vision-Language Models#Pipeline#Data Pipeline#Digital Humanities#Structured Datasets2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation본 논문은 Sampled-Token OPD가 효율적임에도 불구하고 pass@kk 성능이 정체되는 Diversity Distillation Failure 문제를 해결합니다. 기존 방법론들은 무분별하게 엔트로피를 높이거나, 비용이 높은 Forward-KL을 사용하여 효율성을 저해하는 한계가 있습니다.#Review#On-Policy Distillation#Sampled-Token OPD#Diversity Bottleneck#Entropy Influence#First-Order Analysis#Divergence-Adaptive Shrinkage2026년 9월 2일댓글 수 로딩 중
[논문리뷰] Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents본 논문은 전문 분야의 Agent 과제를 설계할 때 발생하는 지식 가용성(Knowledge Availability)과 검증 가능성(Verifiability)의 불균형 문제를 해결하고자 한다.#Review#Synthetic task construction#Knowledge gating#Task validation#Verifiable rewards#Agent evaluation2026년 9월 2일댓글 수 로딩 중
[논문리뷰] HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?본 논문은 Agent가 연구용 프로토타입에서 실제 배포 도구로 진화함에 따라, Agent의 성능을 좌우하는 외부 인프라인 Harness의 중요성을 규명하고 이를 모델이 직접 개발할 수 있는지 검증하고자 합니다.#Review#Agent Harness#LLM#Software Engineering#Benchmark#Agent Evolution#Code Generation2026년 9월 2일댓글 수 로딩 중