[논문리뷰] SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image본 연구는 단일 이미지 기반의 3D 객체 복원 과정에서 발생하는 부품 간 물리적 정합성 부족 문제를 해결하는 데 집중합니다. 기존의 3D reconstruction 방식은 시각적으로는 그럴듯한 결과를 생성하지만, 실제 물리적인 조립이 불가능한 부품 상태로 파편화되거나 공간적인 배치 오류를 범하는 경우가 많습니다.#Review#3D Part Assembly#Physically Grounded#Single-Image Reconstruction#Part Segmentation#Geometric Constraints2026년 9월 13일댓글 수 로딩 중
[논문리뷰] SAS: Simple Attention Sparsification via End-to-End Optimization of Context RankingLong-context inference는 LLM에서 autoregressive generation 시 모든 이전 context token에 대한 dense attention을 요구하며, 이로 인해 cumulative attention cost가 context length에 따라 quadratically 증가하는 심각한 efficiency bottleneck을 초래합니다.#Review#Attention Sparsification#Context Ranking#End-to-End Optimization#Language Modeling Loss#Gated Attention#Long-Context LLMs#Triton Kernel2026년 9월 13일댓글 수 로딩 중
[논문리뷰] PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization본 논문은 기존 DPO가 모든 관측된 preference label을 완벽하게 신뢰할 수 있는 ground truth로 가정함으로써 발생하는 오염 문제를 해결하고자 한다.#Review#Direct Preference Optimization#Label Noise#Preference Alignment#Posterior Label Correction#Robust Optimization2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Online Learning with LLM Experts from Limited FeedbackLarge Language Models (LLMs)는 다양한 태스크에서 널리 활용되고 있지만, 모델마다 비용과 능력이 상이하며, 특정 모델이 모든 태스크에서 다른 모델을 압도하지 못하는 경우가 많습니다.#Review#Online Learning#LLM Experts#Limited Feedback#Contextual Bandit#Regret Minimization#Adaptive Routing2026년 9월 13일댓글 수 로딩 중
[논문리뷰] How Far Can Synthetic Data Take Thai OCR?본 논문은 Synthetic Data가 Thai OCR의 성능을 어디까지 향상시킬 수 있는지 탐구합니다. OCR은 문서 이미지를 기계 판독 가능한 텍스트로 변환하는 핵심 기술이지만, Thai어와 같이 리소스가 부족한 언어는 풍부한 문서에도 불구하고 신뢰할 수 있는 OCR label이 부족한 문제가 있습니다.#Review#Synthetic Data#Thai OCR#Document Reconstruction#Character Error Rate#Vision-Language Models#Handwriting Recognition#PaddleOCR-VL2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models본 논문은 독립적인 연구 팀이 대규모 컴퓨팅 자원 없이도 최첨단(frontier) 수준의 사이버 보안 에이전트 모델을 개발하기 위한 실용적인 방법론을 제안합니다.#Review#Agentic Cyber Capability#Data-centric Framework#Supervised Fine-Tuning#Environment-Grounded Data#Reasoning Fidelity2026년 9월 13일댓글 수 로딩 중
[논문리뷰] DataFlex-RL: An Evaluation Platform for RLVR Data Policies본 논문은 다양한 RLVR 데이터 정책들이 실제로는 일관된 성능 향상을 보장하지 않으며, 기존 연구의 성과들이 과도하게 특정 실험 환경이나 훈련 설정에 의존하고 있다는 문제를 제기합니다.#Review#RLVR#GRPO#Data-Centric Training#Evaluation Platform#Reproducibility2026년 9월 13일댓글 수 로딩 중
[논문리뷰] COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization본 논문은 LLM 에이전트의 skill 최적화 과정에서 발생하는 과도한 평가 비용과 데이터 효율성 문제를 해결하는 데 집중합니다 .#Review#Agent Skills#Contextual Bandit#Evolutionary Optimization#Skill Optimization#Budgeted Optimization#LLM Agents#Procedural Knowledge2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech본 연구는 자원이 제한적인 환경(low-resource)에서 고품질의 Thai TTS를 효율적으로 배포하기 위한 새로운 파이프라인을 구축하고자 합니다.#Review#Thai TTS#Speech Synthesis#Knowledge Distillation#Fixed-Voice#Prosody#Code-switching#Low-resource2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models본 논문은 로봇 파운데이션 모델이 학습 데이터 내에서는 강력한 성능을 보이지만, 카메라 구도나 조명 등 시각적 분포가 변화할 때 성능이 급격히 저하되는 일반화 문제를 해결하고자 합니다.#Review#Robot Foundation Models#Vision-Action Shortcuts#Latent Interface Training#Generalization#Spatial-Goal-Conditioned#VLA#WAM2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents본 논문은 대규모 skill registry에서 LLM agent가 복합적인 작업을 수행할 때 발생하는 skill 선택의 비효율성 문제를 해결합니다. 기존의 skill router는 각 skill을 독립적인 relevance score에 따라 개별적으로 랭킹을 매기는 pointwise 방식을 주로 사용합니다.#Review#LLM Agents#Skill Routing#Determinantal Point Process#Diversity-Aware Selection#Query-Residual Kernel#Information Retrieval2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and EvaluationAI 연구원 및 Large Language Models (LLMs) 개발자들은 관련 Evaluation을 찾아내고, Benchmark Dataset 및 Code를 확보하며, Reported Score 뒤에 숨겨진 Evaluation Settings를 이해하는 데 상당한 어려움을 겪고 있다.#Review#AI Benchmarks#LLM Evaluation#Living Database#Search Engine#Daily Discovery#Score Histories#Benchmark Saturation#Model Reports2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision본 연구는 Proactive wearable assistant의 핵심 과제인 '적절한 시점의 개입 결정' 문제를 해결하고자 합니다. 기존 연구들은 이를 자유 형식의 대화 생성(dialogue generation)으로 접근하여, 개입 여부와 발화 내용이 결합됨으로써 학습 효율이 떨어지는 한계를 보였습니다.#Review#Egocentric video#Proactive assistance#Streaming video-language models#Synthetic data#Agentic annotation#Verbalizer reformulation2026년 9월 13일댓글 수 로딩 중
[논문리뷰] Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model본 논문은 제한된 컴퓨팅 자원 내에서 긴 Egocentric video를 이해하는 효율적인 모델을 구축하는 문제를 해결하고자 합니다. 기존의 강력한 Agentic pipeline은 높은 성능을 보이지만, 추론 시 너무 많은 시간과 파라미터가 요구되어 ≤2B 파라미터 제약이 있는 환경에서는 적용이 불가능합니다.#Review#Egocentric video#Video question answering#Knowledge distillation#Model compression#Dataset bias#EgoLongQA2026년 9월 13일댓글 수 로딩 중
[논문리뷰] ActionSplice: In-Flight Action Editing for Interactive World Models본 논문은 interactive video world model에서 생성 과정 중 발생하는 액션 수정 요구사항에 대해 즉각적인 응답성을 제공하기 위해 ActionSplice를 제안합니다.#Review#Interactive World Models#Diffusion Models#In-Flight Action Editing#Counterfactual State Transport#Action-Conditioned Generation2026년 9월 13일댓글 수 로딩 중
[논문리뷰] X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation본 연구는 Speech LLM의 추론 효율성을 높이기 위해 오디오 인코더의 레이어를 직접 제거할 때 발생하는 성능 저하 문제를 해결하고자 합니다.#Review#Speech LLMs#Audio-Encoder Compression#Cross-Scale Distillation#Progressive Pruning#Transcript-Consistency Filtering#LoRA Finetuning2026년 9월 10일댓글 수 로딩 중
[논문리뷰] World in World: Explore the World with World Models본 논문은 사전 학습된(frozen) Causal Video World Models를 재학습 없이도 자유롭고 정교하게 제어할 수 있는 통합 인터페이스를 제공하고자 합니다 .#Review#World Models#Video Generation#Visual-Evidence Interface#Training-free#Self-Attention#Camera Control2026년 9월 10일댓글 수 로딩 중
[논문리뷰] UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration본 논문은 기존 All-in-One 의료 영상 복원(MedIR) 모델들이 데이터의 이질성(heterogeneity) 모델링에만 치중하고 의료 영상 내의 구조적 동질성(homogeneity)을 간과하고 있다는 문제점을 지적합니다.#Review#Medical Image Restoration#All-in-One#Universal Model#Hierarchical Homogeneity#Hierarchical Heterogeneity#Memory Network#Attention Mechanism2026년 9월 10일댓글 수 로딩 중
[논문리뷰] SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem본 논문은 LVLM(Large Vision-Language Model)이 2D 이미지로부터 3D 구조를 추론하는 공간 지능이 부족하다는 점을 해결하고자 한다. 기존 연구들은 실제 장면(real-scene)에 대한 dense한 기하학적 주석(annotation)에 의존하여 확장성이 낮고 비용이 많이 드는 한계가 있다 .#Review#Spatial Intelligence#LVLM#Block-Stacking#Synthetic Dataset#Spatial Reasoning#Anchor-based Reasoning#Reinforcement Learning2026년 9월 10일댓글 수 로딩 중
[논문리뷰] SenseNova-U1.5: Towards Native Unified Visual Intelligence본 논문은 기존의 복합적 비전 생성 시스템이 가진 아키텍처 분리 문제를 해결하고자 합니다. 대부분의 모델은 Vision Encoder와 VAE를 각각 사용하여 인지 영역과 생성 영역을 분리하는데, 이는 정보의 왜곡과 효율성 저하를 초래합니다.#Review#Native Unified Multimodal Model#Encoder-free#VAE-free#Flow Matching#On-Policy Distillation#Mixture-of-Transformers#Spatial Reconstruction2026년 9월 10일댓글 수 로딩 중