[논문리뷰] SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization본 논문은 현대 LLM에서 요구되는 긴 컨텍스트 처리를 위해 Gated DeltaNet의 컨텍스트 윈도우를 확장하는 과정에서의 근본적인 한계를 해결하고자 한다.#Review#Gated DeltaNet#Long-Context Extension#Spectral Reparameterization#Continual Pretraining#Linear Attention#Spectral Dynamics#Information Decay2026년 9월 16일댓글 수 로딩 중
[논문리뷰] Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM본 논문은 Gated DeltaNet (GDN)과 같은 선형 Attention 레이어를 포함하는 Hybrid Large Language Models (LLMs)에 NVFP4 W4A4 4-bit 양자화를 적용하는 과정에서의 근본적인 문제를 다룬다.#Review#NVFP4#W4A4#Gated DeltaNet#LLM Quantization#Hybrid LLM#Recurrent Networks#KV-cache#Post-training Quantization2026년 9월 3일댓글 수 로딩 중
[논문리뷰] On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability본 논문은 대규모 Sparse Mixture-of-Experts (MoE) 모델인 Qwen3.8-Flash-Next를 설계하며, 기존 모델 대비 컴퓨팅 예산을 절감하면서도 동등 이상의 성능을 유지하는 것을 목표로 한다.#Review#Sparse Mixture-of-Experts#Gated DeltaNet#Qwen Sparse Attention#Gated Residual#Muon Optimizer#Long-context Inference2026년 8월 31일댓글 수 로딩 중
[논문리뷰] Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation본 논문은 하이브리드 모델로의 전환 시 발생하는 부적절한 재귀적 파라미터 초기화 문제를 해결하고자 합니다. 기존 연구들은 Transformer의 가중치를 복사하는 데 집중하지만, 새롭게 도입되는 GDN의 동역학(decay, gate 등)을 고려하지 않아 초기 모델이 최적화되지 않은 상태에서 학습을 시작하게 됩니다 .#Review#Hybrid Linear Attention#Gated DeltaNet#Model Distillation#Initialization#Softmax Attention#Knowledge Distillation#Recurrent Dynamics2026년 6월 18일댓글 수 로딩 중
[논문리뷰] InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models본 연구는 기존 VLM의 이차적인 계산 복잡성과 증가하는 KV 캐시로 인한 장기 컨텍스트 이해 능력 및 배포 제약 문제를 해결하는 것을 목표로 합니다. 특히, 선형 어텐션의 정보 집약적 작업에서의 저조한 성능과 윈도우 기반 어텐션의 장기 기억 유지 부족이라는 한계를 극복하고자 합니다.#Review#Vision-Language Models#Linear Attention#Sliding Window Attention#Gated DeltaNet#Long-Context Understanding#Efficiency#Hybrid Architecture#Multimodal Learning2025년 12월 10일댓글 수 로딩 중