[논문리뷰] Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM본 논문은 Gated DeltaNet (GDN)과 같은 선형 Attention 레이어를 포함하는 Hybrid Large Language Models (LLMs)에 NVFP4 W4A4 4-bit 양자화를 적용하는 과정에서의 근본적인 문제를 다룬다.#Review#NVFP4#W4A4#Gated DeltaNet#LLM Quantization#Hybrid LLM#Recurrent Networks#KV-cache#Post-training Quantization2026년 9월 3일댓글 수 로딩 중
[논문리뷰] Hybrid Architectures for Language Models: Systematic Analysis and Design Insights기존 대규모 언어 모델(LLM)에서 Transformer 의 quadratic 복잡성과 Mamba 의 장문 컨텍스트 처리 한계를 극복하고자 합니다.#Review#Hybrid LLM#Transformer Architecture#Mamba#State Space Models (SSM)#Computational Efficiency#Long-Context#Language Model Architectures#Scaling Laws2025년 10월 7일댓글 수 로딩 중