[논문리뷰] Unlocking Lossless Speedups in LLMs via Discrete Diffusion본 논문은 Autoregressive (AR) Large Language Models (LLMs)의 핵심 bottle-neck인 느린 순차적 토큰 생성을 해결하고자 한다. 기존 AR LLMs는 한 번에 하나의 토큰만 생성하므로 추론 시 높은 Latency를 유발하고 GPU utilization을 저하시킨다.#Review#Diffusion-augmented LLMs#Discrete Diffusion#Lossless Acceleration#Ψ-Spec Sampler#Speculative Decoding#Throughput#Tokens-Per-Forward-pass (TPF)#LoRA2026년 9월 7일댓글 수 로딩 중