[논문리뷰] Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training본 연구는 RL Post-Training을 거친 모델들이 기존의 고정된 Draft Model과 호환되지 않아 Speculative Decoding의 성능이 급격히 하락하는 문제를 해결합니다.#Review#Speculative Decoding#RL Post-Training#Draft Model#Co-Training#Long-Context#Inference Latency#Large-Scale2026년 9월 8일댓글 수 로딩 중
[논문리뷰] Verification-Aware Training for Speculative Decoding본 논문은 기존 Draft Model의 학습 방식이 실제 Speculative Decoding의 검증(verification) 과정과 정렬(alignment)되지 않는 문제를 해결한다.#Review#Speculative Decoding#Large Language Models#Verification-Aware Training#Draft Model#Inference Acceleration#Verification Head#Verification-Adaptive Weighting2026년 8월 31일댓글 수 로딩 중
[SGLang] EAGLE: 은닉 상태 기반 드래프트 모델SGLang의 EAGLE 구현을 분석한다. 타겟 모델의 은닉 상태를 활용한 드래프트 생성, 기존 독립 드래프트 모델 대비 정확도 향상, 트리 기반 검증을 코드와 함께 살펴본다.#sglang#EAGLE#Draft Model#Hidden States#Speculative2026년 4월 12일댓글 수 로딩 중
[논문리뷰] ConFu: Contemplate the Future for Better Speculative Sampling본 논문은 기존의 speculative decoding 드래프트 모델들이 현재 prefix에만 의존하여 예측하는 방식 때문에 발생하는 오류 누적 문제 를 해결하고자 합니다.#Review#Speculative Decoding#LLM Inference Acceleration#Draft Model#Future Prediction#Contemplate Tokens#Mixture-of-Experts#Token Acceptance Rate#Speedup Ratio2026년 3월 10일댓글 수 로딩 중
[논문리뷰] AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders본 논문은 대규모 언어 모델(LLM) 추론 속도 향상을 위한 Speculative Decoding (SD) 과정에서 드래프트 모델과 타겟 모델 간의 불일치 문제를 해결하는 것을 목표로 합니다.#Review#Speculative Decoding#Knowledge Distillation#LLM Inference#Model Acceleration#Token Filtering#Draft Model#Acceptance Rate2025년 10월 24일댓글 수 로딩 중