[논문리뷰] Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel본 논문은 대규모 언어 모델의 성능 향상을 위해 수행되는 재학습(retraining)과 매번 전체 문맥을 재연산해야 하는 추론 과정의 막대한 비용 문제를 해결하고자 합니다. 기존 방식은 모델 가중치를 변경하거나 매번 동일한 문맥을 반복해서 Prefill하는 비효율적인 자원 소모를 동반합니다.#Review#KV-State Grafting#Byte-Exactness#Inference-time Learning#Flywheel#Galahad#KV Cache#Model Efficiency2026년 7월 16일댓글 수 로딩 중