본문으로 건너뛰기

[논문리뷰] LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

링크: 논문 PDF로 바로 열기

저자: Simon P. Villani

1. Key Terms & Definitions

  • Hybrid Language Models: Full-attention layers와 recurrent 메커니즘(예: Gated DeltaNet (GDN) layers)을 결합하여 사용하는 언어 모델을 지칭합니다.
  • KV Cache: Transformer attention 메커니즘에서 사용되는 Key 및 Value 상태를 의미하며, 이는 토큰 시퀀스의 contextual 정보를 저장합니다.
  • Gated DeltaNet (GDN): Hybrid LLM 내에서 recurrent matrices와 convolution history를 유지하는 persistent state를 갖는 특정 recurrent layer type입니다.
  • Persistent State: Hybrid LLM의 완전한 inference state를 구성하는 요소로, attention KV뿐만 아니라 GDN recurrent matrices 및 convolution history를 모두 포함합니다.
  • Negative Log-Likelihood (NLL): 관측된 다음 토큰에 할당된 평균 negative log-probability를 나타내는 성능 지표입니다. NLL 값이 낮을수록 예측 성능이 우수함을 의미합니다.

2. Motivation & Problem Statement

기존 언어 모델(LLM) 스위칭 방식은 수신 모델이 historical prefix를 재인식하고 inference state를 재구축해야 하는 비효율성을 내포한다. 기존의 cross-model transfer 연구들은 주로 full-attention models의 KV cache translation에 집중해 왔으며, Gated DeltaNet (GDN) recurrent matrices와 convolution history를 포함하는 hybrid models의 persistent memory 전송 문제는 충분히 다루지 않았다. 이러한 한계는 source model이 생성한 중요한 persistent state가 target model의 상이한 내부 projection 및 gate 구조로 인해 효율적으로 재활용되지 못함으로써, hybrid 모델 간의 성능 gap을 야기하고 새로운 접근 방식의 필요성을 증대시킨다. [Figure 1]

3. Method & Key Results

본 연구는 Qwen3.5 4B Base 모델에서 9B Base 모델로의 Hybrid-State Handoff 메커니즘을 제안하며, 이는 zero target prefix replay를 통해 이루어진다. 제안하는 방법론은 translated Attention KV와 함께 Gated DeltaNet (GDN) persistent-state package의 직접적인 재사용을 포함한다. Experiment 1에서, 고정된 translated KV에 GDN persistent-state package를 추가함으로써 NLL이 0.7473 nats/token 감소했으며 (95% CI [0.6921, 0.8047]), 64개 PG19 문서 전체에서 개선을 보였다. 이는 KV-only transfer 대비 79.7%의 excess NLL 감소를 의미한다. Experiment 2에서는 GDN 컴포넌트에 대한 direct reuse가 learned maps보다 우수함을 확인하였고, 이 결과를 바탕으로 translated KV와 direct recurrent 및 convolution state의 조합 (TDD)을 기본 handoff 메커니즘으로 선정하였다. [Figure 4] 여기에 추가적인 rank-4 correction (434,176 trainable parameters)을 적용한 결과, 64개의 FineWeb-Edu 문서에서 corrected 9B handoff는 continued 4B inference 대비 0.0521 nats/token (95% CI [0.0185, 0.0843]) 더 낮은 NLL을 달성하였다. 이 corrected handoff는 native 9B의 prefix-derived NLL 개선 중 91.8%를 회복하는 Native Context Recovery (NCR)를 보였으며, Jensen–Shannon (JS) divergence는 0.022로 native 9B에 근접한 출력 분포를 나타냈다. [Figure 2]

4. Conclusion & Impact

본 연구는 KV cache가 hybrid LLM의 전체 transferable inference state가 아니며, persistent GDN state가 source-specific 정보를 전달하여 receiver가 historical prefix를 다시 읽지 않고도 활용할 수 있음을 입증하였다. 이는 differently sized hybrid language models 간에 target prefix replay 없이 persistent recurrent inference state를 cross-model handoff하는 최초의 시연으로, 기존 연구의 한계를 넘어선다. Direct recurrent and convolution reuse의 예상치 못한 강점은 persistent-state coordinates의 부분적인 functional compatibility를 시사하며, 이는 향후 모델 패밀리를 persistent-state compatibility를 위해 의도적으로 훈련하는 설계 방향을 제안한다. 이 연구는 LLM의 효율적인 cross-model switching 가능성을 보여주며, 특히 hybrid architectures에서의 inference latency 및 computational overhead 감소에 기여할 잠재력이 크다.

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글