본문으로 건너뛰기

[논문리뷰] Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

링크: 논문 PDF로 바로 열기

The browsing result is quite long. I need to carefully go through it and extract the required information following the specified structure and constraints.

Part 1: Markdown Summary

Metadata:

  • Authors: Julien Siems, Riccardo Grazzi, Korbinian Pöppel, Jaisidh Singh, Arber Zela, Timur Carstensen, Jenia Jitsev, Frank Hutter, Volkan Cevher, Antonio Orvieto, Aaron Klein
  • Keywords: I need to pick 5-8 relevant keywords. From the abstract and introduction, I can see: Complex KDA, Kimi Delta Attention, Linear RNNs, Expressivity, State Tracking, Rotation, Diagonal-plus-rank-one, Language Modeling.

## 1. Key Terms & Definitions

  • Kimi Delta Attention (KDA): 채널별 게이트를 허용하는 Gated delta-rule Linear RNN 모델.
  • Complex KDA (CKDA): KDA의 확장 버전으로, 게이트 엔트리를 [-1, 1] 범위로, Delta-rule 계수 β를 [0, 2] 범위로 확장하여 2D 회전을 구현하고 Expressivity를 강화한다.
  • Diagonal-plus-rank-one (DPR1): State-transition 행렬의 구조로, Diagonal 행렬에 Rank-one correction 항이 추가된 형태. 효율적인 계산과 채널 간 정보 혼합을 가능하게 한다.
  • State Tracking: 시간이 지남에 따라 입력 의존적 업데이트를 구성하여 시스템의 상태를 추적하는 문제. 패리티, 모듈러 덧셈, 일반적인 순열 합성 등을 포함한다.
  • Non-expansiveness: Linear RNN의 State-transition 행렬이 Norm ||A||_2 <= 1을 만족하여, 시간이 지나도 Hidden state의 크기가 발산하지 않도록 하는 특성.

## 2. Motivation & Problem Statement Linear RNN은 시퀀스 길이에 선형적으로 확장되는 효율적인 시퀀스 모델링이 가능하지만, 낮은 Rank correction을 포함하는 선형 업데이트로 인해 Expressivity가 제한된다. 기존 연구에서는 두 개의 Delta-rule transition을 단일 Recurrent update에 결합하여 2D Rotation을 모델링할 수 있었으나, 이는 단일 transition 대비 Rank와 연산 비용을 증가시켰다. GDN (Gated DeltaNet)의 Scalar gate는 Symmetry를 깰 수 없어 Real Spectrum에 머무는 반면, KDA의 Channel-wise gate는 Householder-diagonal transition의 Symmetry를 깨고 Complex eigenvalues를 가능하게 하지만, 표준 Nonnegative gate는 여전히 Real spectrum으로 제한된다. 저자들은 기존 KDA의 제한된 Expressivity를 극복하고, 단일 Recurrent update 내에서 효율적인 2D Rotation을 구현하기 위한 새로운 접근 방식의 필요성을 제기한다.

## 3. Method & Key Results 본 논문은 KDA의 Channel-wise gate와 β 계수 범위를 확장하여 Complex KDA (CKDA)를 제안한다. CKDA는 게이트 엔트리를 [-1, 1]로, β를 [0, 2]로 확장하여, 단일 Householder transformation과 Channel-wise gate를 결합함으로써 2D Rotation을 구현한다. 이는 Negative gate entries가 Coordinate reflection을 제공하고, β=2의 Householder reflection과 결합하여 Planar rotation을 형성하는 원리이다 [Figure 1, 2]. CKDA는 Diagonal-plus-rank-one 구조와 Non-expansiveness를 유지하면서도, Orthogonal DPR1 행렬을 정확히 실현할 수 있음을 이론적으로 증명한다 [cite: 1, Theorem 1].

핵심 결과로, 단일 CKDA Layer는 SO(3)의 모든 Finite subgroup (예: S3, S4, A5 [cite: 1, Theorem 3])을 추적할 수 있으며, 이는 다른 Diagonal-plus-rank-one Linear RNN (예: Gated DeltaNet)에 비해 요구되는 Layer 수를 하나 줄이는 효과를 가져온다 [cite: 1, Table 1]. 특히, S3 및 S4에서 기존 KDA 설정 대비 가장 강력한 Length extrapolation 성능을 보였고, Periodic waveform continuation task에서도 다른 KDA 변형 및 Transformer를 능가하는 정확도를 달성했다 [cite: 1, Figure 6, 8]. 언어 모델링(1.3B parameters, 100B tokens)에서 CKDA는 Transformer 및 다른 Linear RNN보다 우수한 성능을 보이며, KDA baseline과 유사한 결과를 달성했고, 유망한 Scaling behavior를 나타냈다 [cite: 1, Table 2, Figure 9]. 구현 측면에서, Signed gate를 기존 KDA recurrence kernels에 효율적으로 통합하여 표준 KDA Throughput의 96-97%를 유지한다 [cite: 1, Figure 5].

## 4. Conclusion & Impact 본 연구는 Signed Channel-wise gate와 β∈[0, 2] 범위 확장이 단일 Non-expansive diagonal-plus-rank-one KDA transition 내에서 Planar rotation을 가능하게 함을 보였다. Complex KDA (CKDA)는 이러한 형태의 모든 Orthogonal matrix를 포괄하며, State-tracking Expressivity를 강화한다. 실험적으로, CKDA는 S3, S4 및 Periodic waveforms에서 가장 뛰어난 Length extrapolation 능력을 입증했으며, 1.3B Parameters 모델에서 유사한 Language modeling Downstream accuracy와 Baseline에 가까운 Throughput을 달성했다. 학습된 Transition들은 State tracking 및 Language modeling 모두에서 이론적으로 예측된 메커니즘을 나타낸다 [cite: 1, Figure 7, 10]. 이 연구는 Linear RNN의 Expressivity 한계를 극복하고 장거리 의존성 학습 및 일반화 능력을 향상시키는 데 중요한 기여를 한다.


Part 2: 중요 Figure 정보 (JSON)

I need to find image URLs and captions. Looking for <img> tags in the browsed content.

  • Figure 1: Visualization of CKDA applying a rotation to a vector p within a single recurrent update (I - beta k k^T) Diag(-1,1) by combining a signed diagonal gate with a Householder reflection (beta=2). k=(1,1)^T/sqrt(2) allows a 90 degree rotation, while varying k allows any planar rotation angle.
    • Image URL: The browsed content for Figure 1 does not contain an <img> tag or direct image link. It only has MathML-like symbols. Let me re-check.
    • Correction: I need to check the HTML source for <img> tags, not just the text content. The browsing result shows Figure 1: Visualization of CKDA applying a rotation to a vector p .... This implies it is a figure.
    • Checking the structure: The browsing result contains a list of sections. I'll search for <img> tags in the raw content.
    • Figure 1 does not have a direct image source, it's a visualization described with math and text. This suggests it might be directly rendered in HTML or a complex combination of elements, not a simple <img> tag.
    • Figure 2 is also described with text.
    • Figure 3: Figure 3: Intuition for Theorem 1 in 2D. A single reflection aligns a 2D orthogonal frame with the coordinate axes, giving H A=S and hence A=H S. No direct <img> tag is provided in the browsed content for Figure 3 either.
    • Figure 4: Figure 4: Example of the S3S_{3} permutation task, analogous to the shell game with 3 hidden objects. Again, no <img> tag.
    • Figure 5: Figure 5: Forward–backward kernel throughput on an H100 in BF16, with 16 heads, dk=dv=128d_{k}=d_{v}=128, and 32k tokens per step. DeltaProduct_2_ uses no forget gate. Implementation and timing details are in Appendix E. This seems like a graph. Still no <img> tag.
    • Figure 6: Figure 6: One-layer KDA variants on the S3S_{3}, S4S_{4}, and A5A_{5} word problems, by the eigenvalue range allowed for the diagonal and for the Householder component. Combining a signed gate with a reflection-enabled Householder update gives the strongest extrapolation on S3S_{3} and S4S_{4} among the tested KDA range settings; S4S_{4} accuracy still declines at long lengths. The A5A_{5} result uses the separate theory-initialized setup described in the text (see Figure 16 for baseline results). No <img> tag.
    • Figure 7: ![Refer to caption](2609.24797v1/s3_signed_kda_interpretability.png) - FOUND ONE!
      • figure_id: "Figure 7"
      • image_url: https://arxiv.org/html/2609.24797v1/s3_signed_kda_interpretability.png
      • caption_kr: "학습된 CKDA S3 S3 워드 문제 전환"
    • Figure 8: Figure 8: Periodic waveform continuation with single-layer models. (a) Predictions shortly after the half-bar cue and far beyond the training horizon; gold marks the cue, grey the target, and dark blue the prediction. (b) Waveform MSE versus sequence length; the dashed line marks the maximum training length (136). CKDA extrapolates accurately where the other KDA variants and causal Transformer fail, while the GRU achieves the lowest error. No <img> tag.
    • Figure 9: Figure 9: Advantage of CKDA recurrent and hybrid variants against the Transformer baseline (including QK-Norm, similar to [Ajroldi et al. (2026)]) in nats across scales. Grey area is the approximate noise floor. See Appendix H for a detailed scaling analysis of these numbers. No <img> tag.
    • Figure 10: ![Refer to caption](2609.24797v1/gate_spectrum_three_panel.png) - FOUND ANOTHER ONE!
      • figure_id: "Figure 10"
      • image_url: https://arxiv.org/html/2609.24797v1/gate_spectrum_three_panel.png
      • caption_kr: "언어 모델링 학습 중 확장된 범위 사용"
    • Figure 11: Figure 11: Extension of Figure 2 varying both beta and alpha. The figure demonstrates that both components need to be extended from their standard ranges to obtain complex eigenvalues. No <img> tag.
    • Figure 12: Figure 12: Smallest symmetry-preserving rotations (excluding half-turns) for the cube, grouped by axis. The red point marks the axis viewed end-on; visible edges are solid and hidden edges are dashed. No <img> tag.
    • Figure 13: Figure 13: Smallest symmetry-preserving rotations (excluding half-turns) for the icosahedron, grouped by axis. The red point marks the axis viewed end-on; visible edges are solid and hidden edges are dashed. No <img> tag.
    • Figure 14: Figure 14: Many-to-one tracking, illustrated by h=(123)h=(123), k=(124)k=(124), and ℓ=(k h)−1=(14)(23)\ell=(kh)^{-1}=(14)(23). The cups show identities at successive positions; object 55 is fixed and omitted. Each box groups the two labels F(Qi)=(Li,Ri)F({\bm{Q}}\_{i})=(L\_{i},R\_{i}) of one hidden matrix, where Li=f(Qi)L\_{i}=f({\bm{Q}}\_{i}) and Ri=f−(Qi)R\_{i}=f\_{-}({\bm{Q}}\_{i}). An input a gives Q↦g(a)Q{\bm{Q}}\mapsto g(a){\bm{Q}} and (L,R)↦(aL,a−1R)(L,R)\mapsto(aL,a^{-1}R). The cups and LL return to their initial values, but R3=(12)(34)≠R0=eR\_{3}=(12)(34)\neq R\_{0}=e; hence Q3≠Q0{\bm{Q}}\_{3}\neq{\bm{Q}}\_{0}, although both decode to ee. No <img> tag.
    • Figure 15: Figure 15: Standard and spread initializations for the signed extended-range layer. No <img> tag.
    • Figure 16: Figure 16: Baseline results for Figure 6 ![Refer to caption](2609.24797v1/s3_signed_kda_interpretability_baselines.png) - FOUND A THIRD ONE!
      • figure_id: "Figure 16"
      • image_url: https://arxiv.org/html/2609.24797v1/s3_signed_kda_interpretability_baselines.png
      • caption_kr: "S3, S4, A5 워드 문제 베이스라인 결과"

It's odd that most figures don't have <img> tags directly. It might be due to the ar5iv rendering. However, I found three with explicit <img> tags, so I will use them.

I need to make sure I cite Figure 7, 10, and 16 at least once in the main body.

  • Figure 7 is cited in Section 6, Figure 7.
  • Figure 10 is cited in Section 6, Figure 10.
  • Figure 16 is cited in Section 6, Figure 16 and in the caption for Figure 6. I need to make sure I cite it explicitly in the text. I'll add a citation in the Method & Key Results or Conclusion section if not already present.

Re-checking the citations:

  • Figure 7 is cited implicitly in "A successful CKDA head learns β≈2, nearly sign-valued gates, and complex-conjugate eigenvalues near the unit circle (Figure 7), recovering the mechanism of Sections 3 and 2."
  • Figure 10 is cited implicitly in "To test whether the mechanism of Section 3 also emerges in language modeling, we track the fraction of negative gate entries, fraction of β>1, and of complex transitions in every layer of CKDA during training of the 1.3B parameter models (Figure 10)."
  • Figure 16 is cited implicitly in "CKDA (𝜶∈[−1,1]n, β∈) extrapolates well on S3S3 and S4S4 while other settings of KDA fail (Figure 6; Figure 16 in Appendix I), consistent with Theorem 3."

These are all good. I can now compile the output.저자: Julien Siems, Riccardo Grazzi, Korbinian Pöppel, et al.

1. Key Terms & Definitions

  • Kimi Delta Attention (KDA): 채널별(channel-wise) 게이트를 허용하는 Gated delta-rule Linear RNN 모델로, 시퀀스 모델링에서 효율적인 계산과 채널 간 정보 혼합(channel mixing)을 제공한다.
  • Complex KDA (CKDA): KDA의 확장된 버전으로, 게이트 엔트리를 [-1, 1] 범위로, Delta-rule 계수 β를 [0, 2] 범위로 확장하여 단일 Recurrent update 내에서 2D Rotation을 구현하고 모델의 Expressivity를 강화한다.
  • Diagonal-plus-rank-one (DPR1): State-transition 행렬의 구조 중 하나로, Diagonal 행렬에 Rank-one correction 항이 추가된 형태이다. 이는 Householder reflection을 통해 특정 변환을 가능하게 한다.
  • State Tracking: 시간이 지남에 따라 입력 의존적 업데이트를 구성하여 시스템의 Hidden state를 추적하는 문제이다. 이는 패리티(parity), 모듈러 덧셈(modular addition), 순열 합성(permutation composition)과 같은 복잡한 연산을 포함한다.
  • Non-expansiveness: Linear RNN의 State-transition 행렬 A가 Norm ||A||_2 <= 1을 만족하는 특성으로, Hidden state의 크기가 무한히 발산하지 않아 모델의 안정성을 보장한다.

2. Motivation & Problem Statement

Linear RNN은 시퀀스 길이에 선형적으로 스케일링되는(linear scaling) 효율적인 시퀀스 모델링(sequence modeling)이 가능하지만, 낮은 Rank correction을 포함하는 선형 업데이트(linear updates)로 인해 Expressivity가 제한된다. 기존 연구에서는 두 개의 Delta-rule transition을 단일 Recurrent update에 결합하여 2D Rotation을 모델링할 수 있었으나, 이는 단일 transition 대비 Rank와 연산 비용(computational cost)을 증가시켰다. GDN (Gated DeltaNet)의 Scalar gate는 Symmetry를 깰 수 없어 Real Spectrum에 머무는 반면, KDA의 Channel-wise gate는 Householder-diagonal transition의 Symmetry를 깨고 Complex eigenvalues를 가능하게 하지만, 표준 Nonnegative gate는 여전히 Real spectrum으로 제한되는 한계가 있었다. 저자들은 이러한 기존 KDA의 제한된 Expressivity를 극복하고, 단일 Recurrent update 내에서 효율적이고 확장된 2D Rotation을 구현하기 위한 새로운 접근 방식의 필요성을 제기한다.

3. Method & Key Results

본 논문은 KDA의 Channel-wise gate와 β 계수 범위를 확장하여 Complex KDA (CKDA)를 제안한다. CKDA는 게이트 엔트리를 [-1, 1]로, β를 [0, 2]로 확장하여, 단일 Householder transformation과 Channel-wise gate를 결합함으로써 2D Rotation을 구현한다. 이는 Negative gate entries가 Coordinate reflection을 제공하고, β=2의 Householder reflection과 결합하여 Planar rotation을 형성하는 메커니즘을 따른다 [cite: 1, Figure 1, 2]. CKDA는 Diagonal-plus-rank-one 구조와 Non-expansiveness를 유지하면서도, 모든 Orthogonal DPR1 행렬(matrix)을 정확히 실현할 수 있음을 이론적으로 증명한다 [cite: 1, Theorem 1].

핵심 결과로, 단일 CKDA Layer는 SO(3)의 모든 Finite subgroup (예: S3, S4, A5)을 추적할 수 있으며 [cite: 1, Theorem 3], 이는 다른 Diagonal-plus-rank-one Linear RNN (예: Gated DeltaNet)에 비해 요구되는 Layer 수를 하나 줄이는 효과를 가져온다 [cite: 1, Table 1]. 특히, S3 및 S4 그룹 워드 문제(group word problems)에서 기존 KDA 설정 대비 가장 강력한 Length extrapolation 성능을 보였고 [cite: 1, Figure 6, Figure 16], Periodic waveform continuation task에서도 다른 KDA 변형 및 Transformer를 능가하는 정확도를 달성했다 [cite: 1, Figure 8]. 언어 모델링(1.3B parameters, 100B tokens)에서 CKDA는 Transformer 및 다른 Linear RNN보다 우수한 성능을 보이며, KDA baseline과 유사한 결과를 달성했고, 유망한 Scaling behavior를 나타냈다 [cite: 1, Table 2, Figure 9]. 학습된 Transition들은 β≈2, 거의 Sign-valued gate, 그리고 Unit circle 근처의 Complex-conjugate eigenvalues를 보여주어 이론적 예측 메커니즘을 회복했다 [cite: 1, Figure 7]. 구현 측면에서, Signed gate를 기존 KDA recurrence kernels에 효율적으로 통합하여 표준 KDA Throughput의 96-97%를 유지한다 [cite: 1, Figure 5]. 언어 모델링 학습 중 Negative gate entries, β>1, Complex transitions의 출현이 관찰되어 제안 메커니즘의 학습 가능성이 입증되었다 [cite: 1, Figure 10].

4. Conclusion & Impact

본 연구는 Signed Channel-wise gate와 β∈[0, 2] 범위 확장이 단일 Non-expansive Diagonal-plus-rank-one KDA transition 내에서 Planar rotation을 가능하게 함을 보였다. Complex KDA (CKDA)는 이러한 형태의 모든 Orthogonal matrix를 포괄하며, State-tracking Expressivity를 강화한다. 실험적으로, CKDA는 S3, S4 및 Periodic waveforms에서 가장 뛰어난 Length extrapolation 능력을 입증했으며, 1.3B Parameters 모델에서 유사한 Language modeling Downstream accuracy와 Baseline에 가까운 Throughput을 달성했다. 학습된 Transition들은 State tracking 및 Language modeling 모두에서 이론적으로 예측된 메커니즘을 나타낸다. 이 연구는 Linear RNN의 Expressivity 한계를 극복하고 장거리 의존성 학습 및 일반화 능력을 향상시키는 데 중요한 기여를 하며, DeltaProduct_2와 유사한 Expressivity를 제공하는 동시에 기존 KDA 커널을 재사용하는 효율적인 구현 방안을 제시한다.

Figure 7: S3 워드 문제 학습된 CKDA 전환

Figure 7 — S3 워드 문제 학습된 CKDA 전환

Figure 10: 언어 모델링 학습 중 확장된 범위 사용

Figure 10 — 언어 모델링 학습 중 확장된 범위 사용

⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.

댓글

관련 포스트

Review 의 다른글