[논문리뷰] Parts-of-Speech as Emergent Categories in SAE Latent Space
링크: 논문 PDF로 바로 열기
저자: Alessandro Bondielli, Lucia Passaro, Serena Auriemma, Alessandro Lenci
1. Key Terms & Definitions (핵심 용어 및 정의)
- Sparse AutoEncoders (SAEs): Dense model activations를 고차원의 sparse representations로 분해하여, 각 latent dimension이 더욱 interpretable한 방향의 variations를 포착하도록 설계된 모델입니다.
- Part-of-Speech (PoS): 단어의 문법적 기능(예: 명사, 동사, 형용사)을 나타내는 linguistic category입니다. 본 연구에서는 morpho-syntactic information 분석을 위한 controlled testbed로 활용됩니다.
- Latent Space: SAEs 내부에서 sparse activation vectors가 존재하는 고차원의 추상적인 공간을 의미합니다.
- Monosemanticity: SAE latents가 단일하고 명확하게 해석 가능한 개념이나 feature를 나타내는 이상적인 특성을 지칭합니다.
- Distributed Representations: 언어 정보가 단일한 localization된 단위가 아닌, 여러 차원에 걸쳐 분산된 형태로 인코딩되는 방식을 의미합니다.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
Large Language Models (LLMs)는 lexical, syntactic, semantic 정보를 내부 representations에 encode하지만, 이러한 정보가 어떻게 organize되는지에 대한 명확한 이해는 부족합니다. 특히, linguistic categories가 localized하고 interpretable한 units에 상응하는지, 아니면 representation space의 다양한 dimensions에 걸쳐 distributed되어 있는지 불분명합니다. Sparse AutoEncoders (SAEs)는 LLM의 interpretability를 위한 유망한 도구로 주목받고 있으나, SAE latents와 linguistic categories 간의 관계는 여전히 명확하지 않습니다. 기존 probing methods는 정보의 존재 유무만을 밝힐 뿐, 정보가 latent space 내에서 어떻게 organization되는지에 대한 깊이 있는 설명을 제공하지 못하는 한계가 있습니다. 본 연구는 Part-of-Speech (PoS) categories를 controlled test case로 사용하여, morpho-syntactic information이 개별 latents에 의해 encode되는지, 혹은 structured groups of features에 의해 encode되는지를 체계적으로 분석하여 이러한 interpretability gap을 해소하고자 합니다. 전체 연구 워크플로우는 [Figure 1]에서 개괄적으로 확인할 수 있습니다.

Figure 1 — 제안 방법론의 워크플로우
3. Method & Key Results (제안 방법론 및 핵심 결과)
본 논문은 PoS information의 recoverability, organization, stability를 평가하기 위해 3단계 interpretability pipeline을 제안합니다. 이 pipeline은 LLaMA-3-8B 모델의 Layer 30에서 추출한 SAE activations를 GUM Corpus Treebank 데이터셋과 소규모 controlled dataset에 적용하여 분석합니다.
첫째, RQ1: Recoverability of PoS Information을 평가하기 위해 one-vs-rest binary probing classifiers를 활용하여 PoS 구분이 SAE activations로부터 선형적으로 recoverability한지 테스트합니다. 실험 결과, 대부분의 PoS category에서 높은 F1 scores를 달성하여 morpho-syntactic information이 sparse latent space에 명시적으로 encode되어 있음을 입증합니다. 특히 Closed-class categories (예: CCONJ, DET, PRON)는 Open-class categories (예: NOUN, VERB)보다 높은 F1 score를 기록하며 recoverability가 더 용이했습니다 [Figure 2].

Figure 2 — PoS별 F1 스코어
둘째, RQ2: PoS Organization in Latent Space를 분석하기 위해 feature-salience, coverage, compactness analyses를 수행합니다. 각 PoS category에 대해 L1-regularized logistic regression classifier의 계수를 기반으로 salient latents를 식별하고, 95% coverage에 도달하는 데 필요한 latents의 수(kc⋆k^{\star}_{c})를 정량화합니다. 이 kc⋆k^{\star}_{c} 값은 PoS category별로 상당한 차이를 보이며, Closed-class categories는 더 적은 latents로 compact한 activation patterns을 보이는 반면, Open-class categories는 더 넓은 latent groups를 필요로 함을 나타냅니다 [Figure 4]. 이는 PoS 정보가 개별 monosemantic latents가 아닌, structured groups of latents 수준에서 localized된다는 것을 시사합니다. 또한, 이러한 분석을 통해 선별된 compact feature set(총 498개의 latents)으로 훈련된 multi-class PoS classifier는 full SAE representation으로 훈련된 classifier와 유사한 0.89 Accuracy 및 0.78 Macro F1 성능을 달성하여, 선택된 latents가 대부분의 PoS discrimination 정보를 보존함을 보여줍니다.

Figure 4 — PoS별 95% 커버리지 Latent 수
셋째, RQ3: Validation on Held-Out Data를 통해 식별된 latent groups의 stability와 systematicity를 평가합니다. Held-out test split과 controlled dataset 모두에서, salient latents는 target category의 unseen instances에 대해 일관되게 active한 상태를 유지하여 stability를 확인합니다. Co-activation analysis는 높은 recall (대부분 ≥0.95)을 보여주면서도 categories 간의 중복(overlap)이 존재함을 나타냅니다 [Figure 6]. 이는 latent groups가 부분적으로 category-specific하지만, 완전히 배타적이지는 않다는 점을 시사하며, 특히 Open-class categories에서 더 높은 spurious co-activation rates를 보였습니다.
4. Conclusion & Impact (결론 및 시사점)
본 연구는 Part-of-Speech (PoS) categories가 Sparse AutoEncoders (SAEs)의 latent space에서 localizable하지만 distributed features의 emergent sets으로 internal하게 represent된다는 것을 입증합니다. PoS distinctions은 sparse activations에서 일관되게 recoverability하며, 각 category는 그 linguistic nature에 따라 크기가 상이한 compact group of sparse features에 의해 support됩니다. Closed-class PoS는 더 적은 active latents를 가지는 반면, Open-class PoS는 더 확산된(diffuse) 패턴을 보입니다. 이러한 category-specific latents들의 작은 union은 강력한 multi-class classification performance를 유지하며, 식별된 groups는 held-out treebank data에서 안정적이고 controlled examples에서는 대체로 additive한 특성을 나타냅니다.
본 연구 결과는 latent level에서의 interpretability 주장이 단순히 top-activating examples에 의존하기보다는, theoretically grounded category inventories에 대해 엄격하게 평가되어야 함을 시사합니다. 동시에, cross-category co-activations은 식별된 latents가 lexical, positional, 그리고 annotation-driven regularities를 부분적으로 추적한다는 것을 보여주며, 이는 향후 연구에서 더 풍부한 linguistic levels과 typologically diverse languages에 대한 분석의 필요성을 제기합니다. 본 연구는 LLM의 내부 workings에 대한 이해를 심화하고, interpretability 연구의 방법론적 토대를 강화하는 데 중요한 기여를 합니다.
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
- [논문리뷰] NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
- [논문리뷰] V-RAE: Rethinking Video Latent Spaces for Generation
- [논문리뷰] Beyond Pixels: From Video Priors to 4D Worlds
- [논문리뷰] GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
Review 의 다른글
- 이전글 [논문리뷰] PUBG Ally: A Conversational Embodied Agent as an AI Teammate
- 현재글 : [논문리뷰] Parts-of-Speech as Emergent Categories in SAE Latent Space
- 다음글 [논문리뷰] Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
댓글