[논문리뷰] Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations본 논문은 현재의 NLA가 모델의 내부 상태를 해석하는 데 있어 '충실도(Faithfulness)'를 보장하지 못한다는 구조적 결함을 해결하고자 합니다. 기존 연구들은 Reconstruction 점수를 신뢰성 지표로 사용하지만, 저자들은 이 점수가 실제 모델의 내부 정보를 정확히 반영하지 않는다는 점을 증명합니다.#Review#Natural-Language Autoencoders#Activation Explanations#RECAP#Faithfulness#Decodability Supervision#AI Safety2026년 7월 22일댓글 수 로딩 중