[논문리뷰] SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?본 논문은 Recursive Self-Improvement (RSI) 시스템의 신뢰성과 제어 가능성을 확보하기 위해 필수적인 내부 표현 모니터링 및 감사(auditing)의 부재 문제를 해결하고자 합니다.#Review#AI Agents#Mechanistic Interpretability#Sparse Autoencoders (SAEs)#Recursive Self-Improvement (RSI)#Autonomous Discovery#Contrastive Probes#Causal Steering#Benchmark2026년 9월 9일댓글 수 로딩 중