[논문리뷰] Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions본 논문은 LLM이 structurally unanswerable questions, 예를 들어 cot(−540∘) 계산이나 (1).startswith('1') 평가와 같이 유효한 답이 없는 질문에 대해 abstention 대신 answer-like outputs을 제공하는 문제를 다룹니다.#Review#Recognition–Refusal Misalignment#LLMs#Unanswerable Questions#Internal Representations#Safety Refusal#Activation Steering#Pretraining2026년 9월 8일댓글 수 로딩 중