[논문리뷰] A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal본 논문은 LLM이 내부적으로 알고 있는 지식을 외부에 드러내지 않거나, 의도적으로 잘못된 답변을 생성하는 문제를 해결하고자 한다.#Review#Concealed Information Test#Sandbagging#Unlearning Verification#Probe of Internal Recognition (PIR)#Internal Recognition#Activation Probing#Deception Detection2026년 9월 21일댓글 수 로딩 중