[논문리뷰] A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal본 논문은 LLM이 내부적으로 알고 있는 지식을 외부에 드러내지 않거나, 의도적으로 잘못된 답변을 생성하는 문제를 해결하고자 한다.#Review#Concealed Information Test#Sandbagging#Unlearning Verification#Probe of Internal Recognition (PIR)#Internal Recognition#Activation Probing#Deception Detection2026년 9월 21일댓글 수 로딩 중
[논문리뷰] Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations본 연구는 LLM의 deception detection을 위해 사용되는 Linear Probes가 실전 환경에서 보이는 극심한 성능 저하의 원인을 규명하고자 합니다.#Review#LLM#Deception Detection#Linear Probes#Scaling Laws#Robustness#Geometric Analysis#Activation Engineering2026년 6월 2일댓글 수 로딩 중