[논문리뷰] APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport본 논문은 기존의 LLM 안전성 벤치마크가 대부분 '모델의 거부 여부'에만 초점을 맞추어, 실제 배포 환경에서의 잔여 위험(Residual Risk)을 적절히 측정하지 못한다는 한계를 지적합니다.#Review#AI Agents#Payment Authorization#Open Agent Passport#Adversarial Benchmarking#Tool-using Agents#LLM Security2026년 9월 20일댓글 수 로딩 중
[논문리뷰] LMSM: LLM Security Framework Inspired by Linux Security Modules본 논문은 LLM의 내적 상태를 해석하는 다양한 방법들이 실제 서비스 환경에서 통합된 보안 제어로 작동하지 못하고 파편화되어 있다는 문제를 해결합니다.#Review#LLM Security#Runtime Mediation#Linux Security Modules#Interpretability#Sparse Autoencoders#Adversarial Robustness#Model Serving2026년 8월 30일댓글 수 로딩 중
[논문리뷰] When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse본 논문은 RAG 시스템이 LLM의 지식 부족 문제를 효과적으로 해결하지만, Adversarial Documents 주입을 통한 Poisoning Attack에 취약하다는 핵심 문제를 다룹니다.#Review#Retrieval-augmented Generation#RAG Poisoning#Attention Collapse#Mechanistic Interpretability#LLM Security#Attack Detection#Document-Level Attention2026년 8월 17일댓글 수 로딩 중
[논문리뷰] Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming본 논문은 LLM 기반 에이전트의 취약점을 탐지하기 위한 기존 레드팀 방식의 낮은 효율성과 제한적인 전이 가능성을 해결하고자 합니다. 기존의 RL-based 방식은 특정 타겟 모델에 대한 방대한 학습 데이터와 수만 번의 쿼리가 필요하여 비용이 높고 새로운 모델에 대한 적응력이 떨어집니다.#Review#Prompt Injection#Red Teaming#Agentic System#Hierarchical Memory#Strategy Library#LLM Security2026년 8월 5일댓글 수 로딩 중
[논문리뷰] SkillJack: Persistent Skill Backdoors in Self-Evolving Agents본 연구는 자가 진화형 에이전트가 경험을 Skill로 변환하는 과정에서 발생하는 새로운 보안 위협을 다룬다.#Review#Self-evolving Agents#Skill Extraction#Backdoor Attack#Experience Poisoning#LLM Security#Transformation-Resilient Payload#Persistence Isolation2026년 8월 4일댓글 수 로딩 중
[논문리뷰] Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection본 논문은 HuggingFace와 같은 공공 모델 허브에서 배포되는 LoRA 어댑터가 데이터 오염(Poisoning)을 통해 치명적인 백도어에 취약할 수 있다는 점을 지적합니다.#Review#LoRA Adapter#Backdoor Attack#Data Poisoning#Behavioral Detection#Weight-Level Detection#LLM Security2026년 5월 28일댓글 수 로딩 중
[논문리뷰] On the Evidentiary Limits of Membership Inference for Copyright Auditing본 논문은 LLM(Large Language Model) 학습 데이터의 저작권 감사에서 MIA(Membership Inference Attack) 가 신뢰할 수 있는 기술적 증거로 사용될 수 있는지 여부를 조사합니다.#Review#Membership Inference Attacks#Copyright Auditing#Large Language Models#Adversarial Robustness#Paraphrasing#Sparse Autoencoders#Semantic Preservation#LLM Security2026년 1월 20일댓글 수 로딩 중
[논문리뷰] Soft Instruction De-escalation Defense본 논문은 외부 환경과 상호작용하는 LLM 기반 에이전트 시스템 이 겪는 프롬프트 인젝션 공격에 대한 취약성을 해결하는 것을 목표로 합니다. 특히, 신뢰할 수 없는 데이터 내의 악의적인 명령을 효과적으로 무력화하면서도 에이전트의 유용성을 저해하지 않는 방어 메커니즘을 제안합니다.#Review#Prompt Injection#LLM Security#Agentic Systems#Iterative Sanitization#Instruction Control#Adversarial Robustness#Large Language Models2025년 10월 27일댓글 수 로딩 중
[논문리뷰] Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated본 논문은 대규모 언어 모델(LLM)의 단계별 추론(Chain-of-Thought, CoT) 능력 이 새로운 유형의 데이터 포이즈닝 공격 기회를 제공함과 동시에, 이러한 공격을 최종 답변으로 유도하는 것을 복잡하게 만드는 예상치 못한 견고성을 생성한다는 점을 탐구합니다.#Review#LLM Security#Data Poisoning#Chain-of-Thought#Reasoning Models#Backdoor Attacks#CoT Unfaithfulness#Emergent Robustness2025년 9월 12일댓글 수 로딩 중