[논문리뷰] From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World본 논문은 기존의 사이버 보안 벤치마크가 지나치게 제한된 환경(예: Capture-the-Flag)에 국한되어 있어, 실제 환경에서의 복잡한 공격 표면과 전략적 탐색 능력을 평가하지 못하는 한계를 해결하고자 한다 .#Review#AI Pentesting Agents#Vulnerability Discovery#Evaluation Protocol#Ground-Truth Matching#Stochasticity#Agentic Workflow2026년 7월 15일댓글 수 로딩 중
[논문리뷰] T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search기존 LLM red-teaming 연구는 주로 모델에서 유해한 텍스트 출력(harmful text outputs)을 유도하는 데 초점을 맞추었으나, 이는 Model Context Protocol (MCP)과 같은 통합 표준을 통해 다단계 도구 실행(multi-step tool execution)이 가능한 LLM Agents의 새로운 안전 위험을 간과하고 있습니다.#Review#LLM Agents#Red-Teaming#Vulnerability Discovery#Trajectory-aware Search#MAP-Elites#Tool Call Graph#Attack Realization Rate2026년 3월 25일댓글 수 로딩 중