[논문리뷰] StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing본 논문은 LLM 기반 에이전트의 실질적인 위험(파일 변조, 정보 유출 등)을 사전에 차단하기 위한 Step-level 안전 감독 기술의 부재를 해결하고자 합니다.#Review#LLM Agents#Guardrails#Step-level Supervision#Safety-Utility Balancing#Balance-GRPO#Synthetic Data2026년 8월 30일댓글 수 로딩 중
[논문리뷰] LiSA: Lifelong Safety Adaptation via Conservative Policy Induction본 논문은 배포된 AI 에이전트의 안전 가드레일이 고정된 사전 정의(pre-deployment definition)만으로는 변화하는 환경과 개별적인 로컬 맥락의 안전 위험을 효과적으로 제어하지 못하는 문제를 해결합니다.#Review#Lifelong Safety Adaptation#Guardrails#Conservative Policy Induction#Structured Policy Memory#Confidence-gated Reuse#Conflict-aware Local Refinement#Sparse Feedback2026년 5월 14일댓글 수 로딩 중
[논문리뷰] VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation본 논문은 자율 AI 에이전트, 특히 LLM 기반 에이전트의 배포로 인해 발생하는 안전, 보안, 프라이버시 위험을 해결하고자 합니다.#Review#LLM Agents#Safety#Formal Verification#Code Generation#Runtime Monitoring#Security#Guardrails#Policy Enforcement2025년 10월 8일댓글 수 로딩 중