#Training Data Provenance

1개의 포스트

[논문리뷰] TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs

이 논문은 대규모 언어 모델(LLM)이 왜 안전하지 않거나 정책을 위반하는 출력을 생성하는 '정렬 드리프트(alignment drift)'를 겪는지에 대한 근본적인 원인을 밝히는 것을 목표로 합니다.

#Review #LLM Alignment #Alignment Drift #Training Data Provenance #Belief Conflict Index (BCI)#Suffix Array #Safety Interventions #Reinforcement Learning from Human Feedback #Explainable AI

2025년 8월 6일