[논문리뷰] PUBG Ally: A Conversational Embodied Agent as an AI Teammate
링크: 논문 PDF로 바로 열기
저자: Beomsoo Kim, Byeongju Kim, Dohyun Kim, et al.
1. Key Terms & Definitions
본 논문에서 다루는 핵심 기술 용어 및 개념은 다음과 같다:
- Co-Playable Character (CPC): 인간 플레이어와 함께 공유된 게임 월드에서 음성 기반으로 소통하고 협력하며 자율적으로 행동하는 embodied game agent를 지칭한다.
- System 1:
System 2에서 전달된high-level action을real-time으로movement,combat,recovery등latency-critical한behavior로 변환하여 실행하는deterministic behavior-tree layer를 의미한다. - System 2: 플레이어의
intent를 해석하고, 플레이어와 협력하며speech를 생성하고,high-level action을 선택하는language-model agent로, 고정된clock이 아닌event에 의해invoke된다. - Agent Loop:
tools및game으로부터의feedback이 모델의 다음 응답에 영향을 미치는closed loop과정으로, 들어오는match events를 처리하며 가변적인 수의turn에 걸쳐 진행될 수 있다. - Reactivity Prior: 각
event type에 대해speech response와action response를 얼마나 강력하게 유도할지를 독립적으로 명시하는 사전 정의된 값으로, 모델 재훈련 없이Ally의response behavior를 제어할 수 있게 한다.
2. Motivation & Problem Statement
본 연구는 PUBG: BATTLEGROUNDS와 같은 real-time 게임 환경에서 변화하는 game world에 latency 제약 하에 반응하고 플레이어와 자연스럽게 voice interaction하는 AI teammate를 구축하는 핵심 문제를 해결하고자 한다. 기존 game agents는 주로 autonomous gameplay에 중점을 두어 개발되었으며, Ally와 같이 speech와 actions를 동기화하며 지속적인 voice interaction을 통해 인간 teammate와 소통하고 협력하는 능력은 부족했다. 특히, 에이전트의 발언이 게임 내 행동과 일치하지 않을 경우 플레이어를 오도할 수 있으므로, speech와 action의 synchronization이 중요하다는 점에서 이러한 challenges가 복합적으로 작용한다. 또한, 기존 연구들은 fixed set of observations에 의존하는 경향이 있어, Ally는 현재 상황과 관련된 game information을 동적으로 observe해야 하는 필요성이 제기된다. 이러한 한계를 극복하기 위해 deliberate LM reasoning과 fast control을 분리하는 새로운 접근 방식이 요구되었다 [cite: 1, Figure 2].
3. Method & Key Results
저자들은 language-model reasoning, voice interaction, autonomous gameplay를 통합하는 conversational embodied agent architecture인 PUBG Ally를 제안한다 [cite: 1, Figure 2]. 이 architecture는 bounded tool interface와 System 1–System 2 control hierarchy를 활용한다. System 2인 language-model agent는 observation tools, speech tools, action tools를 포함하는 controlled interface를 통해 관련 game information을 검사하고, player speech를 해석하며, context를 유지하고, 발언할 내용을 결정하며, high-level action choices를 발행한다 [cite: 1, Table 1]. 이러한 high-level action choices는 System 1인 deterministic behavior-tree layer에 의해 game-tick rate로 movement, combat, recovery 등 latency-critical behavior로 전환되어 실행된다 [cite: 1, Figure 2].
Ally는 31B teacher model과 teacher-corrected student rollouts를 활용하여 39,000개에 가까운 실제 플레이어와의 게임 세션 데이터로 on-device model을 훈련했다 [cite: 1, Figure 4]. Capability evaluation 결과, GEPA로 prompt optimization된 31B teacher model은 Game-event response에서 0.917에서 0.952로, Trajectory quality에서 0.884에서 0.910으로 향상되어, Claude Opus 4.8과 같은 외부 quality reference에 근접한 성능을 보였다 [cite: 1, Figure 6(a)]. Distillation 과정을 거친 2B student model은 Fact 0.932 (teacher: 0.944), Event 0.954 (teacher: 0.952), Intent 0.953 (teacher: 0.960)의 capability scores를 달성하며 teacher model에 근접한 성능을 보였다 [cite: 1, Figure 6(b)]. Runtime efficiency 측면에서 on-device SLM은 LM inference에 중앙값 0.80초의 latency를 기록하여 cloud LLM의 2.54초 대비 3.2배 빨랐으며, end-to-end spoken exchange는 on-device에서 약 1.6초가 소요되었다 [cite: 1, Table 8]. Safety evaluation에서는 Final models가 broad-coverage benchmark sets에서 98.3–99.4%, production-representative dialogue sets에서 84.1–99.0%의 harmless response rates를 달성했으며, over-refusal rates는 2.2–6.1%로 낮게 유지되었다 [cite: 1, Table 6, Table 7]. Live-beta player surveys에서는 Ally에 대한 net recommendation이 +25.1 percentage points를 기록했으며, 응답자의 50.0%가 Ally를 teammate 또는 companion으로 인식했다 [cite: 1, Figure 15(a)].
4. Conclusion & Impact
PUBG Ally는 reasoning, autonomous gameplay, voice-based coordination을 상업용 live battle-royale game에 통합한 최초의 conversational embodied teammate로, on-device에서 language 및 speech models를 실행하는 성공적인 사례를 제시한다. real-player interaction data와 player-centered evaluation에 기반한 iterative development process는 AI의 행동을 플레이어 선호도에 맞추고 개선점을 파악하는 데 결정적인 역할을 했다. 이 연구는 real-time, interactive environments에서 인간 사용자와 효과적으로 소통하고 행동하는 embodied agents 개발, 훈련 및 배포를 위한 실용적인 기반을 마련했다. 특히 System 1–System 2 architecture와 bounded tool interface는 복잡하고 동적인 환경에서 deliberative reasoning과 latency-critical control 간의 균형을 관리하는 강력한 framework를 제공한다. Ally의 architectural influence는 이미 physical robotics 분야로 확장되어, Ludi 0.1이 Ally의 agentic design principles를 차용하고 있어, socially intelligent agents 개발에 광범위한 영향을 미 미치고 있음을 보여준다.

Figure 1 — Ally의 역할

Figure 2 — Ally 런타임 파이프라인
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
- [논문리뷰] SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- [논문리뷰] World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
- [논문리뷰] RedVox: Safety and Fairness Gaps in Speech Models Across Languages
- [논문리뷰] SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Review 의 다른글
- 이전글 [논문리뷰] OmniEcho: Spatial Audio Understanding for Embodied Agents
- 현재글 : [논문리뷰] PUBG Ally: A Conversational Embodied Agent as an AI Teammate
- 다음글 [논문리뷰] Parts-of-Speech as Emergent Categories in SAE Latent Space
댓글