[논문리뷰] Diffusion Reward Models
링크: 논문 PDF로 바로 열기
The paper "Diffusion Reward Models" by Wang et al. introduces DRM, a Diffusion Reward Model that addresses the limitations of existing reward models which typically reduce human preference to a point estimate or a fixed parametric distribution. Human preference is inherently multimodal, and DRM aims to capture this complexity by recasting reward modeling as conditional density estimation over p(𝐫∣x,y).
I have browsed the paper. Now I will proceed to extract the required information and format the output according to the instructions.
Part 1: Summary
- Metadata: Authors are clearly listed. I'll extract the first two and use "et al.". Keywords need to be identified from the abstract/introduction/conclusion.
Diffusion Reward Model,Conditional Density Estimation,Multimodal Reward,RLHF,Diffusion Transformer,Human Preference,Reward Modelingseem appropriate. - Key Terms & Definitions:
- Reward Models (RMs): Models that learn to predict a reward score for a given prompt-response pair, used in RLHF.
- Multimodal Preference: Human preference that exhibits multiple distinct modes or local peaks, indicating diverse valid judgments for the same response.
- Diffusion Reward Model (DRM): The proposed model that uses a Diffusion Transformer to estimate the conditional density of reward vectors without parametric assumptions.
- Diffusion Transformer (DiT): A lightweight transformer architecture used in DRM to denoise Gaussian noise into reward vectors.
- Reinforcement Learning from Human Feedback (RLHF): A paradigm for post-training LLMs where an RM provides an optimization signal based on human preferences.
- Motivation & Problem Statement: The core problem is that dominant RM designs (discriminative and generative) collapse human preference into a single scalar or fixed parametric distribution, which is at odds with the inherently multimodal nature of human feedback. Existing methods like multi-objective RMs or parametric distributional heads still commit to specific output distribution families, limiting their expressiveness for multimodal rewards. This loss of information hinders effective LLM alignment.
- Method & Key Results: DRM replaces the conventional scalar value head with a Diffusion Reward Head, implicitly modeling p(𝐫∣x,y) through an iterative denoising process. A frozen LLM encoder provides a semantic representation of the prompt-response pair, which conditions a lightweight Diffusion Transformer (DiT) to denoise Gaussian noise into a K-dimensional reward vector. The model can be trained with both multi-attribute regression (using a masked denoising loss) and pairwise preference data (using an additional Bradley-Terry objective). At inference, N samples form an empirical reward distribution, which can be aggregated into a scalar mean, variance, or quantiles.
- Key results:
- On five benchmarks (RewardBench v2, PPE, RMB, RM-Bench, JudgeBench), DRM-Multi-8B achieved an average of 66.2, outperforming the scalar multi-attribute head ArmoRM (62.3) and the parametric-quantile head QRM (64.1) under matched data and backbone conditions.
- DRM demonstrated competitive performance against larger discriminative, distributional, and generative RMs, such as DeepSeek-GRM-27B (65.6) and GPT-4o (67.7), despite its modest training scale.
- Distributional analysis showed that DRM captures multimodal reward structures associated with human disagreement, with its multimodal ratio increasing as human disagreement grows. For helpfulness, the multimodal ratio increased from 37.6% (rating range <= 1) to 63.2% (low-high polarized examples).
- Distribution-aware decision-making strategies like uncertainty-aware rejection and Lower-Confidence-Bound (LCB) aggregation improved performance. For example, rejecting 30% most uncertain decisions on PPE Correctness improved accuracy by an average of 2.81 percentage points. LCB aggregation consistently outperformed mean aggregation across PPE and RMB Best-of-N settings.
- DRM uniquely offers "reward-axis scaling" at test-time; increasing the number of diffusion samples N improves scoring precision without retraining. For example, DRM-Multi-8B's performance on RewardBench v2 rose from 56.5% (N=1) to 65.6% (N=32).
- Key results:
- Conclusion & Impact: DRM effectively models multimodal human preferences by replacing conventional reward heads with a non-parametric Diffusion Transformer head. This approach not only matches or surpasses existing baselines but also provides richer distributional information that improves decision-making through uncertainty awareness and risk-sensitive aggregation. The research demonstrates the practical benefits of diffusion-based reward modeling for downstream RLHF, where DRM leads to improved policy performance, thereby advancing the alignment of LLMs.
Part 2: Important Figures
I will look for figures that show the overall architecture, key quantitative results, or methodological explanations.
- Figure 1: Overview of DRM. This is a crucial architecture diagram.
- Figure 2: DRM multimodality tracks human disagreement. This shows a key result validating the core hypothesis.
- Figure 3 or Figure 4: Both show quantitative results for test-time scaling. Figure 4 shows reward-axis scaling, which is unique to DRM. Figure 3 shows Best-of-NN scaling which is also important. I will pick Figure 4 because it highlights a unique capability of DRM.
Extract image_url for these figures.
- Figure 1:
2609.33803v1/0523_DRM_overview.png->https://arxiv.org/html/2609.33803v1/0523_DRM_overview.png - Figure 2:
2609.33803v1/Figures/drm_multimodality_tracks_disagreement.png->https://arxiv.org/html/2609.33803v1/Figures/drm_multimodality_tracks_disagreement.png - Figure 4:
2609.33803v1/Figures/tts-b_result.png->https://arxiv.org/html/2609.33803v1/Figures/tts-b_result.png저자: Xiangyang Wang, Bingxiang He, et al.
1. Key Terms & Definitions (핵심 용어 및 정의)
- Reward Models (RMs): 대규모 언어 모델(LLM)의 Alignment를 위해 인간 선호도를 예측하는 모델로, Reinforcement Learning from Human Feedback (RLHF)에서 Optimization Signal을 제공합니다.
- Multimodal Preference: 동일한 응답에 대해 다양한 합리적인 판단이 가능하여, 분포가 여러 Mode나 Local Peak를 가지는 인간 선호도의 본질적인 특성을 의미합니다.
- Diffusion Reward Model (DRM): 본 논문에서 제안하는 새로운 Reward Modeling Paradigm으로, Reward Head를 Diffusion-based Head로 대체하여 Output Distribution에 대한 Parametric Assumption 없이 Conditional Density Estimation을 수행합니다.
- Diffusion Transformer (DiT): DRM의 핵심 구성 요소로, Frozen LLM Encoder에 의해 Conditioning되어 Gaussian Noise를 Reward Vector로 Denoise하는 경량 Transformer 구조입니다.
- Classifier-Free Guidance (CFG): Diffusion Model의 Sampling 과정에서 Conditional Vector와 Unconditional Vector를 활용하여 생성 품질을 제어하는 기법으로, DRM Inference 시 Guidance Scale ω를 통해 Reward Estimate의 Precision을 조절할 수 있습니다.
2. Motivation & Problem Statement (연구 배경 및 문제 정의)
본 논문은 기존 Reward Model(RM)들이 인간 선호도의 본질적인 Multimodality를 제대로 반영하지 못하는 문제점을 지적합니다. 현재 지배적인 Discriminative RM 및 Generative RM은 Prompt-Response Pair에 대해 단일 Scalar Score를 예측하거나, 고정된 Parametric Family 내의 분포를 가정합니다. 이러한 접근 방식은 인간 선호도가 본질적으로 Multimodal하다는 점과 상충되며, Annotator 간의 체계적인 Disagreement, Uncertainty, 그리고 Multimodal Structure를 무시하게 됩니다. 예를 들어, Anthropic-HH 데이터셋의 Annotator Inter-agreement는 약 63%에 불과하며, HelpSteer3-Preference 데이터셋에서도 상당한 Within-rubric Disagreement가 관찰됩니다.
기존 연구들은 이러한 Scalar Bottleneck을 해결하고자 했지만, Multi-objective RM은 고정된 Attribute Schema에 의존하고, Parametric Distributional RM (예: URM의 Gaussian, QRM의 Quantile Grid)은 특정 Output Distribution Family에 제한되어 Multimodal Distribution을 충분히 표현하지 못합니다. 이러한 한계점들은 LLM Alignment 품질에 직접적인 영향을 미치며 Reward Hacking과 같은 문제로 이어질 수 있습니다. 따라서 저자들은 Output Distribution에 대한 어떠한 Parametric 가정 없이 Multimodal Reward Distribution을 표현할 수 있는 새로운 Reward Head의 필요성을 강조합니다.
3. Method & Key Results (제안 방법론 및 핵심 결과)
저자들은 기존 Reward Model의 Value Head를 Diffusion-based Head로 대체하는 DRM(Diffusion Reward Model)을 제안합니다. DRM은 Prompt-Response Pair (x,y)를 Frozen LLM Encoder가 Hidden State 𝐡로 Mapping하고, 이 𝐡에 Conditioning된 Lightweight Diffusion Transformer (DiT)가 Gaussian Noise를 K-dimensional Reward Vector로 Denoise하여 p(𝐫∣x,y)를 직접 Modeling하는 방식을 채택합니다. 이 접근 방식은 Output Distribution에 Parametric Assumption을 부과하지 않아, 인간 선호가 유도하는 Multimodal Reward Distribution을 자연스럽게 나타낼 수 있습니다. 하나의 아키텍처로 Multi-attribute Regression (Masked Denoising Loss ℒMSE)과 Pairwise Preference Data (Bradley-Terry Loss ℒBT)를 모두 처리할 수 있습니다. Inference 시에는 N개의 Sample을 통해 Empirical Reward Distribution을 형성하며, 이를 Scalar, Variance 또는 Quantile로 Aggregation하여 활용합니다.
DRM은 다섯 가지 표준 RM Benchmark (RewardBench v2, PPE, RMB, RM-Bench, JudgeBench)에서 평가되었습니다. 동일한 Training Data와 Backbone을 사용했을 때, DRM-Multi-8B는 평균 66.2의 성능을 달성하며 Scalar Head인 ArmoRM (62.3)과 Parametric-quantile Head인 QRM (64.1)을 능가했습니다. 특히 RMB에서는 78.0으로 가장 큰 성능 향상을 보였습니다. 또한, DRM은 Modest Training Scale에도 불구하고 DeepSeek-GRM-27B (65.6)나 GPT-4o (67.7)와 같은 훨씬 큰 규모의 Discriminative, Distributional, Generative RM들과 경쟁력 있는 성능을 유지했습니다.
DRM의 핵심 강점은 Multimodal Reward Structure를 포착하는 능력입니다. 인간 Disagreement가 강한 Sample에서 DRM의 Multimodal Ratio가 증가하는 것이 확인되었습니다 [Figure 2]. 예를 들어, Helpfulness Metric에서 Human Rating Range가 1 이하일 때 Multimodal Ratio가 37.6%였으나, Low-High Polarized Examples에서는 63.2%까지 상승했습니다. Distributional Uncertainty를 활용한 Uncertainty-aware Rejection 기법은 PPE Correctness에서 Coverage를 70%로 줄였을 때, 정확도를 평균 2.81 Percentage Point 향상시켰습니다. 또한, Lower-Confidence-Bound (LCB) Aggregation은 Mean Aggregation 대비 Best-of-N Candidate Selection에서 일관된 성능 향상을 보였습니다. DRM은 Reward Axis Scaling이라는 고유한 기능을 통해 Inference 시 Diffusion Sample 수를 늘려 Scoring Precision을 향상시킬 수 있으며, RewardBench v2에서 DRM-Multi-8B의 성능은 N=1일 때 56.5%에서 N=32일 때 65.6%로 증가했습니다 [Figure 4].

Figure 2 — DRM 멀티모달성 및 인간 불일치

Figure 4 — Reward-axis 스케일링 결과
4. Conclusion & Impact (결론 및 시사점)
본 논문은 Diffusion Reward Model(DRM)을 통해 Reward Modeling 패러다임을 혁신하며, 인간 선호의 본질적인 Multimodality를 효과적으로 포착할 수 있음을 입증했습니다. DRM은 Parametric Assumption 없이 조건부 Reward Distribution p(𝐫∣x,y)를 모델링하여, 기존 Scalar 또는 Parametric RM이 놓쳤던 풍부한 Distributional Information을 제공합니다. 이는 단일 아키텍처로 Multi-attribute 및 Preference Data를 모두 처리하며, Inference 시 Distributional Statistics와 Reward-axis Scaling을 가능하게 합니다.
DRM은 다양한 Benchmark에서 기존 Baseline들을 능가하거나 경쟁력 있는 성능을 보였으며, 인간 Disagreement와 관련된 Multimodal Reward Structure를 성공적으로 학습했습니다. 특히 Uncertainty-aware Rejection 및 Risk-sensitive Aggregation 실험을 통해 학습된 Reward Distribution이 단순히 평균 이상의 유용한 Decision Signal을 제공한다는 것을 입증했습니다. 더 나아가, DRM이 Downstream RLHF 학습의 Reward Model로 사용될 때 정책 성능 향상으로 이어진다는 점은 Diffusion-based Reward Modeling의 실질적인 이점을 강조합니다. 이 연구는 LLM Alignment 분야에서 보다 정교하고 신뢰할 수 있는 Feedback Mechanism을 구축하는 데 중요한 기여를 하며, LLM의 안전하고 유용한 발전을 위한 새로운 방향을 제시합니다.

Figure 1 — DRM의 전체 개요
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] Qwen-Image-2.0-RL Technical Report
- [논문리뷰] Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
- [논문리뷰] Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
- [논문리뷰] Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- [논문리뷰] LongCat-Video Technical Report
Review 의 다른글
- 이전글 [논문리뷰] DepthBench: Measuring How Residual Connections Enable More Computational Depth
- 현재글 : [논문리뷰] Diffusion Reward Models
- 다음글 [논문리뷰] Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
댓글