[논문리뷰] FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
링크: 논문 PDF로 바로 열기
The browsing was successful. I have the content of the paper. Now I will proceed to extract the information as per the user's request, following all constraints.
Part 1: Summary Body (Markdown)
Metadata Extraction:
- Authors: Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
- Keywords: Based on the abstract and introduction, good keywords would be:
Perceptual Distance,Image Quality Assessment (IQA),Diffusion Models,Forking Moment (FoMo),Generative Trajectory,Automated Annotation,Rank-based Learning.
Section Drafting:
## 1. Key Terms & Definitions
- Reference-based Image Quality Assessment (IQA): Reference image와 distorted counterpart 간의 perceptual difference를 정량화하는 metric.
- Mean Opinion Score (MOS): Human annotator들이 distorted image에 absolute quality score를 부여하는 방식으로, global ordering을 제공하지만 collection 비용이 높고 inconsistent judgments에 취약함.
- Two-Alternative Forced Choice (2AFC): Human annotator들이 두 개의 distorted image 중 reference에 더 유사한 것을 선택하는 pairwise preference labeling protocol로, MOS보다 reliability가 높으나 global ordering을 직접적으로 인코딩하지 못함.
- Forking Moment (FoMo): Diffusion model의 generative trajectory에서 두 이미지가 동일한 denoising process를 따르다가 독립적으로 branching하는 시점. 이 시점이 빠를수록 이미지는 perceptual하게 멀고, 늦을수록 perceptual하게 가깝다.
- Generative Trajectory: Diffusion model에서 noise image가 clean image로 denoising되는 일련의 과정으로, timestep에 따라 coarse structure부터 fine details까지 정보가 복구되는 계층적 특성을 가짐.
## 2. Motivation & Problem Statement 본 논문은 Reference-based IQA metric 학습 시 필요한 human-annotated data 수집의 한계점을 해결하고자 한다. 기존 MOS 기반 pointwise scoring은 global ordering을 제공하지만, 대규모 수집이 어렵고 human judgment의 inconsistency로 인해 noisy하다. 대안으로 2AFC pairwise labels가 신뢰성과 효율성 때문에 널리 사용되지만, 이는 이미지 쌍 간의 relative comparison만을 포착하며 globally consistent ordering을 직접적으로 인코딩하지 못하는 구조적 한계를 가진다. 이러한 문제점들은 IQA metric이 인간의 perceptual judgment와 긴밀하게 정렬되기 어렵게 만들며, 특히 다양한 distortion type과 severity level에 걸쳐 일관된 ordering을 유도하는 데 장애가 된다. 따라서 저자들은 인간의 개입 없이 확장 가능하며, pointwise perceptual distance labels을 자동으로 생성하는 새로운 접근 방식의 필요성을 제기한다.
## 3. Method & Key Results 저자들은 diffusion model의 generative dynamics를 perceptual distance의 proxy로 활용하는 FoMo (Forking Moment) 기반의 완전 자동화된 데이터 생성 파이프라인을 제안한다. 이 방법론은 diffusion process에서 coarse image structure가 early timesteps에, fine details가 later timesteps에 생성된다는 점에 착안한다. 두 이미지가 생성 과정에서 early fork할수록 perceptual하게 멀고, late fork할수록 perceptual하게 가깝다는 원리를 이용한다. 구체적으로, reference image $x_0$에 sampled timestep $s$에 해당하는 noise를 주입하여 $x_{t_s}$를 얻고, 이를 다시 denoising하여 perturbed variant $x_0^s$를 생성한다. 이때 $t_s$ (interpolation factor)를 distance label로 사용하여 $(x_0, x_0^s, t_s)$ 쌍을 구성하며, $t_s$ 값이 클수록 reference image로부터 perceptual deviation이 크다고 정의한다.
저자들은 FoMo가 인간의 perceptual judgment와 잘 정렬됨을 empirical하게 검증했다. Single-reference validation에서는 forking timestep이 늦을수록 reference image에 더 유사하게 인지된다는 점을 Spearman rank correlation 0.970으로 확인했다. Cross-reference validation에서는 forking timesteps가 다른 reference image 쌍들 간에도 일관된 유사성 순위를 유도하며, 인간의 선택과 90.2% 일치하고 Fleiss’ $\kappa$ 0.82를 달성했다.
학습 시에는 FoMo에서 생성된 pointwise distance label을 활용하여 RankNet-style objective function을 사용한다. 이는 기존 2AFC 방식의 triplet-wise comparison과 달리, 배치 내 모든 $B \times B$ 쌍 간의 global ordering을 감독함으로써 더 풍부한 training objective를 가능하게 한다. 실험 결과, FoMo (Ours)는 PIPAL 벤치마크에서 LPIPS-Alex 기준 0.733 SROCC를 달성하여, human-annotated dataset인 BAPPS (0.622), PieAPP (0.602), NIGHTS (0.577), KADID-10K (0.577)보다 우수한 성능을 보였다 [cite: 1, Table 1]. 특히 Transformer-based backbones에서는 더욱 큰 성능 향상을 나타내어, PIPAL에서 DINOv3 기준으로 FoMo는 0.699 SROCC를 기록하며 BAPPS (0.287), PieAPP (0.325), NIGHTS (0.234), KADID-10K (0.298)를 크게 상회했다 [cite: 1, Table 1]. 이 결과는 FoMo가 human annotation 없이도 기존 human-annotated dataset을 능가하는 강력하고 신뢰할 수 있는 training signal을 제공함을 입증한다.
## 4. Conclusion & Impact 본 연구는 diffusion denoising trajectory의 forking moment를 perceptual distance로 재해석하는 혁신적인 데이터 생성 파이프라인을 제시한다. 인간의 개입 없이 자동으로 훈련 쌍과 보정된 perceptual distance labels를 합성함으로써, 기존 human annotation 방식의 한계를 극복한다. FoMo 기반의 RankNet-style supervision은 2AFC 방식보다 더 풍부한 gradient signal과 내재적인 transitivity enforcement를 제공하여, 다양한 backbone과 평가 벤치마크에서 기존 strong baselines 대비 일관된 성능 향상을 입증하였다. 이 연구는 IQA 분야에서 human-aligned metric 구축의 접근 장벽을 낮추고, image restoration 및 synthesis task의 평가 파이프라인을 개선하는 데 긍정적인 영향을 미칠 것이다. 또한, 의료 영상, 이미지 압축, 접근성 등 다양한 응용 분야에서 더 나은 metric 개발에 기여할 잠재력을 가진다.
Part 2: Important Figure Information (JSON)
I need to find img tags and their src attributes.
Looking at the HTML content:
- Figure 1:
2609.25716v1/single_ref_human_study.png(a) and2609.25716v1/figures_raw/cross_ref_human_study_samples.png(b). This figure explains human study setup. - Figure 2:
2609.25716v1/main_figure_fomo.png. This is the overall framework visualization. - Figure 3:
2609.25716v1/visualized_samples.png. This shows data sample image pairs.
Figure 2 is the most important as it visualizes the overall framework. Figure 1 is also important as it details the human studies validating the hypothesis. Figure 3 shows examples of the generated data.
I'll pick Figure 1, 2, and 3. I need to make sure to prepend https://arxiv.org/html/ to the relative paths.
Figure 1:
single_ref_human_study.png(a)figures_raw/cross_ref_human_study_samples.png(b) The figure caption is: "Figure 1: (a) Single-reference study: participants rank synthesized variants to verify whether the divergence timestep in diffusion generative trajectory reflects the degree of perceptual similarity to the reference image. (b) Cross-reference study: participants are asked which of two pairs, built from different references, holds the two more similar images. Further examples in the Appendix." I will combine these into one entry for Figure 1. I will use the path for (a) as the primary image if they are separate, or choose the one that represents the overall figure better if the figure is composed of multiple images. In this case,single_ref_human_study.pngseems to be the main visual part. Wait, the HTML shows two separateimgtags for (a) and (b) under the sameFigure 1caption. I need to select up to 3 figures, not parts of a figure. The guidelines say "Figure/Table", which implies a single entity. I will choose Figure 2 as the architecture, Figure 1 as the empirical grounding, and Figure 3 for data samples. For Figure 1, I will choose the first image listed, as it represents the single-reference validation clearly. Or, I can check if there's a single image that combines (a) and (b). No, they are presented as separate<img>tags. I'll takemain_figure_fomo.pngas the architecture,single_ref_human_study.pngas the human study validation, andvisualized_samples.pngfor data samples. These are the three strongest.
Let's refine the captions and URLs.
Figure 1 (original a): https://arxiv.org/html/2609.25716v1/single_ref_human_study.png - "인간 지각 정렬 연구"
Figure 2: https://arxiv.org/html/2609.25716v1/main_figure_fomo.png - "제안 프레임워크 시각화"
Figure 3: https://arxiv.org/html/2609.25716v1/visualized_samples.png - "생성된 데이터 샘플"
This matches the criteria. Final check on markdown formatting and constraints.
- Technical terminology in English: Checked.
- Natural Korean Phrasing: Checked.
- Mandatory Metadata: Checked.
- Output format (Part 1,

Figure 1 — 인간 지각 정렬 연구

Figure 2 — 제안 프레임워크 시각화

Figure 3 — 생성된 데이터 샘플
⚠️ 알림: 이 리뷰는 AI로 작성되었습니다.
관련 포스트
- [논문리뷰] LLMs4All: A Review on Large Language Models for Research and Applications in Academic Disciplines
- [논문리뷰] ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
- [논문리뷰] VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- [논문리뷰] TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
- [논문리뷰] Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors
Review 의 다른글
- 이전글 [논문리뷰] Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
- 현재글 : [논문리뷰] FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
- 다음글 [논문리뷰] FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
댓글