[논문리뷰] Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning본 논문은 MLLM의 추론 성능을 향상하기 위해 이미지 편집 도구를 활용하는 MLLM → Image Editor → MLLM 파이프라인의 실질적인 유용성을 진단하는 데 목적이 있다. 기존 연구들은 시각적 중간 상태가 추론에 도움이 된다고 주장하지만, 모든 태스크가 시각적 편집으로부터 이득을 얻는 것은 아니다.#Review#Multimodal Reasoning#Image Editing#Visual Workspace#MLLM#Task-Conditioned Utility#Diagnostic Framework2026년 8월 27일댓글 수 로딩 중