[flashinfer] FlashInfer의 FP8 양자화 AllReduce를 통한 통신 대역폭 최적화FP8 양자화를 통해 AllReduce 통신량을 절반으로 줄여 대규모 모델 학습의 NVLink 병목을 해결합니다.#FlashInfer#AllReduce#FP8#Quantization#NVLink#Triton2026년 7월 24일댓글 수 로딩 중