[ray] Ray RDT NIXL 메모리 풀 최적화: 불필요한 복사 제거와 전송 효율 극대화Ray의 RDT NIXL 메모리 풀 개선을 통해 텐서 뷰 복사 비용을 줄이고, 전송 디스크립터 병합으로 대규모 텐서 전송 성능을 100배 향상시킨 사례를 분석합니다.#Ray#PyTorch#Memory Management#Performance Optimization#Distributed Computing2026년 9월 1일댓글 수 로딩 중
[sglang] [SGLang] MoE Prefill의 혁신: DWDP(Distributed Weight Data Parallelism) 도입 분석MoE 모델의 Prefill 단계에서 All-to-All 통신 병목을 제거하고 NVLink를 통한 가중치 프리페칭으로 성능을 극대화하는 DWDP 기법을 살펴봅니다.#LLM#MoE#Distributed Computing#CUDA VMM#SGLang#Performance Optimization2026년 7월 21일댓글 수 로딩 중
[flashinfer] FlashInfer 분산 오토튜닝 동기화: NCCL 데드락 해결을 위한 전략적 접근분산 환경에서 오토튜닝 시 발생하는 GPU 타이밍 오차로 인한 NCCL 데드락 문제를 ProcessGroup 동기화로 해결합니다.#FlashInfer#Distributed Computing#NCCL#AutoTuning#LLM2026년 7월 10일댓글 수 로딩 중
[sglang] [HunyuanVideo] Sequence Parallelism 최적화: Text Token Sharding으로 성능 한계 돌파하기HunyuanVideo 모델에서 텍스트 토큰을 분산 처리하여 중복 연산을 제거하고 추론 속도를 최대 5.7% 향상시킨 기법을 분석합니다.#SGLang#HunyuanVideo#Sequence Parallelism#DeepSpeed Ulysses#Distributed Computing2026년 6월 20일댓글 수 로딩 중
[flashinfer] FlashInfer의 고성능 분산 연산: All-Gather Matmul 최적화 분석FlashInfer에 추가된 All-gather Matmul 연산은 Push-Wait 알고리즘을 통해 분산 환경에서 GEMM 성능을 극대화합니다.#FlashInfer#Distributed Computing#CUDA#GEMM#Performance Optimization2026년 4월 24일댓글 수 로딩 중
[Ray] StreamingRepartition과 MapBatches 연산자 퓨전으로 스케줄링 오버헤드 제거Ray Data의 StreamingRepartition과 MapBatches를 퓨전하여 불필요한 스케줄링 오버헤드를 줄이고 collate 성능을 개선한 분석.#Ray#Python#Performance#Operator Fusion#Distributed Computing2025년 12월 3일댓글 수 로딩 중