[논문리뷰] The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction본 연구는 대규모 MoE 모델을 단일 GPU에서 구동할 때 발생하는 극심한 Memory Wall 문제를 해결하는 데 집중합니다. 기존 연구들은 주로 Weight Quantization이나 Weight Pruning을 통해 모델 크기를 줄여 VRAM 내에 적재하려 시도했으나, 이는 모델의 Accuracy 저하를 초래합니다.#Review#Mixture-of-Experts#SSD#Memory Wall#Routing Prediction#LLM Inference#Throughput2026년 9월 16일댓글 수 로딩 중
[논문리뷰] Unlocking Lossless Speedups in LLMs via Discrete Diffusion본 논문은 Autoregressive (AR) Large Language Models (LLMs)의 핵심 bottle-neck인 느린 순차적 토큰 생성을 해결하고자 한다. 기존 AR LLMs는 한 번에 하나의 토큰만 생성하므로 추론 시 높은 Latency를 유발하고 GPU utilization을 저하시킨다.#Review#Diffusion-augmented LLMs#Discrete Diffusion#Lossless Acceleration#Ψ-Spec Sampler#Speculative Decoding#Throughput#Tokens-Per-Forward-pass (TPF)#LoRA2026년 9월 7일댓글 수 로딩 중
[논문리뷰] Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism본 논문은 기존 Speculative Decoding의 핵심인 다중 토큰 예측(Multi-token prediction) 방식이 갖는 구조적 한계를 극복하고자 합니다.#Review#Speculative Decoding#Pipeline Parallelism#LLM Inference#Feature Aggregation#Latency Hiding#Throughput2026년 6월 1일댓글 수 로딩 중
[논문리뷰] Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving본 논문은 End-to-End Autonomous Driving을 위한 Vision-Language-Action (VLA) 모델이 직면한 High-Fidelity Trajectory Planning과 Efficient Inference 간의 상충 관계 문제를 해결하고자 합니다.#Review#Autonomous Driving#VLM#Block-Diffusion#Inference Efficiency#Trajectory Planning#Scaffold Speculative Decoding#Latency#Throughput2026년 5월 27일댓글 수 로딩 중
[sglang] run_eval에 latency 및 throughput 메트릭 추가평가 프레임워크에 completion token 기반 output throughput과 latency 메트릭을 추가하여 성능 추적 가능#SGLang#Evaluation#Metrics#Throughput2026년 4월 1일댓글 수 로딩 중
[논문리뷰] ECoLAD: Deployment-Oriented Evaluation for Automotive Time-Series Anomaly Detection기존의 Time-Series Anomaly Detection(TSAD) 연구들은 주로 workstation-class hardware에서 unconstrained execution 환경 하에 detection quality(주로 accuracy)만을 비교하고 최적화했습니다.#Review#Time-series anomaly detection#Deployment-oriented evaluation#Compute reduction#CPU parallelism#Throughput#Latency#Automotive telemetry#AUC-PR2026년 3월 15일댓글 수 로딩 중
[Ray] 파이프라인 최적 처리량 계산 유틸리티 함수 추가Ray Data에 파이프라인 연산자별 처리 속도와 리소스 제약을 기반으로 최적 처리량과 리소스 할당을 계산하는 유틸리티 함수를 추가한 PR 분석.#Ray#Ray Data#Resource Allocation#Pipeline Optimization#Throughput#Performance2026년 2월 27일댓글 수 로딩 중
[논문리뷰] SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs본 논문은 대규모 언어 모델(LLM)의 장문맥(long-context) 추론 시 발생하는 Key-Value (KV) 캐시 관련 문제를 해결하는 것을 목표로 합니다.#Review#LLMs#Long-context Reasoning#KV Cache Optimization#Speculative Sparsity#Knowledge Distillation#Adaptive Memory Management#Throughput2025년 12월 1일댓글 수 로딩 중