Presentation Information
[B-7-25]Evaluation of the Impact of Link Bandwidth on PD-Disaggregated LLM Inference Across AI-DC
◎Masaki Sugiyama1, Shunsuke Higuchi1, Kazuaki Ueda1, Ryo Inohara1 (1. KDDI Research, Inc.)
Keywords:
AI Data Center,LLM Inference,PD Disaggregation
This paper evaluates the impact of inter-data-center link bandwidth on Prefill/Decode-disaggregated LLM inference across AI data centers. Using LLMServingSim with Qwen3-14B, we show that 10 Gbps link bandwidth causes TTFT and E2E Latency degradation due to KV Cache transfer delay, while bandwidths of 100 Gbps or higher significantly suppress the degradation. The results also indicate that PD disaggregation improves TPOT and can potentially improve E2E Latency when sufficient link bandwidth is available.
