Inference Traffic Patterns

Inference traffic patterns can vary depending on model size, serving architecture, parallelism strategy, and deployment model. Some inference deployments may include GPU-to-GPU communication, KV cache-related traffic, storage access, or other inference scenarios. This JVD focuses specifically on the frontend inference path, where traffic is primarily composed of client requests, optional load balancer distribution, and response delivery.

In this JVD, inference traffic enters the frontend fabric from benchmark or client systems and is sent either directly to an inference server or to an Envoy load balancer. When Envoy is used, it provides a single frontend endpoint and distributes requests across multiple model-serving endpoints running on AMD Instinct™ MI300X GPU systems.

Table 3: AI Training and AI Inference Traffic Comparison

Characteristic AI Training AI Inference
Primary traffic pattern GPU-to-GPU communication. Client/API-to-inference-server communication.
Main fabric focus GPU backend fabric. Frontend fabric.
Common communication model Collective operations such as all-reduce or all-to-all. Request/response flows between clients, load balancers, and model-serving endpoints.
Typical network technologies RoCEv2, backend congestion control, rail optimization, and lossless or near-lossless design. IP Ethernet frontend connectivity, request distribution, latency, throughput, and observability.
Common performance indicators

Measure how efficiently the system trains a model, including how efficiently the system uses compute, network, and storage resources to train a model.

These may include overall job or training time, collective communication performance, and workload throughput.

Evaluate how efficiently the system responds to user or application queries.

These metrics are more closely tied to user experience because they measure response time, consistency, and request-handling efficiency. (for example, Time to first response/output, Request Latency).

The specific metrics included in this JVD will be described in the Benchmarking Testing Methodology section.