Validated Inference flows

The inference traffic flows include two primary modes, as shown in Figure 4:

  • Single-node inference
  • Multi-node (Load balanced) inference

The following sections describe each in more detail.

Note: These workflows assume that each validated model instance fits a single GPU. More complex serving models, including distributed inference, multi-GPU model parallelism, may be explored in future updates of this JVD.

Figure 4: GenAI-Perf, Envoy, and SGLang Inference Traffic Flow