Create AI Cluster Template

You can create blueprint templates based on L3 Clos AI fabrics or Rail-Collapsed AI fabrics using the GUI.

In an AI data center template, we organize the data center by rails, stripes, and spine elements. AI Clusters are supported on Juniper Networks devices only. This is due to the many variables and changing nature of AI fabrics, and the complexity of supporting advanced protocols (such as DCQCN).

To create an AI Cluster template using the GUI:

  1. From the left navigation menu of the GUI, navigate to Design > Templates and click Create AI Cluster Template.
    HPE Apstra interface showing the Design section under Templates with navigation menu, banner explaining relationships, buttons to create templates, and a table listing template details like Name, Type, Overlay Control Protocol, and Actions.
    The Create AI Cluster Template dialog opens.
  2. Enter specifications (described below) for your design.
    Note: After configuring your template you'll see a detailed preview. If you want to change any aspect of it, you can always return to this parameters page and change any details.
    User interface for creating an AI cluster template with input parameters and results. Key features: Fabric Connectivity Design options L3 Clos with spine or Collapsed Rail Design without spine, 8 GPUs per server, 16 total servers, 8 servers per stripe, GPU NIC speed 400 Gbps, 1:1 oversubscription, 128 total GPUs, Advanced Calculator Settings, Find Matching Configurations button.
    Table 1: AI Cluster Template Parameters

    Parameter

    Description

    Fabric Connectivity Design

    Depends on size of design:

    • L3 Clos (with spine) - for larger designs that justify having a spine

    • Collapsed Rail Design (no spine) - for most small AI cluster designs

    GPUs per Server

    All GPU servers must have the same number of GPUs. Changing the number of GPUs per server automatically changes the total number of GPUs (the big number on the right).

    If you need to design with a different number of GPUs per server, consider creating the design with the rack type builder.

    Total Number of Servers

    Enter the total number of servers to include. Changing the number of servers automatically changes the total number of GPUs (the big number on the right).

    Servers per Stripe (L3 Clos Only)

    Enter the number of servers to include for each stripe.

    GPU NIC Speed (L3 Clos Only)

    Enter the speed of the NIC. Speed affects what devices can be used in the cluster.

    Oversubsription (L3 Clos Only)

    This is the target oversubscription allowed between the spine and the leaf layer of the design. It's calculated as a ratio of southbound bandwidth to northbound bandwidth measured on the leaf switches layer.

    For example, a 2:1 oversubscription ratio with 8 GPUs per server means for each leaf in the design there would be 8 ports facing the server GPUs and 4 ports facing the spine layer.

  3. For optional settings, click Advanced Calculator Settings.
    The advanced settings appear.
  4. Configure advanced settings, as applicable.
    AI cluster template interface with tabs: Input Parameters, Select Calculated Result, and Preview & Create. Advanced Calculator Settings section shows device models selection options: Juniper Recommended or Manually Selected. Device profiles listed include Juniper PTX10008 8x36CD MDP, PTX10016 16x36CD MDP, QFX5220-32CD, QFX5230-64CD, QFX5240-64OD, QFX5240-64QD. Dropdown for Leaf-Spine Link Speed and button labeled Find Matching Configurations are at the bottom.
    1. The device profiles listed are the ones that are appropriate for the template. If you have specific device profiles in mind, you can use the Search function (magnifying glass) to narrow down the list. (Click the name of one of the device profiles to see its details.) If you've created your own device profile you can select the Manually Selected radio button, then search for it and select it.
    2. From the Leaf-Spine Leaf Speed drop-down list, you can specify the speed between the leaf switches and the spine switches.
      If you leave it blank, then the speed between the leaf switches and spine switches is the same as the speed between the leaf switches and the GPUs.
  5. Click Find Matching Configurations.
    A table of suitable logical devices appear.
  6. From the Leaf-Spine Leaf Speed drop-down list, you can specify the speed between the leaf switches and the spine switches.
    If you leave it blank, then the speed between the leaf switches and spine switches is the same as the speed between the leaf switches and the GPUs.
  7. Select one of the options.

    The tables often have many proposals. You can filter the results based on leaf and/or spine logical devices. For example, you may decide to select only leaf logical devices with 400G interfaces excluding those with 200G. Or, you may want to select spines that are fixed-form factors only while excluding those where the port count indicates chassis.

    Create AI Cluster Template interface showing step 2 highlighted with user inputs summary and table of stripe configurations. Includes buttons for recalculating, loading more results, and returning to previous step.
    The template preview opens for your review. This preview renders a topology of your AI cluster, and provides port and logical structure information. Hover over elements of the rendered topology for more information.
  8. Click Create Template (bottom-right).
    User interface for creating AI cluster template. Current step: Preview & Create. Template name: L3 Clos 16x400G Servers 8 per Stripe. Topology Preview shows nodes spine1, leaf_6_1, leaf_7_1 with connections. Options for Expanded View, Compact View, Expand Nodes, and Show Links are checked. Navigation buttons: Back to Calculator Results and Create Template.
The template is created and you're returned to the Templates table view.
After creating your AI Cluster template, you can create a blueprint using the template. For information about creating a blueprint, see Create Blueprint.