Device Health Analytics Report

This topic provides an overview of the device health reports you can generate in the HPE Networking Apstra Data Center Director GUI. To learn how to generate this report, see Generate Analytics Report.

The Device Health report analyzes the health of the device. In this report, you can view the health of all managed devices in your network.

Inventory Overview

The Inventory Overview (Figure 1) shows the number of managed devices in your network. This report shows the device hardware model, system ID, label, network operating system (NOS), average/max memory usage, and average/max CPU usage.

Figure 1: Inventory Overview Table showing system performance metrics: System ID, Label, Hardware, NOS Version, Avg Mem Usage %, Max Mem Usage %, Avg CPU Usage %, Max CPU Usage %.

Memory Usage Analysis

The Memory Usage Analysis reports show detailed memory usage for all devices and includes charts used to identify memory leaks. These reports are useful for capacity planning and identifying usage patterns to predict demand.

Note:

It can be normal for a device to use more memory over time due to increased workload or new features being enabled, however it might also suggest memory leaks.

Device Memory Usage

The Device Memory Usage chart (Figure 2) shows the top 10 devices with the highest memory usage over time.

Figure 2: Device Memory Usage Chart Line graph of memory utilization from April 1-7, 2025. Six lines show stability: green ID 525400FB11B0 increases then stabilizes; blue ID 52540028379E at 28 percent; purple ID 52540044CDE4 at 27 percent; yellow averages at 22 percent; orange ID 505400C5C6F1 at 16 percent; brown ID 50540097D3D9 at 15 percent; red ID 505400088253 at 14 percent.

Memory Trending Charts

The Memory Trending charts show devices with the highest memory increments rate. A device might use more memory over time due to increased workload or new features.

For example, consider three devices. Device A's memory usage increases 5-fold from 10M to 50M. Device B's usage increases 10 percent from 100M to 110M. Device C's memory usage decreases from 500M to 490M. Accordingly, we rank device A before Device B and eliminate Device C.

The following examples show the memory trending charts for two leaf devices and one spine device.

Figure 3: Device Memory Trending Chart - Leaf1 Line graph of memory usage from April 1-7, 2025. Y-axis: memory usage percentage. X-axis: time. Green circles show gradual increase; blue line shows specific pattern with fluctuations around April 3, 2025.
Figure 4: Device Memory Trending Chart - Leaf2 Graph showing memory usage from April 1-7, 2025. Y-axis: 28%-29.3%. Green dotted line: gradual increase. Blue line: stable, spike on April 3.
Figure 5: Device Memory Trending Chart - Spine2 Graph showing memory usage from April 1 to April 7, 2025. Memory utilization remains steady at 27 percent. Green line represents Utilization trend. Blue label 52540044CDE4 indicates monitored system.

CPU Usage Analysis

The CPU analysis section provides on-device CPU usage and telemetry collectors for each device.

Device CPU Usage

The Device CPU usage chart displays the CPU usage for all your devices.

Figure 6: Device CPU Usage Chart Line graph showing CPU utilization from April 1 to April 7, 2025, for multiple systems. Most devices fluctuate between 18 to 24 percent, while two remain below 4 percent. Yellow dotted line marks average utilization.

Device CPU Usage Analysis

The measurement of some metrics, such as CPU usage, might exhibit periodical behaviors due to certain operations repeated at fixed intervals. For example, device telemetry collectors may issue CLI commands to collect traffic counters every few seconds, causing device CPU usage to increase. We use FFT (Fast Fourier Transformation) to convert a time domain measurement to the frequency domain. This helps identify associations between various frequencies and known periodical operations.

For example, in the frequency domain chart (Figure 7) for a leaf device, the x-axis represents frequency in number per hour. Value 0 means a constant signal, while value 120 means the signal repeats 120 times per hour, with a 30-second period.

Figure 7: CPU Usage Frequency Chart Graph of FFT analysis on CPU usage data; X-axis shows rate per hour, Y-axis shows CPU usage amplitude. Significant spike near 0 rate indicates high CPU activity.
Figure 8: Telemetry Service Collector Execution Time Box plot of execution times in seconds for telemetry service collectors in Apstra system. Collectors on x-axis include lldp, resource_util, arp, disk_util, interface_counters, hostname, bgp, route, interface, xcvr. Y-axis shows execution times. Each plot displays median, interquartile range, whiskers, and outliers.