Device System Health Probe

Introduction

The Device Dystem Health probe alerts if the system health parameters (CPU, memory and disk usage) exceed their specified thresholds for the specified duration.

For more information about this probe, from the blueprint, navigate to Analytics > Probes, click Create Probe, then select Instantiate Predefined Probe from the drop-down list. Select the probe from the Predefined Probe drop-down list to see details specific to the probe.

In HPE Networking Apstra Data Center Director 6.0, the Device System Health Probe is enhanced to monitor generic systems, or GPU servers where a compute agent is installed. This means you can monitor metrics like CPU, memory, and disk utilization for both network devices and generic systems (GPU servers) using the same probe.

The following image is an example of system CPU utilization data collected from multiple generic systems:

Dashboard showing system CPU utilization: Header 'Processor: System cpu utilization data', Real-Time source, 2% utilization, updated seconds ago.
The predefined dashboard has been updated to show differentiated gauges for switches and servers:

Device Health Summary dashboard showing system utilization metrics for CPU, memory, and disk usage with gauge charts and navigation links for further data exploration.
The following image shows polling frequency for both resource_util and disk_util telemetry services which have been extended to include generic systems where a compute agent is installed:

Telemetry collection statistics for device IP 10.28.253.8. Table shows services like RESOURCE UTIL and DISK UTIL with polling details. All services successfully polled with no failures or timeouts. Last run timestamp is 2024-10-07, 14:13:15.

Probe Settings


Configuration interface for creating a Device System Health probe to monitor CPU, memory, and disk utilization, with thresholds, durations, and historical data settings.
  • Thresholds:

    • System Type: Set different thresholds for servers and switches.

    • Set Thresholds: Define CPU, memory, and disk utilization threshold alert levels for each system type. If the threshold is exceeded, an anomaly is generated.

  • Duration: Specify how long utilization must exceed thresholds before an alert is triggered. The default value is 6 minutes.

  • Threshold Duration: Minimum time that utilization must remain high to trigger an alert. The default value is 2 minutes.