Interface Flapping (Fabric Interfaces) Probe

The Interface Flapping (Fabric Interfaces) probe determines if any fabric interfaces are flapping and raises anomalies accordingly.

Probe Overview

If the number of times that the operational state of any fabric interface changes is greater than a specified number (Threshold) over a specified amount of time (Duration), then the interface is flapping and an anomaly is raised. Also, if the percentage of flapping interfaces exceeds the specified percentage (Max Flapping Interfaces Percentage), then an anomaly is raised for that device.

Probe Parameters

When you instantiate the predefined Interface Flapping (Fabric Interfaces) probe you can customize the above mentioned parameters or leave default values as is, as shown in the screenshot below.

Configuration interface for setting a probe to monitor fabric interfaces for flapping. Probe type: Interface Flapping Fabric Interfaces. Threshold: 5 state transitions. Max Flapping Interfaces Percentage: 10. Duration: 1 Minute. Description explains detection and anomaly conditions. Create button finalizes probe setup.

Probe Stages

The following stages are used in the interface flapping probe for all fabric interfaces:

Hierarchical flow diagram showing network monitoring stages: leaf fabric interface status, history, flapping, flapping percentage, and system anomalous flapping.

The probe performs the following tasks in stages:

  1. Identify fabric interfaces (leafs facing spines) on all leafs (using the leaf fab int status processor).

  2. For each interface, count the number of times that the operational state changes and create a time series from it (using the leaf fabric interface status history processor).

  3. If the number of state changes during the specified duration is more than the specified upper limit, then generate an anomaly (using the leaf fabric interface flapping processor

  4. If an anomaly was created based on the previous stage, create a time series for it (using the

  5. Calculate the percentage of interfaces on devices with the anomaly. If the percentage is higher than the specified threshold, then raise device level anomaly(so that recent history of existence and clearing of anomaly can be inspected).

  6. Create time series for the anomaly so that recent history can be inspected. the last "Anomaly History Count" anomaly state-changes are stored for observation.

The sections below discuss the stages in detailed.

Leaf Fabric Interface Status Processor

Leaf Fabric Interface Status Service Collector (Input)

We begin with the source processor (no inputs), which is a Service Collector configured to collect interface status telemetry for all fabric interfaces on leaf devices, as shown in the screenshot below.

User interface for a network processor titled leaf fab int status with sections for Service Collector, graph query script, Query Tag Filter, Telemetry parameters, and Advanced options for configuring network systems.

This processor keeps track of when the operational state (up, down) of the interface changes. It collects and outputs this information as leaf interface status (leaf_if_status). Each interface is identified by its system ID and interface name. Operational state and a few other details are included in the output as shown in the list below:

  • System ID - ID of the leaf device, usually the serial number

  • Interface - name of the interface

  • Remote Interface - interface name on the other end

  • Remote System Label - the device name on the other end

  • Value - operational state of the device (up, down)

  • Updated - when the state was last updated

The following screenshot is an example of the details that are collected from this processor.

Table displaying network interface status with columns: System ID, Interface, Remote Interface, Remote System Label, Value, and Updated. All connections are up.

leaf fabric interface status history

The output from the prevous stage (leaf_if_status) becomes the input for this one (leaf_fab_int_status_accumulate). The leaf fabric interface status history is a set of interface status time series (for each spine facing interface on each leaf). Each set member has the following keys to identify it: system_id (id of the leaf system, usually serial number), interface (name of the interface).

Purpose: create recent history time series for each interface status In terms of the number of samples, the time series will hold the smaller of: 1024 samples or samples collected during the last 'total_duration' seconds (facade parameter).

For this stage, the Accumulate processor is configured for collecting leaf fabric status history as shown in the screenshot below. The defaults are shown in the screenshot below. It states if the status changes more than 5 times within one minute

User interface for processor titled leaf fabric interface status history with sections for Inputs Processing and Advanced. Inputs include Input Name: in, Stage Name: leaf_if_status, Column Name: value. Processing shows Total Duration: 1 minute, Max Samples: 6. Advanced section lists Graph Query: Empty, Enable Streaming: False. Related to data processing or monitoring.

This processor collects and outputs leaf fabric interface status accumulate (leaf_fab_int_status_accumulate). Each interface is identified by its system ID and interface name. The inputs here are the same as the outputs from the previous processor with the addition of Count, which counts the transition states as shown below

  • System ID - ID of the leaf device, usually the serial number

  • Interface - name of the interface

  • Remote Interface - interface name on the other end

  • Remote System Label - the device name on the other end

  • Count -

  • Value - operational state of the device (up, down)

  • Updated - when the state was last updated

Table showing leaf and spine switch connections with system IDs, interfaces, status up, and updated 12 days ago.

leaf fabric interface flapping

We begin with the source processor (no inputs), which is a Service Collector configured to collect interface status telemetry for all fabric interfaces on leaf devices, as shown in the screenshot below.

The output from the prevous stage (leaf_if_status) becomes the input for this one (leaf_fab_int_status_accumulate). The leaf fabric interface status history is a set of interface status time series (for each spine facing interface on each leaf). Each set member has the following keys to identify it: system_id (id of the leaf system, usually serial number), interface (name of the interface).

Purpose: create recent history time series for each interface status In terms of the number of samples, the time series will hold the smaller of: 1024 samples or samples collected during the last 'total_duration' seconds (facade parameter).

For this stage, the Accumulate Processor is configured for collecting leaf fabric status history as shown in the screenshot below. The defaults are shown in the screenshot below. It states if the status changes more than 5 times within one minute

leaf fabric interface flapping (Range)

Purpose: Count the number of state changes in the leaf_fab_int_status_accumulate ("up" to "down" and "down" to "up"). If the count is higher than 'threshold' facade parameter return "true", otherwise "false".

Input Stage: leaf_fab_int_status_accumulate

Output Stage: if_status_flapping

Set of statuses (for each spine facing interface on each leaf), indicating if the interface has been flapping or not. Each set member has the following keys to identify it: system_id (id of the leaf system, usually serial number), interface (name of the interface).

For this stage, the Range processor is configured for (collecting leaf fabric interface flapping) as shown in the screenshot below.

Configuration interface for processor leaf fabric interface flapping with inputs, processing, and advanced settings for anomaly detection.

Network interfaces monitoring dashboard showing no interface flapping. All systems report no anomalies. Data is real-time.

percentage flapping per device interfaces

percentage flapping per device interfaces (MatchPercentage)

Input Stage: if_status_flapping

Output Stage: flapping_fab_int_perc

The Match Percentage Processor is configured for collecting flapping per device interfaces.

For this stage, the Match Percentage processor is configured for (collecting leaf fabric interface flapping) as shown in the screenshot below.

Configuration interface for processor labeled percentage flapping per device interfaces with sections for processor name, inputs, processing, and advanced settings.

Dashboard showing flapping percentage metrics for systems leaf1, leaf2, leaf3, all at 0 percent flapping. Updated 12 days ago.

system anomalous flapping

system anomalous flapping (Range)

Input Stage: flapping_fab_int_perc

Output Stage: system_flapping

Set of statuses for each leaf, indicating if the leaf has higher then acceptable percentage of flapping interfaces. Each set member has the following key to identify it: system_id (id of the leaf system, usually serial number).

The Range Processor is configured for collecting system anomalous flapping.

For this stage, the Range processor is configured for (collecting leaf fabric interface flapping) as shown in the screenshot below.

Configuration screen for processor system anomalous flapping with sections for Inputs, Processing, and Advanced settings. Inputs include input name in, stage name flapping_fab_int_perc, column name value. Processing includes anomalous range 11 or more, property value, raise anomaly true. Advanced includes empty graph query, raise on NaN false, anomaly metric logging false, retention duration 1 day, retention size 1073741824, enable streaming false.

Monitoring dashboard for system_flapping showing three system IDs: leaf1, leaf3, and leaf2. All show no anomaly, 11 percent gauge value, last updated 12 days ago. Options for filtering anomalies and showing context with data source set to Real Time.