# A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

Intel White Paper Closed-Loop Automation - Telemetry-Aware Scheduler for Service Repair and Platform Recovery

Summary

Maximize platform resiliency and ensure high Quality of Experience (QoE) with Closed Loop Automation. This guide presents an architectural overview of a resilient prototyping system that uses real-time telemetry, analytics, and Intel solutions to optimize workload placement. By automating self-healing actions before service disruption occurs, it helps Service Providers meet critical Service Level Agreements (SLAs). Ideal for network architects and operators implementing advanced MANO, cloud, or telecommunications infrastructure seeking proactive operational resilience.

📄 Preview 📖 Table of Contents Contents PAGE OF 10

Page 1 Text Content

White Paper Intel Corporation

Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience

Authors 1 Introduction John Browne Closed loop automation is the process of continually monitoring real-time network

conditions, workload requirements, and resource capabilities and availability to

Emma Collins

determine the optimum workload placement for faster and more efficient delivery of Krzysztof Kepka services. Service Providers are adopting Closed Loop Automation as a strategic priority to address their business objectives [1.]. One such objective is to ensure customer Service

Sunku Ranganath

Level Agreements (SLAs) are met to provide an optimum Quality of Experience (Qo E), Jabir Kanhira Kadavathu which is delivered according to operator defined policies. Network Services rely on the infrastructure they run on and so this is a crucial place to start with when trying to ensure

Swati Sehgal

an optimal Qo E.

Killian Muldoon

This document provides an overview of a platform resiliency prototype that showcases Michal Kobylinski the integration of key Closed Loop Automation components, leveraging a mix of solutions provided by Intel and the open source community. Traditional Telecom Services Assurance is an Operations Support System (OSS) function carried out offline [1.] without automated processes and methods for self-optimization of the network. This paper outlines ‘near real-time’ automation at a granular platform level. This allows an orchestration system to detect and respond to an issue before a customer is impacted, including integration with Management and Network Orchestration (MANO), using platform telemetry, and analytics. Closed loop solutions integrate components that include platform, telemetry, analytics processing, MANO and polices to create automated processes. The Intel platform has a wealth of features and resources that provide a rich set of data and control points that can be used for configuration, reporting, monitoring and managing workloads in a network infrastructure. This data can be used as part of a Closed Loop Solution to provide granular insights into the behavior of the system. When these metrics provide actionable alerts, the configuration controls can be used to provide platform resiliency support and ensure minimal unplanned operational downtime. By collecting Intel® Architecture specific feature telemetry, this prototype will provide host insights to a new Kubernetes* (K8s*) scheduling extension, called Telemetry Aware Scheduling (TAS), to trigger corrective healing actions on a workload and influence workload placement decisions, depending on the reliability of the underlying infrastructure. By gaining these insights into the health of the platform, the services they are running on, and having the control to be able to remediate against faults and errors in an automated way, this prototype provides the methodology to reduce risk of unplanned outages and enable scheduled downtime for maintenance, therefore maximizing service availability. Along with the video showing this use case in action included in the release, this document is part of the Network Transformation Experience Kit, which is available at: https://networkbuilders.intel.com/network-technologies/network-transformation-exp- kits

Page Summary Contents For Intel White Paper Closed-Loop Automation - Telemetry-Aware Scheduler for Service Repair and Platform Recovery

Page 1 White Paper Intel Corporation Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience Authors 1 Introduction John Browne Closed loop automation is the process of...
Page 2 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience Table of Contents 3.2.1 Host Health Indicator Calculation Logic ..............................
Page 3 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience 2 Document Overview This document describes the components of a closed loop solution for au...
Page 4 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience 3 Components of the Intel® Closed Loop Automation Solution This section includes the ETSI N...
Page 5 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience All of these features have an associated collectd* plugin that can be found as part of coll...
Page 6 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience The following figure is an example of mapping the prototype Host Health Indications to the ...
Page 7 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience The TAS telemetry policy takes the form: api Version: telemetry.ie/v1 kind: Telemetry Polic...
Page 8 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience Once a workload is placed TAS continues to monitor metrics specified in the telemetry polic...
Page 9 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience 4.1 Method of Operation The stress-ng application is used in this prototype both as applica...
Page 10 White Paper | Closed Loop Automation - Telemetry Aware Scheduler for Service Healing and Platform Resilience 2. Amber/Yellow Scenario: An amber indicator will identify that a minor fault has been foun...

Manual Details

Brand Intel
Pages 10
File Size 630.52 KB
Published June 03, 2026
32 views

Enter the captcha to get the download link:

captcha

Frequently Asked Questions

What is Closed Loop Automation used for in networking?

It continually monitors real-time network conditions to determine the optimum workload placement for faster and more efficient service delivery.

How does the system achieve high service availability?

By collecting platform telemetry, it can trigger corrective healing actions and remediate faults automatically, minimizing unplanned outages.

What is the role of Telemetry Aware Scheduling (TAS)?

TAS is a Kubernetes extension that uses host insights to influence workload placement decisions based on underlying infrastructure reliability.

Which components are integrated for closed loop automation?

The process requires platform telemetry, analytics processing, Management and Network Orchestration (MANO), and defined policies.