Before joining Red Hat, senior product manager for edge computing Luke Thompson spent 7 years of his career working at a large industrial automation company, where he had considerable opportunity to talk with customers and learn about the operational challenges they face every day. Whether the conversation was about infrastructure management, workforce shortages, supply chain disruptions, or production efficiency, one concern remained consistent across every manufacturing industry: minimizing unplanned downtime.

The challenge of minimizing unplanned downtime

In manufacturing, unplanned downtime translates directly to lost production, which is ultimately lost revenue. If a production line goes down, the plant stops producing what it makes, whether that’s laundry detergent, pharmaceuticals, automotive parts, or food products. Plant downtime can result in anywhere from $10,000 to $2 million in lost revenue per hour. And there are many causes of downtime, because plants operate like small cities with a lot of moving parts: equipment failures, operator error, material shortages, and safety incidents are all common causes. 

In an increasingly digitized manufacturing landscape, another risk factor has emerged. Manufacturers see the benefits of virtualized infrastructure, industrial IoT, machine learning applications, and numerous other technologies that have evolved productivity. But what do these things have in common? They run on software that will eventually need updates at some point. While these updates are essential, they don’t always go as planned, and when critical infrastructure fails to update successfully, the result can be the one thing manufacturers work to avoid: unplanned downtime.

Manufacturers plan for this. They don’t just send updates to machines during production; they carve out planned maintenance windows to repair equipment, replace worn components, apply security patches, and upgrade critical infrastructure. Ideally, these maintenance windows are as quick and efficient as possible, because every minute the plant is not operating is a minute more where the plant is not generating revenue.

However, there’s a critical distinction between performing an update and recovering from a failed update. Manufacturers accept planned maintenance as part of operating a plant. What they don’t accept is a 4-hour maintenance window turning into an 8-hour production outage because an infrastructure update failed and recovery required manual intervention. What if infrastructure updates came with an automatic path back to a known-good state, reducing the risk of this happening?

Demonstration of an example scenario

To explore this capability, let’s first examine a typical update scenario in an industrial environment. To do so, imagine that you work at a brewery that is responsible for running multiple batch lines. As a member of the team responsible for keeping devices up to date and secure, you need to adhere to a plan.

You have devices that require Red Hat Enterprise Linux (RHEL) OS updates. These updates could be security fixes, new application updates, etc. Your plan is as follows:

  1. Gain preapproval for a planned outage time.
  2. During that outage, stop the process that each device controls and put it into a safe state.
  3. Perform the update to the OS and reload any pertinent software or applications.
  4. Get the greenlight that the OS installed correctly and the application has been validated as working.
  5. For the remaining devices requiring updates, start the process again, watching to make sure everything looks good.
  6. If it fails to start for whatever reason, manually revert to the previous (working) version. Depending on how much time is left in your outage window, either troubleshoot and try again, or gain approval for a new outage window.

You’ll notice that the rollback steps here require manual intervention, consuming valuable time within the planned outage window. RHEL includes GreenBoot, a health-check framework that can detect when an updated system fails predefined checks and trigger an automatic rollback to the previous working state. Using RHEL and GreenBoot automatic rollback capabilities, we can automate these rollback steps, reducing recovery time and the risk of exceeding the planned outage window.

When RHEL boots off of an updated OS image, GreenBoot must make a determination between possibilities, as illustrated in Figure 1:

  • The system is working properly and should remain running on the updated image.
  • There is an issue and GreenBoot should roll back the system to the previous image.

GreenBoot relies on health check scripts to make this determination. These scripts are coordinated ahead of time to verify that requirements are met.

Diagram showing RHEL's update flow

 Figure 1. RHEL OS update flow

With the demo below, you can explore how a health check for our brewing process control application is used to check the health of an OS update. In this demo, you will also be using Red Hat Edge Manager to orchestrate the management of our brewing device fleets. You can mouse over the full-screen icon at the top right of the demo to expand the window, or click here to open it in a new browser window.

 

 

Final thoughts: Reducing operational risk through automated recovery

As we just saw, this capability is deceptively simple but critically important to preventing a situation where operators are unable to quickly revert back to a known working state. Instead of spending valuable maintenance time diagnosing and manually reversing a failed update, operators can rely on an automated recovery path that helps keep planned outages from becoming costly production disruptions Minimizing downtime through this type of automated lifecycle management is at the heart of the advanced compute platform from Red Hat—a unified, open source framework designed to bring enterprise-grade resilience and zero-touch recovery directly to the plant floor edge. By combining declarative updates with automated health checks, manufacturers can modernize their operations without risking critical production availability.

In our next article in this series, we'll explore how the Red Hat Edge portfolio helps manufacturers strengthen security at the edge without compromising operational integrity.

Ready to get started? 

Thanks to Tim Mirth, Josh Swanson, and Sean O’Keeffe for their reviews on this article.

Hatville: miniature city where edge computing comes to life

Hatville. Explore this tiny metropolis and experience how edge computing can revolutionize industries all around us.

About the authors

Ken Osborn is a Principal Technical Marketing Manager at Red Hat focused on edge technologies. With more than 20 years of experience in the IT industry, Ken has held roles spanning pre-sales engineering, product and technical marketing, with a strong emphasis on helping organizations bridge traditional infrastructure with modern cloud-native platforms.

Luke Thompson is a Senior Product Manager for the Edge Computing portfolio at Red Hat, focused on delivering platform capabilities that simplify how edge technologies are adopted and used in real-world environments.

He brings a strong background in industrial automation and IT/OT convergence, with experience helping organizations integrate complex systems, streamline data flows, and turn operational data into actionable insights in manufacturing and industrial settings.

Today, Luke applies this expertise to bridge the gap between sophisticated IT infrastructure and the practical needs of users at the edge, making advanced technologies more accessible, usable, and impactful.

UI_Icon-Red_Hat-Close-A-Black-RGB

Keep exploring

Browse by channel

automation icon

Automation

The latest on IT automation for tech, teams, and environments

AI icon

Artificial intelligence

Updates on the platforms that free customers to run AI workloads anywhere

open hybrid cloud icon

Open hybrid cloud

Explore how we build a more flexible future with hybrid cloud

security icon

Security

The latest on how we reduce risks across environments and technologies

edge icon

Edge computing

Updates on the platforms that simplify operations at the edge

Infrastructure icon

Infrastructure

The latest on the world’s leading enterprise Linux platform

application development icon

Applications

Inside our solutions to the toughest application challenges

Virtualization icon

Virtualization

The future of enterprise virtualization for your workloads on-premise or across clouds