Before joining Red Hat, senior product manager for edge computing Luke Thompson spent 7 years of his career working at a large industrial automation company, where he had considerable opportunity to talk with customers and learn about the operational challenges they face every day. Whether the conversation was about infrastructure management, workforce shortages, supply chain disruptions, or production efficiency, one concern remained consistent across every manufacturing industry: minimizing unplanned downtime.
The challenge of minimizing unplanned downtime
In manufacturing, unplanned downtime translates directly to lost production, which is ultimately lost revenue. If a production line goes down, the plant stops producing what it makes, whether that’s laundry detergent, pharmaceuticals, automotive parts, or food products. Plant downtime can result in anywhere from $10,000 to $2 million in lost revenue per hour. And there are many causes of downtime, because plants operate like small cities with a lot of moving parts: equipment failures, operator error, material shortages, and safety incidents are all common causes.
In an increasingly digitized manufacturing landscape, another risk factor has emerged. Manufacturers see the benefits of virtualized infrastructure, industrial IoT, machine learning applications, and numerous other technologies that have evolved productivity. But what do these things have in common? They run on software that will eventually need updates at some point. While these updates are essential, they don’t always go as planned, and when critical infrastructure fails to update successfully, the result can be the one thing manufacturers work to avoid: unplanned downtime.
Manufacturers plan for this. They don’t just send updates to machines during production; they carve out planned maintenance windows to repair equipment, replace worn components, apply security patches, and upgrade critical infrastructure. Ideally, these maintenance windows are as quick and efficient as possible, because every minute the plant is not operating is a minute more where the plant is not generating revenue.
However, there’s a critical distinction between performing an update and recovering from a failed update. Manufacturers accept planned maintenance as part of operating a plant. What they don’t accept is a 4-hour maintenance window turning into an 8-hour production outage because an infrastructure update failed and recovery required manual intervention. What if infrastructure updates came with an automatic path back to a known-good state, reducing the risk of this happening?
Demonstration of an example scenario
To explore this capability, let’s first examine a typical update scenario in an industrial environment. To do so, imagine that you work at a brewery that is responsible for running multiple batch lines. As a member of the team responsible for keeping devices up to date and secure, you need to adhere to a plan.
You have devices that require Red Hat Enterprise Linux (RHEL) OS updates. These updates could be security fixes, new application updates, etc. Your plan is as follows:
- Gain preapproval for a planned outage time.
- During that outage, stop the process that each device controls and put it into a safe state.
- Perform the update to the OS and reload any pertinent software or applications.
- Get the greenlight that the OS installed correctly and the application has been validated as working.
- For the remaining devices requiring updates, start the process again, watching to make sure everything looks good.
- If it fails to start for whatever reason, manually revert to the previous (working) version. Depending on how much time is left in your outage window, either troubleshoot and try again, or gain approval for a new outage window.
You’ll notice that the rollback steps here require manual intervention, consuming valuable time within the planned outage window. RHEL includes GreenBoot, a health-check framework that can detect when an updated system fails predefined checks and trigger an automatic rollback to the previous working state. Using RHEL and GreenBoot automatic rollback capabilities, we can automate these rollback steps, reducing recovery time and the risk of exceeding the planned outage window.
When RHEL boots off of an updated OS image, GreenBoot must make a determination between possibilities, as illustrated in Figure 1:
- The system is working properly and should remain running on the updated image.
- There is an issue and GreenBoot should roll back the system to the previous image.
GreenBoot relies on health check scripts to make this determination. These scripts are coordinated ahead of time to verify that requirements are met.
Figure 1. RHEL OS update flow
With the demo below, you can explore how a health check for our brewing process control application is used to check the health of an OS update. In this demo, you will also be using Red Hat Edge Manager to orchestrate the management of our brewing device fleets. You can mouse over the full-screen icon at the top right of the demo to expand the window, or click here to open it in a new browser window.
Final thoughts: Reducing operational risk through automated recovery
As we just saw, this capability is deceptively simple but critically important to preventing a situation where operators are unable to quickly revert back to a known working state. Instead of spending valuable maintenance time diagnosing and manually reversing a failed update, operators can rely on an automated recovery path that helps keep planned outages from becoming costly production disruptions Minimizing downtime through this type of automated lifecycle management is at the heart of the advanced compute platform from Red Hat—a unified, open source framework designed to bring enterprise-grade resilience and zero-touch recovery directly to the plant floor edge. By combining declarative updates with automated health checks, manufacturers can modernize their operations without risking critical production availability.
In our next article in this series, we'll explore how the Red Hat Edge portfolio helps manufacturers strengthen security at the edge without compromising operational integrity.
Ready to get started?
- Learn more about how Red Hat supports digital transformation in manufacturing.
- Download the e-book “How to unite modern and traditional at the industrial edge.”
- Contact your Red Hat representative to discuss your edge deployment scenarios.
Thanks to Tim Mirth, Josh Swanson, and Sean O’Keeffe for their reviews on this article.
Hatville: miniature city where edge computing comes to life
About the authors
Ken Osborn is a Principal Technical Marketing Manager at Red Hat focused on edge technologies. With more than 20 years of experience in the IT industry, Ken has held roles spanning pre-sales engineering, product and technical marketing, with a strong emphasis on helping organizations bridge traditional infrastructure with modern cloud-native platforms.
Luke Thompson is a Senior Product Manager for the Edge Computing portfolio at Red Hat, focused on delivering platform capabilities that simplify how edge technologies are adopted and used in real-world environments.
He brings a strong background in industrial automation and IT/OT convergence, with experience helping organizations integrate complex systems, streamline data flows, and turn operational data into actionable insights in manufacturing and industrial settings.
Today, Luke applies this expertise to bridge the gap between sophisticated IT infrastructure and the practical needs of users at the edge, making advanced technologies more accessible, usable, and impactful.
More like this
Insights-client updates standardize RHEL package management
Better together: RHEL system roles and Red Hat Ansible Automation Platform
Infrastructure At The Edge | Compiler
Operating System Management | Compiler
Keep exploring
- Boost security, flexibility, and scale at the edge with Red Hat Enterprise LinuxE-book
- Red Hat and Cisco: Extending network automation at the edge
Video - Samsung propels 5G and edge networks with Red Hat OpenShift
Case study
Browse by channel
Automation
The latest on IT automation for tech, teams, and environments
Artificial intelligence
Updates on the platforms that free customers to run AI workloads anywhere
Open hybrid cloud
Explore how we build a more flexible future with hybrid cloud
Security
The latest on how we reduce risks across environments and technologies
Edge computing
Updates on the platforms that simplify operations at the edge
Infrastructure
The latest on the world’s leading enterprise Linux platform
Applications
Inside our solutions to the toughest application challenges
Virtualization
The future of enterprise virtualization for your workloads on-premise or across clouds