AI clouds have advanced beyond initial pilot testing, and the challenge now isn’t simply configuring the physical hardware. Organizations must focus on operating an efficient shared platform that provides predictable operating costs, access to the latest computer chips, and ongoing platform updates without relying on fragile custom code or complex software adjustments.

By integrating NVIDIA DSX OS™ software, part of the DSX platform, with Red Hat AI, we are co-engineering a deployment framework for AI clouds that accelerates innovation while delivering the reliability, scalability, and operational excellence required by NVIDIA Cloud Partners (NCPs). The integration of NVIDIA Infra Controller (NICo), a software component of DSX OS, enables operators to ease and simplify lifecycle management by leveraging automated bare-metal provisioning, and to quickly establish secure, multi-tenant revenue streams.

The era of generative AI has evolved into the era of the agentic AI Factory, where data centers operate as industrial engines designed to continuously manufacture intelligence in the form of tokens. As AI clouds and infrastructure teams scale these environments from megawatts to gigawatts, success no longer depends on experimentation and pilots but on operational efficiency, tokens-per-watt, and total cost per token.

As a core component of the NVIDIA DSX Platform, open-source DSX OS software is designed to efficiently operate and scale multi-tenant AI clouds.  As an inaugural NVIDIA AI Cloud Ready validated partner, Red Hat combines Red Hat AI with DSX to bring business-critical capabilities required by AI clouds, enabling them to deploy and manage at scale.

By combining NVIDIA accelerated hardware and reference designs with a certified, validated, and trusted platform in Red Hat AI, we are establishing the benchmark for fully integrated AI factories and Model Context Protocol (MCP) tool execution. Here is a look at the co-engineering work that makes this "Better Together" story a reality for AI clouds.

The foundation: Zero-day architecture with Project Voyager

An AI factory is only as optimized as the foundational operating layer supporting it. To eliminate custom, risky development that can delay infrastructure deployments by months, Red Hat and NVIDIA have embedded Day 0 platform support directly into the operating system through Red Hat Enterprise Linux for NVIDIA.

Through a dedicated CentOS Special Interest Group (the Accelerated Infrastructure Enablement SIG), both companies co-engineer deep kernel-level integration, automated driver signing, and upstream code paths for advanced NVIDIA platforms such as Grace Blackwell and Vera Rubin. This co-engineered pipeline pairs immediate access to new silicon performance with the production stability, hardened security and predictable lifecycles customers expect from Red Hat Enterprise Linux.

Driving efficiency: Operational and lifecycle excellence

Operating a gigawatt-scale AI factory requires continuous, auditable, and secure management across thousands of multi-tenant nodes. Red Hat AI addresses these AI cloud infrastructure demands by integrating directly with core DSX OS lifecycle and provisioning components:

  • NVIDIA Infra Controller™ (NICo): Managing DPU deployments at scale presents a major operational challenge. NICo includes NVIDIA DOCA™ Platform Framework (DPF), which delivers end-to-end NVIDIA BlueField DPU provisioning and lifecycle management. By integrating NICo with Red Hat OpenShift, bare-metal lifecycle management becomes fully API-driven. With NVIDIA BlueField-3, our joint architecture enforces secure, hardware-isolated, and multi-tenant IaaS boundaries managed directly from the Red Hat OpenShift control plane.
  • NVIDIA AI Cluster Runtime (AICR): Configuration drift is the leading cause of unexpected job failures. AICR prevents this by enforcing versioned, validated configuration recipes. Through close, upstream collaboration, Red Hat delivers in-tree Helm template-style integration for AICR components (including Node Feature Discovery, NVIDIA GPU Operators, and Network Operators) on Red Hat OpenShift, ensuring software parity across global deployments.

Automated resiliency: Eliminating downtime in high-density GPU fleets

In an environment featuring tens of thousands of GPUs, continuous health visibility across every node is a daily operational necessity. Manual alerting loops delay recovery times and burn power and compute cycles. By combining DSX OS health automation with Red Hat OpenShift, AI cloud operators can automatically detect, quarantine, and remediate system anomalies in real time without human intervention.

  • NVIDIA NVSentinel: Red Hat and NVIDIA validate NVSentinel deployment on Red Hat OpenShift. Utilizing custom OpenShift manifests and automated scripting, NVSentinel provides Kubernetes-native GPU fault detection. Upon detecting critical hardware or driver events, it automatically isolates compromised nodes and reassigns workloads to healthy nodes to protect jobs.
  • NVIDIA Fleet Intelligence: Running Fleet Intelligence agents on Red Hat OpenShift nodes enables continuous scraping and aggregation of cluster telemetry into Grafana dashboards, giving operations teams real-time visibility across hybrid-cloud environments.

Pre-deployment validation: Building digital twins with NVIDIA DSX Air

Before deploying physical hardware, architects can model and validate operational workflows using NVIDIA DSX Air, a SaaS-based platform that simulates data center computing, networking, and switch configurations. By integrating lightweight Red Hat footprints like Single Node OpenShift  into DSX Air, AI cloud engineering teams can safely simulate multi-tenant network topologies, test BGP routing, and validate Day 2 software upgrade scripts on virtual nodes before pushing changes to production.  

Agentic operations: Securing Model Context Protocol (MCP) in a multi-tenant AI cloud

NVIDIA DSX OS surfaces data center telemetry and operational controls through domain-specific Model Context Protocol (MCP) servers, providing AI agents with a unified tool catalog. Red Hat AI governs this agentic layer by establishing namespace-scoped access controls and zero-trust security boundaries within OpenShift. Using OpenShift’s native RBAC, multi-tenancy, and audit logging, validated MCP servers allow autonomous agents to safely query cluster metrics, inspect logs, and assist with root-cause analysis, such as correlating thermal spikes with network drops—without exceeding administrative privileges or introducing operational risk.

Conclusion: Solutions for scalable AI clouds and AI factories

Building an AI factory shouldn’t feel like a science project held together by fragile glue code and custom driver plumbing. The true value of the Red Hat and NVIDIA partnership lies in turning hardware innovation into predictable, hardened, pre-validated solutions for both enterprises and NVIDIA Cloud Partners (NCPs).

By combining NVIDIA DSX OS architectures with the validated security and operational boundaries of Red Hat AI, production teams gain a production-ready platform for deploying AI clouds at scale. From zero-day hardware support to automated fleet remediation with NVSentinel, this AI cloud platform gives NCPs the speed to innovate, without the overhead of managing underlying platform complexity.

Red Hat AI Factory with NVIDIA is an enterprise-grade AI solution for building, deploying, and managing AI factories with NVIDIA’s accelerated computing, networking, and software. Red Hat AI Factory with NVIDIA combines the capabilities of Red Hat AI Enterprise with NVIDIA AI Enterprise software to deploy repeatable, scalable, and secure enterprise AI factories.

Discover how Red Hat and NVIDIA are co-engineering the standard for modern, secure, and fully automated AI factories for both enterprise and cloud-scale deployments. Explore our strategic partnership and joint solutions here.

Learn more about Red Hat AI Factory with NVIDIA

Resource

The adaptable enterprise: Why AI readiness is disruption readiness

This e-book, written by Michael Ferris, Red Hat COO and CSO, navigates the pace of change and technological disruption with AI that faces IT leaders today.

About the author

Ami Zvieli is the Director of Engineering for the Ecosystem Engineering team at Red Hat. In this leadership capacity, he serves as the Nvidia Technical Partnership lead, focusing on bridging the gap between cutting-edge hardware acceleration technologies and world-class open-source software.

UI_Icon-Red_Hat-Close-A-Black-RGB

Browse by channel

automation icon

Automation

The latest on IT automation for tech, teams, and environments

AI icon

Artificial intelligence

Updates on the platforms that free customers to run AI workloads anywhere

open hybrid cloud icon

Open hybrid cloud

Explore how we build a more flexible future with hybrid cloud

security icon

Security

The latest on how we reduce risks across environments and technologies

edge icon

Edge computing

Updates on the platforms that simplify operations at the edge

Infrastructure icon

Infrastructure

The latest on the world’s leading enterprise Linux platform

application development icon

Applications

Inside our solutions to the toughest application challenges

Virtualization icon

Virtualization

The future of enterprise virtualization for your workloads on-premise or across clouds