IT leaders face a frustrating dilemma when scaling big data: Should they force engineering teams to rewrite applications running on traditional virtual machines (VMs) for containers, or maintain separate, expensive infrastructure environments? Red Hat OpenShift lets organizations modernize VMs alongside containers—without forcing immediate application refactoring.
To help organizations facing this dilemma, Cloudera has validated Cloudera Data Platform (CDP) on Red Hat OpenShift Virtualization at scale—testing deployments across more than 100 virtualized nodes. On the data side, organizations need consistent tools to govern workloads across hybrid cloud environments. Cloudera Data Platform provides this unified foundation, managing, analyzing, and governing data across hybrid cloud environments. Traditionally deployed on bare-metal servers alongside Red Hat OpenShift, Cloudera Data Platform can now run on Red Hat OpenShift Virtualization, allowing enterprises to consolidate application and data workloads on a unified application platform.
Here is a breakdown of how Cloudera validated this architecture on Red Hat OpenShift Virtualization, along with practical deployment options for your datacenter.
Red Hat OpenShift Virtualization at a glance
Red Hat OpenShift Virtualization extends Red Hat OpenShift by allowing VMs to run side by side with containers on the same platform. OpenShift Virtualization relies on Kernel-based Virtual Machine (KVM) and KubeVirt—proven open source technologies that enterprise IT teams have trusted in production for more than a decade.
By running OpenShift Virtualization, organizations are able to:
- Run traditional VM workloads within Red Hat OpenShift.
- Modernize applications at their own pace.
- Simplify infrastructure management.
- Consolidate virtualization and container platforms.
- Use Kubernetes-native automation and operations.
Organizations can now host Cloudera Data Platform infrastructure components to benefit from Red Hat OpenShift’s operational model, while preserving compatibility with existing deployment requirements.
Recently, in collaboration with Red Hat, Cloudera tested the deployment of Cloudera Base on premises with OpenShift Virtualization. The validation results were published in an article on Cloudera’s blog and are now listed in the Red Hat Ecosystem Catalog.
Let’s take a closer look at the benefits of this deployment.
Cloudera Data Platform with OpenShift benefits
Cloudera Data Platform on OpenShift Virtualization creates a unified, modernized architecture for on-premise data infrastructure.
Key benefits include:
- Validated scale and architecture: The platform provides a tested blueprint for production-grade deployments, successfully running Cloudera Base on premises on over 100 virtualized nodes.
- Unified operational model: The converged platform allows organizations to run both VMs (for Cloudera Base) and containerized workloads on a single Kubernetes platform. This simplifies lifecycle management, eliminates split operational models, and automates resource scaling and VM provisioning through GitOps and Red Hat Ansible Automation Platform.
- Flexible hybrid deployments: This architecture allows Cloudera Base with OpenShift Virtualization to connect directly to Cloudera Data Services running on bare metal or VMs, giving teams consistent operational control across environments.
- Technical foundation: The solution uses Red Hat Enterprise Linux (RHEL), Red Hat OpenShift, and OpenShift Virtualization for Cloudera Base on premises, while Cloudera Data Services on premises uses Red Hat OpenShift. Always consult the Cloudera Product Support Matrix to determine the versions of RHEL and Red Hat OpenShift that are supported with Cloudera Data Platform. Because Cloudera's support matrix specifies supported operating systems rather than hypervisors, running Cloudera Base with OpenShift Virtualization relies on supported RHEL guest images.
- Robust security and governance: The environment enforces central identity and access policies across both VMs and containers using Kerberos, LDAP, and FreeIPA, with granular access control via Apache Ranger and gateways through Knox.
- Real-world validation: Functional testing confirmed end-to-end support for core Cloudera services—including Cloudera Data Warehouse (CDW), Cloudera Data Engineering (CDE), and Cloudera AI (CAI). To test real-world limits, Cloudera’s engineering team collaborated with Red Hat to run a bank branch performance analytics pipeline—pushing heavy data streams through the lifecycle from ingestion to visualization.
- Improved infrastructure utilization: By running VMs with OpenShift Virtualization, teams pack more workloads onto fewer physical servers—cutting hardware costs while simplifying management.
Let’s take a look at the environment we used for our lab deployment.
Cloudera Data Platform with OpenShift reference architecture
The Cloudera Data Platform is a hybrid and multi-cloud data solution that manages the entire lifecycle of data, analytics, and AI—from edge ingestion to generative AI deployment.
For enterprise architects evaluating sizing, the following host and node breakdowns illustrate how to structure a large deployment. Architects typically deploy Cloudera Base within VMs to simplify migration and lifecycle management for traditional stateful infrastructure, while deploying Cloudera Data Services directly on bare-metal Red Hat OpenShift worker nodes to maximize containerized analytics throughput. In this example, we used 2 separate clusters to run Cloudera on-premise using Red Hat OpenShift: 1 virtualized host cluster for Cloudera Base, and 1 bare-metal cluster for Cloudera Data Services.
1. Cloudera Base on premises with OpenShift Virtualization (Cloudera Base VM versus bare metal)
This architecture, illustrated in Figure 1, runs Cloudera Base on premises with VMs for scalability. To prove enterprise-grade scalability, Red Hat and Cloudera validated a massive 110-node benchmark cluster. While your organization's environment will vary based on analytics density, this tested architecture provides a proven baseline for sizing your deployment.
Host cluster: 110 physical bare-metal server nodes running the Red Hat OpenShift host cluster.
- Nodes: Three Red Hat OpenShift control plane nodes, 24 Red Hat OpenShift storage/worker nodes, and 83 Red Hat OpenShift worker nodes.
- Platform layer: Includes Red Hat OpenShift Data Foundation, local storage operator, and OpenShift Virtualization to host VMs.
Cloudera Base nodes (VMs): Cloudera Base on premises is deployed to 100 Red Hat Enterprise Linux VMs.
- Nodes: Three Cloudera Base master nodes, 95 Cloudera Base worker nodes, 1 Cloudera Manager node, and 1 Cloudera utility node.
- Other: An IPA server provides identity management (IdM), DNS, LDAP, and REPO services.
Figure 1: Cloudera Base on premises with OpenShift Virtualization reference architecture.
2. Cloudera Data Services on premises
This architecture, illustrated in Figure 2, focuses on running Cloudera Data Services on premises on Red Hat OpenShift bare-metal worker nodes.
Host Cluster: Seven physical bare-metal server nodes running the Red Hat OpenShift host cluster.
- Nodes: Three Red Hat OpenShift control plane nodes and 4 OpenShift bare-metal storage/worker nodes.
- Platform components: OpenShift Data Foundation and local storage operator are used for storage.
- Other: An IPA server for IdM, DNS, LDAP, and REPO services.
Figure 2: Cloudera Data Services on premises with Red Hat OpenShift reference architecture.
Deployment considerations
Before deploying OpenShift Virtualization for Cloudera Data Platform workloads, plan for key technical requirements:
- Account for hypervisor memory reserves (typically 5–10% per host node) and size storage controllers for high-concurrency I/O throughput before migrating stateful Cloudera Base VMs.
- To achieve peak analytics performance, size your storage throughput for heavy concurrent I/O and account for KVM memory overhead during initial host cluster planning. Review the complete Cloudera validation results, or connect with a Red Hat architect to evaluate OpenShift Virtualization in your datacenter.
This validation demonstrates that Cloudera Base environments with 100 nodes or more can maintain sub-second query responsiveness on OpenShift Virtualization, giving IT leaders a tested blueprint to eliminate redundant infrastructure. Organizations gain cloud-native virtualization for their data lakehouse, hybrid deployment flexibility, and unified security and governance. Whether you need to consolidate aging hypervisors or scale a containerized analytics pipeline, this benchmarked architecture gives you a validated baseline to build on in your own datacenter.
Ready to evaluate OpenShift Virtualization for your data lakehouse? Read the full Cloudera validation results or schedule a sizing assessment with a Red Hat architect.
Prueba del producto
Red Hat OpenShift Virtualization Engine | Versión de prueba
Sobre los autores
Patrick is a Technical Sales Strategist with the Ecosystem Center of Excellence team at Red Hat. He joined Red Hat in 2019 and currently works with our OEM and ISV partner ecosystem. Patrick is passionate about creating AI/ML, infrastructure and platform solutions with OpenShift.
Firas Yasin is a distinguished technology leader and award-winning author, currently serving as the Global Alliance Manager for AI/ML partnerships at Red Hat. With an impressive career journey, Firas has excelled in various roles, from being a Global Sales Leader at IBM to a skilled software engineer and a visionary Global Lead Architect.
With a keen eye for transformation, Firas navigated through various roles, from software engineer to the strategic position of a Global Lead Architect. He also served as a Sales Director at Atos for Hybrid Cloud. Throughout his career, Firas has demonstrated a remarkable ability to adapt to the ever-evolving technology landscape. Prior to his AI/ML focus, he focused on Cybersecurity partnerships at Red Hat. In summary, Firas is a dynamic professional, a thought leader in technology, and a driving force in the AI/ML partnerships.
Kuldeep Sahu is a Partner Solutions Engineer at Cloudera, passionate about helping partners turn Cloudera's data and AI capabilities into real, production-ready solutions for their customers. Based in Indore, India, he works closely with the partner ecosystem to position Cloudera's value proposition and drive successful platform adoption across the region. With a strong background in DevOps, cloud infrastructure, and automation, Kuldeep brings an engineering-first perspective to every partner engagement — helping teams design architectures that are scalable, resilient, and built for real-world production. He's always exploring emerging technologies in cloud, Data and AI, and enjoys tackling complex technical challenges through smart automation and thoughtful system design.
Más como éste
Gestión de máquinas virtuales en Red Hat OpenShift con Service Mesh
Red Hat Enterprise Linux CoreOS 10 llega pronto a Red Hat OpenShift
Virtualization Is (Still) King | Compiler
Transforming Your Secrets Management | Code Comments
Navegar por canal
Automatización
Las últimas novedades en la automatización de la TI para los equipos, la tecnología y los entornos
Inteligencia artificial
Descubra las actualizaciones en las plataformas que permiten a los clientes ejecutar cargas de trabajo de inteligecia artificial en cualquier lugar
Nube híbrida abierta
Vea como construimos un futuro flexible con la nube híbrida
Seguridad
Vea las últimas novedades sobre cómo reducimos los riesgos en entornos y tecnologías
Edge computing
Conozca las actualizaciones en las plataformas que simplifican las operaciones en el edge
Infraestructura
Vea las últimas novedades sobre la plataforma Linux empresarial líder en el mundo
Aplicaciones
Conozca nuestras soluciones para abordar los desafíos más complejos de las aplicaciones
Virtualización
El futuro de la virtualización empresarial para tus cargas de trabajo locales o en la nube