OVERVIEW
AI the open way
With open source, no one faces the uncertainty and complexity of AI alone. Navigate the guesswork, mitigate the risks, and stay in control with AI-enabled tech you already trust.
Enterprises around the world already trust and rely on our hybrid cloud platforms to work anywhere. And since real-world AI success requires technology that adapts to your business needs and new advancements, those same platforms are at the foundation of our AI offerings—and are now even easier to use. Red Hat AI is a platform that accelerates AI innovation and reduces the operational cost of developing and delivering AI solutions across hybrid cloud environments.
We can't wait to meet you at Ai4!
Visit us at booth #906. Stop by to discuss open source AI, check out a demo, and grab some swag.
SPEAKING SESSIONS
Check out Red Hat's experts during our speaking sessions.
Tuesday, August 4
3:50 - 4:10 p.m.
Inference at Scale: Deploying Open-Source LLMs in Production with vLLM and llm-d
Training gets the headlines, but inference is where generative AI meets production — and where the costs, latency, and scaling challenges actually live. This session demystifies LLM inference from the ground up: how models generate tokens, why the KV cache and the prefill/decode split matter, and what an inference server actually does to make GPUs go "brrrr." From there we turn to the hard part — running inference at scale. As reasoning and agentic workloads drive token demand 20x higher (and compute 150x higher), single-node serving stops being enough.
We'll show how Red Hat approaches distributed inference with llm-d, a Kubernetes-native, open-source initiative built jointly with Google, NVIDIA, Hugging Face, and others. You'll see how disaggregated prefill and decode, intelligent request scheduling, KV-cache-aware routing, and hardware-agnostic serving across NVIDIA, AMD, Intel, and Google TPUs come together to maximize GPU utilization and hit demanding SLOs — across edge, private, and public cloud.
We'll connect the dots from vLLM as the de facto inference engine to llm-d as the scaling layer, rounding out the picture with validated and quantized models that stretch more tokens out of fixed hardware.
Attendees will leave understanding not just what inference is, but how to deploy it efficiently and reliably at enterprise scale.
Speaker: Rob Greenberg, Principal Product Manager, Red Hat AI
Wednesday, August 5
11:20 - 11:40 a.m.
AI Agents You Can Actually Trust: Platform Best Practices for Safe AI
Your AI agent works in testing. In production, it makes unpredictable calls and breaks things. Four platform practices make agents safe: isolation, Identity, Observability, and Governance. This talk shows how to add them, and why infrastructure, not just prompting, is how you control AI at scale.
Speakers: Sawyer Bowerman, Associate Technical Marketing Manager, Red Hat
Additional Resources
Connecting data to models
Join the conversation
OPPORTUNITIES ARE OPEN
Open unlocks the world's potential
At Red Hat, our commitment to open source extends beyond technology into virtually everything we do. We collaborate and share ideas, create inclusive communities, and welcome diverse perspectives from all Red Hatters, no matter their role. It’s what makes us who we are.