Red Hat OpenShift

Introduction to llm-d

Interactive demo2 minsMay 20, 2026

  Christopher Nuland

Click through for a high-level tour of Red Hat’s LLM-D distributed inference demo, showcasing how KV cache–aware routing accelerates responses and improves efficiency.

3D graphic of OpenShift icon, a cloud, and a cursor aimed at a purple target
What you'll learn Next steps Recursos Feedback

What you'll learn in this interactive demo

This experience highlights:

  • How LLM-D separates prefill and decode stages to scale each independently
  • Testing prompts and seeing cache hit rates in action
  • Comparing repeated and unique prompts to show the impact on latency and load distribution
  • Using the inference gateway for load testing and measuring session stickiness
  • Tracking per-pod cache hit rates, token throughput, and latency in Grafana

Next steps

Red Hat LLM-D helps scale distributed inference across environments, improving efficiency and reducing latency for generative AI applications.

About the author of this page

Christopher Nuland headshot

Christopher Nuland

Technical Marketing Manager for AI Inference

Christopher Nuland is a Principal Technical Marketing Manager for AI at Red Hat and has been with the company for over six years. Before Red Hat, he focused on machine learning and big data analytics for companies in the finance and agriculture sectors. Once coming to Red Hat, he specialized in clou...