Introduction to llm-d
Click through for a high-level tour of Red Hat’s LLM-D distributed inference demo, showcasing how KV cache–aware routing accelerates responses and improves efficiency.
快速跳转
快速跳转
What you'll learn in this interactive demo
This experience highlights:
- How LLM-D separates prefill and decode stages to scale each independently
- Testing prompts and seeing cache hit rates in action
- Comparing repeated and unique prompts to show the impact on latency and load distribution
- Using the inference gateway for load testing and measuring session stickiness
- Tracking per-pod cache hit rates, token throughput, and latency in Grafana
Next steps
Red Hat LLM-D helps scale distributed inference across environments, improving efficiency and reducing latency for generative AI applications.
- Read the press release: Red Hat launches LLM-D community
- Explore the broader Red Hat AI portfolio
- Try it with a Red Hat AI trial
Related resources
About the author of this page
Note: This demo may contain AI-generated content and/or media. All AI-generated content was reviewed or edited by a human before being made available to you.
平台
工具
试用购买与出售
联系我们
关于红帽
红帽是开放混合云技术的领导者,为企业变革性 IT 和人工智能 (AI) 应用提供一致、全面的基础。作为深受《财富》500 强企业信赖的顾问,红帽提供云、开发人员、Linux、自动化和应用平台技术,以及屡获殊荣的服务。