Introduction to llm-d
게시됨 May 20, 2026
Click through for a high-level tour of Red Hat’s LLM-D distributed inference demo, showcasing how KV cache–aware routing accelerates responses and improves efficiency.
Note: This demo may contain AI-generated content and/or media. All AI-generated content was reviewed or edited by a human before being made available to you.
바로 가기
바로 가기
What you'll learn in this interactive demo
This experience highlights:
- How LLM-D separates prefill and decode stages to scale each independently
- Testing prompts and seeing cache hit rates in action
- Comparing repeated and unique prompts to show the impact on latency and load distribution
- Using the inference gateway for load testing and measuring session stickiness
- Tracking per-pod cache hit rates, token throughput, and latency in Grafana
Next steps
Red Hat LLM-D helps scale distributed inference across environments, improving efficiency and reducing latency for generative AI applications.
- Read the press release:
Red Hat launches LLM-D community영어 (English) 버전으로 제공됩니다 (한국어 미지원) - Explore the broader Red Hat AI portfolio
- Try it with a Red Hat AI trial
About the author of this page
플랫폼
툴
체험, 구매 & 영업
커뮤니케이션
Red Hat 소개
Red Hat은 Fortune 선정 500대 기업이 신뢰하는 어드바이저이며, 클라우드, 개발자, Linux, 자동화, 애플리케이션 플랫폼 기술 분야에서 전문성은 물론 수상 경력을 갖춘 서비스를 제공합니다.