Introduction to llm-d
公開日 May 20, 2026
Click through for a high-level tour of Red Hat’s LLM-D distributed inference demo, showcasing how KV cache–aware routing accelerates responses and improves efficiency.
Note: This demo may contain AI-generated content and/or media. All AI-generated content was reviewed or edited by a human before being made available to you.
セクションを選択
セクションを選択
What you'll learn in this interactive demo
This experience highlights:
- How LLM-D separates prefill and decode stages to scale each independently
- Testing prompts and seeing cache hit rates in action
- Comparing repeated and unique prompts to show the impact on latency and load distribution
- Using the inference gateway for load testing and measuring session stickiness
- Tracking per-pod cache hit rates, token throughput, and latency in Grafana
Next steps
Red Hat LLM-D helps scale distributed inference across environments, improving efficiency and reducing latency for generative AI applications.
- Read the press release: Red Hat launches LLM-D community
- Explore the broader Red Hat AI portfolio
- Try it with a Red Hat AI trial
About the author of this page
プラットフォーム
ツール
試用、購入、販売
コミュニケーション
Red Hat について
Red Hat は、オープン・ハイブリッドクラウド・テクノロジーのリーダーであり、エンタープライズにおける革新的な IT および人工知能 (AI) アプリケーションのための一貫性のある包括的な基盤を提供しています。フォーチュン 500 企業に信頼されるアドバイザーとして、Red Hat はクラウド、開発者向け、Linux、自動化、アプリケーション・プラットフォームといったテクノロジーと、受賞歴のあるさまざまなサービスを提供しています。
ページの言語を選択してください
Red Hat legal and privacy links
- Red Hat について
- 採用情報
- イベント
- 各国のオフィス
- Red Hat へのお問い合わせ
- Red Hat ブログ
- Red Hat におけるインクルージョン
- Cool Stuff Store
- Red Hat Summit