What is llm-d?

已發表 2026 年 2 月 18 日•1-minute read

llm-d is a Kubernetes-native, open source framework that speeds up distributed large language model (LLM) inference at scale.

This means when an AI model receives complicated queries with a lot of data, llm-d provides a framework that makes processing faster.

llm-d was created by Google, NVIDIA, IBM Research, and CoreWeave. Its open source community contributes updates to improve the technology.

How Red Hat AI speeds up inference

LLM prompts can be complex and nonuniform. They typically require extensive computational resources and storage to process large amounts of data.

llm-d has a modular architecture that can support the increasing resource demands of sophisticated and larger reasoning models like LLMs.

A modular architecture allows all the different parts of the AI workload to work either together or separately, depending on the model's needs. This helps the model inference faster.

Imagine llm-d is like a marathon race: Each runner is in control of their own pace. You may cross the finish line at a different time than others, but everyone finishes when they’re ready. If everyone had to cross the finish line at the same time, you’d be tied to various unique needs of other runners, like endurance, water breaks, or time spent training. That would make things complicated.

A modular architecture lets pieces of the inference process work at their own pace to reach the best result as quickly as possible. It makes it easier to fix or update specific processes independently, too.

This specific way of processing models allows llm-d to handle the demands of LLM inference at scale. It also empowers users to go beyond single-server deployments and use generative AI (gen AI) inference across the enterprise.

How does distributed inference work?

The llm-d modular architecture is made up of:

Kubernetes: an open source container-orchestration platform that automates many of the manual processes involved in deploying, managing, and scaling containerized applications.
vLLM: an open source inference server that speeds up the outputs of gen AI applications.
Inference Gateway (IGW): a Kubernetes Gateway API extension that hosts features like model routing, serving priority, and “smart” load-balancing capabilities.

This accessible, modular architecture makes llm-d an ideal platform for distributed LLM inference at scale.

What is operationalized AI?

繼續閱讀

Why choose Red Hat for Linux?

Workloads need to be portable and scalable across environments. Red Hat Enterprise Linux is your consistent, stable foundation across hybrid cloud deployments.

什么是配置管理

配置管理是指将计算机系统、服务器和软件维持在理想、一致状态的过程。它可以通过自动化进行管理。

Artificial intelligence resources

特色產品

Red Hat AI

我們的方法

我們的產品組合

互動與學習

平台解決方案

案例

按產業分類的解決方案

探索雲技術

平台產品

精選

試用並購買

服務與支援

培訓及認證

精選

服務

培養技能

更多學習方式

對於開發人員

顧客專屬

合作夥伴專屬

構建由值得信賴的合作夥伴提供支援的解決方案

What is llm-d?

What is llm-d?

How does llm-d work?

繼續閱讀

Why choose Red Hat for Linux?

什么是配置管理

Artificial intelligence resources

特色產品

Red Hat AI

平台

Tools

Try, buy, & sell

Communicate

About Red Hat

Change page language

Red Hat legal and privacy links

Red Hat legal and privacy links