The AI industry is shifting its attention from training models to running models more efficiently. Enterprise AI applications generate millions of inference requests as they coordinate multiple models, tools, and agents. Inference is the process where a trained AI model generates a response to a user prompt. Every request consumes computing capacity, making inference efficiency one of the primary drivers of both AI performance and infrastructure cost.

The challenge is both acquiring enough capacity and using that capacity intelligently. As model parameters have grown exponentially in size and complexity, traditional deployment approaches have struggled to keep up. Many organizations still rely on standard Kubernetes load balancing to distribute inference requests,  which was originally designed for stateless applications. 

However, large language model (LLM) inference behaves differently. LLMs depend heavily on the key-value (KV) cache, the model's short-term memory, to avoid repeating work. When conventional load balancers blindly distribute related requests evenly across different inference servers, cache locality is lost. The result is redundant computation, unnecessary GPU utilization, and unnecessary infrastructure costs as the model is forced to recompute work it has already performed.

From better models to better inference

Scaling AI isn't simply a matter of adding more inference servers, inference infrastructure needs to be LLM-aware. Open source has played a central role in making AI models more accessible, and 2 open source projects, vLLM and llm-d, are doing the same for inference.

  • vLLM is an open source inference engine optimized for running models more efficiently on a single node or server.
  • llm-d acts as the overarching control plane for distributed inference. While vLLM increases single-node efficiency, llm-d spans multiple vLLM instances to orchestrate and optimize the entire cluster.

Using intelligent, KV cache-aware scheduling, llm-d understands where computed prompts already exist and routes inference requests accordingly. By preserving cache locality across multiple inference instances, it minimizes redundant computation, improves GPU utilization, reduces latency, and helps lower the cost of serving AI workloads. Capabilities such as intelligent request routing and inference disaggregation also help organizations to extract more value from the infrastructure they already have instead of simply adding more hardware.

Looking ahead

Open source projects like vLLM and llm-d demonstrate how collaborative innovation is helping organizations overcome the cost and compute barriers that stand between AI experimentation and enterprise production at scale.

In Part II of this series, we'll explore why efficient inference is only one piece of the puzzle, and why infrastructure independence is that path to sovereign AI. 

To learn more, watch Red Hat's Brian Stevens and Robert Shaw discuss the fundamental challenges of scaling AI inference.

리소스

AI 추론 시작하기

더욱 스마트하고 효율적인 AI 추론 시스템을 구축하는 방법을 알아보세요. Red Hat AI를 통해 양자화와 희소성, 그리고 vLLM 같은 고급 기술에 대해 배울 수 있습니다.

저자 소개

Pete serves as the Principal Community Architect for AI within Red Hat’s Open Source AI Program Office (OSAIPO). In this role, he drives community development and engagement for Red Hat’s open source AI initiatives, including key projects like llm-d. Pete helps scale Red Hat’s contributions to AI by supporting the open source communities and developers working to advance these technologies.

UI_Icon-Red_Hat-Close-A-Black-RGB

채널별 검색

automation icon

오토메이션

기술, 팀, 인프라를 위한 IT 자동화 최신 동향

AI icon

인공지능

고객이 어디서나 AI 워크로드를 실행할 수 있도록 지원하는 플랫폼 업데이트

open hybrid cloud icon

오픈 하이브리드 클라우드

하이브리드 클라우드로 더욱 유연한 미래를 구축하는 방법을 알아보세요

security icon

보안

환경과 기술 전반에 걸쳐 리스크를 감소하는 방법에 대한 최신 정보

edge icon

엣지 컴퓨팅

엣지에서의 운영을 단순화하는 플랫폼 업데이트

Infrastructure icon

인프라

세계적으로 인정받은 기업용 Linux 플랫폼에 대한 최신 정보

application development icon

애플리케이션

복잡한 애플리케이션에 대한 솔루션 더 보기

Virtualization icon

가상화

온프레미스와 클라우드 환경에서 워크로드를 유연하게 운영하기 위한 엔터프라이즈 가상화의 미래