Kubernetes introduces Gateway API extension for LLM inference routing
The Kubernetes project released the Gateway API Inference Extension to standardize routing for generative AI workloads, reducing latency and improving GPU utilization.
In a post on the Kubernetes Blog in June 2025, contributors from Solo.io, Google, and Bytedance introduced the Gateway API Inference Extension. This new standardized extension addresses the specific traffic-routing challenges of self-hosted large language models (LLMs) by adding inference-aware capabilities to the existing Gateway API framework.
What happened
Modern generative AI services differ significantly from traditional web applications. While standard HTTP requests are often short-lived and stateless, LLM inference sessions are long-running, resource-intensive, and partially stateful. A single GPU-backed model server might maintain active inference sessions and keep token caches in memory. Traditional load balancers, which typically rely on simple HTTP path matching or round-robin distribution, lack the specialized logic required to handle these complex workloads effectively. They do not account for model identity or the criticality of a request, such as distinguishing between an interactive chat session and a background batch job.
To solve this, the community developed the Gateway API Inference Extension. This project builds upon the familiar Gateway API model, allowing platform engineers to transform a standard gateway into an "Inference Gateway." The goal is to provide a standardized approach to routing inference workloads across the ecosystem, moving away from ad-hoc, custom solutions. By enabling model-aware routing and supporting per-request criticalities, the extension aims to reduce latency and optimize the utilization of accelerators like GPUs.
How it works
The architecture introduces two new Custom Resource Definitions (CRDs) that separate concerns between platform operators and AI/ML owners. The first, InferencePool, defines a group of pods running model servers on shared compute resources. Platform administrators use this resource to configure deployment, scaling, and balancing policies, ensuring consistent resource usage across the cluster. It functions similarly to a Kubernetes Service but is specifically aware of model-serving protocols.

The second resource, InferenceModel, is managed by AI/ML teams. It maps a public endpoint name, such as "gpt-4-chat," to a specific model within an InferencePool. This allows workload owners to define which models are served, including any fine-tuning variants, and to set traffic-splitting or prioritization policies. This separation ensures that platform teams manage the infrastructure while application teams manage the model exposure.
When a client sends a request, the Gateway examines the HTTPRoute to identify the target InferencePool. Instead of forwarding traffic to any available pod, the Gateway consults an Endpoint Selection Extension. This component analyzes live metrics from the pods, such as queue lengths, memory usage, and loaded adapters. It then selects the optimal pod based on real-time conditions, ensuring the request is handled with the lowest possible latency or highest efficiency. This process remains transparent to the client, appearing as a standard single request.
Key details
- Two new CRDs:
InferencePoolfor platform-level resource management andInferenceModelfor user-facing model endpoints. - Endpoint Selection Extension: Replaces simple round-robin with metric-aware routing that considers queue depth and memory state.
- Benchmark environment: Tests used vLLM version 1 on H100 (80 GB) GPU pods with 10 Llama2 replicas.
- Latency improvements: The extension showed significantly lower p90 latency at high loads (500+ QPS) compared to standard Kubernetes Services.
- Throughput parity: Throughput remained comparable to standard services across the tested range of 100 to 1000 Queries per Second.
- Extensible design: The framework supports additional extensions for new routing strategies or specialized hardware needs.
Why it matters
For engineering teams building AI products, this extension offers a path to more reliable and efficient self-hosted model serving. By routing requests based on real-time pod metrics rather than static rules, organizations can avoid hotspots that occur when GPU memory approaches saturation. This leads to more predictable tail latencies, which is critical for maintaining a smooth user experience in interactive applications. The ability to prioritize traffic based on criticality also ensures that high-value interactions are not blocked by lower-priority batch processes.

From an operational perspective, the standardization reduces the maintenance burden of custom routing logic. Platform teams can enforce consistent policies across different models and teams using native Kubernetes tools. As the project moves toward general availability, features like prefix-cache aware load balancing and support for heterogeneous accelerators will further enhance its utility. This alignment with Kubernetes-native tooling simplifies the integration of GenAI services into existing infrastructure.
What you can do
- Review the official project documentation to understand the API specifications and installation requirements.
- Deploy a test Inference Gateway in a non-production environment to evaluate the Endpoint Selection Extension.
- Map existing model services to
InferencePoolandInferenceModelresources to test the migration path. - Monitor p90 latency and GPU utilization metrics during load testing to quantify performance gains.
- Contribute to the project by developing new extensions for specific routing strategies or hardware types.
- Provide feedback on the roadmap items, such as LoRA adapter pipelines and disaggregated serving support.


