Optimizing long-context embedding inference on Cloud TPU with vLLM
Google details how native TPU support in vLLM enables elastic scaling and high-precision embedding inference for Qwen3 models.
In a post on the Google Developers Blog in August 2026, engineers described how they integrated native Tensor Processing Unit (TPU) support into vLLM. This update allows AI infrastructure teams to run large-scale embedding models with enterprise-grade precision while leveraging the elastic compute capabilities of Google Kubernetes Engine (GKE).
What happened
Embedding models are critical components in modern AI systems, translating unstructured data like text, images, and audio into dense vector representations. These vectors power semantic search, recommendation engines, and content clustering by capturing mathematical relationships between concepts. While prototyping with small models is manageable, scaling these pipelines to handle millions of queries introduces significant bottlenecks in capacity management and cost efficiency.
To address these challenges, Google Cloud added native TPU support to vLLM, the popular open-source serving engine for large language models. This integration enables true elasticity, allowing engineering teams to scale serving capacity dynamically. By provisioning TPU nodes alongside other accelerator instances, organizations can manage traffic fluctuations more effectively. If primary TPU reservations are fully utilized, the infrastructure can automatically fall back to secondary GPU spot or on-demand pools without interrupting inference traffic.
The implementation relies on GKE primitives like Custom Compute Classes to automate node autoscaling based on strict priority rules. This ensures that workloads scale up across different capacity types or accelerators if the preferred resource is unavailable. The goal is to maintain high availability and performance while optimizing costs during peak demand periods.
How it works
Serving next-generation embedding models requires processing ultra-long sequence contexts, ranging from over 4,000 tokens for text to more than 15,000 tokens for multimodal inputs. Maintaining strict mathematical parity across different hardware backends is essential for enterprise applications. The team focused on optimizing the Qwen3 Embedding model series on TPU hardware, addressing three specific engineering challenges.

First, TPU Matrix Execution Units impose strict divisibility constraints when sharding vocabulary matrices. The engineers implemented a hardware-safe vocabulary padding strategy to ensure exact tensor alignment during execution. Second, they hardened the materialization process by introducing attribute promotion in the unquantization pipeline. This change made weight loading compatible with vLLM’s TPU lazy-loader, eliminating initialization failures. They also implemented sharding-aware pre-warming to lock compilation caches before inference, reducing runtime latency.
Third, handling ultra-long contexts required a new architecture to prevent memory exhaustion. The team engineered a hybrid StepPool mechanism that migrates metadata to a cached request state. This ensures that pooling states accumulate correctly across steps and survive request preemptions. These optimizations allow the system to handle chunked prefill operations efficiently without losing state integrity.
Key details
- Native TPU support in vLLM enables dynamic scaling of embedding workloads alongside other accelerator types.
- The solution targets Qwen3-Embedding-8B and Qwen3-VL-Embedding-8B models for text and multimodal tasks.
- Hardware-safe vocabulary padding ensures tensor alignment across TPU topology meshes during parallel execution.
- Sharding-aware pre-warming locks JAX/XLA compilation caches to stabilize production pipelines and reduce latency.
- Cosine similarity scores reached ≥0.999 for text and ≥0.995 for multimodal inputs compared to reference baselines.
- Serving Qwen3-Embedding-8B on TPU Ironwood achieved 83,996 total tokens per second and 5.13 requests per second.
Why it matters
For software engineers building AI products, the ability to scale embedding inference elastically is crucial for maintaining service level agreements during traffic spikes. Traditional static provisioning often leads to over-provisioning or service degradation. By integrating TPUs with vLLM and GKE, teams can automate resource allocation based on real-time demand. This reduces operational overhead and improves cost efficiency without sacrificing performance.

Precision is another critical factor. In enterprise settings, even minor numerical discrepancies between hardware backends can degrade search quality or recommendation accuracy. The rigorous parity evaluations described in the post confirm that TPU-based inference maintains near-perfect alignment with reference implementations. This gives developers confidence that migrating workloads to TPUs will not compromise the quality of their AI features.
Furthermore, the support for long-context multimodal inputs opens new possibilities for complex retrieval tasks. Applications that need to understand relationships between large documents and associated images can now process these inputs efficiently. The hybrid StepPool architecture ensures that these heavy workloads remain stable and responsive, even under high load.
What you can do
- Review the official AI-Hypercomputer Qwen3-Embedding-8B Recipes on GitHub for setup scripts and deployment steps.
- Test the provided Python example to initialize Qwen3-Embedding-8B on Cloud TPU using vLLM’s native pooling runner.
- Evaluate cosine similarity between your current embedding outputs and TPU-generated vectors to verify numerical parity.
- Configure Custom Compute Classes in GKE to define priority rules for autoscaling across TPU and GPU resources.
- Implement sharding-aware pre-warming in your deployment pipeline to minimize cold-start latencies for TPU workloads.
- Explore the AI-Hypercomputer Public Repository for multimodal embedding recipes if your application processes images and text.


