Running Gemma LLMs on Raspberry Pi 5 with LiteRT
Google's 2026 guide details using LiteRT to run Gemma models on Raspberry Pi 5, achieving real-time inference for robotics and edge agents without cloud dependency.
In a post on the Google Developers Blog in August 2026, engineers described how to deploy large language models directly on compact hardware. The guide focused on running Google’s Gemma family of models on the Raspberry Pi 5 using the LiteRT inference runtime. This approach enables fully offline, low-latency AI applications for robotics and local agents.
What happened
The article demonstrated that developers can build autonomous systems, such as the Reachy Mini robot, that perceive and react to their environment in real time without internet connectivity. By combining LiteRT with Gemma models, the system achieves total data privacy and ultra-low latency. The setup relies entirely on the Raspberry Pi 5, eliminating the need for cloud servers or external accelerators for basic tasks.
Google highlighted specific performance metrics for this configuration. On the Raspberry Pi 5, the LiteRT-LM orchestration layer enabled the Gemma 4 E2B model to process text at speeds sufficient for real-time interaction. The system maintained a low memory footprint while delivering responsive generative capabilities. This demonstration served as a practical reference for engineers looking to implement edge AI in resource-constrained environments.
The guide also outlined the broader ecosystem supporting this deployment. It pointed to available tools for model conversion, quantization, and benchmarking. By providing ready-to-use models via the LiteRT Hugging Face Community, Google aimed to reduce the friction typically associated with deploying large models on edge devices. The focus remained on making high-performance on-device inference accessible through streamlined workflows.
How it works
LiteRT serves as the core inference engine, optimized for both CPU and GPU execution on ARM architectures. For the Raspberry Pi 5, the system leverages the quad-core ARM Cortex-A76 CPU and the Broadcom VideoCore VII GPU. LiteRT uses XNNPACK for CPU acceleration, ensuring efficient memory usage and low-latency execution. This allows the device to handle complex mathematical operations required by large language models without overheating or stalling.

The architecture employs a heterogeneous parallel execution strategy. Instead of forcing all tasks onto the CPU, the system delegates specific workloads to the GPU. For instance, continuous computer vision tasks like object detection run on the GPU, freeing up CPU cycles for language processing and system orchestration. This division of labor maximizes the thermal and computational efficiency of the single-board computer.
LiteRT-LM acts as a specialized layer on top of the base runtime, simplifying the deployment of language models. It handles tokenization and generation loops efficiently. The Gemma 4 E2B model, designed for mobile and edge environments, uses memory-mapped per-layer embeddings to minimize RAM usage. This design choice is critical for maintaining performance on devices with limited memory resources.
Key details
- Hardware: The demonstration used a Raspberry Pi 5 with an ARM Cortex-A76 CPU and VideoCore VII GPU.
- Model: Gemma 4 E2B was selected for its balance of reasoning capability and compact size, suitable for edge deployment.
- Performance: The system achieved 99 tokens/sec for prefill and 9 tokens/sec for decode, with a peak memory footprint of 1432 MB.
- Speed: End-to-end generation reached approximately 27.3 characters per second, or about 300 words per minute.
- Tools: LiteRT CLI provides a unified command set for converting, quantizing, and running models from the Hugging Face Community.
- Pipeline: The Reachy Mini demo split tasks, running Ultralytics YOLO on the GPU and Moonshine ASR plus Gemma on the CPU.
Why it matters
For software engineers building IoT or robotics products, this development removes a major barrier: the need for constant cloud connectivity. Running models locally ensures data privacy and reduces latency, which is crucial for real-time interactions. It also lowers operational costs by eliminating server fees for inference. Developers can now prototype and deploy sophisticated AI agents on inexpensive, widely available hardware.

The ability to offload vision tasks to the GPU while keeping language processing on the CPU offers a blueprint for efficient edge architecture. This pattern allows devices to handle multi-modal inputs—such as video and audio—simultaneously without bottlenecking the main processor. It demonstrates that modern single-board computers are capable of handling workloads previously reserved for more powerful, expensive systems.
Furthermore, the availability of optimized models and tools like LiteRT CLI accelerates the development cycle. Engineers no longer need to manually manage complex dependencies or write custom optimization code from scratch. This standardization encourages innovation in edge AI, allowing teams to focus on application logic rather than infrastructure tuning.
What you can do
- Install the LiteRT CLI using pip in a virtual environment to access unified edge AI tools.
- Download the Gemma 4 E2B model from the LiteRT Hugging Face Community for testing on your device.
- Use the provided code snippets to run inference with image attachments and text prompts on Raspberry Pi 5.
- Explore the LiteRT-Samples GitHub repo to find reference implementations for voice translators and object detection.
- Experiment with offloading computer vision models like YOLO to the GPU to preserve CPU resources for LLM tasks.
- Monitor memory usage and thermal performance to ensure your specific application stays within the device's limits.



