Open-source FPGA accelerator runs modern LLMs with full toolchain
The openTPU project releases a complete AI accelerator design in one repository, running models like Qwen3 and LFM2.5 on a Kintex-7 FPGA card with bit-exact simulation.
GitHub user FeSens has released openTPU, an open-source AI accelerator that includes its own hardware description, instruction set architecture, simulator, compiler, and profiler in a single repository. Published on October 6, 2026, the project demonstrates that modern large language models can run efficiently on field-programmable gate arrays using a fully transparent software and hardware stack.
What happened
The openTPU project answers two fundamental questions about the current state of AI hardware design: how far autonomous agents can go in designing hardware, and whether they can build the very chips that run their own inference. The entire accelerator is contained within one small monorepo, allowing developers to read the code end-to-end. This includes the SystemVerilog hardware design, the instruction set definition, a bit-exact simulator, a kernel language with its compiler, and the host software required to drive a physical PCIe card.
The design has been tested on an Inspur YPCB-00338 card, which features a Xilinx Kintex-7 xc7k480t FPGA and two DDR3 memory channels. It successfully runs ten modern models with their real weights, producing tokens that match the simulator output bit for bit. The project provides detailed benchmarks for models such as LFM2.5-230M, Qwen3-0.6B, Qwen3.5-0.8B, Gemma 4 E2B, LFM2-2.6B, SmolLM3-3B, Phi-4-mini, and Qwen3.5-2B and 4B. These tests cover both int8 and 4-bit quantization schemes, showing decode speeds ranging from 3.75 tokens per second for larger models up to 85.8 tokens per second for smaller ones.
A significant update, referred to as Build B, improved decode performance by 8-9% for several models compared to previous production images. This build also increased DRAM bandwidth utilization to between 91% and 94% of the peak DDR3-1066 speed. The system supports mixture-of-experts models larger than the card’s 4 GiB memory by streaming experts from host storage, maintaining bit-exact accuracy with the simulator even during these complex operations.
How it works
The architecture is deliberately simple to ensure transparency and ease of debugging. A sequencer issues one instruction per cycle to a few specialized units: a DMA engine for data movement, a matrix unit for int8 weight multiplication streamed from DRAM, a vector unit for fp32 math, and a quantizer to convert results back to int8. There are no caches or hidden scheduling mechanisms. Every data movement is explicitly defined as an instruction, meaning a trace of execution reveals exactly where cycles are spent.
Developers write kernels in a Python-like language called ol, which uses decorators like @ol.jit to compile high-level operations into the custom instruction set. The compiler handles layouts, affine loop addressing, and fusion. The resulting instructions are executed on the FPGA or verified against the Python-based ISA simulator. Because the simulator and the RTL (Register Transfer Level) hardware design process the same bits, they are checked by tests to ensure consistency. This allows developers to prototype and debug on a laptop before deploying to the physical hardware.
The system includes a profiler named Lens, which records runs from the RTL, simulator, or card and displays them in a browser. Lens provides a roofline model, a timeline, and per-instruction tables, coloring each cycle to show if a unit is busy, waiting on DRAM, or waiting on another instruction. This visibility helps engineers understand performance bottlenecks, such as DRAM bandwidth limits, which currently constrain decode speed to 82-85% of the theoretical peak.
Key details
- The project runs on a Xilinx Kintex-7 xc7k480t FPGA with two DDR3 channels, achieving up to 17.1 GB/s peak bandwidth.
- Benchmarks show decode speeds from 3.75 tok/s for Qwen3.5-4B to 85.8 tok/s for LFM2.5-230M with 4-bit quantization.
- The system supports 4-bit weights using FP4 values with two-level block scales, reducing bytes per token by about a third.
- Mixture-of-experts models like LFM2.5-8B-A1B run by streaming experts from host storage, achieving 10.6 tok/s with 98.5% slot hit rate.
- All configurations match the simulator token for token, ensuring bit-exact reproducibility between software simulation and hardware execution.
- The host software overhead is minimal, adding only 0.17 to 0.30 ms per token on optimized systems, keeping the accelerator largely independent.
Why it matters
For software engineers and ML practitioners, openTPU demystifies the black box of AI accelerators. By providing a complete stack from Python kernels down to SystemVerilog RTL, it offers a rare opportunity to understand how matrix multiplications and attention mechanisms translate into physical wire movements and clock cycles. This level of transparency is invaluable for educational purposes and for developers who need to optimize models for specific hardware constraints without relying on proprietary tools.
The project also highlights the practical limits of current FPGA-based inference. The detailed benchmarks show that decode performance is heavily bound by DRAM bandwidth rather than compute power. This insight guides engineers toward optimizing memory access patterns and quantization strategies rather than just increasing computational throughput. The ability to run modern models like Qwen3 and Gemma 4 on relatively older hardware like the Kintex-7 suggests that efficient software and architecture design can extend the life of existing infrastructure.
Furthermore, the integration of AI agents in the design process, as suggested by the "auto-arch-tournament" approach, points to a future where hardware design itself may be automated. The fact that the system can build and run the chip that executes its own inference loops creates a closed-loop development environment that could accelerate innovation in specialized AI hardware.
What you can do
- Install the openTPU package via pip and run the ISA simulator on your laptop to test models like LFM2.5-230M without hardware.
- Study the instruction set documentation in
docs/isa.mdto understand the low-level operations that drive the accelerator. - Use the Lens profiler to visualize execution traces and identify DRAM bottlenecks in your own kernel implementations.
- Experiment with 4-bit quantization schemes to improve decode speeds, noting the trade-offs in perplexity reported in the quantization docs.
- If you have a compatible Kintex-7 PCIe card, build the bitstream and load it via JTAG to run real-time inference with
otpu-chat. - Contribute to the project by exploring the
rtl/directory and helping to improve timing margins and area efficiency in future builds.



