Magnitude: An open source inference engine that tunes itself to your hardware
Magnitude compiles and tunes kernels on-device to run open models up to 2x faster than llama.cpp across Apple Silicon, NVIDIA, AMD, and CPU setups.
GitHub has introduced Magnitude, an open source inference engine designed specifically for AI agents. Released in late September 2026, this tool optimizes itself for the exact hardware it runs on, compiling and tuning its kernels directly on the user's device. The project claims to deliver performance gains of up to twice the speed of llama.cpp while maintaining strict local privacy.
What happened
Magnitude arrives as a desktop application available for macOS, Windows, and Linux. It functions as both a standalone runner for open-weight models and a backend for popular AI coding agents. The installation process is streamlined: users download the app, select a recommended model from the built-in Discover section, and connect their preferred agent through the Connections tab. The desktop package also includes the magnitude command-line interface, eliminating the need for separate tool installations.
The core distinction of Magnitude lies in its approach to kernel optimization. Unlike generalist engines that ship with precompiled kernels designed for broad hardware classes, Magnitude performs just-in-time compilation and tuning on the specific chip inside the user’s machine. This process happens before the model begins running, ensuring that the mathematical operations are tailored to the exact capabilities and constraints of the local hardware. This method allows it to support a wide range of configurations, from high-end NVIDIA or AMD GPUs to Apple Silicon and even CPU-only setups without fixed minimum requirements.
How it works
The performance advantage comes from hand-optimized kernels written for popular open-weight model families. By focusing on specific architectures rather than trying to be a universal solver for every possible model structure, the engine reduces overhead. When a user starts a session, Magnitude compiles these kernels on the device. This means the software adapts to the unique characteristics of the user’s processor, whether it is an M-series Mac, a CUDA-enabled NVIDIA card, or an AMD GPU.
Memory management is another key mechanical improvement. The engine uses 27% less memory per agent compared to previous standards. It achieves this by freeing resources when agents stop working and by sharing prefix caches across concurrent sessions. This caching mechanism prevents slowdowns when multiple tasks are running simultaneously, as the system does not need to reprocess identical initial prompts or context windows for each new request.
Key details
- Performance: Claims up to 2x faster inference than llama.cpp, with specific benchmarks showing 92% faster decode on Metal and 19% on CUDA.
- Hardware support: Runs on Apple Silicon, NVIDIA GPUs, AMD GPUs, or CPU-only systems with no fixed minimum specs.
- Agent integration: One-click connections for Pi, OpenCode, Hermes, Codex, Claude Code, Oh My Pi, Cline, and OpenClaw.
- Compatibility: Any other agent can connect via the provided OpenAI-compatible API endpoint.
- Privacy: All prompts, files, and models remain on the local machine with no internet requirement after initial download.
- License: The project is open source under the Apache 2.0 license.
Why it matters
For developers building local AI tools, the bottleneck has often been the trade-off between ease of use and performance. General-purpose engines like Ollama or LM Studio offer convenience but may leave performance on the table because their kernels are compiled for average hardware profiles. Magnitude addresses this by shifting the compilation step to the end-user’s device. This allows teams to deploy heavier open-weight models on existing workstations without needing cloud infrastructure, reducing latency and cost.
The memory efficiency improvements also have practical implications for multitasking. Engineers often run multiple agent instances for different tasks, such as code review, documentation generation, and unit test writing. By reducing memory usage per agent by 27% and sharing prefix caches, Magnitude enables more concurrent workflows on the same machine. This makes local development environments more responsive and capable of handling complex, multi-step agent chains without crashing or slowing down due to resource exhaustion.
What you can do
- Download the Magnitude desktop app for macOS, Windows, or Linux from the official repository.
- Install the app and use the Discover tab to download a recommended open-weight model.
- Connect your existing AI agent, such as Codex or Hermes, using the one-click integration in the Connections tab.
- Use the included magnitude CLI if you prefer terminal-based workflows over the graphical interface.
- Test the performance difference by running benchmarks against your current setup, particularly if you use Apple Silicon or NVIDIA hardware.
- Star the GitHub repository to support the community and track updates to the open source project.


