AI agents

Building a local voice agent in Rust with Voxlocal

Voxlocal is a minimal, open-source voice agent written in Rust that runs locally on macOS. It demonstrates how to chain speech recognition, vector search, and small language models for low-latency int

Rust gears interacting with digital audio waves representing local voice processing
Illustration generated for this article

Sam Khawase has released Voxlocal, an open-source command-line tool written in Rust that functions as a fully local voice agent for macOS. Published in October 2026, this project strips away cloud dependencies to demonstrate the inner workings of voice-driven AI systems using off-the-shelf micro models.

What happened

Voice agents typically rely on complex cloud infrastructure to handle audio streaming, speech-to-text conversion, and large language model inference. Khawase created Voxlocal to demystify this stack by building a bare-bones implementation that runs entirely on a local machine. The project focuses on a narrow automotive-service use case, allowing users to ask questions like service pricing via English voice commands.

The tool deliberately avoids advanced features like asynchronous streaming or Voice Activity Detection to keep the pipeline transparent and inspectable. Instead of reinventing core components, Voxlocal integrates existing open-source models such as Whisper for transcription and Piper for speech synthesis. This approach allows developers to study each stage of the conversation loop without the opacity of proprietary APIs or heavy abstraction layers.

By running locally, the agent eliminates network latency associated with remote servers, though it requires sufficient local compute resources. The source code is available publicly, serving as an educational resource for engineers interested in the mechanics of real-time audio processing and local AI inference.

How it works

The agent follows a linear pipeline starting with audio capture. The microphone records air pressure changes, converting them into digital audio samples at 16,000 samples per second in mono format. These samples are passed to Whisper, a neural network trained for speech recognition. Voxlocal uses greedy decoding with a temperature of zero, meaning the model selects the most likely next token rather than sampling from a probability distribution. This ensures deterministic output suitable for the limited context of the demo.

Once transcribed, the text undergoes normalization to remove filler words like "um" and standardize formatting. The cleaned text is then converted into a numerical vector using MiniLM, an embedding model. This process involves masked mean pooling, where token representations are averaged to create a single sentence vector, which is then L2-normalized to have a length of one. This mathematical representation allows the system to compare semantic meaning rather than just keyword matching.

The system performs retrieval-augmented generation by comparing the query vector against pre-computed vectors of service documents. It calculates cosine similarity via dot products to find relevant context, discarding matches below a threshold of 0.35. The top matches are fed into SmolLM2, a small language model running via Candle. The model is prompted to output a structured JSON tool call, such as checking a price, rather than generating free-form text. Finally, Piper synthesizes the response text back into audio for playback.

Key details

  • Voxlocal is a CLI tool built with Rust, designed specifically for macOS environments.
  • The speech-to-text component uses Whisper with greedy decoding and zero temperature for deterministic results.
  • Text normalization relies on regular expressions to strip filler words and standardize spacing before embedding.
  • Semantic search uses MiniLM embeddings and cosine similarity, with a minimum similarity threshold of 0.35 for document retrieval.
  • The language model, SmolLM2, is constrained to output structured JSON tool calls, which are parsed and validated by Rust code before execution.
  • Speech synthesis is handled by Piper, converting the final text response into audio waves using pre-trained voice models.

Why it matters

For software engineers building conversational interfaces, Voxlocal offers a rare look under the hood of a functional voice agent. Most production systems abstract these steps behind managed services, making it difficult to understand where latency bottlenecks or accuracy issues originate. By implementing the pipeline in Rust, Khawase highlights the importance of type safety and performance in real-time audio applications, where milliseconds matter.

The project also demonstrates the viability of small, local models for specific domains. Using SmolLM2 and MiniLM shows that specialized tasks do not always require massive cloud-hosted models. This approach can reduce costs and improve privacy, as data never leaves the user's device. However, it also reveals the fragility of small models, requiring careful parsing and repair logic to handle malformed JSON outputs.

Understanding the transition from analog sound waves to digital vectors and back to audio helps developers make better architectural decisions. It clarifies why jitter buffering and audio upsampling are critical in telephony providers and why normalization is essential before embedding. This knowledge enables teams to debug issues more effectively when integrating third-party voice services.

What you can do

  • Clone the Voxlocal repository from GitHub to inspect the Rust source code and understand the pipeline structure.
  • Experiment with the greedy decoding parameters in Whisper to see how temperature settings affect transcription accuracy.
  • Modify the regular expressions in the normalization stage to handle different filler words or speech patterns.
  • Adjust the cosine similarity threshold in the retrieval step to balance between precision and recall for your specific dataset.
  • Test the tool call parsing logic by feeding malformed JSON outputs to see how the repair function handles errors.
  • Replace the mock executor with a real API call to integrate actual booking or pricing services into the workflow.

Tools from the Bytechap store

Keep reading

All stories