Bespoke Labs releases Nimble, a fast 9B decision model for local inference
Bespoke Labs has released Nimble, an open-weight 9B decision model fine-tuned from Qwen3.5-9B that provides probabilistic answers to structured questions in under 100ms on local hardware.
Bespoke Labs has released Nimble, a new open-weight decision model designed for fast, local inference. Fine-tuned from the Qwen3.5-9B architecture, this 9-billion parameter model allows developers to send text and a set of structured questions, receiving probabilistic answers without a traditional reasoning step. The release includes clear API documentation and performance metrics, targeting engineers who need efficient edge processing capabilities.
What happened
Nimble represents a shift toward specialized decision models that prioritize speed and structured output over generative creativity. Unlike standard large language models that generate text token by token through a reasoning process, Nimble is optimized to pick answers from a defined set of options. It reads the input prompt once per question and scores the answer tokens directly. This architectural choice eliminates the latency associated with chain-of-thought reasoning, enabling the model to return results in under 100 milliseconds on a MacBook Pro equipped with an M5 Max chip.
The model is available under an Apache 2.0 license, making it suitable for commercial and open-source projects alike. Bespoke Labs trained Nimble on contrastive pairs, a dataset construction method where two examples differ by only one specific fact that flips the correct answer. This training strategy helps the model focus on precise factual distinctions rather than general language patterns. Developers can pull the model directly using Ollama with the command ollama pull nimble, integrating it into existing local workflows with minimal setup.
How it works
Nimble operates through Ollama’s /v1/systemone endpoint, which adheres to TypeSafe’s Jev API standard. The interaction model is distinct from typical chat interfaces. Users provide the text to be evaluated in a field called state, which can be a simple string or a structured JSON object. Alongside the state, users submit up to 64 named questions in a single request. The model processes these questions against the provided text and returns answers in the same order they were asked.
The system supports three primary question types: choice, binary truth verification, and rubric scoring. For choice questions, the user defines a list of options, and the model returns the selected choice along with probabilities for all options. Binary questions, referred to as noul in the API, return a true or false verdict with an associated probability. Score questions allow users to define ordered levels with specific criteria, returning a probability-weighted score. In all cases, the model provides a confidence metric between 0 and 1, indicating how concentrated the probabilities are, though this is not a direct measure of correctness.
Key details
- Model size and base: Nimble is a 9-billion parameter model fine-tuned from Qwen3.5-9B.
- Performance: It delivers answers in under 100ms on a MacBook Pro with an M5 Max chip.
- Batching capability: A single API call can handle up to 64 questions about the same text.
- Training data: The model was trained on contrastive pairs where a single fact change flips the answer.
- License: Released under the Apache 2.0 license, permitting broad commercial and private use.
- Output format: Returns structured data including choices, scores, probabilities, and confidence metrics.
Why it matters
For software engineers building production systems, the elimination of the reasoning step offers significant advantages in latency and cost. Traditional large language models often struggle with consistent structured output, requiring complex prompt engineering or post-processing to extract reliable decisions. Nimble’s design ensures that every response fits a predefined schema, reducing the need for validation layers and error handling in downstream applications. This reliability is crucial for tasks like routing customer support tickets, applying compliance policies, or rating content quality, where deterministic behavior is preferred over creative generation.
The ability to run this model locally at high speeds also addresses growing concerns around data privacy and operational costs. By processing sensitive text on-device without sending it to external APIs, organizations can maintain stricter control over their data. Furthermore, the low latency enables real-time decision-making in user-facing applications, such as instant form validation or dynamic content filtering, without introducing noticeable delays. The open-weight nature of the model allows teams to audit its behavior and fine-tune it further for specific domain requirements.
What you can do
- Install Ollama and pull the Nimble model using the command
ollama pull nimbleto test local inference speeds. - Experiment with the
/v1/systemoneendpoint by sending structured JSON states and defining up to 64 questions per request. - Implement a routing system that uses Nimble’s choice type to direct user requests to appropriate departments based on probability scores.
- Build a policy enforcement tool that uses the score type to rate user-generated content against a defined rubric of acceptable levels.
- Use the binary
noulquestion type to create automated fact-checking pipelines that verify specific conditions within documents. - Monitor the confidence metric returned by the model to flag low-certainty decisions for human review, improving overall system reliability.


