Cloudflare releases Clef, a multimodal decision model for structured AI tasks
Cloudflare has launched Clef, a 27B parameter model optimized for document extraction and routing. It processes text and images in a single pass via Ollama.
Cloudflare has introduced Clef, a new multimodal decision model designed to handle structured tasks such as document field extraction and request routing. Released in early October 2026, this 27-billion parameter model is fine-tuned from Qwen3.8-27B and integrates directly into existing AI engineering workflows through standard APIs.
What happened
Clef represents a shift toward specialized decision-making models rather than general-purpose chatbots. The model is capable of processing text, JSON structures, and images simultaneously. This allows it to analyze screenshots, receipts, forms, and photos alongside textual data to make deterministic decisions. Cloudflare states that Clef currently leads the Decision Index, their internal leaderboard for decision models, and outperforms its predecessor, Jev, on most benchmark suites.
The release includes full compatibility with the Jev and System One APIs. This means developers can deploy Clef using Ollama’s /v1/systemone endpoint without rewriting their integration logic. A smaller, faster variant called clef-flash is also available for use cases where latency is more critical than maximum accuracy. Both versions are released under the Apache 2.0 license, allowing for broad commercial and open-source adoption.
How it works
Unlike traditional autoregressive language models that generate text token by token, Clef operates in a single non-autoregressive forward pass. It scores every possible option for every question jointly. This architecture significantly reduces inference time for classification and extraction tasks because the model does not need to generate verbose explanations before arriving at a conclusion. Instead, it outputs structured probabilities for predefined choices.
The model supports a 64K context window, which is double the capacity of Jev’s 32K window. This larger context allows Clef to process longer documents or more complex sets of instructions in one go. When images are included in the request, they are encoded and scored together with the text. This joint scoring ensures that visual cues, such as the layout of a form or a stamp on a receipt, influence the final decision alongside the textual content.
Developers interact with Clef by sending a state (the data to be judged) and a list of questions. The questions define the type of decision required, such as a multiple-choice selection, a binary true/false check, or a scored rating. The model returns the chosen option along with probability distributions and a confidence metric. This confidence score indicates how concentrated the probabilities are, helping engineers determine when to trust the output or escalate to a human reviewer.
Key details
- Model size: 27 billion parameters, fine-tuned from Qwen3.8-27B.
- Context window: 64K tokens, supporting long documents and complex instructions.
- Input types: Text, JSON, and base64-encoded images (PNG, JPEG, WebP).
- Deployment: Compatible with Ollama via the
/v1/systemoneendpoint; installable viaollama pull clef. - License: Apache 2.0, permitting free commercial use and modification.
- Output format: Structured JSON with choices, probabilities, and confidence scores.
Why it matters
For software engineers building AI agents, Clef offers a more reliable alternative to prompting large language models for simple classification tasks. General-purpose LLMs often hallucinate or provide inconsistent formatting when asked to extract specific fields from documents. Clef’s non-autoregressive design and structured output format reduce these errors by forcing the model to choose from predefined options. This makes it easier to integrate AI decisions into deterministic business logic, such as routing support tickets or approving expense reports.
The ability to process images and text in a single pass also simplifies multimodal pipelines. Previously, engineers might have needed separate OCR tools and language models, chaining them together with fragile error handling. Clef unifies this process, allowing a single API call to interpret a scanned invoice and extract key values like date and total amount. The open-source license further lowers the barrier to entry, enabling teams to self-host the model for data privacy compliance without licensing fees.
What you can do
- Install Clef locally using Ollama by running
ollama pull clefin your terminal. - Replace existing classification prompts with Clef’s structured question format to improve accuracy.
- Use the
imagesfield to submit screenshots or scanned documents for joint text-image analysis. - Implement confidence thresholds in your application to flag low-certainty decisions for human review.
- Test the clef-flash variant if your application requires lower latency for high-volume tasks.
- Review the System One API documentation to map your current Jev integrations to Clef.



