Microsoft launches Decision-1 model on Qwen base to rival OpenAI and TypeSafe
Microsoft released Microsoft-Decision-1 in Foundry, built on Alibaba’s Qwen3.5-9B, matching competitor pricing while omitting image support and open weights.
Cet article est disponible uniquement en anglais.
Microsoft launched Microsoft-Decision-1 in its Foundry platform on Friday, October 10, 2026, entering the rapidly growing market for AI decision models. The release comes just three days after OpenAI made its Decisions API available in public beta and weeks after startup TypeSafe introduced its Jev model. By skipping OpenAI’s architecture and instead post-training a model based on Alibaba’s Qwen3.5-9B, Microsoft signals a shift toward independent infrastructure for agent logic, even as it plans to eventually rebase on its own MAI models and OpenAI’s technology.
What happened
The new model arrives at a price point of $0.042 per million input tokens with free output, exactly matching the rate charged by TypeSafe for its Jev model. This pricing move coincides with TypeSafe announcing an $870 million Series A funding round led by a16z, valuing the startup at $7.5 billion. TypeSafe claims that nearly 30% of the Fortune 500 has tried Jev in the three weeks since its launch, a traction metric that likely pressured Microsoft to respond quickly. In the same period, competitors including Upstage, Perplexity, Cloudflare, and AWS have also shipped their own decision-focused models.
Microsoft Chairman and CEO Satya Nadella announced the launch on X, stating that the company is already testing the model internally. Four internal teams are currently using Decision-1: Xbox Research uses it to sort over 10,000 pieces of player feedback, the Copilot team employs it to grade chat and agent responses, on-call engineers use it to pull context during live incidents, and Microsoft Discovery uses it to score experiments before an agent replans. Internal benchmarks suggest the model is more than 14 times faster than GPT-6 Sol for Xbox tasks and 46 times more consistent for Discovery workflows.
How it works
Decision-1 is a specialized model designed for structured decision tasks rather than general generative content. It accepts up to 32,768 tokens of text input and returns outputs in JSON format, making it suitable for integration into automated agents and workflows. Unlike some competitors, this initial version is text-only and does not support image inputs. Microsoft post-trained the model on Alibaba’s Qwen3.5-9B base, a choice that aligns with industry trends where developers are porting open-source models to create cost-effective decision layers. Cognition’s vice president of engineering, Jared Palmer, noted that porting similar models to Qwen3.5 required minimal compute resources, highlighting the efficiency of this approach.
The model is designed to provide calibrated probabilities, meaning a 90% confidence score should correspond to correct decisions nine out of ten times on representative data. However, Microsoft advises customers to validate these calibration scores on their own specific datasets. While the model aims for high consistency, recent research on competing models like Jev shows that adversarial inputs can significantly skew confidence scores. Microsoft tested Decision-1 against eight types of perturbations, such as reordered options, finding that only 1.3% of answers changed on average, though independent adversarial testing has not yet been published.
Key details
- Base Model: Post-trained on Alibaba’s Qwen3.5-9B, with plans to rebase on Microsoft MAI and OpenAI models later.
- Pricing: $0.042 per million input tokens with free output, matching TypeSafe’s Jev pricing.
- Performance: Reported as 14 times faster than GPT-6 Sol for Xbox feedback sorting and 46 times more consistent for Microsoft Discovery.
- Input Limits: Accepts up to 32,768 tokens of text; does not currently support image inputs.
- Availability: Launched in Microsoft Foundry; no open weights have been released yet.
- Internal Use: Currently tested by Xbox Research, Copilot, on-call engineering, and Microsoft Discovery teams.
Why it matters
For software engineers building AI agents, the emergence of dedicated decision models represents a shift away from using large, expensive language models for simple logical routing. Decision-1 allows developers to offload structured choices—such as determining whether a task should run on-device or in the cloud—to a specialized, low-latency model. This separation of concerns can reduce costs and improve response times for complex agent workflows. With major cloud providers and startups all releasing similar tools, the standard for agent architecture is crystallizing around a tiered system where small, fast models handle logic and larger models handle generation.
However, the lack of image support and open weights places Decision-1 behind some competitors like Cloudflare’s Clef, which supports vision encoders and is released under Apache 2.0. Additionally, while Microsoft’s Foundry sample code references a /systemone endpoint, the company has not confirmed full compatibility with the System One API adopted by AWS, Upstage, and Ollama. This uncertainty may affect interoperability for developers who rely on standardized interfaces to switch between providers. The reliance on calibrated probabilities also requires careful validation, as adversarial prompts can still manipulate decision outcomes, a risk highlighted by recent academic research on similar models.
What you can do
- Evaluate Decision-1 in Microsoft Foundry for text-based routing tasks within your existing agent workflows.
- Validate the model’s probability calibration using your own historical data to ensure reliability in production.
- Compare latency and cost against other decision models like Jev, Kev, or Clef for your specific use case.
- Test the model’s robustness by introducing paraphrased inputs or reordered options to check for consistency.
- Monitor Microsoft’s announcements for future updates regarding image support and System One API compatibility.
- Consider using the model for internal triage tasks, such as sorting feedback or grading agent responses, before deploying it to customer-facing features.



