OpenAI launches Decisions API to cut agent monitoring costs
OpenAI introduced a new Decisions API at Dev Day, offering fast, low-cost classification similar to TypeSafe AI's Jev model for securing AI agents.
At its Dev Day event on Tuesday, OpenAI CEO Sam Altman revealed the Decisions API, a new tool designed to streamline decision-making for software agents. This announcement follows the recent release of Jev by TypeSafe AI, signaling a shift toward specialized, high-speed models for automation tasks.
What happened
The Decisions API allows developers to provide the Luna model with a predefined set of options, such as image categories or specific agent behaviors. The model then selects from these choices with high speed and lower computational cost than traditional large language model inference. Altman stated that this approach maintains capabilities like image understanding and safety protections while focusing the model on specific decisions.
TypeSafe AI, the startup behind Jev, did not respond to requests for comment on the new product. However, Diogo Almeida, TypeSafe’s CEO and a former OpenAI engineer, commented on social media about the emergence of similar technologies. He suggested that OpenAI’s move validates the concept of "System One" compatible architectures, which prioritize fast, intuitive processing over deliberate, slow reasoning.
The market is seeing increased activity in this niche, with other startups also releasing models that function as super-powered classifiers. While the exact technical similarities between the Decisions API and Jev remain unclear due to the limited preview status of OpenAI’s tool, industry interest is evident. The core value proposition for both tools is addressing the latency and expense associated with using general-purpose LLMs for routine software automation tasks.
How it works
These decision models operate as specialized classifiers built on top of large language model foundations. Instead of generating open-ended text, they evaluate a given input against a constrained list of possible outputs. The system assigns probabilities to each option, allowing the application to select the most likely correct action or category.
This mechanism differs from standard generative AI by restricting the output space. By limiting the model to choosing from a fixed set of alternatives, the computational load decreases significantly. This results in faster response times and reduced costs per query, making it feasible to run these checks on every single action an autonomous agent takes.
TypeSafe emphasizes that the intelligence of these models comes from the quality of synthetic data used during training. Almeida noted that while achieving speed and low cost is straightforward, maintaining high accuracy and statistical usefulness is the primary challenge. The goal is to optimize the ratio of intelligence to cost, ensuring that the rapid decisions remain reliable and calibrated to real-world scenarios.
Key details
- OpenAI launched the Decisions API as a limited preview during its Dev Day event on Tuesday.
- The API enables the Luna model to choose from predefined options, such as image categories or agent behaviors.
- TypeSafe AI released Jev earlier this month, a model designed for similar high-speed, low-cost software automation.
- Shapor Naghibzadeh demonstrated that monitoring agent actions with Jev costs $2.94, compared to $372 with a frontier LLM.
- Diogo Almeida described the trend as a move toward "System One" thinking, favoring fast, intuitive processing.
- The technology is being explored as a security measure to monitor and block misbehaving AI agents in real time.
Why it matters
For engineers building autonomous agents, cost and latency are critical bottlenecks. Traditional LLMs are often too slow and expensive to serve as real-time monitors for every step an agent takes. The introduction of dedicated decision APIs offers a viable path to implementing continuous oversight without prohibitive expenses. This shift could make it economically feasible to secure agents against unintended behaviors, such as the incidents previously observed on platforms like Hugging Face.
The competition between established labs and specialized startups highlights a broader trend in AI infrastructure. As agents become more complex, the need for lightweight, specialized components grows. Developers can no longer rely solely on monolithic models for all tasks. Instead, they will likely adopt hybrid architectures where heavy reasoning is reserved for complex problems, while fast classifiers handle routine decisions and safety checks.
This evolution also raises questions about calibration and reliability. As more providers enter the space with similar offerings, the differentiator will be how well these models align with real-world outcomes. Engineers must evaluate not just the speed and price, but the statistical robustness of the decisions made by these systems. The ability to trust these fast classifiers will determine their adoption in critical security and automation workflows.
What you can do
- Evaluate whether your current agent architecture requires real-time monitoring for every action.
- Test the OpenAI Decisions API in the limited preview to compare its performance against your existing classifiers.
- Experiment with TypeSafe’s Jev or similar startup offerings to benchmark cost and latency improvements.
- Define clear, constrained sets of options for your agents to reduce ambiguity and improve decision speed.
- Monitor the calibration of these fast models by comparing their outputs against ground truth data in your domain.
- Consider implementing a hybrid approach where fast decision models handle routine checks and larger LLMs handle complex exceptions.
