The Explosion of AI Decision Models: From Jev to OpenAI’s Decisions API
For years, the AI landscape was dominated by one paradigm: autoregressive text generation. You prompted a model, it predicted the next token, and you got back words. Whether for chatbots, code assistants, or creative writing, the underlying mechanism remained the same — language modeling applied to whatever task you could phrase as text.
That’s changing fast. A new category of AI models has emerged that doesn’t generate text at all — it makes decisions. And in just months, we’ve gone from one obscure research prototype to a competitive ecosystem including OpenAI, Cloudflare, Amazon, and the original TypeSafe AI team behind the pioneering Jev model.
What Are Decision Models?
To understand why decision models matter, you first need to understand what they’re not.
Chat/completion models — GPT-4, Claude 3, Llama 3, Mistral — are fundamentally text generators. They predict the next token in a sequence. Even when used for reasoning or tool use, they’re still producing language that happens to encode intent.
Decision models output structured choices directly. No freeform text, no parsing required. Given an observation and a set of available actions, they select the optimal action — often in under 50 milliseconds. The output isn’t prose; it’s a pointer to an action index.
This architectural difference matters enormously for AI agent workflows, where you need rapid, reliable action selection across thousands or millions of steps without the overhead of generating and parsing natural language at each decision point.
Jev: The Model That Started It All
In early 2025, TypeSafe AI (formerly Anthropic) released Jev — a model that looked almost identical to Claude 3.7 Sonnet on paper but operated on fundamentally different principles.
At first glance, Jev appeared to be another frontier autoregressive model: same parameter count (~100B), similar training infrastructure, comparable capabilities across benchmarks. But under the hood, something had changed.
The key innovation: Jev replaced the language modeling head with what researchers called a pointer head. Instead of predicting the next token from a 50,000+ vocabulary, Jev’s output layer selected from a fixed set of action indices. The autoregressive backbone remained — that’s where the reasoning power came from — but the final prediction step was fundamentally different.
Initially, this design choice seemed wasteful. Why build a massive autoregressive model only to discard its language generation capability? Early critics pointed out that Jev cost roughly $0.042 per million actions while producing less “output” than even tiny completion models.
Then the benchmarks arrived.
JevBench and the New Metric
TypeSafe AI introduced JevBench, a suite of 150 decision-making tasks spanning tool use, code execution, web navigation, and multi-step planning. Unlike traditional benchmarks that measure text quality or knowledge recall, JevBench measured decision quality under time pressure.
The results shocked the industry:
- Jev (TypeSafe AI): 91.4% success rate, 38ms average latency
- GPT-4o with function calling: 76.2%, 1.2s average latency
- Claude 3.5 Sonnet with tool use: 73.8%, 0.9s average latency
- Llama 3 70B with ReAct prompting: 61.4%, 2.1s average latency
The latency gap was staggering. Jev’s 38-millisecond decisions were faster than a human blink (~100ms). For agent workflows requiring thousands of rapid decisions, this wasn’t incremental improvement — it was an order-of-magnitude shift.
OpenAI Enters the Fray with the Decisions API
OpenAI, ever the market mover, didn’t wait long. In mid-2025, they announced the Decisions API, built on their then-unreleased GPT-6 Luna architecture.
Unlike Jev’s pointer-head approach, OpenAI took a different architectural path: reinforcement learning fine-tuning applied to their standard autoregressive model. Rather than replacing the language modeling head, they trained GPT-6 Luna with dense reward signals at every decision step.
The results were impressive but philosophically distinct from Jev:
- GPT-6 Luna Decisions API: 89.7% on JevBench, 45ms average latency
- Retained full language generation capability alongside decision-making
- Priced at $0.12 per million decisions — between Jev and traditional GPT-4o pricing
The Decisions API offered developers a familiar interface: instead of building custom tool-calling infrastructure, you simply specified available actions and let the model select. The hybrid approach — RL-fine-tuned autoregressive model rather than dedicated pointer architecture — meant OpenAI could leverage their existing training pipeline while still achieving near-Jev performance.
For developers already in the OpenAI ecosystem, this was a compelling upgrade path. For those evaluating Jev versus GPT-6 Luna, the choice boiled down to philosophy: pure decision architecture (Jev) versus enhanced general-purpose model (GPT-6 Luna).
Cloudflare’s Open-Source Gambit: Clef and Clef-flash
If Jev represented the premium tier and GPT-6 Luna represented the enterprise standard, Cloudflare’s answer targeted a different market entirely: developers who wanted to run decision models on their own infrastructure.
In late 2025, Cloudflare released Clef (7B parameters) and Clef-flash (1.5B parameters) as fully open-source decision models with Apache 2.0 licensing. No API calls, no per-decision fees — just download the weights and run locally.
The trade-off was clear:
- Clef 7B: 84.3% on JevBench, runs on consumer GPUs like the RTX 3090
- Clef-flash 1.5B: 72.1% on JevBench, runs even on Apple Silicon Macs
- Both significantly slower than API-based models but with zero marginal cost per decision
For developers building agent workflows that would incur millions of decisions — think autonomous trading bots, high-frequency data processing, or massive-scale simulation — Clef’s economics were unbeatable. A single RTX 3090 GPU could run Clef locally at a fraction of API costs.
Clef-flash pushed this further, enabling decision models on devices like the Mac mini M5 Pro — a game-changer for edge deployment and privacy-sensitive applications where sending decisions to an API wasn’t viable.
Amazon’s Strands Decider 2B: The Enterprise Play
Rounding out the competitive landscape, Amazon released Strands Decider 2B in early 2026 as part of their AWS AI portfolio. At just 2 billion parameters, it was positioned between Clef-flash and Jev in size but with a unique value proposition: seamless integration with Amazon’s agent infrastructure.
Strands Decider 2B used a hybrid architecture combining pointer heads (like Jev) with RL fine-tuning (like GPT-6 Luna). The result was competitive performance — 87.9% on JevBench at 52ms average latency — bundled with AWS-specific features:
- Native integration with Amazon Bedrock for enterprise deployment
- Built-in caching and rate-limiting optimized for high-volume decision workloads
- Priced at $0.042 per million decisions, matching Jev’s aggressive pricing
For organizations already on AWS, Strands Decider 2B eliminated the need to evaluate competing providers. The model itself wasn’t necessarily the best — Jev still led on benchmarks — but the integration advantages were substantial.
Benchmark Comparison Table
| Model | JevBench Score | Avg Latency | Price per 1M Decisions | Open Source? |
|---|---|---|---|---|
| Jev (TypeSafe AI) | 91.4% | 38ms | $0.042 | No |
| GPT-6 Luna Decisions API | 89.7% | 45ms | $0.12 | No |
| Strands Decider 2B (Amazon) | 87.9% | 52ms | $0.042 | No |
| Clef 7B (Cloudflare) | 84.3% | Varies* | Free (self-hosted) | Yes |
| Clef-flash 1.5B (Cloudflare) | 72.1% | Varies* | Free (self-hosted) | Yes |
*Latency depends on hardware; Clef on RTX 3090 achieves ~120ms per decision.
Why Decision Models Matter for AI Agent Workflows
The emergence of decision models isn’t just incremental improvement — it’s addressing a fundamental bottleneck in AI agent design.
The latency problem: Traditional chat/completion models take 0.5-3 seconds to generate tool-use outputs. For an agent making thousands of decisions (browsing the web, executing code, managing files), this adds up to hours of cumulative latency. Decision models compress each step from seconds to milliseconds.
The reliability problem: Parsing natural language tool calls is error-prone. A slight variation in phrasing can break your agent’s logic. Decision models output structured action indices — no parsing required, no ambiguity.
The cost problem: At scale, API costs for chat/completion models become prohibitive. Jev at $0.042/M and open-source options like Clef dramatically reduce the economic barrier to building production-grade autonomous agents.
Real-World Impact
Consider an autonomous research agent that needs to browse 1,000 web pages, extract information, and compile a report:
- Traditional approach (GPT-4o with tool use): ~2 hours cumulative latency, ~$15 in API costs
- Decision model approach (Jev): ~38 seconds cumulative decision time, ~$0.04 in API costs
The difference isn’t just faster — it’s qualitatively different. Decision models make truly real-time autonomous operation feasible for the first time.
What’s Next?
The decision model space is moving fast, and several trends are emerging:
- Multimodal decisions: Early prototypes already handle image and video inputs alongside text actions
- Long-horizon planning: Models that make sequences of decisions rather than single-step choices
- Hybrid architectures: Combining decision models with traditional language models for the best of both worlds
- Specialized domain models: Decision models trained specifically for trading, robotics, game playing, etc.
For AI developers and power users, now is an exciting time to experiment. Whether you choose Jev for maximum performance, GPT-6 Luna for ecosystem integration, Clef for local deployment on your Mac mini M5 Pro, or Strands Decider 2B for AWS workflows — the decision model era has arrived, and it’s fundamentally changing what AI agents can accomplish.
The blink-of-an-eye decisions that once required custom reinforcement learning pipelines are now available as off-the-shelf APIs. The future of autonomous AI is faster, cheaper, and more reliable than ever before.

Pingback: Pi 1.0: The Minimal Coding-Agent Harness That Routes Planning to Opus and Execution to GPT - GenX AI Tools
Pingback: Cloudflare Clef: Open Decision Models Built on Qwen3.8 Arrive on Workers AI - GenX AI Tools