Cloudflare Clef: 39ms AI Decision Models for Autonomous Agents

Alex da Cruz
Alex da Cruz is a full-stack developer based in São Paulo, Brazil. He works with React, TypeScript and automation, and uses AI daily to solve real problems in code and operations — not as a demo. He has run an e-commerce operation end to end, and now builds and maintains the automation pipeline behind this blog. He writes about what he actually tests.
According to Cloudflare, the newly launched Clef and Clef-flash decision models eliminate the need for human intervention in routine AI agent tasks by returning structured probabilities instead of raw text. Built on Qwen architectures, these models assign likelihoods to predefined options, allowing downstream systems to execute decisions instantly.
How fast are Cloudflare's Clef models in practice?
Cloudflare reports a median response time of 39 milliseconds for Clef-flash and 209 milliseconds for the larger Clef model. In comparative tests, this performance outpaces TypeSafe AI's Jev model, which logs a median latency of over 524 milliseconds.
What changes for engineering teams on Monday morning?
Engineering teams can route support tickets, filter abuse reports, or classify incoming data streams without waiting for slow general-purpose language models. Processing an entire website classification task takes 2.2 seconds with Clef, compared to 4.7 seconds for standard LLMs, cutting wait times in half for automated pipelines.
Both models run directly on Cloudflare Workers AI and are available under the Apache 2.0 license, allowing developers to integrate structured decision-making into existing edge applications immediately.
Sources
Frequently asked questions
- What is a decision model like Cloudflare Clef?
- A decision model assigns probabilities to a set of predefined answers rather than generating free-form text. This structure lets downstream code act on the results automatically.
- How does Clef-flash compare to competitors in speed?
- Clef-flash delivers median response times of 39 milliseconds, making it significantly faster than comparable decision models like TypeSafe AI's Jev.
Comments
0 comments
Be the first to comment.
Continue Lendo

Microsoft MAI Models: Voice Agents with 100ms Latency
Microsoft released new transcription and voice models with 100ms latency and voice cloning capabilities for AI agents.

Manus 2.0: mobile control, video editing and AI agents
Manus 2.0 adds mobile remote control, video editing, multiplayer game hosting, and cuts agent running costs by 32 percent.

Shopify WebMCP Checkout: AI Agents Complete Purchases
Shopify's new WebMCP checkout support allows browser-based AI agents to complete purchases natively using structured APIs.