machine learning

Clef Gives AI Agents a Fast Decision Layer

Clef Gives AI Agents a Fast Decision Layer

At 9:02 a.m., a support ticket lands: “Checkout has failed for every customer for the last hour.” A general Large Language Model (LLM)—a model built to generate open-ended text—can summarize the incident and propose a reply. Production code needs a different handful of facts: Is this urgent? Which team owns it? Should the system page someone, route the ticket, or pause for human review? Once prose has to drive an action, parsing and uncertainty become part of the engineering problem.

Cloudflare’s Clef models are aimed at that middle layer. A decision model reads a state—the text, records, screenshots, or other context describing a situation—and answers a bounded set of typed questions. Its output is structured data, meaning machine-readable fields, with estimated probabilities: numbers between 0 and 1 that represent the model’s likelihood for each answer. An agentic workflow, meaning software that gathers context, chooses a next step, and calls tools, can use those answers directly. (developers.cloudflare.com)

The practical search question—what is a decision model, and why not ask a general LLM to return JSON?—has a concrete answer: the contract is part of the model’s job, rather than a formatting request added after free-form generation. That makes the boundary between model output and application logic much cleaner.

A decision model speaks in types

The interface looks like a small decision table. You provide a state and a schema, meaning a machine-readable description of the questions and allowed answers:

const decision = await env.AI.run('@cf/cloudflare/clef', {
 model: 'clef',
 state: {
 subject: 'Checkout failures',
 message: 'Every customer has received a payment error for the last hour.'
 },
 questions: {
 urgent: {
 type: 'noul',
 instructions: 'Is this incident urgent?'
 },
 team: {
 type: 'choice',
 instructions: 'Which team should handle it?',
 criteria: {
 billing: 'Payments, invoices, and refunds',
 technical: 'Outages, errors, and configuration',
 sales: 'Plans and upgrades'
 }
 },
 severity: {
 type: 'score',
 instructions: 'How severe is the customer impact?',
 criteria: ['No impact', 'Minor', 'Major', 'Critical']
 }
 }
});

noul is Clef’s name for a yes-or-no question. choice selects one option from a set, while score evaluates an ordered scale. The returned answer can include the chosen value and a probability for each permitted option, so application code can set rules such as “route automatically above 90%, otherwise send to review.” There is no sentence parser waiting for a model to spell technical correctly. (developers.cloudflare.com)

That probability deserves a careful reading. It is an estimate of likelihood, not a guarantee, and it becomes useful only when tested against real outcomes. A model that says “urgent: 0.92” is giving your policy engine a control signal; your code still decides what 0.92 means.

Why Clef can be fast without becoming shallow

Most chat-oriented LLMs are autoregressive: they generate one token, or small piece of text, after another. Clef takes a different path. Cloudflare describes a prefill-only pass over the input—an initial read that loads the situation into the model’s internal state—followed by parallel scoring of the valid choices in the supplied schema. Because the decision step is non-autoregressive, the model does not need to compose an intermediate answer token by token before returning its probabilities. (blog.cloudflare.com)

The two released versions make the tradeoff visible. Clef is a 27-billion-parameter model, where parameters are the learned numerical settings inside the network. Clef-flash is a 9-billion-parameter variant aimed at latency-sensitive paths. Both are described as multimodal, meaning they can work with more than one kind of input, and both have a 64K-token context window—the amount of model-readable text they can consider at once. Images can be included too, which makes a screenshot or rendered webpage part of the decision rather than a separate preprocessing step.

Cloudflare’s published 43-benchmark comparison reports median model latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, compared with 524.1 milliseconds for Jev, another decision-model system. Those figures are useful for understanding the design goal, not a promise for every deployment; network distance, hardware, input size, and application code still shape the result. In a threat-intelligence example, Cloudflare also reported that fetching, rendering, and classifying a domain took 2.2 seconds with Clef versus 4.7 seconds with the general-purpose gpt-oss-120b model in the same workflow.

Open weights, managed inference

Clef is available through Workers AI, Cloudflare’s managed service for running models on its network of serverless GPUs. That route removes the need to operate GPU machines for an early prototype or a production endpoint. The model weights are also published under the Apache 2.0 license, a permissive open-source license, so teams can download them for local experiments, evaluation, or a deployment they control.

This gives the architecture two useful shapes. A fast decision model can sit in the hot path—the time-sensitive path before an agent calls a tool—while a larger language model handles explanation, planning, or the final customer-facing response. Or Clef can make the whole choice locally and hand only a compact, typed result to the rest of the system. Either way, the generative model no longer has to pretend that every small routing decision deserves a page of prose.

Where reinforcement learning fits

Fine-tuning means adapting a general model to a narrower workload. Reinforcement learning (RL) adds a reward signal: instead of asking only whether an answer matches a label, you score how useful the complete decision was. That matters when the right result depends on several fields, downstream actions, or a graded scale.

Cloudflare’s announced RL service starts as a hands-on engagement with a forward-deployed engineer—an engineer who works alongside a customer on the real workload—with a self-serve platform planned later. The proposed loop connects several pieces: AI Gateway, Cloudflare’s traffic and logging layer, captures request and response data; Workers AI generates rollouts, meaning candidate decisions produced from the base model; Cloudflare Containers provide a sandbox for scoring and replaying actions; a Trainer updates the fine-tuned weights; and the result can be redeployed on Workers AI.

The training objective can also respect near misses. Cloudflare describes Reinforcement Learning for Calibrated Decisions, or RLCD, as giving partial credit to adjacent ordinal choices, rewarding precise structured records, and penalizing drift from a reference model. Calibration means that a predicted 70% should behave roughly like a 70% outcome rate across similar cases. Confusing “major” with “critical” is still an error, but it is a different error from assigning a payment outage to the sales team.

Domain fine-tuning will not automatically improve every benchmark. A model tuned for bot classification may become better at that task while losing some broad, general-purpose behavior. The practical work is to keep a clean evaluation set, measure incorrect safe-case flags (false positives) and missed risky cases (false negatives), and check whether the model’s probabilities stay trustworthy as traffic changes.

Confidence is a control signal, not a verdict

A useful Clef integration treats confidence as part of a policy. High-confidence, reversible actions can run automatically. Ambiguous cases can go to a human or a slower reasoning model. High-impact decisions need stronger safeguards, audit logs, and a clear escape hatch even when the probability looks impressive.

That is the larger idea behind decision models. Large language models remain valuable for open-ended reasoning and communication. Clef adds a compact decision layer for the moments when an agent needs to choose from known options, expose its uncertainty, and move on. The result feels less like asking a chatbot to run a business process and more like giving the process a well-defined instrument panel.

ahsan

ahsan

Hello! I am Mr Ahsan, the writer of the Website. I am from Netherland. I like to write about technology and the news around it.

Comments (0)

No comments yet. Be the first to respond!

Leave a Comment

Your comment will be visible after review.