Skip to content

Cloudflare Ships Open-Source Clef Models for Automated Support and Ops Triage

Short answer

Cloudflare released Clef and Clef-flash, open-source decision models hosted on Workers AI that return typed probability outputs (like ticket urgency or routing team) in one call, plus a new RL fine-tuning platform. This lets teams automate classification-heavy decisions like support ticket routing faster and cheaper than using a general LLM.

What this means for operators

The example Cloudflare publishes is the exact workflow most support teams already run by hand or with brittle rules: a message comes in, someone decides if it's urgent, which team owns it, and how severe it is. Clef does all three in one API call with typed, probability-scored outputs your existing code can act on directly, no LLM prompt engineering or free-text parsing required. For a 10-200 person company this means ticket routing, escalation flags, and severity scoring can run as a deterministic step inside an existing support stack, with humans only pulled in when the model defers. The new RL fine-tuning option (via Cloudflare's hands-on team first, self-serve later) also means a company with a backlog of labeled support tickets or sales qualification outcomes can tune the model on its own data rather than relying on a generic classifier.

Cloudflare released two open-source "decision models," Clef and Clef-flash, hosted on its Workers AI platform, alongside a new reinforcement learning (RL) fine-tuning product.

Unlike large language models, which generate open-ended text, decision models are built to produce bounded, typed, probability-scored outputs for a fixed set of questions, the kind of output a workflow can act on directly without parsing free text. Cloudflare's example: pass in a support message and ask if it's urgent, which team should own it, and how severe the impact is, and get back structured, probability-weighted answers your code can use to route the ticket or escalate to a human.

Cloudflare says Clef currently leads the Jev Decision Index benchmark and is fully API-compatible with Typesafe AI's Jev models, so existing integrations can swap in Clef without rework. Both models are released on Hugging Face under an Apache 2.0 license for local use, and are also hosted on Workers AI for immediate use via API.

On speed, Cloudflare reports that using Clef to classify a website domain (fetching, rendering, and categorizing it) took 2.2 seconds and returned more category labels, versus 4.7 seconds for its general-purpose LLM gpt-oss-120b on the same task, which returned only two classifications. Clef also adds a vision encoder for classifying images, a capability Jev does not currently have, and a 64k context window versus Jev's 32k.

Alongside the models, Cloudflare is launching an RL fine-tuning service. Initially this runs as a hands-on engagement with Cloudflare's forward-deployed engineer (FDE) team; a self-serve version is planned later, letting customers capture their own request data via Cloudflare's AI Gateway, generate training rollouts on Workers AI, and redeploy a fine-tuned Clef model on Cloudflare's infrastructure. Cloudflare states it does not read, store, or train on customer requests unless they opt into this fine-tuning product.

Source: Cloudflare Blog · In the Atlas: Cloudflare OS →

Next step

Visibility Analyzer

This is what the Visibility Analyzer measures on a real site: which answers cite you, which pages an engine cannot retrieve, and what to fix first. Free to run.

Run a free visibility audit

Free to run. No card.

Fee
Free
Length
One run, minutes

Free tier: two analyses a day, no card required.