Skip to main content
株式会社オブライト
AI2026-10-029 min read

Cloudflare Clef: 41GB VRAM, $0.24/M, Jev Alternative

Cloudflare Clef and Clef-flash: open-weight decision models with a Jev-compatible API. Apache 2.0, 41GB minimum VRAM, $0.24/M on Workers AI. With benchmarks.


Cloudflare Clef is an open-weight "decision model" that Cloudflare released on October 1, 2026. It comes in two sizes, Clef and Clef-flash. Instead of generating prose, it takes a state (text or images) plus typed questions, and returns calibrated probabilities for yes/no, multiple-choice, and ranking or score questions. It is positioned directly against TypeSafe AI's Jev System One: the API is described as Jev-compatible, the license is Apache 2.0, and the weights are on Hugging Face. The announcement earned 512 points on Hacker News, and Cloudflare also unveiled a reinforcement-learning fine-tuning platform. This article covers how Clef works, pricing, required VRAM, API usage, benchmarks, and how to choose between it and Jev.

What Cloudflare Clef is

A decision model does not write free-form text. It returns only the judgment: Is this request urgent? Which team should handle it? How severe is the impact on a four-level scale? Clef is post-trained from existing backbones: Clef is built on Qwen3.8-27B and Clef-flash on Qwen3.5-9B. The license is Apache 2.0, and the Hugging Face repositories are Cloudflare/clef and Cloudflare/clef-flash. Only the weights are public, not the training datasets, so "open weight" is the accurate term rather than "open source".

Its operation differs from an LLM as well. Clef does not generate an answer token by token. It scores all choices in parallel in a single prefill-only pass (non-autoregressive), so responses are fast and the output is a set of probabilities. On the training side, Cloudflare describes a two-stage attention routing scheme, rank-256 low-rank adapters, and a label-smoothed cross-entropy plus Brier loss to keep probabilities calibrated, followed by RLCD (Reinforcement Learning for Calibrated Decisions). That is the same method name used in Jev's announcement, and the design goal of calibrated decisions is shared.

What it can do: inputs and question types

The context window is 64k tokens, and inputs include text and images. Clef has a vision encoder, which Cloudflare presents as something Jev lacks. The Register also reports video support, so check the official documentation for modalities beyond images. The question types follow the same three-way design as Jev, which we covered in the Ollaya article.

noul: a yes/no question that returns the probability of true (for example, is this request urgent?)
- choice: pick from options, with a probability for each (for example, which team should handle it?)
- score: graded evaluation, returning a rank or score distribution (for example, how severe is the customer impact?)

Typical uses include support ticket triage, email and inquiry routing, invoice processing, security incident severity, evaluation of agent traces, and guardrails for LLM output. The basic pattern is to embed it as a "smart if statement" and set probability thresholds that split work between automatic handling and human review.

Specs, pricing, and required VRAM

ItemClefClef-flash
BackboneQwen3.8-27B (post-trained)Qwen3.5-9B (post-trained)
LicenseApache 2.0 (open weight)Apache 2.0 (open weight)
Context window64k tokens64k tokens
InputsText, images (video per The Register)Text, images
Latency median / p95209.3 ms / 238.6 ms38.8 ms / 122.4 ms
Minimum VRAM to self-host85 GB41 GB
Workers AI model ID@cf/cloudflare/clef@cf/cloudflare/clef-flash
Hugging FaceCloudflare/clefCloudflare/clef-flash

Pricing on Workers AI is $0.24 per million tokens. The Register notes that this is nearly 6x Jev's $0.042 per million. On latency, Cloudflare reports Jev at a 524.1 ms median and 536.0 ms p95, against 209.3 ms and 238.6 ms for Clef and 38.8 ms and 122.4 ms for Clef-flash. The trade-off is therefore a higher price for lower latency. Cloudflare says it does not read, store, or train on requests.

Cloudflare-reported latency: Clef-flash 38.8 ms median (122.4 ms p95), Clef 209.3 ms (238.6 ms), Jev 524.1 ms (536.0 ms).

For self-hosting, The Register gives minimum VRAM of 41 GB for Clef-flash and 85 GB for Clef, assuming single concurrency and a 64k context. As a rough guide, Clef-flash fits a 48 GB-class GPU or a Mac with large unified memory, while Clef needs a 96 GB-class setup, such as multiple GPUs or a Mac with 128 GB of unified memory. More concurrency or KV-cache headroom needs more. Whether local runtimes such as Ollaya can run Clef weights is unconfirmed, and we have seen no official support announcement.

How to use it: a Workers AI API example

On Workers AI you call the model by ID, @cf/cloudflare/clef or @cf/cloudflare/clef-flash. A request consists of a state and a set of questions, and one call can carry several questions. The example below judges a checkout outage support request on urgency (noul), responsible team (choice), and impact (score) at the same time.

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{
    "model": "clef",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
      "urgent": {
        "type": "noul",
        "instructions": "Is this support request urgent?"
      },
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
          "billing": "Payments, invoices, and refunds",
          "technical": "Outages, errors, and configuration",
          "sales": "Plans and upgrades"
        }
      },
      "severity": {
        "type": "score",
        "instructions": "How severe is the customer impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"]
      }
    }
  }'

Three questions about one state go out together, and each returns typed probabilities. For choice, criteria holds each option with a description; for score, criteria holds ordered level labels. The instructions text and option descriptions drive accuracy, so tune the wording on a small sample of real data. Because the API is described as Jev-compatible, Cloudflare says code already using Jev can be tried by swapping the endpoint and model name.

Benchmark results

These are the main benchmarks Cloudflare published, in the order Clef / Clef-flash / Jev. All comparisons were run by Cloudflare itself.

BenchmarkMetricClefClef-flashJev
BFCLcase exact98.4798.7695.75
ToolRetnDCG@1069.1966.4365.28
API-Bankacc91.9393.1188.19
Home appliancescase exact82.9597.7352.27
When2Callacc72.3765.5880.97
BANKING77macro-F194.2090.9379.74
CLINC150+OOSmacro-F197.4366.7789.27
BRIGHTnDCG@1045.9139.2647.52
Amazon ESCImacro-F157.4857.3955.21
PhishNChipsacc79.6075.0562.55

Of the ten benchmarks, Jev wins two: When2Call (deciding whether to call a tool or ask a clarifying question) and BRIGHT (reasoning-heavy retrieval). Clef or Clef-flash wins the rest. The gap between the two Clef models matters: the small Clef-flash beats Clef on BFCL, API-Bank, and Home appliances, but falls far behind on CLINC150+OOS (intent classification with out-of-scope detection), scoring 66.77 against Clef's 97.43. For classification tasks where out-of-scope inputs matter, be careful about choosing the small model.

Cloudflare also published workflow-oriented scores across Jev Decision Index areas, in the order Clef / Clef-flash / Jev: invoice processing 64.7 / 57.1 / 61.8, customer service 76.3 / 77 / 76.0, security incidents 62.9 / 61.7 / 61.7, and agent trace observability 68.5 / 69.8 / 71.6. Clef beats Jev in three of four areas and loses only on agent trace observability. These results are self-reported and have not yet been reproduced on the official Decision Index.

Comparison with Jev and how to choose

AspectCloudflare Clef / Clef-flashTypeSafe Jev
InputsText and imagesText only
WeightsOpen weight (Apache 2.0)Not released (hosted API)
Price$0.24/M tokens (Workers AI)$0.042/M tokens
Median latency209.3 ms (Clef) / 38.8 ms (flash)524.1 ms
Self-hostingPossible (41 GB / 85 GB minimum VRAM)Not possible
StrengthsClassification and tool-selection benchmarks, speed, imagesWhen2Call, BRIGHT, low unit price

As a rule of thumb, if cost is the top priority, inputs are text only, and Jev's accuracy is enough, staying with Jev at roughly one-sixth the unit price is well justified. If latency matters (real-time processing), you want images in the decision, or you must keep weights in your own environment because data cannot leave, Clef becomes a strong candidate, and Clef-flash offers a 38.8 ms median. For Jev's background and basics, see our practical guide to Jev. The reliable approach is to measure accuracy, latency, and cost on your own data.

The RL fine-tuning platform

Cloudflare also announced a platform for reinforcement-learning fine-tuning of decision models on your own data. The pipeline captures datasets from real requests in AI Gateway, generates rollouts in Workers AI, evaluates them in an RL sandbox on Containers, updates weights with a Trainer, and redeploys through Workers AI BYO Model. For now it is offered with hands-on support from Cloudflare's FDE (forward deployed engineer) team, with self-serve planned for later. Pricing has not been disclosed. It targets teams that want to train a model on their own decision criteria, but it is not something anyone can use immediately.

Caveats

The benchmarks are Cloudflare's own figures with no third-party reproduction confirmed; the Decision Index area results in particular are self-reported
- It is open weight with undisclosed training data. Apache 2.0 allows commercial use, but you need to assess data-provenance risk yourself
- Workers AI costs about 6x Jev per token; estimate the cost difference before adopting it for high-volume workloads
- The minimum VRAM (41 GB / 85 GB) assumes single concurrency and a 64k context; higher concurrency needs more headroom
- Operation on local runtimes such as Ollaya is unconfirmed; start by considering a setup that loads the weights directly
- Video support comes from The Register's report; check the official specification
- The RL fine-tuning platform is hands-on support only, with undisclosed pricing

FAQ

Should I choose Clef or Clef-flash?

Pick Clef-flash if speed and VRAM matter most (38.8 ms median, 41 GB minimum VRAM), and Clef if accuracy matters most (209.3 ms median, 85 GB minimum VRAM). In Cloudflare's benchmarks Clef-flash wins on BFCL and Home appliances, while Clef wins by a wide margin on CLINC150+OOS, so the ranking flips by task. Testing both on your own data is the safest approach.

Is migrating from Jev easy?

Cloudflare says the API is Jev-compatible and works as a drop-in replacement. However, Workers AI costs $0.24 per million tokens, about 6x Jev's $0.042. Jev also wins on When2Call and BRIGHT, so compare accuracy and cost on your own decision tasks before switching.

Can I run Clef locally?

The weights are published under Apache 2.0 on Hugging Face (Cloudflare/clef and Cloudflare/clef-flash), so self-hosting is possible. According to The Register, the minimum VRAM is 41 GB for Clef-flash and 85 GB for Clef, assuming single concurrency and a 64k context. Support in Ollaya is unconfirmed for now.

Are the training datasets public?

No. Only the model weights are released, not the training datasets. That makes it open-weight rather than open source in the strict sense. Apache 2.0 still allows commercial use and modification.

Feel free to contact us

Contact Us