Cloudflare27B~17 GB VRAM at Q4_K_M

Clef 27B — VRAM & /v1/systemone setup

Written by Jakub Rusinowski · Last updated

Cloudflare's 27B decision model on Qwen3.8-27B, with a vision encoder. In Ollama as clef (Apache-2.0) and called at /v1/systemone; Cloudflare has published no image benchmark.

Clef 27B needs about 17 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Call Clef 27B

Pull it with Ollama 0.35.1 or newer, then POST to /v1/systemone. Do not use ollama chat commands — this is not a chat model.

ollama pull clef
curl http://localhost:11434/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "clef",
    "state": {
      "ticket": "I was charged twice. Please refund the extra payment."
    },
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this ticket?",
        "criteria": {
          "billing": "Payments and refunds",
          "technical": "Bugs and integrations",
          "other": "None of the above"
        }
      }
    }
  }'
Weights on Hugging Face: Cloudflare/clef ↗

Build a request for this model Decision models guide

The builder writes text requests only; it does not add an images field. See the images guide to send photos.

Hosted: Cloudflare Workers AI (@cf/cloudflare/clef): $0.24 per million input tokens, as of 2026-10-03. Informational; not a recommendation. Source ↗

Related guides: Decision models on images · Image request errors · Ollama setup · /v1/systemone errors

Other decision models: Nimble 9B · Tev1 4B · Tev1 0.8B · Winnow 12B · Winnow E4B · Decider 2B · Decider 4B · Decider 35B-A3B (NVFP4) · JevK5 4B · Intern-Decision 4B · AutoJev 27B · Laya · Clef-Flash 9B

Hardware fit

Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown. Image tokens are not modelled: each image adds to the prompt, so a request with images needs more memory than these rows.

QuantVRAMFit
Ollama build
Publisher build · 17.99 GB file
20.5 GBFits
Q4_K_M
4.83 bpw · modelled quant
19.1 GBFits
Q6_K
6.56 bpw · modelled quant
25 GBOffload
Q8_0
8.50 bpw · modelled quant
31.6 GBOffload

Published file size: Ollama build 17.99 GB. A download size from the model publisher — not a VRAM requirement.

How it was scored

Two different suites on two different scales. Never compare the numbers across the two cards.

Author's public benchmark

Cloudflare's internal run of the Decision Index 0.2.1 suite (vendor-reported)
BFCL (case exact accuracy)98.5%
API-Bank (accuracy)
91.9%
BANKING77 (macro-F1)
94.2%
CLINC150+OOS (macro-F1)
97.4%
RAGTruth (hallucination F1)
79.4%
Home appliance simulator (case exact accuracy)
83.0%

Raw per-benchmark scores from the vendor's own internal run, not chance-corrected and not independently reproduced. They are not the Decision Index's balanced skill, so do not compare them with the card beside this one. Six of the published rows are shown; the rest are on the Hugging Face model card. Cloudflare has published no image benchmark for either Clef model.

Source ↗

Decision Index 0.2.1 (snapshot 2026-09-28)

Decision Index 0.2.1
Balanced skill—

Clef and Clef-Flash were released on 1 October 2026, after the 28 September snapshot. Community-maintained; not affiliated with the model authors.

Source ↗

Decision models are not ranked on chat, creative or coding scores.

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source. Still unconfirmed: contextWindow, baseModel.

Parameters
27.36B
Context window
64K (documented)
Architecture
Fine-tune of Qwen/Qwen3.8-27B
Provider
Cloudflare
Licence
Apache-2.0
Specified at
Q4_K_M
System RAM
36 GB
Record updated
2026-10-03
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Other Clef sizes

Clef 27B — frequently asked questions

What is Clef 27B?

Clef 27B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.

How do I run Clef 27B locally?

Install Ollama 0.35.1 or newer, run `ollama pull clef`, then POST your state and questions to http://localhost:11434/v1/systemone. It is called through the API, not a chat session.

How much memory does Clef 27B need?

About 17 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.

Should I use Clef or Clef-Flash?

On Cloudflare's own text benchmarks (vendor-reported) they differ. Clef-Flash scores higher on BFCL (98.8 vs 98.5), API-Bank (93.1 vs 91.9) and the home appliance simulator (97.7 vs 83.0), which suit decisions over a small set of options. Clef scores higher on CLINC150+OOS (97.4 vs 66.8), a many-class intent task with out-of-scope inputs, and on RAGTruth hallucination detection (79.4 vs 35.6). Neither has a published image result, so measure both on your own examples.

Does Clef accept images?

Ollama's API reference lists Clef and Clef Flash for image input: base64 PNG, JPEG or WebP in an images array, Ollama 0.35.1 or later. Cloudflare has published no image benchmark for either model, so measure it on your own images before relying on it. The images guide shows how.

Why does Ollama list MLX tags for Clef?

Ollama lists MLX tags for both Clef models, but its System One API reference says MLX models are not supported on /v1/systemone. Use the default tag (ollama pull clef or ollama pull clef-flash).