Clef-Flash 9B — VRAM & /v1/systemone setup
Written by Jakub Rusinowski · Last updated
Cloudflare's 9B decision model on Qwen3.5-9B, with a vision encoder. In Ollama as clef-flash (Apache-2.0) and called at /v1/systemone; Cloudflare has published no image benchmark.
Clef-Flash 9B needs about 6 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
Call Clef-Flash 9B
Pull it with Ollama 0.35.1 or newer, then POST to /v1/systemone. Do not use ollama chat commands — this is not a chat model.
ollama pull clef-flashcurl http://localhost:11434/v1/systemone \
-H 'Content-Type: application/json' \
-d '{
"model": "clef-flash",
"state": {
"ticket": "I was charged twice. Please refund the extra payment."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above"
}
}
}
}'Build a request for this model Decision models guide
The builder writes text requests only; it does not add an images field. See the images guide to send photos.
Hosted: Cloudflare Workers AI (@cf/cloudflare/clef-flash): $0.09 per million input tokens, as of 2026-10-03. Informational; not a recommendation. Source ↗
Related guides: Decision models on images · Image request errors · Ollama setup · /v1/systemone errors
Other decision models: Nimble 9B · Tev1 4B · Tev1 0.8B · Winnow 12B · Winnow E4B · Decider 2B · Decider 4B · Decider 35B-A3B (NVFP4) · JevK5 4B · Intern-Decision 4B · AutoJev 27B · Laya · Clef 27B
Hardware fit
Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown. Image tokens are not modelled: each image adds to the prompt, so a request with images needs more memory than these rows.
| Quant | Memory | VRAM | Fit |
|---|---|---|---|
| Ollama build Publisher build · 10.93 GB file | 12.9 GB | Fits | |
| Q4_K_M 4.83 bpw · modelled quant | 7.7 GB | Fits | |
| Q6_K 6.56 bpw · modelled quant | 9.7 GB | Fits | |
| Q8_0 8.50 bpw · modelled quant | 12 GB | Fits |
Published file size: Ollama build 10.93 GB. A download size from the model publisher — not a VRAM requirement.
How it was scored
Two different suites on two different scales. Never compare the numbers across the two cards.
Author's public benchmark
Cloudflare's internal run of the Decision Index 0.2.1 suite (vendor-reported)- API-Bank (accuracy)
- 93.1%
- BANKING77 (macro-F1)
- 90.9%
- CLINC150+OOS (macro-F1)
- 66.8%
- RAGTruth (hallucination F1)
- 35.6%
- Home appliance simulator (case exact accuracy)
- 97.7%
Raw per-benchmark scores from the vendor's own internal run, not chance-corrected and not independently reproduced. They are not the Decision Index's balanced skill, so do not compare them with the card beside this one. Six of the published rows are shown; the rest are on the Hugging Face model card. Cloudflare has published no image benchmark for either Clef model.
Source ↗Decision Index 0.2.1 (snapshot 2026-09-28)
Decision Index 0.2.1Clef and Clef-Flash were released on 1 October 2026, after the 28 September snapshot. Community-maintained; not affiliated with the model authors.
Source ↗Decision models are not ranked on chat, creative or coding scores.
Specifications
Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source. Still unconfirmed: contextWindow, baseModel.
- Parameters
- 9.41B
- Context window
- 64K (documented)
- Architecture
- Fine-tune of Qwen/Qwen3.5-9B
- Provider
- Cloudflare
- Licence
- Apache-2.0
- Specified at
- Q4_K_M
- System RAM
- 14 GB
- Record updated
- 2026-10-03
Commercial use permitted. No usage restrictions beyond attribution.
Other Clef sizes
Clef-Flash 9B — frequently asked questions
What is Clef-Flash 9B?
Clef-Flash 9B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.
How do I run Clef-Flash 9B locally?
Install Ollama 0.35.1 or newer, run `ollama pull clef-flash`, then POST your state and questions to http://localhost:11434/v1/systemone. It is called through the API, not a chat session.
How much memory does Clef-Flash 9B need?
About 6 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.
Should I use Clef or Clef-Flash?
On Cloudflare's own text benchmarks (vendor-reported) they differ. Clef-Flash scores higher on BFCL (98.8 vs 98.5), API-Bank (93.1 vs 91.9) and the home appliance simulator (97.7 vs 83.0), which suit decisions over a small set of options. Clef scores higher on CLINC150+OOS (97.4 vs 66.8), a many-class intent task with out-of-scope inputs, and on RAGTruth hallucination detection (79.4 vs 35.6). Neither has a published image result, so measure both on your own examples.
Does Clef accept images?
Ollama's API reference lists Clef and Clef Flash for image input: base64 PNG, JPEG or WebP in an images array, Ollama 0.35.1 or later. Cloudflare has published no image benchmark for either model, so measure it on your own images before relying on it. The images guide shows how.
Why does Ollama list MLX tags for Clef?
Ollama lists MLX tags for both Clef models, but its System One API reference says MLX models are not supported on /v1/systemone. Use the default tag (ollama pull clef or ollama pull clef-flash).