Clef — local AI model by Cloudflare

Written by Jakub Rusinowski · Last updated

Cloudflare's decision models, Clef 27B and Clef-Flash 9B: typed questions in, a probability per option out, in one scoring pass with no text generation. Both are in Ollama's library and accept images.

Variants

The smallest Clef variant needs about 6 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.

ModelVRAM
Clef 27B →
27B
~17.3 GB
Clef-Flash 9B →
9B
~6.5 GB

Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.

How to run Clef locally

Install Ollama, then pull the tag.

ollama pull clef

Pick a size above for its own VRAM figure, speed estimate and install command.

Licence

Apache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Applies to: Clef 27B, Clef-Flash 9B

Recommended GPU

The cheapest catalogued GPU that runs Clef locally (min 6 GB VRAM) is the Intel Arc B570 (10 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Intel Arc B570 10GB
10 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.

Clef — frequently asked questions

How much VRAM does Clef need?

Clef needs about 6 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Clef 27B (17 GB, Q4_K_M); Clef-Flash 9B (6 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.

Can I run Clef on an RTX 4090 (24 GB)?

Yes — Clef runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.

What quantization should I use for Clef?

Q4_K_M is the best balance of quality and VRAM for Clef in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.

How do I run Clef locally?

Clef is a decision model: it is called over an HTTP endpoint (/v1/systemone), not chatted with. Where a variant is available in Ollama 0.35 or newer, pull it with `ollama pull` and send requests to the local server; the others ship their own server. Each variant page shows the exact commands for that model.