Decision models on images: classify receipts, screenshots and forms locally with Clef
Written by Jakub Rusinowski · Last updated
To send an image to a local decision model, pull clef or clef-flash (Ollama 0.35.1 or later), base64-encode a PNG, JPEG or WebP file, and put it in the images array of a POST /v1/systemone request next to a short text state and your typed questions. Ollama's API reference names these two models as the ones that accept images. The image is scored together with the text, and every question in the request sees the same images.
This guide gives you a request that works, a Python helper that turns any phone photo into something the endpoint accepts, a receipt-triage example with routing, and a way to measure accuracy on your own pictures. That last step matters: Cloudflare has published no image benchmark for either model.
What you need
| Item | Detail |
|---|---|
| Ollama | 0.35.1 or later. The Clef and Clef-Flash pages in Ollama's library state this minimum; the setup guide covers installing and upgrading. |
| A vision-capable decision model | clef (27B, listed by Ollama at about 18 GB) or clef-flash (9B, about 11 GB for the Q8_0 GGUF). Nimble and Tev1 are not listed for image input. |
| Memory | Enough to hold the weights plus the image tokens. Check your machine on the model page before you pull. |
| Python 3.9+ and Pillow | For Steps 3 to 5 only: pip install pillow pillow-heif. pillow-heif is optional and lets Pillow open iPhone HEIC photos. |
What an image adds, and what it cannot do
A decision model returns probabilities, never text. It cannot read a receipt back to you or extract the total as a number. It can answer narrow questions about what it sees:
- Is the total on this receipt fully legible, or should the employee retake the photo?
- Is this a receipt, an invoice or a boarding pass?
- Does this screenshot show an error dialog, the login form, or neither?
- Is a signature present in the signature box?
If you need the values printed on the page, extract them with OCR or a vision-language model and let the decision model make the judgement call around them. Your code stays in charge of the logic; the model scores the messy input.
Cloudflare's launch post says Clef has "a vision encoder so it's able to take in images and classify visual content", but every benchmark row it publishes is a text task. Until you have measured it on your own images (Step 5), treat image accuracy as unknown.
Step 1: Pull the model and run a text-only request first
Pull the model the same way as the other decision models. Do not use ollama run; decision models are called through the API.
ollama -v # 0.35.1 or later
ollama pull clef # or: ollama pull clef-flashBefore adding an image, prove the model answers a text-only request. If this fails, the problem is your install, not your image, and the errors guide will sort it out.
curl http://localhost:11434/v1/systemone \
-H 'Content-Type: application/json' \
-d '{
"model": "clef",
"state": "Receipt from a cafe: two coffees.",
"questions": {
"meal": {"type": "noul", "instructions": "Is this a food or drink purchase?"}
}
}'Did it work? You get an answers.meal.noul number close to 1. The first request is slow while the model loads into memory; later requests are fast.
Step 2: Send your first image
Use a screenshot, or a photo already shrunk to about 1,600 pixels on its long side, for this first test. A full-resolution phone photo can be larger than the loaded context allows, and Step 3 shrinks photos for you.
Build the request body in a file. Even a modest photo encodes to hundreds of kilobytes of base64, far more than a command line can hold, so passing it with -d '...' fails with "Argument list too long". Linux caps a single argument at 131,072 bytes; macOS caps the whole command line at about 1 MiB, which a full-size photo exceeds.
macOS and Linux (bash or zsh):
# request.json: the image goes in "images", the text goes in "state"
{
printf '{"model":"clef","state":"Expense receipt photographed by an employee.","images":["'
base64 < receipt.jpg | tr -d '\n'
printf '"],"questions":{"legible":{"type":"noul","instructions":"Is the receipt total fully legible?"},"category":{"type":"choice","instructions":"What kind of purchase is this?","criteria":{"meals":"Restaurants, cafes, food delivery","travel":"Transport, fuel, parking, accommodation","office":"Stationery, equipment, software","other":"Anything else"}}}}'
} > request.json
curl http://localhost:11434/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary @request.jsontr -d '\n' matters: Linux base64 wraps its output every 76 characters, and a raw line break cannot appear inside a JSON string.
Windows (PowerShell 5.1 or 7+):
# "$PWD\receipt.jpg": .NET resolves relative paths against the process folder, not your PowerShell location
$image = [Convert]::ToBase64String([IO.File]::ReadAllBytes("$PWD\receipt.jpg"))
$body = [ordered]@{
model = "clef"
state = "Expense receipt photographed by an employee."
images = @($image)
questions = [ordered]@{
legible = [ordered]@{ type = "noul"; instructions = "Is the receipt total fully legible?" }
category = [ordered]@{
type = "choice"
instructions = "What kind of purchase is this?"
criteria = [ordered]@{
meals = "Restaurants, cafes, food delivery"
travel = "Transport, fuel, parking, accommodation"
office = "Stationery, equipment, software"
other = "Anything else"
}
}
}
} | ConvertTo-Json -Depth 8
Invoke-RestMethod -Uri "http://localhost:11434/v1/systemone" -Method Post -ContentType "application/json" -Body $body | ConvertTo-Json -Depth 8The response has the same shape as a text-only one. This one is illustrative, not a measurement, and it omits the usage block:
{
"model": "clef",
"answers": {
"legible": {"type": "noul", "noul": 0.93},
"category": {
"type": "choice",
"choice": "meals",
"probabilities": {"meals": 0.88, "travel": 0.06, "office": 0.04, "other": 0.02},
"confidence": 0.648
}
}
}noul is the probability that the statement is true, so 0.93 means "very likely legible", not "legible". confidence only says how peaked the distribution is; it is not the chance of being right (see the confidence guide). usage.input_tokens counts the image positions as well as the text, so it tells you what an image costs on your setup.
Did it work? You get an answers object, not an error. A 400 or 413 here usually points at the image itself: see Ollama /v1/systemone image errors.
Step 3: Make any photo safe to send
Ollama accepts base64 PNG, JPEG and WebP only. Real photos arrive as HEIC from iPhones, with the rotation stored in metadata instead of in the pixels, at 12 megapixels or more, sometimes with transparency. This helper fixes all of that, and decide() refuses to send a body over the documented limit instead of uploading it first.
Save it as decision_images.py:
"""decision_images.py - helpers for sending images to Ollama's /v1/systemone (Clef, Clef-Flash).
Needs Python 3.9+ and Pillow: pip install pillow pillow-heif
(pillow-heif is optional; it lets Pillow open iPhone HEIC photos.)
"""
import base64
import io
import json
import os
import urllib.error
import urllib.request
from PIL import Image, ImageOps
try:
from pillow_heif import register_heif_opener
register_heif_opener() # teaches Pillow to open .heic / .heif files
except ImportError:
pass # without it, HEIC files raise an error in encode_image()
BASE_URL = os.environ.get("OLLAMA_HOST", "127.0.0.1:11434").replace("0.0.0.0", "127.0.0.1")
if "://" not in BASE_URL:
BASE_URL = "http://" + BASE_URL
TEXT_LIMIT = 64 * 1024 # request body limit without images (Ollama API reference)
IMAGE_LIMIT = 32 * 1024 * 1024 # request body limit with images, JSON and base64 included
class DecisionError(Exception):
def __init__(self, status, message):
super().__init__(f"HTTP {status}: {message}")
self.status = status
def encode_image(path, max_side=1600, fmt="JPEG", quality=85):
"""Open any image Pillow can read and return base64 text that /v1/systemone accepts.
- applies the phone's EXIF rotation, so the model sees the photo upright
- shrinks the longest side to max_side (never enlarges)
- flattens transparency onto white and converts to RGB
- writes PNG, JPEG or WEBP, the only formats Ollama accepts
- copies no metadata: GPS position and camera details stay behind
"""
with Image.open(path) as im:
im = ImageOps.exif_transpose(im)
im.thumbnail((max_side, max_side))
if im.mode in ("RGBA", "LA") or (im.mode == "P" and "transparency" in im.info):
im = im.convert("RGBA")
flat = Image.new("RGB", im.size, "white")
flat.paste(im, mask=im.getchannel("A"))
im = flat
else:
im = im.convert("RGB")
buf = io.BytesIO()
options = {"quality": quality} if fmt in ("JPEG", "WEBP") else {}
im.save(buf, fmt, **options)
return base64.b64encode(buf.getvalue()).decode("ascii")
def decide(model, state, questions, images=None, keep_alive="10m", timeout=300):
"""POST one decision request. Raises DecisionError on any HTTP error."""
body = {"model": model, "state": state, "questions": questions, "keep_alive": keep_alive}
if images:
body["images"] = images
raw = json.dumps(body).encode("utf-8")
limit = IMAGE_LIMIT if images else TEXT_LIMIT
if len(raw) > limit:
raise DecisionError(413, f"body is {len(raw):,} bytes, over the {limit:,}-byte limit; shrink it before sending")
request = urllib.request.Request(
BASE_URL + "/v1/systemone", data=raw, headers={"Content-Type": "application/json"}
)
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
return json.loads(response.read())
except urllib.error.HTTPError as err:
detail = err.read().decode("utf-8", "replace").strip()
raise DecisionError(err.code, detail) from NoneWhat encode_image() does, and why:
| Step | Why |
|---|---|
| Applies the EXIF rotation | Phones store "rotate 90 degrees" as a tag instead of turning the pixels. If the tag is ignored, the model sees a sideways receipt. |
Shrinks the longest side to max_side (default 1600 px) | Image positions count toward the prompt, which must fit the loaded context, and large images inflate the request body. 1600 is a starting point, not a vendor recommendation; Step 5 shows how to choose. |
| Flattens transparency onto white, converts to RGB | JPEG cannot store transparency, so alpha is flattened onto white; everything else becomes a plain RGB image. |
| Writes JPEG, PNG or WebP | The only formats Ollama accepts. Use fmt="PNG" for screenshots with small text; JPEG at quality 85 is a good default for photos. |
| Copies no metadata | GPS position and camera details stay in the original file. |
decide() also sets keep_alive to ten minutes so a burst of photos does not pay the model load again. Ollama's default is five minutes.
Did it work? python -c "from decision_images import encode_image; print(len(encode_image('receipt.jpg')))" prints a number: the length of the base64 text. If a .heic file raises UnidentifiedImageError, pillow-heif is not installed.
Step 4: A receipt-triage example
Three questions about one photo, then a routing decision your code owns. Save this as triage.py next to the helper:
# triage.py - sort expense-receipt photos with Clef
import sys
from decision_images import decide, encode_image
QUESTIONS = {
"legible": {
"type": "noul",
"instructions": "Is the receipt total fully legible?",
},
"cut_off": {
"type": "noul",
"instructions": "Is any part of the receipt cut off or missing from the photo?",
},
"category": {
"type": "choice",
"instructions": "What kind of purchase is this?",
"criteria": {
"meals": "Restaurants, cafes, food delivery",
"travel": "Transport, fuel, parking, accommodation",
"office": "Stationery, equipment, software",
"other": "Anything else",
},
},
}
def triage(path, model="clef"):
result = decide(
model,
"Expense receipt photographed by an employee.",
QUESTIONS,
images=[encode_image(path)],
)
answers = result["answers"]
legible = answers["legible"]["noul"]
cut_off = answers["cut_off"]["noul"]
category = answers["category"]["choice"]
category_p = answers["category"]["probabilities"][category]
# Starting thresholds, not recommendations. Measure your own (see the calibration guide).
if legible < 0.80 or cut_off > 0.30:
return "retake", {"legible": legible, "cut_off": cut_off}
if category_p < 0.70:
return "human_review", {"category": category, "p": category_p}
return "auto_file", {"category": category, "p": category_p}
if __name__ == "__main__":
for path in sys.argv[1:]:
action, detail = triage(path)
print(f"{path}: {action} {detail}")Run it with python triage.py receipts/*.jpg. The model never decides what happens next; it supplies three probabilities, and your code maps them to retake, human_review or auto_file. The thresholds above are starting values to replace after Step 5.
Two habits make this reliable. Give the choice question an "other" option, because the model will pick one of your options even when none fits. And ask about the picture in the questions, not only in the state: the state is the context ("expense receipt photographed by an employee"), the questions are what you want judged.
Step 5: Measure accuracy on your own images
This is the step to take seriously. Collect 100 to 300 real photos for one question, label them yourself, and include the bad ones: blurry, tilted, partly cut off, taken in a restaurant at night. Production photos are worse than the ones you would pick for a test.
The confidence and calibration guide has an evaluation script. To use it for images, store a path in each example and change one thing in run():
# in eval_decisions.py: add `from decision_images import encode_image`, then replace the requests.post call with
r = requests.post(URL, json={
"model": model,
"state": ex["state"],
"images": [encode_image(ex["image"])],
"questions": {"q": QUESTION},
"keep_alive": "10m",
}, timeout=300)Use examples shaped like {"state": "Expense receipt photographed by an employee.", "image": "receipts/0001.jpg", "label": "meals"}. Then run it four ways, each model at each size:
clefandclef-flash, because Cloudflare gives no image numbers for either.max_sideof 1600 and 1024 (encode_image(path, max_side=1024)). Pick the smallest size where accuracy holds on the small-print cases; it saves memory and time.
Read the accuracy-versus-coverage output the way the calibration guide describes, and set your accept, retake and review thresholds from it.
Limits that matter for images
Ollama /v1/systemone | Cloudflare Workers AI (hosted) | |
|---|---|---|
| Formats | Base64 PNG, JPEG, WebP | Embedded PNG, JPEG, WebP |
| URLs and data URLs | Not supported | Remote URLs not accepted |
| Request body | 64 KiB without images, 32 MiB with images (JSON and base64 included) | 13 MiB for the whole body |
| Image count | Not stated in the API reference | At most 4 |
| Per image | Not stated | 4 MiB and 16 megapixels each, 8 MiB total decoded |
| Context | The whole input must fit the loaded context window; never truncated, so too long is a 400 | 65,536 tokens; a long text state is truncated to fit |
| Image order | Shared by all questions, in request order | Placed before the state |
Workers AI figures are from its Clef model page as read on 3 October 2026. If you might move between local and hosted, stay inside the stricter hosted numbers: at most four images, under 4 MiB and 16 megapixels each.
Several images in one request
Images are shared by all questions, so you cannot point a question at "the second image". For a separate verdict per image, send one request per image. For a question about the whole set, such as "is a signature present on any of these pages?", send them together. Each image adds its own tokens (watch usage.input_tokens), so keep the count small and test the combination on your labelled set before trusting it.
Keep going
- Ollama /v1/systemone image errors: 400, 413, HEIC and sideways photos
- Decision models: the hub
- Confidence, thresholds and calibration
- Decision request builder
- Clef on the model library
Verified against: Ollama System One API reference, Ollama library pages for Clef and Clef-Flash, Cloudflare Workers AI Clef model page, Cloudflare blog "Introducing Clef" (all read 3 October 2026).
Frequently asked questions
Can I send a PDF?
Not directly. Only PNG, JPEG and WebP are accepted. Render the pages to images first, for example pdftoppm -png -r 150 form.pdf page (poppler-utils on Linux, brew install poppler on macOS), then send each page through encode_image().
Can I pass an image URL?
No. Ollama's reference says URLs and data URLs are not supported. Download the file, then send the base64 of its bytes.
Can I send video?
Not through Ollama's endpoint; its reference lists images only. The Hugging Face model card describes video input for Cloudflare's own Python code, which is a different path.
Does the model read the text in the image?
It scores your options using the picture and the text you send, but it never outputs text, so it cannot transcribe or extract values. Use OCR for that and keep the decision model for the judgement.
Does the photo leave my machine?
Not when you call a local Ollama on 127.0.0.1. Re-encoding with encode_image() also drops metadata such as GPS before the file goes anywhere else. If you bind Ollama to other interfaces, remember its API has no authentication.
Which is better for images, Clef or Clef-Flash?
There is no published image data for either, so measure both on your own photos (Step 5). On Cloudflare's own text benchmarks the two have different strengths; the model page lists the figures and says whose they are.
Verified against: Ollama System One API reference, Ollama library pages for Clef and Clef-Flash, Cloudflare Workers AI Clef model page and Cloudflare blog, read 3 October 2026.
Keep going
- Decision models (Jev-style System One): run them locally
- Run Jev-style decision models locally with Ollama (Nimble, Tev1)
- Open Jev reproductions beyond Ollama: Winnow, Decider, JevK5, AutoJev, Laya
- Decision-model confidence, thresholds and calibration — use the probabilities safely
- Ollama /v1/systemone errors and fixes
- Ollama /v1/systemone image errors
- Decision request builder