Home / Guides / Troubleshooting / Windows

Ollama /v1/systemone image errors: 400, 413, HEIC and sideways photos

WindowsLinuxmacOS

Written by Jakub Rusinowski · Last updated October 3, 2026

AI educator & workshop leader on local LLM deployment

The error

HTTP 400 or 413 from POST /v1/systemone as soon as the request has an "images" array, or a 200 whose answers ignore the picture

Which one is it?

If you seeThe cause isGo to
400 on the image request, and the same request without images worksThe file is not a PNG, JPEG or WebP (HEIC from an iPhone is the usual one), the base64 is malformed, or the model is not Clef or Clef-Flash (only those two take images; nimble and tev1 reject them)Fix 1, Fix 2, then Fix 3
413 on a small photo, or 413 on a large oneOn a small photo the image is inside state, so Ollama applies the 64 KiB limit for requests without images; on a large photo the body is over 32 MiB once base64 and JSON are countedFix 3 (small photo) or Fix 5 (large photo)
Argument list too long, or curl fails before sendingBase64 passed as a command-line argumentFix 4
400 on a full-size photo; a screenshot or a small image worksThe image positions do not fit the loaded contextFix 6
200, but answers look random or wrongSideways photo, shrunk too far, or the question never mentions the pictureFix 7
500, or the first image request takes a very long timeThe model did not fit in memory, or it is still loadingFix 8

When you see it

Ollama's reference documents the status codes, and the wording of only the size-limit error; other error text can change between releases. Read the status first, then use the table. Verified against Ollama's System One API reference and FAQ, the Ollama library pages for Clef and Clef-Flash, and Cloudflare's Workers AI Clef model page, all read 3 October 2026.

What's actually going on

Ollama is strict about image requests, and each requirement has its own fix below. It needs a model with vision weights, which its reference names as Clef or Clef Flash. Each entry in images must be base64 text for a PNG, JPEG or WebP file; URLs and data URLs are not supported. The whole body must fit 64 KiB without images or 32 MiB with them. Finally, the complete input, including the image positions, must fit the loaded context window, and Ollama never truncates, so too much input is a 400 rather than a quietly cropped picture.

How to fix it

1. Use Clef or Clef-Flash on Ollama 0.35.1 or later

Confirm the version, and that the model is pulled. Nimble and Tev1 are not listed for image input and reject a request that has images.

bash
ollama -v
ollama list
ollama pull clef
Did it work? The same request with "model": "clef" and no images returns answers. That proves the install; everything below is about the image.

2. Convert the file to PNG, JPEG or WebP

iPhones save photos as HEIC by default. Ollama accepts only PNG, JPEG and WebP, so HEIC, GIF, TIFF, BMP, AVIF and PDF files are out. The extension does not matter: a HEIC file renamed to .jpg is still HEIC. Screenshots from a Mac or an iPhone are PNG and need no conversion. On a Mac, the command below produces a JPEG; or open the photo in Preview and choose File, Export, Format JPEG. On Debian and Ubuntu, sudo apt install libheif-examples provides heif-convert IMG_0001.HEIC IMG_0001.jpg. On Windows, and anywhere else, the Python helper below is the dependable route. To stop it at the source, set the iPhone to Settings, Camera, Formats, Most Compatible, which captures JPEG. The more robust fix is to normalise every image in code with encode_image() from the images guide, which opens HEIC (with pillow-heif), applies rotation, shrinks and writes JPEG. It works the same on macOS, Windows and Linux.

bash
sips -s format jpeg IMG_0001.HEIC --out IMG_0001.jpg
Did it work? file IMG_0001.jpg on macOS or Linux says "JPEG image data", and the diagnostic script in Fix 9 reports JPEG.

3. Put bare base64 in the images array, nowhere else

Three mistakes cause 400s that survive Fix 2, and the first also causes a confusing 413. They are shown below: base64 inside state, an image URL, and a data URL. Only bare base64 of the file's bytes in images is right. A base64 string in state is just text, so Ollama sees a request with no images and applies the 64 KiB limit. That is why a 3 MB photo gets a 413 although 3 MB is far below 32 MiB. Two more traps are platform specific. On Windows, certutil -encode wraps the output in -----BEGIN CERTIFICATE----- lines. That is not base64 of your image. Use [Convert]::ToBase64String([IO.File]::ReadAllBytes("$PWD\photo.jpg")); the $PWD matters because .NET resolves relative paths against the process folder, not your PowerShell location. On Linux, base64 wraps lines at 76 characters, and a raw line break cannot appear inside a JSON string. Use base64 < photo.jpg | tr -d '\n'. macOS and Linux both accept that form.

Text
WRONG  "state": "Receipt: /9j/4AAQSkZJRgABAQ..."          base64 inside state
WRONG  "images": ["https://example.com/receipt.jpg"]       a URL
WRONG  "images": ["data:image/jpeg;base64,/9j/4AAQ..."]    a data URL
RIGHT  "images": ["/9j/4AAQSkZJRgABAQ..."]                 bare base64 of the file's bytes
Did it work? The 413 disappears, and images holds one long unbroken string per image.

4. Build the body in a file, not on the command line

Linux caps a single argument at 131,072 bytes, and macOS caps the whole command line at about 1 MiB, so curl -d "$BODY" fails with Argument list too long for any real photo. Write the JSON to a file and send it with --data-binary @file. The images guide has a copy-paste version that assembles the file with printf and base64; with jq installed you can also do:

bash
base64 < photo.jpg | tr -d '\n' | jq -Rs --arg model clef \
  '{model: $model, state: "Expense receipt.", images: [.],
    questions: {legible: {type: "noul", instructions: "Is the receipt total fully legible?"}}}' > request.json
curl http://localhost:11434/v1/systemone -H 'Content-Type: application/json' --data-binary @request.json
Did it work? curl sends the request, and you get an HTTP status instead of Argument list too long.

5. Keep the body under 32 MiB

The 32 MiB limit includes base64 and JSON, and base64 makes data a third larger, so a raw image file must stay under 24 MiB: a 24 MiB file encodes to exactly 32 MiB. In practice you should be far below it, because a multi-megapixel photo is also what overflows the context in Fix 6. Check the size of the request file with the command for your system. Shrink the image rather than hunting for a bigger limit; the limit is Ollama's, not a setting. encode_image(path, max_side=1600) does it, and the images guide explains how to pick the size from your own accuracy numbers.

bash
wc -c request.json              # macOS and Linux; must be under 33,554,432
PowerShell
(Get-Item request.json).Length  # Windows
Did it work? The file is under 33,554,432 bytes and the 413 is gone.

6. Shrink the image, or raise the context (400 on large photos only)

Image positions count toward the prompt, and the prompt must fit the loaded context. Ollama reports them in usage.input_tokens ("total evaluated input tokens, including image positions"). A full-resolution phone photo is a likely way to overflow it. First, shrink: send the same photo at 1,600 and then 1,024 pixels on the long side, and use the smallest size that still answers correctly. The diagnostic script in Fix 9 does this automatically when it gets a 400. If you need big images, the context can be raised. Ollama's FAQ states that the default context window is 4,096 tokens and that OLLAMA_CONTEXT_LENGTH changes it on the server; a larger context also needs more memory. We could not confirm that the decision runner reads that variable, so treat it as an experiment: change it, restart Ollama, and check whether the same image now returns a 200. How to set server environment variables differs by system; on macOS it is launchctl setenv OLLAMA_CONTEXT_LENGTH 16384 and a restart of the app, and on Linux it goes in the systemd unit, see environment variables ignored on Linux.

Did it work? The request returns 200, and usage.input_tokens is comfortably below the context you set.

7. Fix sideways, over-shrunk or ignored images (answers look wrong)

An error-free answer can still be wrong for three reasons. The photo is sideways. Phones store rotation as an EXIF tag instead of turning the pixels. Ollama's documentation does not say it applies that tag, so apply it yourself: encode_image() does. To see exactly what the model receives, write the encoded bytes back to a file and open it, as in the code below. It was shrunk too far. If the text you care about is unreadable to you in sent.jpg, it will not be reliable for the model either. Raise max_side, or crop to the part that matters before encoding. The question never refers to the picture. The model scores the images together with the text. Say what the picture is in state ("Expense receipt photographed by an employee") and ask about it in the question. Then measure on 100 or more of your own labelled images, because Cloudflare published no image benchmark for Clef or Clef-Flash; the calibration guide shows how.

Python
import base64
from decision_images import encode_image

with open("sent.jpg", "wb") as f:
    f.write(base64.b64decode(encode_image("photo.jpg")))   # open sent.jpg: this is what the model gets
Did it work? sent.jpg is upright and readable, and accuracy on your labelled set is acceptable.

8. 500 or a very slow first image request Most common fix

A 500 means loading, rendering or scoring failed, and memory is the first thing to check: the model plus its image tokens may not fit. Look at ollama ps and the server log. clef-flash is the smaller model: Ollama lists it at roughly 11 GB against about 18 GB for clef. A slow first request is the model loading into memory; Ollama unloads it after five minutes idle unless you send "keep_alive": "30m". For memory problems in general, see model requires more system memory. Server logs: macOS ~/.ollama/logs/server.log; Windows $env:LOCALAPPDATA\Ollama; Linux journalctl -u ollama --no-pager --pager-end.

Did it work? ollama ps lists the model as loaded, and the next image request no longer pays the load time.
Check what fits your hardware — check that Clef fits your memory, then build the request
Open the VRAM checker →

9. Find the cause automatically with a diagnostic script

Save the script below as check_image_request.py and run python3 check_image_request.py photo.heic (or add a model name as a second argument). It uses only the standard library, honors OLLAMA_HOST, and checks in order: the real file format from its first bytes, the base64 size against the 32 MiB limit, Ollama's version, that the model accepts images and is pulled, a text-only request, and then your image exactly as it is. If the image gets a 400 it retries the same picture at 1600, 1024 and 512 pixels (this part needs Pillow) to show whether size is the cause.

Python
#!/usr/bin/env python3
"""Diagnose why an image is rejected by Ollama's /v1/systemone (Clef, Clef-Flash).

Usage:  python3 check_image_request.py photo.heic [model]      (default model: clef)
Standard library only; Pillow is optional and adds dimensions, rotation and a size ladder.
Honors OLLAMA_HOST. Exit code 0 means every check passed.
"""
import base64
import io
import json
import math
import os
import re
import sys
import urllib.error
import urllib.request

MIN_VERSION = (0, 35, 1)           # stated on the Clef and Clef-Flash pages in Ollama's library
IMAGE_LIMIT = 32 * 1024 * 1024     # body limit with images, from Ollama's System One API reference
VISION_MODELS = ("clef", "clef-flash")
MAGIC = [
    (b"\x89PNG\r\n\x1a\n", "PNG", True),
    (b"\xff\xd8\xff", "JPEG", True),
    (b"GIF8", "GIF", False),
    (b"BM", "BMP", False),
    (b"II*\x00", "TIFF", False),
    (b"MM\x00*", "TIFF", False),
    (b"%PDF", "PDF", False),
]
HINTS = {
    400: "Not an accepted image (PNG, JPEG, WebP only), not valid base64, a URL or data URL, "
         "an unsupported model, or a prompt larger than the loaded context.",
    404: "The model is not pulled locally. Run: ollama pull {model}",
    413: "Body over 64 KiB (no images field was recognised) or over 32 MiB with images.",
    500: "The model failed to load or score. Check the server log and memory (ollama ps).",
}


def base_url():
    host = os.environ.get("OLLAMA_HOST", "127.0.0.1:11434").replace("0.0.0.0", "127.0.0.1")
    return host if "://" in host else "http://" + host


def call(path, body=None, timeout=300):
    data = json.dumps(body).encode() if body is not None else None
    request = urllib.request.Request(base_url() + path, data=data, headers={"Content-Type": "application/json"})
    try:
        with urllib.request.urlopen(request, timeout=timeout) as response:
            return response.status, json.loads(response.read())
    except urllib.error.HTTPError as err:
        raw = err.read().decode(errors="replace")
        try:
            return err.code, json.loads(raw)
        except ValueError:
            return err.code, {"error": raw.strip()}
    except urllib.error.URLError as err:
        return None, {"error": str(err.reason)}


def check(label, ok, detail=""):
    print(f"[{'ok' if ok else 'FAIL'}] {label}" + (f": {detail}" if detail else ""))
    return ok


def sniff(raw):
    """Name the real format from the first bytes, whatever the file extension says."""
    for magic, name, accepted in MAGIC:
        if raw.startswith(magic):
            return name, accepted
    if raw[:4] == b"RIFF" and raw[8:12] == b"WEBP":
        return "WebP", True
    if raw[4:8] == b"ftyp":
        brand = raw[8:12].decode("ascii", "replace").strip()
        return ("AVIF" if brand.startswith("avif") else "HEIC/HEIF") + f" ({brand})", False
    return "unknown", False


def request_for(model, images):
    return {
        "model": model,
        "state": "Image diagnostic.",
        "images": images,
        "questions": {"q": {"type": "noul", "instructions": "Does the attached image contain any visible content?"}},
    }


def describe(status, body, model):
    if status == 200:
        tokens = body.get("usage", {}).get("input_tokens")
        return f"HTTP 200, input_tokens={tokens}"
    return f"HTTP {status}: {body.get('error', body)}\n      hint: " + HINTS.get(status, "Unexpected response.").format(model=model)


def shrunk(raw, side):
    from PIL import Image, ImageOps
    with Image.open(io.BytesIO(raw)) as im:
        im = ImageOps.exif_transpose(im)
        im.thumbnail((side, side))
        out = io.BytesIO()
        im.convert("RGB").save(out, "JPEG", quality=85)
    return base64.b64encode(out.getvalue()).decode("ascii")


def run(path, model):
    failures = 0
    if not os.path.isfile(path):
        check("File exists", False, path)
        return 1
    raw = open(path, "rb").read()
    name, accepted = sniff(raw)
    ext = os.path.splitext(path)[1].lower()
    failures += not check("File format", accepted, f"{name} ({len(raw):,} bytes, extension {ext or 'none'})")
    if not accepted:
        print("      fix: convert it to PNG, JPEG or WebP first (see Fix 2). Ollama will not accept " + name + ".")

    b64_len = 4 * math.ceil(len(raw) / 3)
    failures += not check("Base64 body fits the 32 MiB limit", b64_len < IMAGE_LIMIT - 4096,
                          f"base64 is about {b64_len:,} bytes of {IMAGE_LIMIT:,}")
    try:
        from PIL import Image, ImageOps
        with Image.open(io.BytesIO(raw)) as im:
            orientation = im.getexif().get(0x0112, 1)
            print(f"      {im.width} x {im.height} px ({im.width * im.height / 1e6:.1f} megapixels), EXIF orientation {orientation}")
            if orientation not in (0, 1):
                print("      note: stored rotated; apply the EXIF rotation before sending or the model sees it sideways (Fix 7).")
    except Exception as err:  # Pillow missing, or it cannot open the format
        print(f"      (no dimensions: {type(err).__name__})")

    status, body = call("/api/version", timeout=10)
    if status is None:
        check("Ollama reachable", False, f"{body['error']} (is the server running?)")
        return 1
    version = body.get("version", "")
    match = re.match(r"(\d+)\.(\d+)\.(\d+)", version)
    parsed = tuple(int(p) for p in match.groups()) if match else None
    needed = ".".join(map(str, MIN_VERSION))
    failures += not check("Ollama version", parsed is not None and parsed >= MIN_VERSION, f"{version or 'unknown'} (need {needed} or later)")

    failures += not check("Model accepts images", model.split(":")[0] in VISION_MODELS,
                          f"{model}" + ("" if model.split(":")[0] in VISION_MODELS else " is not clef or clef-flash; Ollama's reference lists only those for image input"))
    status, body = call("/api/tags", timeout=10)
    names = [m.get("name", "") for m in body.get("models", [])] if status == 200 else []
    wanted = model if ":" in model else model + ":latest"
    if not check(f"Model {wanted} is pulled", wanted in names, "" if wanted in names else f"run: ollama pull {model}"):
        return 1

    text_only = {"model": model, "state": "Hello.", "questions": {"q": {"type": "noul", "instructions": "Is this a greeting?"}}}
    status, body = call("/v1/systemone", text_only)
    if not check("Text-only request works", status == 200, "" if status == 200 else describe(status, body, model)):
        print("      The problem is the install or the model, not your image. Stop here and fix that first.")
        return 1

    if not accepted or b64_len >= IMAGE_LIMIT - 4096:
        print("      Skipping the image request until the file problems above are fixed.")
        return 1
    status, body = call("/v1/systemone", request_for(model, [base64.b64encode(raw).decode("ascii")]))
    ok = check("Image request, your file as it is", status == 200, describe(status, body, model))
    if ok:
        return 1 if failures else 0
    failures += 1
    if status == 400:
        for side in (1600, 1024, 512):
            try:
                small = shrunk(raw, side)
            except Exception as err:
                print(f"      (cannot test smaller sizes: {type(err).__name__})")
                break
            s2, b2 = call("/v1/systemone", request_for(model, [small]))
            if check(f"Same image shrunk to {side} px", s2 == 200, describe(s2, b2, model).splitlines()[0]):
                print(f"      It works at {side} px, so size was the problem: the image positions did not fit the loaded context (Fix 6).")
                break
    return 1 if failures else 0


if __name__ == "__main__":
    if len(sys.argv) < 2:
        sys.exit(__doc__)
    sys.exit(run(sys.argv[1], sys.argv[2] if len(sys.argv) > 2 else "clef"))

On a HEIC photo that was renamed to .jpg, the first lines read:

Text
[FAIL] File format: HEIC/HEIF (heic) (460,688 bytes, extension .jpg)
      fix: convert it to PNG, JPEG or WebP first (see Fix 2). Ollama will not accept HEIC/HEIF (heic).

To test with a tiny known-good image:

bash
python3 -c "from PIL import Image; Image.new('RGB', (256, 256), 'white').save('tiny.png')"
python3 check_image_request.py tiny.png

If none of this worked

Send a tiny known-good image in the same request (the last two commands in Fix 9). If it works, the problem is your photo; if it fails, the problem is the request or the install. Ollama's API has no authentication, so do not expose it to your network to share a failing case.

Include this when you report it

Also affects: Linux · Apple

Related

A model that fits most setups:
View model & requirements →

Frequently asked questions

Why does a 3 MB photo return 413 when the limit is 32 MiB?

Because the image is almost certainly not in the images array. Base64 inside state is plain text, and a request with no images is limited to 64 KiB. Move it to images (Fix 3).

Can I raise the 32 MiB limit?

No. It is a documented limit, not a setting. Shrink the image; a smaller image also saves memory and time.

Does converting to JPEG hurt accuracy?

There is no published data on that for Clef, so measure it. For photos, JPEG at quality 85 is a reasonable default; for screenshots with small text, use PNG. The images guide shows how to compare formats and sizes on your own examples.

My request works on Cloudflare's hosted Workers AI but fails locally. Why?

The two have different rules. Workers AI limits a request to at most four images, 4 MiB and 16 megapixels each, 13 MiB for the whole body, and it truncates a long text state; Ollama allows 32 MiB, does not state a count limit, and never truncates. A request tuned for one may break on the other.

Can I send video?

Not through Ollama's endpoint; its reference lists images only.