Home / Guides / Troubleshooting / Windows

Ollama /v1/systemone errors: 400, 404, 413 explained

WindowsLinuxmacOS

Written by Jakub Rusinowski · Last updated September 30, 2026

AI educator & workshop leader on local LLM deployment

The error

{"error": "request body must not exceed 64 KiB"}

Which one is it?

If you seeThe cause isGo to
404 on every request, even with a correct model name — or a 404 that names the modelOllama is older than 0.35 (or an old server is still running), or the model was never pulled or its tag is misspelledFix 1, then Fix 2
400 with a chat model, a :cloud tag or an MLX modelThat model isn't a local System One modelFix 3: use a System One model
400 after adding more options or questions, or on long input onlyA limit: 1–64 questions and 2–26 options, or the rendered prompt doesn't fit the loaded contextFix 4 or Fix 5
400 from PowerShell only; the same JSON works in bashConvertTo-Json depth, or curl being an alias for Invoke-WebRequestFix 6: PowerShell builds a different JSON
413The request body is over 64 KiBFix 7: keep the body under 64 KiB
Works in curl, fails from a web pageCORS: the page’s origin is not allowedFix 8: allow your page’s origin

When you see it

Ollama's decision-model endpoint, POST /v1/systemone, arrived in Ollama 0.35. It is strict by design: it never truncates your input, never guesses at a malformed schema and never falls back to a chat model. That strictness means most failures come back as a clear HTTP status — if you know what each one means. Only the size limit has a documented error body; everything else is a status code plus a JSON error field whose wording can change between releases, so read the status first: 400 is an invalid request, an unsupported model or a prompt that is too long; 404 is a model that isn't pulled or an endpoint that doesn't exist; 413 is a body over 64 KiB; 500 is a model that failed to load, render or score.

What's actually going on

A decision model doesn't generate an answer; the runner scores each allowed option against the full prompt. For that to work, Ollama needs three things to be exactly right: a model whose weights were trained for this scoring (today nimble, tev1 and tev1:0.8b), a schema it can render into a prompt, and a prompt that fits in the model's loaded context window with two token positions spare for the scoring itself. Each question is rendered and scored separately against the whole state, so a long state is paid for once per question. Ollama checks all of this before loading anything and refuses with 400 rather than silently cutting your text — a truncated contract would give you a confident answer about half a contract.

How to fix it

1. Upgrade Ollama and restart the server

You need 0.35.0 or later. On macOS and Windows, quit Ollama from the menu bar or system tray before installing the new version — otherwise the old server keeps answering on port 11434 and the new endpoint seems missing. On Linux, the commands below upgrade it and restart the service.

bash
ollama -v

# Linux upgrade
curl -fsSL https://ollama.com/install.sh | sh
sudo systemctl restart ollama
Did it work? ollama -v shows 0.35.0 or later and the same request no longer returns 404. If ollama -v warns that the client and server versions differ, an old server is still running — quit it fully and start Ollama again.

2. Pull the model

Model names must match exactly — tev1:0.8b, not tev1-0.8b or tev1:0.8B.

bash
ollama pull nimble
ollama list
Did it work? ollama list shows the model, and the request returns answers.

3. Use a System One model

The endpoint accepts only local models trained for System One, packaged as GGUF with a scoring-capable runner. It rejects chat models (llama3, qwen3, gemma4 …), which weren't trained for this scoring; anything with a :cloud tag, because cloud models are refused on this endpoint; and MLX or Safetensors models imported on a Mac. Switch "model" to nimble, tev1 or tev1:0.8b. For open models beyond these, use the model's own server.

Did it work? The same request with "model": "tev1" returns answers.

4. Stay inside the schema limits

The limits are: 1 to 64 questions per request; a choice criteria object with 2 to 26 keys, none blank; a score criteria array of 2 to 26 strings, lowest first; an optional noul criteria with only the keys "false" and "true"; a type of exactly choice, noul or score, lowercase; and a non-empty state and instructions (a string, or a JSON object or array). A choice with 40 labels needs splitting: first route to a group, then choose within the group, in two requests — or use a runtime that supports up to 255 options (Decider, JevK5).

Did it work? Remove questions or options until the request succeeds, then add back one at a time.

5. Shorten the state Most common fix

Every question's rendered prompt must fit the loaded context with 2 tokens to spare. Nimble's must fit 8,192 tokens; Tev1's card recommends keeping inputs to roughly 2,000 tokens. First, send only the part of the document the question is about — extract the relevant section in code, which also improves accuracy because irrelevant detail distracts decision models. Second, split one long request into several with shorter states. Third, check usage.input_tokens on a request that works: it is the total across all questions, so divide by your question count to see the size of each prompt.

Did it work? The request succeeds and usage.input_tokens ÷ the number of questions is comfortably under the model’s limit.
Check what fits your hardware — check that the decision model fits, then build a valid request
Open the VRAM checker →

6. PowerShell builds a different JSON than you think

Two traps. First, in Windows PowerShell 5.1, curl is an alias for Invoke-WebRequest, which doesn't understand curl's flags — call curl.exe explicitly. Second, ConvertTo-Json stops at a depth of 2 by default and turns anything deeper into a type name. Your criteria object sits three levels below the root, so it is silently replaced; PowerShell 7 at least warns that the resulting JSON is truncated as serialization has exceeded the set depth of 2. Always pass -Depth 8, and use [ordered]@{} so options keep their order.

PowerShell
$body | ConvertTo-Json -Depth 8 | Set-Content -Encoding utf8 request.json
Get-Content request.json   # inspect before sending
Did it work? request.json shows your full criteria objects, not System.Collections.Hashtable.

7. Keep the body under 64 KiB

Large states are usually whole documents or long chat histories. Trim them in code, or split the work across requests. Remember that 64 KiB is the whole JSON body, including instructions and option descriptions.

Did it work? The file is under 65,536 bytes. In PowerShell the equivalent check is (Get-Item request.json).Length.
wc -c request.json

8. Allow your web page’s origin

A browser calling http://localhost:11434 from a page on another origin gets a console message saying the request "has been blocked by CORS policy". Ollama allows extra browser origins through the OLLAMA_ORIGINS environment variable, set on the server, then restart Ollama. How to set server environment variables differs by OS — on Linux it goes in the systemd unit, not your shell. Allow only the origins you need, not *.

Did it work? The browser request succeeds and the console shows no CORS error.

If none of this worked

Try the smallest model, tev1:0.8b, with a one-question request copied from the setup guide. If that works, the problem is your request, not your install. One more thing before you expose the endpoint: Ollama's API has no authentication. Binding it to 0.0.0.0 so another computer can send decision requests also lets anyone who can reach that port pull models and use your hardware. Prefer, in order, an SSH tunnel (ssh -N -L 11434:127.0.0.1:11434 user@your-box), a reverse proxy with authentication in front of Ollama, or a firewall rule restricted to one source address.

Include this when you report it

Also affects: Linux · Apple

Related

A model that fits most setups:
View model & requirements →

Frequently asked questions

Why does Ollama refuse long input instead of truncating it?

Because a decision about half a document is worse than no decision. Ollama's spec states that input is never truncated. If the prompt doesn't fit, you get a 400 and can decide yourself what to cut.

Can I raise the context limit to fit longer documents?

The limit comes from the loaded model’s context window. The decision models in Ollama were trained on short prompts — Nimble up to 8,192 tokens — so accuracy on much longer inputs is untested even where the window allows it. Extract the relevant part instead.

My request works but the answers look random. Is that an error?

Not one Ollama can detect. Check that your option descriptions are specific, that the state contains the evidence the question needs, and that you are not asking for arithmetic or date comparisons — known weak spots for decision models. The confidence and thresholds guide shows how to measure accuracy on your own examples.

Does a 500 mean my GPU is broken?

Usually not. It means loading, rendering or scoring failed — most often the model didn’t fit in memory. Check ollama ps and the server log, and try tev1:0.8b to confirm the endpoint itself works.