Home / Guides / Troubleshooting / Windows
Ollama /v1/systemone errors: 400, 404, 413 explained
WindowsLinuxmacOS
Written by Jakub Rusinowski · Last updated September 30, 2026
AI educator & workshop leader on local LLM deployment
The error
{"error": "request body must not exceed 64 KiB"}
Which one is it?
| If you see | The cause is | Go to |
|---|---|---|
| 404 on every request, even with a correct model name — or a 404 that names the model | Ollama is older than 0.35 (or an old server is still running), or the model was never pulled or its tag is misspelled | Fix 1, then Fix 2 |
400 with a chat model, a :cloud tag or an MLX model | That model isn't a local System One model | Fix 3: use a System One model |
| 400 after adding more options or questions, or on long input only | A limit: 1–64 questions and 2–26 options, or the rendered prompt doesn't fit the loaded context | Fix 4 or Fix 5 |
| 400 from PowerShell only; the same JSON works in bash | ConvertTo-Json depth, or curl being an alias for Invoke-WebRequest | Fix 6: PowerShell builds a different JSON |
| 413 | The request body is over 64 KiB | Fix 7: keep the body under 64 KiB |
| Works in curl, fails from a web page | CORS: the page’s origin is not allowed | Fix 8: allow your page’s origin |
When you see it
Ollama's decision-model endpoint, POST /v1/systemone, arrived in Ollama 0.35. It is strict by design: it never truncates your input, never guesses at a malformed schema and never falls back to a chat model. That strictness means most failures come back as a clear HTTP status — if you know what each one means. Only the size limit has a documented error body; everything else is a status code plus a JSON error field whose wording can change between releases, so read the status first: 400 is an invalid request, an unsupported model or a prompt that is too long; 404 is a model that isn't pulled or an endpoint that doesn't exist; 413 is a body over 64 KiB; 500 is a model that failed to load, render or score.
- Following a tutorial written for TypeSafe's hosted Jev and pointing it at localhost.
- After upgrading Ollama on Windows or macOS while the old server kept running in the tray.
- Moving from a short test string to real documents — contracts, email threads, logs.
- Building the request in PowerShell instead of pasting JSON into curl.
- Trying a chat model you already have (
llama3,qwen3) instead of a decision model.
What's actually going on
A decision model doesn't generate an answer; the runner scores each allowed option against the full prompt. For that to work, Ollama needs three things to be exactly right: a model whose weights were trained for this scoring (today nimble, tev1 and tev1:0.8b), a schema it can render into a prompt, and a prompt that fits in the model's loaded context window with two token positions spare for the scoring itself. Each question is rendered and scored separately against the whole state, so a long state is paid for once per question. Ollama checks all of this before loading anything and refuses with 400 rather than silently cutting your text — a truncated contract would give you a confident answer about half a contract.
How to fix it
1. Upgrade Ollama and restart the server
You need 0.35.0 or later. On macOS and Windows, quit Ollama from the menu bar or system tray before installing the new version — otherwise the old server keeps answering on port 11434 and the new endpoint seems missing. On Linux, the commands below upgrade it and restart the service.
ollama -v
# Linux upgrade
curl -fsSL https://ollama.com/install.sh | sh
sudo systemctl restart ollamaollama -v shows 0.35.0 or later and the same request no longer returns 404. If ollama -v warns that the client and server versions differ, an old server is still running — quit it fully and start Ollama again.2. Pull the model
Model names must match exactly — tev1:0.8b, not tev1-0.8b or tev1:0.8B.
ollama pull nimble
ollama listollama list shows the model, and the request returns answers.3. Use a System One model
The endpoint accepts only local models trained for System One, packaged as GGUF with a scoring-capable runner. It rejects chat models (llama3, qwen3, gemma4 …), which weren't trained for this scoring; anything with a :cloud tag, because cloud models are refused on this endpoint; and MLX or Safetensors models imported on a Mac. Switch "model" to nimble, tev1 or tev1:0.8b. For open models beyond these, use the model's own server.
"model": "tev1" returns answers.4. Stay inside the schema limits
The limits are: 1 to 64 questions per request; a choice criteria object with 2 to 26 keys, none blank; a score criteria array of 2 to 26 strings, lowest first; an optional noul criteria with only the keys "false" and "true"; a type of exactly choice, noul or score, lowercase; and a non-empty state and instructions (a string, or a JSON object or array). A choice with 40 labels needs splitting: first route to a group, then choose within the group, in two requests — or use a runtime that supports up to 255 options (Decider, JevK5).
5. Shorten the state Most common fix
Every question's rendered prompt must fit the loaded context with 2 tokens to spare. Nimble's must fit 8,192 tokens; Tev1's card recommends keeping inputs to roughly 2,000 tokens. First, send only the part of the document the question is about — extract the relevant section in code, which also improves accuracy because irrelevant detail distracts decision models. Second, split one long request into several with shorter states. Third, check usage.input_tokens on a request that works: it is the total across all questions, so divide by your question count to see the size of each prompt.
usage.input_tokens ÷ the number of questions is comfortably under the model’s limit.Open the VRAM checker →
6. PowerShell builds a different JSON than you think
Two traps. First, in Windows PowerShell 5.1, curl is an alias for Invoke-WebRequest, which doesn't understand curl's flags — call curl.exe explicitly. Second, ConvertTo-Json stops at a depth of 2 by default and turns anything deeper into a type name. Your criteria object sits three levels below the root, so it is silently replaced; PowerShell 7 at least warns that the resulting JSON is truncated as serialization has exceeded the set depth of 2. Always pass -Depth 8, and use [ordered]@{} so options keep their order.
$body | ConvertTo-Json -Depth 8 | Set-Content -Encoding utf8 request.json
Get-Content request.json # inspect before sendingrequest.json shows your full criteria objects, not System.Collections.Hashtable.7. Keep the body under 64 KiB
Large states are usually whole documents or long chat histories. Trim them in code, or split the work across requests. Remember that 64 KiB is the whole JSON body, including instructions and option descriptions.
(Get-Item request.json).Length.wc -c request.json8. Allow your web page’s origin
A browser calling http://localhost:11434 from a page on another origin gets a console message saying the request "has been blocked by CORS policy". Ollama allows extra browser origins through the OLLAMA_ORIGINS environment variable, set on the server, then restart Ollama. How to set server environment variables differs by OS — on Linux it goes in the systemd unit, not your shell. Allow only the origins you need, not *.
If none of this worked
Try the smallest model, tev1:0.8b, with a one-question request copied from the setup guide. If that works, the problem is your request, not your install. One more thing before you expose the endpoint: Ollama's API has no authentication. Binding it to 0.0.0.0 so another computer can send decision requests also lets anyone who can reach that port pull models and use your hardware. Prefer, in order, an SSH tunnel (ssh -N -L 11434:127.0.0.1:11434 user@your-box), a reverse proxy with authentication in front of Ollama, or a firewall rule restricted to one source address.
- Decision models setup with Ollama (the one-question request)
- Connection refused
- Ollama not reachable from another machine (Windows firewall)
- Headless Linux server: disk and remote access
- Report it upstream
Include this when you report it
ollama -v, your OS, GPU model and driver version- The model tag and the request body with private data removed
- The full error response
- The server log: macOS
~/.ollama/logs/server.log, Windows$env:LOCALAPPDATA\Ollama, Linuxjournalctl -u ollama --no-pager --pager-end
Related
- Decision models: the hub
- Confidence, thresholds and calibration
- Open Jev reproductions beyond Ollama
- Decision request builder
- Check whether the model fits your hardware
Frequently asked questions
Why does Ollama refuse long input instead of truncating it?
Because a decision about half a document is worse than no decision. Ollama's spec states that input is never truncated. If the prompt doesn't fit, you get a 400 and can decide yourself what to cut.
Can I raise the context limit to fit longer documents?
The limit comes from the loaded model’s context window. The decision models in Ollama were trained on short prompts — Nimble up to 8,192 tokens — so accuracy on much longer inputs is untested even where the window allows it. Extract the relevant part instead.
My request works but the answers look random. Is that an error?
Not one Ollama can detect. Check that your option descriptions are specific, that the state contains the evidence the question needs, and that you are not asking for arithmetic or date comparisons — known weak spots for decision models. The confidence and thresholds guide shows how to measure accuracy on your own examples.
Does a 500 mean my GPU is broken?
Usually not. It means loading, rendering or scoring failed — most often the model didn’t fit in memory. Check ollama ps and the server log, and try tev1:0.8b to confirm the endpoint itself works.