Run Jev-style decision models locally with Ollama
Written by Jakub Rusinowski · Last updated
To run a decision model locally: install Ollama 0.35 or newer, run ollama pull nimble (or tev1 / tev1:0.8b), then POST a state and named questions to http://localhost:11434/v1/systemone. You get back one answer per question with a probability for every option. No API key, no network round-trip, no per-token bill.
This guide takes about 15 minutes, most of it the download.
What you need
| Minimum | |
|---|---|
| Ollama | 0.35.0 or later — the /v1/systemone endpoint does not exist before that |
| OS | macOS, Windows or Linux (anything Ollama supports) |
| Disk | 812 MB (tev1:0.8b) · 4.5 GB (tev1) · 9.5 GB (nimble) |
| Memory | Enough to hold the model — check your machine in the analyzer |
Decision requests are short (Nimble's prompts must fit 8,192 tokens), so the context cache stays small. The weight file is most of what has to fit.
Step 1 — Install or upgrade Ollama
Check what you have:
# bash / zsh / PowerShell — same command everywhere
ollama -vIf it prints anything below 0.35.0, upgrade:
- macOS and Windows: download the latest installer from ollama.com/download and run it over the old install. Quit Ollama from the menu bar or system tray first.
- Linux: re-run the official install script, which upgrades in place:
# bash
curl -fsSL https://ollama.com/install.sh | shRun ollama -v again and confirm 0.35.0 or higher before moving on. An old server still running in the background is the most common reason this step "didn't work" — see the troubleshooting page.
Step 2 — Pull a decision model
Pick one. You can pull all three and compare.
ollama pull nimble # 9B, Bespoke Labs, 9.5 GB — the most accurate of the three
ollama pull tev1 # 4B, Together AI, experimental, 4.5 GB
ollama pull tev1:0.8b # 0.8B, Together AI, experimental, 812 MBDon't use ollama run with these. Decision models are called through the API. They are not built to chat, and a chat session won't give you probabilities.
Step 3 — Send your first request
The request has three parts: which model, the state (the text you want judged — a string or a JSON object), and questions (named, typed questions about that state).
macOS / Linux (bash or zsh):
curl http://localhost:11434/v1/systemone \
-H 'Content-Type: application/json' \
-d '{
"model": "nimble",
"state": {"ticket": "I was charged twice. Please refund the extra payment."},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above"
}
},
"refund": {
"type": "noul",
"instructions": "Does the customer explicitly ask for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Routine", "Soon", "Urgent"]
}
}
}'Windows (PowerShell): build the body as objects and let PowerShell serialise it. Two details matter: use [ordered] so your options keep their order (ties go to the first option in request order), and pass -Depth 8, because ConvertTo-Json silently flattens anything deeper than two levels by default.
# PowerShell 5.1 or 7+
$body = [ordered]@{
model = "nimble"
state = [ordered]@{ ticket = "I was charged twice. Please refund the extra payment." }
questions = [ordered]@{
team = [ordered]@{
type = "choice"
instructions = "Which team should handle this ticket?"
criteria = [ordered]@{
billing = "Payments and refunds"
technical = "Bugs and integrations"
other = "None of the above"
}
}
refund = [ordered]@{
type = "noul"
instructions = "Does the customer explicitly ask for a refund?"
}
urgency = [ordered]@{
type = "score"
instructions = "How urgent is this ticket?"
criteria = @("Routine", "Soon", "Urgent")
}
}
} | ConvertTo-Json -Depth 8
Invoke-RestMethod -Uri "http://localhost:11434/v1/systemone" -Method Post `
-ContentType "application/json" -Body $body | ConvertTo-Json -Depth 8Step 4 — Read the answer
Ollama's published example response for exactly this request:
{
"model": "nimble",
"answers": {
"team": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003},
"confidence": 0.922
},
"refund": {"type": "noul", "noul": 0.997},
"urgency": {
"type": "score",
"score": 0.815,
"legend": {"0": "Routine", "1": "Soon", "2": "Urgent"},
"probabilities": {"0": 0.378, "1": 0.429, "2": 0.193},
"confidence": 0.046
}
},
"usage": {"input_tokens": 841, "output_tokens": 4}
}Your numbers will differ slightly. What each field means:
| Field | Meaning | Common mistake |
|---|---|---|
choice | The option key with the highest probability | Ignoring probabilities — a 0.51 / 0.49 split is not a confident answer |
noul | Probability that the statement is true, from 0 to 1 | Treating it as a boolean. 0.997 is a number; you pick the threshold |
score | Probability-weighted average of the level indices — here 0 to 2 | Reading it as 0–1. 0.815 means "between Routine and Soon", not "81.5%" |
confidence | How concentrated the distribution is: 0 = all options equal, near 1 = one option dominates | Reading it as "chance I'm right". Ollama's spec says it is not calibrated correctness |
usage.input_tokens | Full prompt length summed over every question | Thinking the state was sent once. Each question is scored against the whole state |
Look at urgency: the model thinks "Soon" (0.429) is marginally more likely than "Routine" (0.378), and its confidence of 0.046 says the distribution is nearly flat. That is the model telling you it doesn't really know — exactly the case your code should send to a human or a bigger model. The confidence and thresholds guide covers how.
Step 5 — Use it from code
Python, with TypeSafe's official SDK
The SDK written for TypeSafe's hosted Jev works against Ollama unchanged — point it at localhost:
# bash / zsh
pip install typesafe-sdk # or: uv add typesafe-sdk
export TYPESAFE_BASE_URL=http://localhost:11434
export TYPESAFE_API_KEY=ollama # any non-empty value; local Ollama doesn't check it
export TYPESAFE_DEFAULT_MODEL=nimble# PowerShell equivalent
pip install typesafe-sdk
$env:TYPESAFE_BASE_URL = "http://localhost:11434"
$env:TYPESAFE_API_KEY = "ollama"
$env:TYPESAFE_DEFAULT_MODEL = "nimble"from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
questions = {
"team": Choice(
instructions="Which team should handle this ticket?",
criteria={
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above",
},
),
"refund": Noul(instructions="Does the customer explicitly ask for a refund?"),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["Routine", "Soon", "Urgent"],
),
}
with TypeSafeClient(timeout=120) as client: # generous timeout: the first call loads the model
result = client.system_one(
state={"ticket": "I was charged twice. Please refund the extra payment."},
questions=questions,
)
print(result.choices["team"].choice) # billing
print(result.nouls["refund"].noul) # ~0.997
print(result.scores["urgency"].score) # ~0.8The same code talks to TypeSafe's hosted Jev if you change the three environment variables — handy for comparing local and hosted answers on your own data.
JavaScript / TypeScript, with plain fetch
// Node 18+ or any modern browser context that can reach localhost
const res = await fetch("http://localhost:11434/v1/systemone", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "tev1",
state: "Our checkout has returned 500 errors since 9am.",
questions: {
label: {
type: "choice",
instructions: "Which label fits this ticket?",
criteria: {
billing: "Payments and refunds",
bug: "Software errors",
account: "Login and account access",
},
},
},
}),
});
if (!res.ok) throw new Error(`${res.status}: ${(await res.json()).error}`);
const { answers } = await res.json();
console.log(answers.label.choice, answers.label.probabilities);Calling from a web page on another origin also needs OLLAMA_ORIGINS set on the server — see troubleshooting.
Step 6 — Check it's running on your GPU
Send one request, then:
ollama psThe Processor column should read 100% GPU on an NVIDIA/AMD card or Apple Silicon. 100% CPU works but is slower; a split like 40%/60% CPU/GPU means the model didn't fit and spilled. If you expected GPU and got CPU, the fixes are in our Windows, Linux and Apple troubleshooting hubs.
To keep the model loaded between bursts of requests, add "keep_alive": "30m" to the request body (or -1 to keep it loaded until Ollama restarts). The default unloads after 5 minutes idle, and the next request pays the load time again.
The limits you'll hit first
| Limit | Value (Ollama 0.35) | What happens |
|---|---|---|
| Questions per request | 1–64 | 400 error above 64 |
Options per choice / levels per score | 2–26 | 400 error outside the range |
| Request body | 64 KiB | 413 error |
| Prompt length | Must fit the loaded context with 2 tokens spare; never truncated | 400 error — shorten the state |
| Model type | Local, System One–trained, GGUF | Chat, cloud or MLX models get a 400 |
| Output | One JSON response | No streaming, images, tools or sampling options |
Tev1's model card recommends keeping inputs to roughly 2,000 tokens. Nimble's prompts must fit 8,192 tokens.
Which of the three should you use?
tev1:0.8b | tev1 | nimble | |
|---|---|---|---|
| Maker | Together AI | Together AI | Bespoke Labs |
| Size | 812 MB | 4.5 GB | 9.5 GB |
| Bespoke public benchmark (13 datasets) | 63.5% | 73.3% | 75.7% |
| Decision Index 0.2.1 | 12.9 | 29.2 | 39.6 (v2 repo) |
| Status | Experimental | Experimental | Released |
| Good for | CPU boxes, smoke tests, edge | Laptops, high volume | Default choice when it fits |
For reference, TypeSafe's hosted Jev 1.13 scores 76.0% on the same Bespoke benchmark and 57.9 on the Decision Index. The gap between the two columns is the subject of the explainer post.
If none of the three is accurate enough on your data, the open reproductions in the next guide go further — at the cost of running their own server.
Frequently asked questions
Does this cost anything?
No. Local requests don't need an API key and nothing is billed. You pay in electricity and in the memory the model occupies while loaded.
Can I run a decision model and a chat model at the same time?
Yes, if both fit. Ollama keeps several models loaded at once (by default up to three per GPU) and unloads the least recently used one when memory runs out. Check with ollama ps.
Is my data sent anywhere?
No. Requests to localhost:11434 stay on your machine. That is one of the main reasons to run decision models locally for moderation or triage of private data.
Why is the first request slow?
The model is loaded into memory on first use. Later requests are fast until the model unloads after the keep-alive period. Set keep_alive to avoid the reload.
Can I call it from another computer on my network?
Yes, but Ollama's API has no authentication. Read the network-exposure section of the troubleshooting page before you open it up. --- Verified against the Ollama blog (29 Sep 2026) and the Ollama API reference and decision guide, read 30 September 2026. Ollama's decision-model support is days old; limits may change in later releases.
Verified against: Ollama blog 29 Sep 2026; docs.ollama.com/api/systemone and /capabilities/decision, read 30 Sep 2026.