Run Jev-style decision models locally with Ollama

Written by Jakub Rusinowski · Last updated

To run a decision model locally: install Ollama 0.35 or newer, run ollama pull nimble (or tev1 / tev1:0.8b), then POST a state and named questions to http://localhost:11434/v1/systemone. You get back one answer per question with a probability for every option. No API key, no network round-trip, no per-token bill.

This guide takes about 15 minutes, most of it the download.

What you need

Minimum
Ollama0.35.0 or later — the /v1/systemone endpoint does not exist before that
OSmacOS, Windows or Linux (anything Ollama supports)
Disk812 MB (tev1:0.8b) · 4.5 GB (tev1) · 9.5 GB (nimble)
MemoryEnough to hold the model — check your machine in the analyzer

Decision requests are short (Nimble's prompts must fit 8,192 tokens), so the context cache stays small. The weight file is most of what has to fit.

Step 1 — Install or upgrade Ollama

Check what you have:

bash
# bash / zsh / PowerShell — same command everywhere
ollama -v

If it prints anything below 0.35.0, upgrade:

  • macOS and Windows: download the latest installer from ollama.com/download and run it over the old install. Quit Ollama from the menu bar or system tray first.
  • Linux: re-run the official install script, which upgrades in place:
bash
# bash
curl -fsSL https://ollama.com/install.sh | sh

Run ollama -v again and confirm 0.35.0 or higher before moving on. An old server still running in the background is the most common reason this step "didn't work" — see the troubleshooting page.

Step 2 — Pull a decision model

Pick one. You can pull all three and compare.

bash
ollama pull nimble      # 9B, Bespoke Labs, 9.5 GB — the most accurate of the three
ollama pull tev1        # 4B, Together AI, experimental, 4.5 GB
ollama pull tev1:0.8b   # 0.8B, Together AI, experimental, 812 MB
Don't use ollama run with these. Decision models are called through the API. They are not built to chat, and a chat session won't give you probabilities.

Step 3 — Send your first request

The request has three parts: which model, the state (the text you want judged — a string or a JSON object), and questions (named, typed questions about that state).

macOS / Linux (bash or zsh):

bash
curl http://localhost:11434/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "nimble",
    "state": {"ticket": "I was charged twice. Please refund the extra payment."},
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this ticket?",
        "criteria": {
          "billing": "Payments and refunds",
          "technical": "Bugs and integrations",
          "other": "None of the above"
        }
      },
      "refund": {
        "type": "noul",
        "instructions": "Does the customer explicitly ask for a refund?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this ticket?",
        "criteria": ["Routine", "Soon", "Urgent"]
      }
    }
  }'

Windows (PowerShell): build the body as objects and let PowerShell serialise it. Two details matter: use [ordered] so your options keep their order (ties go to the first option in request order), and pass -Depth 8, because ConvertTo-Json silently flattens anything deeper than two levels by default.

powershell
# PowerShell 5.1 or 7+
$body = [ordered]@{
  model = "nimble"
  state = [ordered]@{ ticket = "I was charged twice. Please refund the extra payment." }
  questions = [ordered]@{
    team = [ordered]@{
      type = "choice"
      instructions = "Which team should handle this ticket?"
      criteria = [ordered]@{
        billing   = "Payments and refunds"
        technical = "Bugs and integrations"
        other     = "None of the above"
      }
    }
    refund = [ordered]@{
      type = "noul"
      instructions = "Does the customer explicitly ask for a refund?"
    }
    urgency = [ordered]@{
      type = "score"
      instructions = "How urgent is this ticket?"
      criteria = @("Routine", "Soon", "Urgent")
    }
  }
} | ConvertTo-Json -Depth 8

Invoke-RestMethod -Uri "http://localhost:11434/v1/systemone" -Method Post `
  -ContentType "application/json" -Body $body | ConvertTo-Json -Depth 8

Step 4 — Read the answer

Ollama's published example response for exactly this request:

json
{
  "model": "nimble",
  "answers": {
    "team": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003},
      "confidence": 0.922
    },
    "refund": {"type": "noul", "noul": 0.997},
    "urgency": {
      "type": "score",
      "score": 0.815,
      "legend": {"0": "Routine", "1": "Soon", "2": "Urgent"},
      "probabilities": {"0": 0.378, "1": 0.429, "2": 0.193},
      "confidence": 0.046
    }
  },
  "usage": {"input_tokens": 841, "output_tokens": 4}
}

Your numbers will differ slightly. What each field means:

FieldMeaningCommon mistake
choiceThe option key with the highest probabilityIgnoring probabilities — a 0.51 / 0.49 split is not a confident answer
noulProbability that the statement is true, from 0 to 1Treating it as a boolean. 0.997 is a number; you pick the threshold
scoreProbability-weighted average of the level indices — here 0 to 2Reading it as 0–1. 0.815 means "between Routine and Soon", not "81.5%"
confidenceHow concentrated the distribution is: 0 = all options equal, near 1 = one option dominatesReading it as "chance I'm right". Ollama's spec says it is not calibrated correctness
usage.input_tokensFull prompt length summed over every questionThinking the state was sent once. Each question is scored against the whole state

Look at urgency: the model thinks "Soon" (0.429) is marginally more likely than "Routine" (0.378), and its confidence of 0.046 says the distribution is nearly flat. That is the model telling you it doesn't really know — exactly the case your code should send to a human or a bigger model. The confidence and thresholds guide covers how.

Step 5 — Use it from code

Python, with TypeSafe's official SDK

The SDK written for TypeSafe's hosted Jev works against Ollama unchanged — point it at localhost:

bash
# bash / zsh
pip install typesafe-sdk          # or: uv add typesafe-sdk
export TYPESAFE_BASE_URL=http://localhost:11434
export TYPESAFE_API_KEY=ollama    # any non-empty value; local Ollama doesn't check it
export TYPESAFE_DEFAULT_MODEL=nimble
powershell
# PowerShell equivalent
pip install typesafe-sdk
$env:TYPESAFE_BASE_URL = "http://localhost:11434"
$env:TYPESAFE_API_KEY = "ollama"
$env:TYPESAFE_DEFAULT_MODEL = "nimble"
python
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

questions = {
    "team": Choice(
        instructions="Which team should handle this ticket?",
        criteria={
            "billing": "Payments and refunds",
            "technical": "Bugs and integrations",
            "other": "None of the above",
        },
    ),
    "refund": Noul(instructions="Does the customer explicitly ask for a refund?"),
    "urgency": Score(
        instructions="How urgent is this ticket?",
        criteria=["Routine", "Soon", "Urgent"],
    ),
}

with TypeSafeClient(timeout=120) as client:  # generous timeout: the first call loads the model
    result = client.system_one(
        state={"ticket": "I was charged twice. Please refund the extra payment."},
        questions=questions,
    )

print(result.choices["team"].choice)   # billing
print(result.nouls["refund"].noul)     # ~0.997
print(result.scores["urgency"].score)  # ~0.8

The same code talks to TypeSafe's hosted Jev if you change the three environment variables — handy for comparing local and hosted answers on your own data.

JavaScript / TypeScript, with plain fetch

javascript
// Node 18+ or any modern browser context that can reach localhost
const res = await fetch("http://localhost:11434/v1/systemone", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    model: "tev1",
    state: "Our checkout has returned 500 errors since 9am.",
    questions: {
      label: {
        type: "choice",
        instructions: "Which label fits this ticket?",
        criteria: {
          billing: "Payments and refunds",
          bug: "Software errors",
          account: "Login and account access",
        },
      },
    },
  }),
});
if (!res.ok) throw new Error(`${res.status}: ${(await res.json()).error}`);
const { answers } = await res.json();
console.log(answers.label.choice, answers.label.probabilities);

Calling from a web page on another origin also needs OLLAMA_ORIGINS set on the server — see troubleshooting.

Step 6 — Check it's running on your GPU

Send one request, then:

bash
ollama ps

The Processor column should read 100% GPU on an NVIDIA/AMD card or Apple Silicon. 100% CPU works but is slower; a split like 40%/60% CPU/GPU means the model didn't fit and spilled. If you expected GPU and got CPU, the fixes are in our Windows, Linux and Apple troubleshooting hubs.

To keep the model loaded between bursts of requests, add "keep_alive": "30m" to the request body (or -1 to keep it loaded until Ollama restarts). The default unloads after 5 minutes idle, and the next request pays the load time again.

The limits you'll hit first

LimitValue (Ollama 0.35)What happens
Questions per request1–64400 error above 64
Options per choice / levels per score2–26400 error outside the range
Request body64 KiB413 error
Prompt lengthMust fit the loaded context with 2 tokens spare; never truncated400 error — shorten the state
Model typeLocal, System One–trained, GGUFChat, cloud or MLX models get a 400
OutputOne JSON responseNo streaming, images, tools or sampling options

Tev1's model card recommends keeping inputs to roughly 2,000 tokens. Nimble's prompts must fit 8,192 tokens.

Which of the three should you use?

tev1:0.8btev1nimble
MakerTogether AITogether AIBespoke Labs
Size812 MB4.5 GB9.5 GB
Bespoke public benchmark (13 datasets)63.5%73.3%75.7%
Decision Index 0.2.112.929.239.6 (v2 repo)
StatusExperimentalExperimentalReleased
Good forCPU boxes, smoke tests, edgeLaptops, high volumeDefault choice when it fits

For reference, TypeSafe's hosted Jev 1.13 scores 76.0% on the same Bespoke benchmark and 57.9 on the Decision Index. The gap between the two columns is the subject of the explainer post.

If none of the three is accurate enough on your data, the open reproductions in the next guide go further — at the cost of running their own server.

Frequently asked questions

Does this cost anything?

No. Local requests don't need an API key and nothing is billed. You pay in electricity and in the memory the model occupies while loaded.

Can I run a decision model and a chat model at the same time?

Yes, if both fit. Ollama keeps several models loaded at once (by default up to three per GPU) and unloads the least recently used one when memory runs out. Check with ollama ps.

Is my data sent anywhere?

No. Requests to localhost:11434 stay on your machine. That is one of the main reasons to run decision models locally for moderation or triage of private data.

Why is the first request slow?

The model is loaded into memory on first use. Later requests are fast until the model unloads after the keep-alive period. Set keep_alive to avoid the reload.

Can I call it from another computer on my network?

Yes, but Ollama's API has no authentication. Read the network-exposure section of the troubleshooting page before you open it up. --- Verified against the Ollama blog (29 Sep 2026) and the Ollama API reference and decision guide, read 30 September 2026. Ollama's decision-model support is days old; limits may change in later releases.

Verified against: Ollama blog 29 Sep 2026; docs.ollama.com/api/systemone and /capabilities/decision, read 30 Sep 2026.

Keep going