Together AI4B~4 GB VRAM at Q4_K_M

Tev1 4B — VRAM & /v1/systemone setup

Written by Jakub Rusinowski · Last updated

Together AI's experimental 4B decision model, a supervised fine-tune of Qwen3.5-4B. Available in Ollama as tev1; keep inputs around 2,000 tokens.

Tev1 4B needs about 4 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Call Tev1 4B

Pull it with Ollama 0.35 or newer, then POST to /v1/systemone. Do not use ollama chat commands — this is not a chat model.

ollama pull tev1
curl http://localhost:11434/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "tev1",
    "state": {
      "ticket": "I was charged twice. Please refund the extra payment."
    },
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this ticket?",
        "criteria": {
          "billing": "Payments and refunds",
          "technical": "Bugs and integrations",
          "other": "None of the above"
        }
      }
    }
  }'

Build a request for this model Decision models guide

Other decision models: Nimble 9B · Tev1 0.8B · Winnow 12B · Winnow E4B · Decider 2B · Decider 4B · Decider 35B-A3B (NVFP4) · JevK5 4B · Intern-Decision 4B · AutoJev 27B · Laya

Hardware fit

Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown.

QuantVRAMFit
Ollama build
Publisher build · 4.5 GB file
6.3 GBFits
Q4_K_M
4.83 bpw · modelled quant
4.6 GBFits
Q6_K
6.56 bpw · modelled quant
5.6 GBFits
Q8_0
8.50 bpw · modelled quant
6.7 GBFits

Published file size: Ollama build 4.5 GB. A download size from the model publisher — not a VRAM requirement.

Model card: a Jev-inspired experiment, not a non-autoregressive Jev runtime; keeps Qwen's LM head. Ollama card: keep inputs ~2,000 tokens; not for high-stakes decisions without review.

How it was scored

Two different suites on two different scales. Never compare the numbers across the two cards.

Author's public benchmark

Bespoke Labs public benchmarks (13 datasets, 3,880 decisions)
Accuracy73.3%
Source ↗

Decision Index 0.2.1 (snapshot 2026-09-28)

Decision Index 0.2.1
Balanced skill29.24
ECE (lower is better)
0.104
Median compute
35.8 ms on 1x NVIDIA RTX PRO 6000 (96 GB)

Community-maintained; not affiliated with the model authors.

Source ↗

Decision models are not ranked on chat, creative or coding scores.

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source. Still unconfirmed: license.

Parameters
4.66B
Context window
Not published
Architecture
Fine-tune of Qwen/Qwen3.5-4B
Provider
Together AI
Licence
Not stated on the model card
Specified at
Q4_K_M
System RAM
8 GB
Record updated
2026-09-30

Other Tev1 (experimental) sizes

Tev1 4B — frequently asked questions

What is Tev1 4B?

Tev1 4B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.

How do I run Tev1 4B locally?

Install Ollama 0.35 or newer, run `ollama pull tev1`, then POST your state and questions to http://localhost:11434/v1/systemone. It is called through the API, not a chat session.

How much memory does Tev1 4B need?

About 4 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.

How accurate is Tev1 4B?

It has two separate published scores on two different suites — the author's own benchmark and the community Decision Index 0.2.1. They are not comparable with each other, and neither is a calibration guarantee.