alibiserikbay4B~4 GB VRAM at Q4_K_M

JevK5 4B — VRAM & /v1/systemone setup

Written by Jakub Rusinowski · Last updated

A 4B decision model on Qwen3.5-4B-Base with a TypeSafe-style server; English only, and inputs over 16,384 tokens are refused rather than cut.

JevK5 4B needs about 4 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Call JevK5 4B

Served by jevk5-serve (default port 8090). This model does not run in Ollama.

Hardware fit

Weights plus overhead plus the KV cache at a 16,384-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown.

QuantVRAMFit
BF16
Publisher build · 9 GB file
11.7 GBFits
Q4_K_M
4.83 bpw · modelled quant
5.6 GBFits
Q6_K
6.56 bpw · modelled quant
6.6 GBFits
Q8_0
8.50 bpw · modelled quant
7.7 GBFits

Published file size: BF16 9 GB. A download size from the model publisher — not a VRAM requirement.

English only. Inputs over 16,384 tokens refused, not cut. Server: jevk5-serve --model alibiserikbay/JevK5 --port 8090. GGUF Q8_0/Q4_K_M in JevK5-GGUF.

How it was scored

Two different suites on two different scales. Never compare the numbers across the two cards.

Decision Index 0.2.1 (snapshot 2026-09-28)

Decision Index 0.2.1
Balanced skill38.81
ECE (lower is better)
0.0268
Median compute
22 ms on 1x NVIDIA RTX PRO 6000 (96 GB)

Community-maintained; not affiliated with the model authors.

Source ↗

Decision models are not ranked on chat, creative or coding scores.

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source. Still unconfirmed: weightsGbByQuant.

Parameters
4.66B
Context window
16K (prompt limit)
Architecture
Fine-tune of Qwen/Qwen3.5-4B-Base
Provider
alibiserikbay
Licence
Apache-2.0
Specified at
Q4_K_M
System RAM
8 GB
Record updated
2026-09-30
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

JevK5 4B — frequently asked questions

What is JevK5 4B?

JevK5 4B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.

How do I run JevK5 4B locally?

JevK5 4B does not run in Ollama. Served by jevk5-serve (default port 8090). This model does not run in Ollama. See the setup guide for servers other than Ollama.

How much memory does JevK5 4B need?

About 4 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.

How accurate is JevK5 4B?

It has two separate published scores on two different suites — the author's own benchmark and the community Decision Index 0.2.1. They are not comparable with each other, and neither is a calibration guarantee.