denis-pplx27B~18 GB VRAM at Q4_K_M

AutoJev 27B — VRAM & /v1/systemone setup

Written by Jakub Rusinowski · Last updated

A 27B multimodal decision model on Qwen3.8-27B, the highest-scoring downloadable model on the Decision Index as of 28 Sep 2026; served by autojev-serve.

AutoJev 27B needs about 18 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Call AutoJev 27B

Served by autojev-serve (default port 8000). This model does not run in Ollama.

Rent a GPU by the hour
Need more VRAM than you own? Spin up a cloud GPU big enough for any model in minutes — pay for the hours you use instead of buying a card.

Affiliate links — we may earn a commission if you sign up, at no extra cost to you.

Vast.ai
Marketplace pricing — the cheapest per-hour rates for spot/interruptible GPUs.
Rent GPUs on Vast.ai →
RunPod
On-demand pods with a simple UI — good for a quick one-off inference or fine-tune.
Rent GPUs on RunPod →

Hardware fit

Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown.

QuantVRAMFit
BF16 (49 GiB)
Publisher build · 52.6 GB file
55.1 GBOffload
Q4_K_M
4.83 bpw · modelled quant
19.3 GBFits
Q6_K
6.56 bpw · modelled quant
25.3 GBOffload
Q8_0
8.50 bpw · modelled quant
32.1 GBOffload

Published file size: BF16 (49 GiB) 52.6 GB. A download size from the model publisher — not a VRAM requirement.

uv run autojev-serve → playground on :8000, POST /v1/systemone with optional base64 images; set AUTOJEV_API_KEY before exposing.

How it was scored

Two different suites on two different scales. Never compare the numbers across the two cards.

Decision Index 0.2.1 (snapshot 2026-09-28)

Decision Index 0.2.1
Balanced skill56.4
ECE (lower is better)
0.0178
Median compute
101.4 ms on 1x NVIDIA RTX PRO 6000 (96 GB)

Community-maintained; not affiliated with the model authors.

Source ↗

Decision models are not ranked on chat, creative or coding scores.

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source. Still unconfirmed: weightsGbByQuant.

Parameters
27.78B
Context window
Not published
Architecture
Fine-tune of Qwen/Qwen3.8-27B
Provider
denis-pplx
Licence
Apache-2.0
Specified at
Q4_K_M
System RAM
36 GB
Record updated
2026-09-30
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

AutoJev 27B — frequently asked questions

What is AutoJev 27B?

AutoJev 27B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.

How do I run AutoJev 27B locally?

AutoJev 27B does not run in Ollama. Served by autojev-serve (default port 8000). This model does not run in Ollama. See the setup guide for servers other than Ollama.

How much memory does AutoJev 27B need?

About 18 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.

How accurate is AutoJev 27B?

It has two separate published scores on two different suites — the author's own benchmark and the community Decision Index 0.2.1. They are not comparable with each other, and neither is a calibration guarantee.