Interactive scatter plot and sortable table of every supported local LLM. Toggle the X and Y axes between VRAM, quality, tok/s, disk size, or context length to find the best model for your hardware.
| VRAM required (GB) | How much GPU memory the model needs at the selected quant |
| Disk size (GB) | Storage required for the GGUF file |
| Model size (params B) | Total parameter count in billions |
| Est. tok/s | Bandwidth-based throughput estimate on selected GPU |
| Quality index | 0–100 aggregate quality score from Artificial Analysis |
| Context length | Maximum context window in tokens |
| Best value | Highlights the Pareto frontier — highest quality per GB of VRAM |
| Fastest | Sorts and colours models by estimated tokens/sec on your GPU |
| Fits my hardware | Hides models that won't run on your selected accelerator |
Click any point or table row to add the model to the Compare tool.