product image
product image
product image
product image
product image

The Local LLM Sizing Guide

$9/ month

The monthly reference that sizes every model for your card, context included.

Other calculators guess your context. This guide reads it from the lab’s config.

A popular rule of thumb budgets 31.8 GB of cache for Qwen3.8 27B at 64K of context. llama.cpp actually needs 4.2 GB. A guess that far off tells you a model will not run on a card that runs it perfectly well.

The Local LLM Sizing Guide is the monthly reference that gets it right.

Download the right model the first time. For every model: the best quantization for your card, and the context it holds with an F16, a Q8_0 or a Q4_0 cache, from 8 GB to 128 GB.

Buy the right card. What each memory size really runs, side by side.

Own a DGX Spark or a 128 GB Mac? You get your own chapter: every model that fits, and why the same 128 GB fits different models on each, with measured Spark speeds.

Get the speed you paid for. Up to twice the speed from a single setting, measured on real cards.

The month’s biggest releases, sized and tested. This issue: DeepSeek V4.1 Flash, 552 billion parameters on two DGX Sparks, plus MiMo-V2.6, Ternary Bonsai 2 27B on an 8 GB card, and EXL3, the format everyone is moving to.

Every model checked against its lab’s own config file. Every figure dated. Every correction counted in the next issue.

Issue 1, September 2026, out 30 September. The October issue is included.

Frequently asked questions