
Run open models on your own machine. Docker-first, no API keys.
A working, pinned Docker Compose stack - not a list of links. Bring it up,
verify it, and reach your first completion.
WHAT YOU GET
WHY THESE THREE
Ollama gives you an OpenAI-compatible endpoint with one-command model pulls.
llama.cpp is the low-level GGUF engine when you want control over quantisation
and context. TGI adds continuous batching and token streaming for production
serving. Together they cover local development through to a real endpoint.
ALL OPEN SOURCE, ALL PERMISSIVE
Every component is MIT or Apache-2.0. Full attribution with pinned commits is
included. You are buying the integration layer and operational knowledge - the
source is freely available upstream, and the bundle tells you where and how.
REQUIREMENTS
Docker Engine 24+ with Compose v2. 16 GB RAM for a 7B model on CPU, or an
NVIDIA GPU with the container toolkit for real throughput.
Instant delivery. One-time purchase, free updates.