Welcome to Atlas Stack Developer Hub
15-Minute Docker Compose Ollama Quickstart
Get a local LLM running on your machine in about 15 minutes. This is the free Atlas Stack starter. Paid stacks add production inference and RAG on top of the same base.
Hardware prerequisites
CPU: 8+ cores recommended
RAM: 16 GB minimum, 32 GB better for 7B–13B models
GPU (optional): NVIDIA with 8 GB+ VRAM for 7B models; 24 GB+ for 13B–34B
Disk: 20 GB free (models are large)
OS: Linux, macOS, or Windows with Docker Desktop
Software: Docker Engine 24+ and Docker Compose v2
If you are CPU-only, start with a quantized 7B model (Q4). GPU is worth it once you care about latency.
Step-by-step: Ollama via Docker Compose
1. Create a project folder
mkdir atlas-ollama && cd atlas-ollama2. Write docker-compose.yml
services:
ollama:
image: ollama/ollama:latest
container_name: atlas-ollama
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
restart: unless-stopped
# Uncomment the next 4 lines if you have an NVIDIA GPU
# deploy:
# resources:
# reservations:
# devices:
# - driver: nvidia
# count: all
# capabilities: [gpu]
volumes:
ollama_data:3. Start the stack
docker compose up -d4. Pull a starter model
docker exec -it atlas-ollama ollama pull llama3.1:8bCPU-only? Use llama3.1:8b-instruct-q4_K_M if you need a smaller footprint.
5. Smoke test
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1:8b",
"prompt": "Say hello in one sentence.",
"stream": false
}'You should get a JSON response with response filled in. OpenAI-compatible chat is at POST /api/chat.
6. Talk to it from the CLI
docker exec -it atlas-ollama ollama run llama3.1:8bWhat you just built
A local inference endpoint on port 11434. No cloud keys. Models stay on disk in the ollama_data volume.
Paid Atlas Stacks (what comes next)
Inference stack (paid): production-ready serving. Multi-model routing, GPU scheduling, OpenAI-compatible gateway, auth, rate limits, and observability so you can ship this to a team instead of a laptop.
RAG stack (paid): retrieval on top of that inference. Document ingest, chunking, embeddings, vector store, and a query pipeline so answers cite your own files instead of generic model memory.
Stay in this forum for hardware sizing, compose issues, and architecture questions. Post your GPU/RAM and the model you pulled if you get stuck.
