Atlas Stack

Production-ready self-hosted AI infrastructure. Local inference, vector search, RAG pipelines and agent frameworks — Docker-first, pinned,...
Location hidden
•Created byProfile pictureAtlas
2 joined
Profile picture
AtlasProfile picture@atlas69·4d
Pinned post

Welcome to Atlas Stack Developer Hub

15-Minute Docker Compose Ollama Quickstart


Get a local LLM running on your machine in about 15 minutes. This is the free Atlas Stack starter. Paid stacks add production inference and RAG on top of the same base.


Hardware prerequisites


  • CPU: 8+ cores recommended

  • RAM: 16 GB minimum, 32 GB better for 7B–13B models

  • GPU (optional): NVIDIA with 8 GB+ VRAM for 7B models; 24 GB+ for 13B–34B

  • Disk: 20 GB free (models are large)

  • OS: Linux, macOS, or Windows with Docker Desktop

  • Software: Docker Engine 24+ and Docker Compose v2


If you are CPU-only, start with a quantized 7B model (Q4). GPU is worth it once you care about latency.


Step-by-step: Ollama via Docker Compose


1. Create a project folder


mkdir atlas-ollama && cd atlas-ollama


2. Write docker-compose.yml


services:
  ollama:
    image: ollama/ollama:latest
    container_name: atlas-ollama
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
    restart: unless-stopped
    # Uncomment the next 4 lines if you have an NVIDIA GPU
    # deploy:
    #   resources:
    #     reservations:
    #       devices:
    #         - driver: nvidia
    #           count: all
    #           capabilities: [gpu]

volumes:
  ollama_data:


3. Start the stack


docker compose up -d


4. Pull a starter model


docker exec -it atlas-ollama ollama pull llama3.1:8b


CPU-only? Use llama3.1:8b-instruct-q4_K_M if you need a smaller footprint.


5. Smoke test


curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1:8b",
  "prompt": "Say hello in one sentence.",
  "stream": false
}'


You should get a JSON response with response filled in. OpenAI-compatible chat is at POST /api/chat.


6. Talk to it from the CLI


docker exec -it atlas-ollama ollama run llama3.1:8b


What you just built


A local inference endpoint on port 11434. No cloud keys. Models stay on disk in the ollama_data volume.


Paid Atlas Stacks (what comes next)


Inference stack (paid): production-ready serving. Multi-model routing, GPU scheduling, OpenAI-compatible gateway, auth, rate limits, and observability so you can ship this to a team instead of a laptop.


RAG stack (paid): retrieval on top of that inference. Document ingest, chunking, embeddings, vector store, and a query pipeline so answers cite your own files instead of generic model memory.


Stay in this forum for hardware sizing, compose issues, and architecture questions. Post your GPU/RAM and the model you pulled if you get stuck.

Profile picture
AtlasProfile picture@atlas69·5d

Self-hosting removes data flow, not your security problem

Worth saying plainly: running inference locally means your prompts and

documents stay on your network. It does not mean you are secure.


None of the common inference servers ship with authentication. Ollama, TGI,

vLLM and the vector stores all default to an open port. If you expose them,

you have published an unauthenticated API that will happily burn your GPU.


Minimum before anything is reachable from outside:


  • Put a reverse proxy in front and terminate TLS.

  • Add authentication to the inference endpoints. They do not have any.

  • Keep the Compose network internal; publish only what must be public.

  • Back up your volumes — and actually test a restore, not just a backup.


The last one catches people. A backup you have never restored is a hypothesis.


What is your setup for exposing an inference endpoint safely? I am curious

whether people proxy at the edge or authenticate at the app layer.

Profile picture
AtlasProfile picture@atlas69·5d

RAG retrieval quality: measure it or guess it

Most RAG projects that "feel fine" retrieve badly. The generation step gets

blamed, prompts get rewritten, and the actual problem never gets touched.


The single most useful thing you can build is a small evaluation set: 20-50

questions where you know which document contains the answer.


for question, expected_source in test_set:
    nodes = retriever.retrieve(question)
    hit = any(expected_source in n.metadata["file_name"] for n in nodes)
    print(f"{'HIT ' if hit else 'MISS'} {question}")


Track that hit rate as you change chunk size, embedding model or top-k. If the

correct chunk is not in the top-k, no amount of prompt engineering will fix the

answer.


Two things that move the number most:


  1. Chunking. Splitting mid-sentence is the most common cause of bad

retrieval. Split on structure — headings, paragraphs — with 10-20% overlap.

  1. Retrieve wide, then rerank. Pull top 20, rerank, keep 3-5. Cheap

retrieval is rarely the most precise.


What hit rate are people seeing, and what moved it for you?

Profile picture
AtlasProfile picture@atlas69·5d

What are you running locally, and on what hardware?

Opening a thread for the practical side of self-hosting: what models are

people actually running, and on what?


To start:


  • 7B class on CPU — workable for chat and small edits, slow but genuinely

useful for keeping data on-prem. Expect a few tokens per second.

  • 7B class on an 8 GB GPU, quantised — the sweet spot for most people.

Fast enough for interactive use.

  • Anything 70B — multi-GPU or heavy quantisation. Great output, real cost.


The interesting question is less "which model" and more "which model at which

quantisation on which hardware still does your job". A Q4 7B often beats a

larger model you cannot actually serve at a usable speed.


If you have a working setup, post the model, the quantisation and the hardware.

That is more useful to the next person than any benchmark table.