If your webhook handler writes to the DB in the HTTP request, you will drop events
Stripe, GitHub, Shopify — they retry. Your process does not.
The typical FastAPI webhook looks like this:
@app.post("/webhooks")
async def handle(payload):
await db.insert(payload)
return {"ok": true}That works until it doesn't. Deploy in the middle of a burst. Worker OOM. DB lock. The provider times out, you return 500, they retry, and now you have duplicates and a gap.
What actually holds:
Ack first. Validate the signature, enqueue, return 200. The HTTP process should not own the side effect.
Idempotency key on every job. Provider event ID as the key. Unique constraint in the DB, not a "check then insert" in Python.
A dead-letter queue you actually look at. Celery retries are not a strategy if failed jobs vanish into Redis.
Crash the worker, not the request. If Redis is down, fail the enqueue loudly. Don't half-write.
I got tired of rebuilding this for every client, so HookEngine Pro is the FastAPI + Redis + Celery kit I now start from: thin receiver, durable queue, worker with retries and DLQ already wired.
If your current handler still does the write inside the request, that's the first line to delete.
