Cut LLM API costs by up to 60% with an OpenTelemetry-native semantic caching proxy and strict budget guardrails.