Every team that gets past "one engineer with an API key" ends up building the same thing: an AI gateway. One endpoint in front of every model, local or cloud. Apps get a virtual key, the gateway holds the real ones, and you finally get budgets, rate limits, routing and an audit trail in one place. It's the right pattern. It also creates the single most valuable secret store you run. On 24 March 2026, LiteLLM, the most widely used open-source gateway, had two PyPI releases (1.82.7 and 1.82.8) published with a credential-stealing payload. They were live for roughly 40 minutes. The attacker got the publishing credentials by compromising a security scanner in the project's own CI. Anyone whose pipeline pulled "latest" in that window shipped a stealer onto the box that holds every model key they own. (Sources: LiteLLM security update, Datadog Security Labs.) That isn't a reason to skip the gateway. It's a reason to build it like a vault instead of a convenience proxy. What a gateway should do (the five jobs): 1. Identity. One virtual key per app or team, never a shared provider key. 2. Budgets and quotas. Dollar caps and tokens per minute per key, enforced before the call, not discovered on the invoice. 3. Routing. Bulk traffic to your local model, the hard cases to a frontier API, with fallback rules you actually chose. 4. Policy. "Confidential" traffic maps to the local model only, with no cloud fallback. Ever. 5. Audit. Who called which model, how many tokens, what it cost, which policy decision fired. How to not turn it into your worst incident: - Pin the version AND the container image digest. No auto-upgrades on the box with every key. - Where the provider supports it, use workload identity instead of static keys (Bedrock through an IAM role, Vertex through a service account). A stolen short-lived token beats a stolen permanent key. - Egress allow-list on the gateway host: your model providers and nothing else. That one rule would have neutered a stealer's exfiltration.