User
Write something
Azure Exchange Clinic is happening in 42 days
Pinned
Start here — how to get the most out of this community
Welcome. If you just joined, here's the 3-minute version of how to use this place. 1. Classroom → pick one lesson. Every lesson is a real build (Terraform drift agent, K8s cost optimizer, AWS anomaly detector) with working code, not slides. Start with whichever problem you actually have this week. 2. Comment on this post with what you're working on right now — cost overruns, drift, K8s chaos, IAM sprawl, anything. I read every one and it shapes what gets built next. 3. Feed gets a new real-world project weekly. Free tier = tips, tutorials, Q&A. Paid tier (Classroom) = full labs, scripts, and templates you keep. This only works if it's not a monologue — so tell me what's breaking in your stack. What are you fighting with right now? Update, Sept 2026: prefer video? Most posts now have a 3-minute explainer. Full playlist: https://www.youtube.com/playlist?list=PLOKaUp_A8PQo Not on Premium yet? The Kubernetes Cost Optimization course in the Classroom is open to every member. Start there.
2
0
BigQuery autoscaling bills what it grants, not what you use. Short queries hold the meter open.
Here's a BigQuery billing rule that surprises nearly everyone who reads it for the first time. When an editions reservation autoscales, Google bills the slots it scaled to, not the slots your query used. Three rules from Google's own autoscaling docs explain the gap: 1. Autoscaling rounds up to a multiple of 50. If a query needs 60 slots, you're billed for 100. 2. You pay for scaled slots even if the job that triggered the scale-up fails. 3. By default, every scale-up is held for at least 60 seconds, and a new peak restarts that window. Rule 3 does the damage. One big ETL job per hour barely notices it. A BI dashboard that fires a 3-second query every 20 seconds never lets the window close, so the reservation sits at 100 slots all business day. Illustrative math (invented org, US Enterprise pay-as-you-go $0.06/slot-hour, list price as of 3 Oct 2026, 220 business hours a month): - Slots the queries actually used: 60 slots x 3 s every 20 s = about 9 slots on average = about $119/month - What the default autoscaler bills: 100 slots held continuously = about $1,320/month - Same workload with fluid scaling on: 100 slots for ~3 s of every 20 = about $198/month That's an 11x gap between the work done and the bill, caused entirely by query shape. The fix list, cheapest first: - Measure before you touch anything. Compare autoscale slot-minutes from INFORMATION_SCHEMA.RESERVATIONS_TIMELINE with the slot time your jobs actually consumed from JOBS_TIMELINE. Google says outright that the JOBS view will not match your bill. Use the billing export as the source of truth. - Turn on BigQuery fluid scaling for the reservation. It's opt-in. It removes the one-minute minimum and keeps per-second billing. It doesn't fix the 50-slot rounding. - Put spiky BI traffic in a different reservation from steady ETL. Idle slots are only shared within the same edition and region, so plan the split on purpose. - Give steady load a baseline and buy commitments for it ($0.048 one year, $0.036 three years on Enterprise as of today). Let autoscale handle real bursts only.
0
0
"Read-only" is not least privilege for AI agents. The credential is the blast radius.
You gave the agent a read-only database account. Security signed off. You're still exposed. Here's why. In July 2025, General Analysis showed an MCP-connected coding assistant reading a support ticket that contained planted instructions. The assistant was running with Supabase's service_role, which bypasses row-level security. It read a table of integration tokens and wrote them into the ticket thread, where the attacker could see them. No permission was violated. The agent did exactly what its credential allowed. Read-only would have blocked that particular write. It would not have blocked the read. And an agent has plenty of other ways to move data once it has it: the chat reply, a Slack tool, a URL fetch, a ticket comment. General Analysis make the same point in their own mitigation notes. Prompt injection is not solved, and I don't expect it to be this year. So stop treating the prompt as the security boundary. The credential is the boundary. Design that instead. The model I'd use for on-prem data: 1. The agent holds no credentials. Tools hold them, next to the data. The agent gets a token that is only valid for the tool server. The current MCP authorization spec (2026-07-28) makes this explicit: servers must reject tokens not issued for them and must not pass a client's token through to upstream APIs. 2. Identity is the human, not "the agent". One shared agent service account means every user gets the union of everyone's access. Run queries with the requesting user's entitlements (OAuth token exchange, or Kerberos constrained delegation if you live in AD). 3. Scope at the data layer. Curated views, not tables. No secret or token columns, ever. Statement timeout, row cap, and short-lived credentials from something like Vault's database secrets engine instead of a password in a .env file. 4. No raw SQL tool on anything that matters. If the agent can run arbitrary SQL, it can SET the session variable your row-level security relies on. Expose parameterised tools, not a database prompt.
0
0
Azure reservation exchanges end Feb 1, 2027. You get one last swap. Don't waste it.
On February 1, 2027, Azure reservations stop being exchangeable for anything savings plans cover: VMs, Dedicated Host, App Service, SQL Database and similar. That's four months away, and it's easy to read as a footnote. It isn't. What changes (Microsoft Learn, pages updated July and September 2026): 1. Reservations bought on or after Feb 1, 2027 for those services can't be exchanged. At all. 2. Reservations bought before that date keep exactly one final exchange. 3. Refunds don't change: cancelled commitment is capped at $50,000 per billing profile or enrollment in a rolling 12 months. Refunds that come out of an exchange don't count against the cap. Microsoft says a 12% early-termination fee might arrive in the future. 4. Trade-in to a savings plan is unchanged. But a savings plan can never be exchanged for anything afterwards. 5. Not affected: Azure VMware Solution and anything else savings plans don't cover. Instance size flexibility stays. The part that bites: today, exchanges are unlimited. After Feb 1 you get one per reservation. So every exchange you already know you'll need (the VM series you're moving off, the region you're consolidating out of, the App Service plan you're retiring) is free to do now and costs you your last move later. Second trap: the "safe" exit. When you trade reservations in for a savings plan, the portal proposes an hourly commitment derived from the money left in the reservations. Microsoft's own doc warns it might not be large enough to cover the same VMs. A savings plan usually gives a smaller discount than a reservation for the same SKU, and the remaining money is spread over a brand-new term. Illustrative numbers (invented, check your own SKUs in the pricing calculator): 20 VMs at $0.40/hr on-demand. Reservation rate $0.24/hr (40% off) = $4.80/hr. Savings plan rate $0.28/hr (30% off). Accept a $4.80/hr savings plan and it covers about 17 VMs; the other 3 run on-demand. You now pay about $5.94/hr. That's roughly $10k a year extra for flexibility you may never use.
0
0
Your AI gateway holds every model key you own. In March, the popular one got backdoored.
Every team that gets past "one engineer with an API key" ends up building the same thing: an AI gateway. One endpoint in front of every model, local or cloud. Apps get a virtual key, the gateway holds the real ones, and you finally get budgets, rate limits, routing and an audit trail in one place. It's the right pattern. It also creates the single most valuable secret store you run. On 24 March 2026, LiteLLM, the most widely used open-source gateway, had two PyPI releases (1.82.7 and 1.82.8) published with a credential-stealing payload. They were live for roughly 40 minutes. The attacker got the publishing credentials by compromising a security scanner in the project's own CI. Anyone whose pipeline pulled "latest" in that window shipped a stealer onto the box that holds every model key they own. (Sources: LiteLLM security update, Datadog Security Labs.) That isn't a reason to skip the gateway. It's a reason to build it like a vault instead of a convenience proxy. What a gateway should do (the five jobs): 1. Identity. One virtual key per app or team, never a shared provider key. 2. Budgets and quotas. Dollar caps and tokens per minute per key, enforced before the call, not discovered on the invoice. 3. Routing. Bulk traffic to your local model, the hard cases to a frontier API, with fallback rules you actually chose. 4. Policy. "Confidential" traffic maps to the local model only, with no cloud fallback. Ever. 5. Audit. Who called which model, how many tokens, what it cost, which policy decision fired. How to not turn it into your worst incident: - Pin the version AND the container image digest. No auto-upgrades on the box with every key. - Where the provider supports it, use workload identity instead of static keys (Bedrock through an IAM role, Vertex through a service account). A stolen short-lived token beats a stolen permanent key. - Egress allow-list on the gateway host: your model providers and nothing else. That one rule would have neutered a stealer's exfiltration.
0
0
1-30 of 57
powered by
AI for Cloud Engineers
skool.com/cloud-cost-optimization-3746
Automate your cloud work with AI. GCP, Azure, VMware. Save hours every week with real workflows.
Build your own community
Bring people together around your passion and get paid.
Powered by