User
Write something
Pinned
Start here — how to get the most out of this community
Welcome. If you just joined, here's the 3-minute version of how to use this place. 1. Classroom → pick one lesson. Every lesson is a real build (Terraform drift agent, K8s cost optimizer, AWS anomaly detector) with working code, not slides. Start with whichever problem you actually have this week. 2. Comment on this post with what you're working on right now — cost overruns, drift, K8s chaos, IAM sprawl, anything. I read every one and it shapes what gets built next. 3. Feed gets a new real-world project weekly. Free tier = tips, tutorials, Q&A. Paid tier (Classroom) = full labs, scripts, and templates you keep. This only works if it's not a monologue — so tell me what's breaking in your stack. What are you fighting with right now?
1
0
Tiering to Glacier can cost more than doing nothing. Here's the math nobody runs first.
A FinOps consultant told a team I know to move their log archive to Glacier Deep Archive. 96% cheaper per GB. The lifecycle rule took four minutes to write. The next bill was $6,000 higher. Nothing was misconfigured. The rule did exactly what it was told. The problem is that three of the four things S3 charges you for during a tiering operation are priced per OBJECT, not per gigabyte — and that bucket had 120 million objects averaging 4 KB. Here is what the "96% cheaper" pitch leaves out. 1. TRANSITION REQUESTS COST $0.05 PER 1,000 OBJECTS. Lifecycle transitions to Glacier Flexible Retrieval or Deep Archive are billed requests. 120 million objects = $6,000, one time, on the first bill. The fee is completely blind to object size. 120 million 4 KB log lines and 120 million 4 GB video files pay exactly the same. 2. GLACIER ADDS 40 KB OF BILLED METADATA PER OBJECT. 32 KB at the archive rate, plus 8 KB billed at S3 STANDARD rates — the hot rate, forever, for the object's name. A 4 KB object in Deep Archive is billed as 44 KB. That is not a rounding error, that is 11x your data. 3. THE IA TIERS HAVE A 128 KB MINIMUM BILLABLE SIZE. Standard-IA, One Zone-IA and Glacier IR round every object up to 128 KB for billing. Your 4 KB object is billed as 128 KB. Same trap, different shape. 4. MINIMUM STORAGE DURATIONS ARE LOCK-IN, NOT A GUIDELINE. Deep Archive is 180 days. Delete on day 90 and you pay for days 90-180 anyway. Azure Archive is 180 days prorated. GCS Archive is 365. The classic own-goal is a transition-at-30-days rule sitting next to an expire-at-90-days rule. You pay for six months of storage on data that stopped existing in month three. THE NUMBER THAT DECIDES IT Work the arithmetic and you get a break-even object size. For Deep Archive, paying back the transition fee inside one year, it is roughly 208 KB. Below that, tiering loses money. Standard-IA lands around 108 KB. That is why the single most valuable line in a lifecycle policy is one most people have never used:
0
0
Tiering to Glacier can cost more than doing nothing. Here's the math nobody runs first.
Your team of 50 is not 50 concurrent users. That number picks your inference server.
Every "which inference server should we run" thread turns into a benchmark fight within about four replies. Someone posts tokens/sec for a model nobody else is running, at a batch size nobody states, on a GPU nobody else has. It is the wrong argument. The number that actually decides this is concurrency. Not team size. Concurrency. "We have 50 engineers" is not 50 concurrent requests. Fifty engineers, working the same timezone, using an internal assistant a handful of times an hour, with each request taking a few seconds, gives you a peak concurrency in the low single digits. I would bet on 2 to 5 for most internal tooling. That is a wildly different engineering problem from 50, and it points at a different server. Go and measure it before you read another benchmark. Peak concurrent in-flight requests, over a normal week. Every gateway, proxy or server you already run can tell you. If you skip this step, everything downstream is guesswork. Now the part people get wrong about the three tools: they are not three competitors at the same layer. llama.cpp is an inference engine. C++, GGUF weights, and by far the widest hardware reach of the three: CPU, Apple Silicon via Metal, CUDA, ROCm, Vulkan. Minimal dependencies. It ships a server binary. If you are on an air-gapped box, odd hardware, no GPU, or a machine where installing a CUDA stack is a six-week change request, this is the one that will actually run. Ollama is llama.cpp with a model manager and a good UX bolted on. Single binary, pull a model, it works in minutes. That ease is why most teams start here, and starting here is usually correct. vLLM is a different category. It is a serving system built for throughput, and its whole design centre - continuous batching, paged KV cache - is about making aggregate throughput scale as concurrency rises. That is the thing it does that the other two do not do as well. Which means the choice is nearly decided by your concurrency number. Low single-digit concurrency, and vLLM's central advantage is mostly dormant while you pay its operational cost: CUDA and driver version alignment, a heavier dependency stack, more to go wrong at 2am. High and sustained concurrency, especially batch work, and that same advantage is the entire reason you would run anything else.
0
0
Your team of 50 is not 50 concurrent users. That number picks your inference server.
Buying GPUs to escape API bills? Run this math first — it usually says don't.
Someone in your org has said "our data can't leave the building." Fair. Then someone else translated that into "so we'll buy GPUs, and it'll be cheaper than the API anyway." That second sentence is where the money gets lost. The first reason is legitimate. The second one usually isn't, and when you staple them together you end up defending a bad cost model in front of a CFO who will eventually check it. Here's the math nobody puts on the slide. The unit that matters is cost per 1M tokens. Not GPU price. Not power. Cost per 1M tokens, because that's the only number that compares to an API invoice. To get there on-prem you need: amortised capital (GPU + chassis + CPU/RAM/NVMe, over 36 months, and be honest that the resale value of a three-year-old accelerator is not what you hope), power at load AND at idle, cooling — multiply your power by your PUE, and if you don't know your PUE, assume 1.5 and flag it, rack or colo, network, spares and the failure domain (one node is not a service, it's a single point of failure with a maintenance window), software and support, and ops headcount. Ops headcount is the line that kills it. A half-FTE to keep an inference node healthy, patched, and on a current model version is, at European loaded cost, comfortably more per year than most teams' entire API spend. I've watched people model a $30k GPU to two decimal places and then write "0.2 FTE" as if that were free. And then the one that actually decides it: utilisation. A GPU is a fixed cost that runs whether or not anyone is talking to it. If your box is genuinely busy 60% of the time, your cost per token is one number. If it's busy 8% of the time — which is what internal tooling with European working hours actually looks like — your effective cost per token is roughly 7x worse. Same hardware. Same invoice. The API, by contrast, charges you nothing at 3am on a Sunday. That asymmetry is the whole argument. Owning hardware is a bet that you will keep it busy. Most internal AI workloads are bursty, business-hours, and nowhere near the volume needed to win that bet.
0
0
Buying GPUs to escape API bills? Run this math first — it usually says don't.
Your GCP inter-zone traffic is not free. Most teams find this in month 14.
Everyone knows internet egress costs money. Almost nobody prices the traffic that never leaves the region. us-central1-a to us-central1-b, over internal IPs, inside your own VPC: roughly $0.01/GB. It looks like LAN traffic. It is billed like WAN traffic. That number sounds trivial until you do the arithmetic on a real cluster. A three-zone GKE setup with no topology awareness, where a service mesh, a Prometheus remote-write, and a chatty API-to-cache path all round-robin across zones, will do 30-60 TB a month of cross-zone chatter without anyone noticing. At ~$0.01/GB that is $300-$600/month for bytes that were never supposed to leave the rack, generated by a HA topology nobody re-examined after the initial deploy. Three more places GCP charges you for traffic that feels free: 1. Cloud NAT is a second meter. ~$0.045/GB processed, on top of egress, billed on data processed in both directions. Your private GKE nodes pulling images from Artifact Registry and shipping logs to Cloud Logging through NAT are paying for traffic Private Google Access carries for $0. One subnet flag and a private DNS zone fixes it. 2. Multi-region buckets quietly convert free reads into billed ones. A VM in us-central1 reading from a us-central1 bucket: free. The same VM reading from the US multi-region bucket: different location, ~$0.02/GB. Multi-region is a durability choice, not a performance one. Plenty of hot paths are pointed at one by default. 3. Cloud VPN gets no egress discount. VPN egress is billed at standard internet rates plus tunnel hours. Only Interconnect gets the cheaper SKU. If you moved to VPN to save on egress, you did not. The reason this hides for years is that none of it shows up as a line item you can attribute. The bill says Network Inter Zone Egress: $487. It does not say which service. VPC Flow Logs do - if you enable them with sampling (0.1-0.5, 30-second aggregation, custom metadata including src_instance, dest_instance and src_gke_details), sink them to BigQuery with --use-partitioned-tables, and classify each flow by the boundary it crosses.
Your GCP inter-zone traffic is not free. Most teams find this in month 14.
1-30 of 50
powered by
AI for Cloud Engineers
skool.com/cloud-cost-optimization-3746
Automate your cloud work with AI. GCP, Azure, VMware. Save hours every week with real workflows.
Build your own community
Bring people together around your passion and get paid.
Powered by