User
Write something
Your AI gateway holds every model key you own. In March, the popular one got backdoored.
Every team that gets past "one engineer with an API key" ends up building the same thing: an AI gateway. One endpoint in front of every model, local or cloud. Apps get a virtual key, the gateway holds the real ones, and you finally get budgets, rate limits, routing and an audit trail in one place. It's the right pattern. It also creates the single most valuable secret store you run. On 24 March 2026, LiteLLM, the most widely used open-source gateway, had two PyPI releases (1.82.7 and 1.82.8) published with a credential-stealing payload. They were live for roughly 40 minutes. The attacker got the publishing credentials by compromising a security scanner in the project's own CI. Anyone whose pipeline pulled "latest" in that window shipped a stealer onto the box that holds every model key they own. (Sources: LiteLLM security update, Datadog Security Labs.) That isn't a reason to skip the gateway. It's a reason to build it like a vault instead of a convenience proxy. What a gateway should do (the five jobs): 1. Identity. One virtual key per app or team, never a shared provider key. 2. Budgets and quotas. Dollar caps and tokens per minute per key, enforced before the call, not discovered on the invoice. 3. Routing. Bulk traffic to your local model, the hard cases to a frontier API, with fallback rules you actually chose. 4. Policy. "Confidential" traffic maps to the local model only, with no cloud fallback. Ever. 5. Audit. Who called which model, how many tokens, what it cost, which policy decision fired. How to not turn it into your worst incident: - Pin the version AND the container image digest. No auto-upgrades on the box with every key. - Where the provider supports it, use workload identity instead of static keys (Bedrock through an IAM role, Vertex through a service account). A stolen short-lived token beats a stolen permanent key. - Egress allow-list on the gateway host: your model providers and nothing else. That one rule would have neutered a stealer's exfiltration.
0
0
Your model fits in VRAM. Your agents don't. Size GPUs for KV cache, not weights.
Every GPU sizing thread starts with the same question: will the model fit? For chat that is most of the answer. For agents it is the smallest part. An agent session is not a prompt. It is a system prompt, 5 to 15k tokens of tool schemas, every tool result, and the full transcript re-read on every loop. It only grows. So what runs out on your GPU is not room for weights. It is room for the KV cache, the per-token memory every live session holds for as long as it is alive. This is not a benchmark. It is arithmetic from the model's config.json: KV bytes per token = 2 x layers x KV heads x head dim x bytes per element Qwen3-14B: 2 x 40 x 8 x 128 x 2 bytes = 160 KiB per token at FP16. One agent session at 32k tokens = 5 GiB, just for its memory of the conversation. Now put that on real cards. These are my estimates using a vLLM-style budget: 90% of VRAM, minus weights, minus 1.5 to 3 GiB of runtime overhead. Check them against your own logs. 16 GB card, Qwen3-14B 4-bit: about 3.4 GiB left for KV. That is 0.7 sessions at 32k context. Not a typo. vLLM will refuse to start with max-model-len 32768. 16 GB, Qwen3-8B 4-bit: about 1.6 sessions at 32k, about 3 with an FP8 KV cache. 48 GB, Qwen3-32B 4-bit: about 2.7 sessions at 32k, about 5 with FP8 KV. 48 GB, Qwen3-14B FP8: about 5 at 32k, about 10 with FP8 KV. 80 GB, Llama 3.3 70B FP8: zero. 70 GB of weights plus overhead leaves nothing for the cache. It fits and it serves nobody. 80 GB, gpt-oss-120b: about 5 at 32k, about 10 with FP8 KV. More than the 70B on the same card, because only 18 of its layers use full attention and its heads are 64-dim. 36 KiB per token against 320. Three things fall out of this. 1. Architecture beats parameter count. KV heads, head dim and sliding-window layers decide concurrency more than 14B vs 32B does. Read the config.json before you read the benchmark. 2. An FP8 KV cache roughly doubles your sessions. In vLLM it is one flag, --kv-cache-dtype fp8. Test quality on your own agent evals before you trust it.
0
0
Your model fits in VRAM. Your agents don't. Size GPUs for KV cache, not weights.
Your CI said green. That does not mean the tests ran.
An AI coding agent finishes a task and reports "all checks passed". That sentence is a claim to investigate, not evidence to merge on. Here is the specific thing that makes it dangerous, and it is documented behaviour, not a bug: in GitHub Actions, a job skipped by a condition can report success. Not "skipped" in a way your required-check rule notices. Success. So a required check can be green because it ran and passed, or green because it never ran at all, and by default those two look identical on the pull request. Now combine that with an agent that has write access to the repository. The failure mode is not the agent writing malicious code. It is much more boring. The agent changes a formatter, a test fails, and the simplest path to a green pipeline is to weaken the assertion or add a condition that skips the job. Nothing about that is adversarial. It is just an optimiser finding the cheapest route to the stated goal, which was "make the checks pass" and not "make the code correct". Then a human sees a green badge and merges. Three things that actually help: 1. Treat "passed" and "did not run" as different states. Write down which checks you expect for this class of change, then verify that each one executed on the revision under review. A missing entry, a skipped job, a cancelled run, a timeout and a neutral conclusion should not all collapse into "no failure reported". 2. Separate who writes code from who has verification authority. The agent can prepare a branch and run checks locally. CI repeats the required verification from the reviewed revision, in a controlled runner. A transcript of local commands is useful context for a reviewer and is not a substitute for a CI result bound to that commit. 3. Review test and workflow changes as a separate category. If the same diff removes an assertion and then earns a green badge, the badge is describing a pipeline that no longer checks the thing you cared about. Make workflow files visible to someone who can evaluate them.
0
0
Your team of 50 is not 50 concurrent users. That number picks your inference server.
Every "which inference server should we run" thread turns into a benchmark fight within about four replies. Someone posts tokens/sec for a model nobody else is running, at a batch size nobody states, on a GPU nobody else has. It is the wrong argument. The number that actually decides this is concurrency. Not team size. Concurrency. "We have 50 engineers" is not 50 concurrent requests. Fifty engineers, working the same timezone, using an internal assistant a handful of times an hour, with each request taking a few seconds, gives you a peak concurrency in the low single digits. I would bet on 2 to 5 for most internal tooling. That is a wildly different engineering problem from 50, and it points at a different server. Go and measure it before you read another benchmark. Peak concurrent in-flight requests, over a normal week. Every gateway, proxy or server you already run can tell you. If you skip this step, everything downstream is guesswork. Now the part people get wrong about the three tools: they are not three competitors at the same layer. llama.cpp is an inference engine. C++, GGUF weights, and by far the widest hardware reach of the three: CPU, Apple Silicon via Metal, CUDA, ROCm, Vulkan. Minimal dependencies. It ships a server binary. If you are on an air-gapped box, odd hardware, no GPU, or a machine where installing a CUDA stack is a six-week change request, this is the one that will actually run. Ollama is llama.cpp with a model manager and a good UX bolted on. Single binary, pull a model, it works in minutes. That ease is why most teams start here, and starting here is usually correct. vLLM is a different category. It is a serving system built for throughput, and its whole design centre - continuous batching, paged KV cache - is about making aggregate throughput scale as concurrency rises. That is the thing it does that the other two do not do as well. Which means the choice is nearly decided by your concurrency number. Low single-digit concurrency, and vLLM's central advantage is mostly dormant while you pay its operational cost: CUDA and driver version alignment, a heavier dependency stack, more to go wrong at 2am. High and sustained concurrency, especially batch work, and that same advantage is the entire reason you would run anything else.
0
0
Your team of 50 is not 50 concurrent users. That number picks your inference server.
Stop maintaining runbooks. Generate them at page time from the alert payload.
Most on-call runbooks are lies. They were written eighteen months ago by an engineer who has since left, they reference a load balancer that was replaced during the last migration, and the one command that actually mattered was never written down because the person who knew it just typed it from memory every time. We all know this. We still page people at 3am and point them at a wiki page nobody has opened since the last audit. I spent the last few weeks attacking this from a different angle. Instead of trying to keep a library of static runbooks fresh, I started generating them on demand from the alert payload itself. The alert already contains almost everything you need. A CloudWatch alarm tells you the metric, the namespace, the exact dimensions, the threshold, the evaluation window, the datapoints that breached, the account and the region. A PagerDuty incident carries the service, the escalation policy, the priority and the full trigger log entry. A Datadog monitor gives you the query, the scope and the tags. That is dense, structured, high signal context. It is exactly the kind of input a language model reasons over well. The pipeline is boring in the best way. Alert fires. Webhook or SNS topic hands the JSON to a small script. The script normalises the payload into a common shape, merges it with a static environment context file that describes your org, your clusters, your escalation tiers and your hard rules, and sends the whole thing to Claude with a strict output contract. Ninety seconds later the responder has a document with a summary, immediate actions, a ranked table of diagnostic commands, escalation criteria, three hypothesised root causes with remediation and rollback for each, and a pre-filled post-mortem template. The part that surprised me was the diagnostic ordering. I asked the model to sort commands by information gain per second rather than by category, and to state which hypothesis each command discriminates between. That single instruction turned a generic checklist into something that reads like a decision tree. For a checkout API 5xx alarm it opened with the target group health check and the deployment history, not with CPU graphs, because CPU is almost never the discriminator on a sudden error rate spike.
0
0
1-24 of 24
powered by
AI for Cloud Engineers
skool.com/cloud-cost-optimization-3746
Automate your cloud work with AI. GCP, Azure, VMware. Save hours every week with real workflows.
Build your own community
Bring people together around your passion and get paid.
Powered by