Before an AI agent can troubleshoot Kubernetes, the platform needs reliable observability, useful runbooks, clear permissions, safe rollback procedures, and engineers who understand the systems underneath.
That is why this upcoming hands-on workshop from Packt caught my attention: The workshop explores how infrastructure agents can combine:
- Live Kubernetes state
- Prometheus metrics
- Runbooks and postmortems
- MCP-based tooling
- Human approval gates
- Dry-run and audit workflows
The practical scenarios include Kubernetes CrashLoopBackOff investigations, GPU scheduling failures, latency troubleshooting, autoscaling inspection, and controlled deployment rollbacks.
Through my work with Kubernetes, GPU infrastructure, observability, and production platforms, I have seen how much time incident response depends on collecting the right context quickly. I am now exploring how AI agents could assist operators with on-call investigations, cluster troubleshooting, and incident diagnosis without removing human judgment or control.
I am particularly interested in how these patterns could evolve into reusable platform capabilities for teams and clients running complex Kubernetes and AI workloads.
My public focus remains firmly grounded in Kubernetes, observability, platform engineering, and operational fundamentals. I chose to share this workshop because it connects those foundations with practical AI tooling rather than treating AI as a substitute for sound engineering.
Packt has provided an exclusive 50% discount for my network. Use code: ERDEN50
๐
August 27, 2026
๐ป Online, hands-on workshop
๐๐ถ๐๐ฐ๐น๐ผ๐๐๐ฟ๐ฒ: Packt invited me to share this workshop and provided a tracked registration link and discount code. I have not attended the workshop, so this is not a review of its quality. I chose to share it because the curriculum closely matches my current focus on infrastructure fundamentals, Kubernetes operations, and practical AI tooling.