Before an AI agent can troubleshoot Kubernetes, the platform needs reliable observability, useful runbooks, clear permissions, safe rollback procedures, and engineers who understand the systems underneath. That is why this upcoming hands-on workshop from Packt caught my attention: Link to my post in LinkedIn The workshop explores how infrastructure agents can combine: - Live Kubernetes state - Prometheus metrics - Runbooks and postmortems - MCP-based tooling - Human approval gates - Dry-run and audit workflows The practical scenarios include Kubernetes CrashLoopBackOff investigations, GPU scheduling failures, latency troubleshooting, autoscaling inspection, and controlled deployment rollbacks. Through my work with Kubernetes, GPU infrastructure, observability, and production platforms, I have seen how much time incident response depends on collecting the right context quickly. I am now exploring how AI agents could assist operators with on-call investigations, cluster troubleshooting, and incident diagnosis without removing human judgment or control. I am particularly interested in how these patterns could evolve into reusable platform capabilities for teams and clients running complex Kubernetes and AI workloads. My public focus remains firmly grounded in Kubernetes, observability, platform engineering, and operational fundamentals. I chose to share this workshop because it connects those foundations with practical AI tooling rather than treating AI as a substitute for sound engineering. Packt has provided an exclusive 50% discount for my network. Use code: ERDEN50 📅 August 27, 2026 💻 Online, hands-on workshop 🔗 Registration: https://lnkd.in/dJxCpnRJ