# Production Monitoring with Codex: Grafana, Kubernetes, & Security

## Executive summary

This session demonstrates how Codex, an AI agent, significantly accelerates production monitoring and incident response by automating the manual process of gathering evidence and proposing fixes. Three practical demos showcase Codex's ability to diagnose checkout failures using Grafana, resolve Kubernetes rollout issues (like OOM kills), and mitigate resource starvation vulnerabilities using the Codex Security plugin. The core benefit is offloading repetitive investigative work to an agentic process, allowing human engineers to approve fixes and restore service availability much faster than traditional manual methods.

## Key takeaways

- Accelerated Incident Response: Codex streamlines the investigation process by gathering evidence from multiple sources (e.g., successful checkouts, response time, service health) and proposing a fix, allowing service restoration in minutes rather than hours.
- Kubernetes Rollout Diagnosis: The 'Kubernetes rollout investigator skill' can diagnose complex failures, such as cascading failures caused by an OOM-killed container, and identify the causal chain to propose a patch and restore the cluster to a healthy baseline.
- Security-Driven Monitoring: Production monitoring can intersect with security. The Codex Security plugin identified a resource-starving request that was taking down the service, generating a patch that successfully blocked the expensive request.

## Technical details

- Production Monitoring Workflow: Engineers monitor key metrics on a Grafana dashboard, including checkout health, current release version (e.g., V2), error rate, and P95 metrics. Codex uses this telemetry, combined with deployment context, to investigate outages.
- Container Failure Analysis: When a service fails (e.g., an Inventory API container is OOM killed), the system uses a dedicated investigator skill to analyze the rollout, identify the root cause, and propose a patch to restore the entire cluster (including Orders API and Edge Gateway).
- Advanced Automation: The system supports hosting custom runners within observability platforms (Kubernetes or Grafana) to monitor alerts above baseline. Furthermore, multi-agent systems can be configured to validate and push fixes without human intervention.

## Practical implications

- Reduces Mean Time To Recovery (MTTR) by automating the manual data gathering phase of incident response.
- Allows engineers to move from reactive troubleshooting to proactive, AI-assisted diagnosis.
- Enables the integration of security findings (resource overuse) directly into production monitoring and patching workflows.

## Topics

Production Monitoring, Site Reliability Engineering (SRE), AI Agents, Kubernetes, Incident Response, Codex, Grafana, Codex Security plugin

Source: https://www.youtube.com/watch?v=RQfg-zo4R-g
