# The 5 Levels of Self-Driving Production — Eric Schwartz, Traversal

## Executive summary

The proliferation of coding agents has accelerated development but has significantly increased the complexity and volume of production issues, shifting engineering time from building features to troubleshooting. Traditional observability tools can identify *what* broke, but they fail to determine the root cause, which the speaker argues is a 'causal problem,' not an 'observability problem.' The solution is 'Self-Driving Production,' a closed-loop system that uses causal AI to automatically diagnose multi-hop failures, propose fixes, and verify resolution, moving organizations up a five-level autonomy spectrum.

## Key takeaways

- The Shift in Engineering Focus: The advent of coding agents (e.g., Cloud Code, Codeex, Cursor) has made development faster, but it has resulted in more complex codebases and increased time spent on troubleshooting, which is a major industry problem.
- Root Cause vs. Observability: Observability tools (like DataDog, Elastic, Splunk) are limited to pointing out correlations and symptoms. They cannot determine the root cause because root cause analysis is fundamentally a causal problem, requiring understanding cause and effect.
- Self-Driving Production Levels: Autonomy is measured on a spectrum (Level 0 to Level 5). Level 5 represents the 'holy grail': a system that can diagnose issues across a full production environment, propose fixes, and verify them autonomously.
- AI SRE Evaluation Criteria: Effective AI SRE solutions must be able to see all production data, search petabytes of data cost-effectively, map relationships between entities, improve autonomously, and find non-obvious, multi-hop root causes quickly.

## Technical details

- Multi-Hop Failure Analysis: Diagnosing an issue often requires tracing five to ten 'hops' across dozens of services and petabytes of data to reach the true root cause (e.g., an expired TLS certificate).
- Causal AI Framework: The core technical approach involves building a 'Production World Model' to map all data relationships, which feeds into a 'Causal Search Engine' to identify cause and effect.
- Self-Driving Production Loop: This is a closed-loop system where the AI system finds the incident, determines the root cause, applies a fix, and verifies the fix, ideally without human paging.
- Use Case: Alert Intelligence: For supply chains (e.g., PepsiCo), the system filters thousands of raw alerts into a prioritized, pre-investigated set, reducing alert fatigue and noise.
- Use Case: Incident RCA: For large enterprises (e.g., American Express), the system acts as the first responder, providing detailed root cause analysis within minutes, drastically reducing the need for large-scale, chaotic paging of engineers.

## Practical implications

- Build teams must shift focus from reactive troubleshooting to proactive design and architecture, as AI agents increase code volume and complexity.
- Organizations should evaluate AI SRE solutions not just on data ingestion, but on their ability to perform multi-hop, causal root cause analysis.
- Implementing self-driving production can drastically reduce Mean Time To Resolution (MTTR) and prevent alert fatigue by prioritizing actionable incidents.

## Topics

Site Reliability Engineering (SRE), AI/ML, DevOps, Production Observability, Causal Inference, Traversal, DataDog, Elastic, Splunk, ServiceNow, PepsiCo, American Express

Source: https://www.youtube.com/watch?v=y-OVWZD4j6U
