AI Engineer

The 5 Levels of Self-Driving Production — Eric Schwartz, Traversal

Published 2026-10-02 · Duration 18:33

Summary

The proliferation of coding agents has accelerated development but has significantly increased the complexity and volume of production issues, shifting engineering time from building features to troubleshooting. Traditional observability tools can identify *what* broke, but they fail to determine the root cause, which the speaker argues is a 'causal problem,' not an 'observability problem.' The solution is 'Self-Driving Production,' a closed-loop system that uses causal AI to automatically diagnose multi-hop failures, propose fixes, and verify resolution, moving organizations up a five-level autonomy spectrum.

Download summary

Key takeaways

  1. The Shift in Engineering Focus 0:22

    The advent of coding agents (e.g., Cloud Code, Codeex, Cursor) has made development faster, but it has resulted in more complex codebases and increased time spent on troubleshooting, which is a major industry problem.

  2. Root Cause vs. Observability 5:12

    Observability tools (like DataDog, Elastic, Splunk) are limited to pointing out correlations and symptoms. They cannot determine the root cause because root cause analysis is fundamentally a causal problem, requiring understanding cause and effect.

  3. Self-Driving Production Levels 11:46

    Autonomy is measured on a spectrum (Level 0 to Level 5). Level 5 represents the 'holy grail': a system that can diagnose issues across a full production environment, propose fixes, and verify them autonomously.

  4. AI SRE Evaluation Criteria

    Effective AI SRE solutions must be able to see all production data, search petabytes of data cost-effectively, map relationships between entities, improve autonomously, and find non-obvious, multi-hop root causes quickly.

Technical details

  • Multi-Hop Failure Analysis 431s

    Diagnosing an issue often requires tracing five to ten 'hops' across dozens of services and petabytes of data to reach the true root cause (e.g., an expired TLS certificate).

  • Causal AI Framework

    The core technical approach involves building a 'Production World Model' to map all data relationships, which feeds into a 'Causal Search Engine' to identify cause and effect.

  • Self-Driving Production Loop 626s

    This is a closed-loop system where the AI system finds the incident, determines the root cause, applies a fix, and verifies the fix, ideally without human paging.

  • Use Case: Alert Intelligence 1105s

    For supply chains (e.g., PepsiCo), the system filters thousands of raw alerts into a prioritized, pre-investigated set, reducing alert fatigue and noise.

  • Use Case: Incident RCA

    For large enterprises (e.g., American Express), the system acts as the first responder, providing detailed root cause analysis within minutes, drastically reducing the need for large-scale, chaotic paging of engineers.

Mentioned resources

  • Traversal (Company/Product)
  • DataDog (Observability Tool)
  • Elastic (Observability Tool)
  • Splunk (Observability Tool)
  • ServiceNow (Observability Tool)
  • PepsiCo (Enterprise Client)
  • American Express (Enterprise Client)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.