AI Engineer

Always-on agents run production without the on-call tax — Justin Smith, Resolve AI

Published 2026-08-09 · Duration 24:56

Summary

The talk introduces the concept of 'always-on agents' designed to automate operational tasks in complex production environments, thereby reducing the burden of manual on-call work. While CI/CD handles baseline checks well, the biggest gap is monitoring non-alerted changes—such as feature flag rollouts or infrastructure updates—that require continuous context understanding. Background agents can run autonomously (on schedules, events, or messages) to perform deep analysis, root cause investigations, and proactive health checks across systems like Kafka pipelines.

Download summary

Key takeaways

  1. The Operational Bottleneck 2:05

    A significant portion of an engineer's time (estimated at 70%) is spent running code in production—maintaining platforms, debugging incidents, and handling alerts—rather than writing it. This complexity increases with the velocity of change driven by AI.

  2. Background Agents vs. Incident Response 10:40

    While on-call agents handle immediate alerts and incidents, background agents address the 'long tail' of operational work—such as routine health checks, summarizing handoffs, or watching for subtle performance drifts (e.g., P99 drift) that don't trigger an alert.

  3. The Importance of Context 12:00

    Execution is easy; production context is hard. The value lies in building knowledge systems that can determine if a metric 'smells wrong' or understand the causal chain impact of a change, rather than just loading a dashboard.

Technical details

  • Agent Architecture 410s

    Agents operate in a cloud sandbox with their own file system, ensuring state persistence regardless of the user's local machine status. They utilize a learning loop to capture knowledge and understand how services interact as the environment evolves.

  • Agent Triggers 820s

    Agents can be triggered in multiple ways: on a schedule (e.g., weekly reports), by event streams (e.g., CI/CD pipeline completion, Slack messages), or via direct message commands.

  • Deployment Monitoring 1045s

    Agents can analyze GitHub release tags to understand the specific changes (e.g., 'checkout replacing currency service'). They then build customized check plans, monitoring targeted KPIs like checkout latency and error rates, following causal chains into systems like Kafka.

  • Agent Interaction 580s

    Agents can proactively watch channels (e.g., Slack) for engineering questions. If they lack confidence in an answer, they can DM the user to confirm before responding publicly.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.