Liz Fong-Jones: 2x the PRs, 1.5x the Incidents
Summary
The rapid adoption of AI coding agents (like those generating pull requests) is accelerating development volume (e.g., 30 to 70 PRs/day) but does not automatically solve systemic reliability issues. The core challenge is scaling human review capacity. The discussion emphasizes that AI amplifies existing organizational practices—magnifying both high ownership and dysfunction. To safely scale, organizations must focus on improving guardrails, implementing automated review classifiers (like Jev), enforcing strong CI/CD patterns, and ensuring human accountability remains paramount, especially during production incidents.
Key takeaways
-
AI amplifies existing organizational practices.
18:14
AI will not fix existing problems; it will magnify them. Organizations must first fix underlying issues (e.g., poor ownership, weak patterns) before introducing AI to accelerate the path.
-
Automated code review is critical for scaling review capacity.
23:21
Tools like Honeycomb's Autobot focus on automating the review process, allowing human developers to focus on complex design patterns rather than trivial bugs. The bottleneck shifts from code generation to code validation.
-
Using classifiers (like Jev) for PR safety.
35:30
A key strategy is classifying PRs into 'safe to merge automatically' versus 'requires human review.' This focuses human attention on the most complex or risky changes, rather than attempting to review every PR.
-
Ownership and accountability remain non-negotiable.
The principle 'If your name's on it, you own it' must extend beyond code to include the responsibility for hardening systems and analyzing failures. 'Claude did it' is not an acceptable excuse.
-
Production maturity dictates AI agent trust.
AI agents can only operate safely if the underlying system has mature practices, including working automatic rollbacks, feature flags, and robust observability. Without these, agents are 'throwing darts at the dartboard.'
Technical details
-
Automated Code Review & PR Management
1401s
Honeycomb's internal bot, Autobot, functions as a 100% reviewer of PRs, helping developers scale their review capacity. The goal is to shift human effort from checking trivial bugs to scrutinizing design patterns and complex changes.
-
Jev Classifier
2130s
Jev is being deployed to classify PRs, determining if they are 'safe to merge' automatically or if they require human intervention. This uses a classification model to assess the complexity and risk level of the code change.
-
Observability and Telemetry
2651s
Observability requires both data (telemetry) and the ability to understand the system. It is a joint responsibility (human and AI) to turn raw data into actionable understanding, which should be integrated into the development cycle (e.g., having LLMs test telemetry in dev environments).
-
System Guardrails and Least Privilege
For AI agents, implementing guardrails like MCP-mediated access and enforcing read-only profiles (even if read-write is available) is crucial. The focus must be on making the tools the agents use safe, rather than trusting the agents' judgment.
-
CI/CD and Open Source Contributions
2400s
Maintaining strong CI/CD patterns, high-quality comments, and testability is foundational. For open source, maintainers should continue accepting bug reports but should not lower standards for AI-created PRs compared to human-created ones.
Mentioned resources
- Honeycomb
- Tessl
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.