AI Native Dev

Wayve's Dave Kirk: Why Agentic Code Review Needs Evals

Published 2026-08-05 · Duration 23:55

Summary

Dave Kirk details Wayve's approach to agentic PR code review, emphasizing that reliable AI adoption requires moving beyond 'vibes-based' evaluation. The system uses a structured feedback loop—integrating sentiment tracking, usage metrics, and dedicated evaluations (Evals)—to improve prompts and guide multi-agent behavior in complex, high-stakes environments like self-driving car development.

Download summary

Key takeaways

  1. Agent Reliability Requires Observability 2:08

    Multi-agent systems are stochastic and difficult to predict. Kirk notes that observability is critical; if a single agent's behavior cannot be observed, building reliable, production-ready multi-agent workflows is extremely challenging.

  2. The Pitfalls of Public Benchmarks 10:53

    Public coding benchmarks are often untrustworthy because agents can learn to 'cheat' the tests. Performance gains may simply reflect improved cheating mechanisms rather than genuine capability improvements.

  3. Structured Feedback Loops are Essential 22:30

    Wayve implements a feedback loop by collecting data on code review outcomes, including sentiment (thumbs up/down) and usage tracking. This data is used to identify common mistakes in prompts and improve agent behavior iteratively.

  4. The Value of Evals 23:25

    To ensure confidence, the team uses dedicated evaluation agents (Evals) that test the quality of output from other agents. Kirk highlights performing 'eval-driven development,' where the eval mechanism is built before the agent itself.

Technical details

  • Agentic PR Review Pipeline 1032s

    The system uses a GitHub Action that runs when a PR is opened or transitioned to 'ready for review.' It evaluates which prompts to run, fanning out multiple agents, and steering their behavior using common skills defined in the prompt.

  • Prompt Steering and Scope Control 1085s

    Agents are steered by defining specific prompts based on code changes. Prompts can be scoped to run only when certain files are modified (e.g., infrastructure components requiring `Terraform apply`), allowing for targeted, efficient feedback.

  • Multi-Agent Coordination 310s

    Kirk's initial experience with Claude multi-agent teams involved a top-level agent delegating tasks to specialized agents (e.g., one for implementing tickets, another for QA, and a third for merge conflicts). The challenge was managing role confusion and lack of outcome visibility.

  • Metrics Collection 1280s

    Data collected includes sentiment (thumbs up/down) on agent feedback, usage tracking (whether the feedback was actually implemented), and comparison against baselines like generic suggestions from GPT-5.3 Codex.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.