Wayve's Dave Kirk: Why Agentic Code Review Needs Evals
Summary
Dave Kirk details Wayve's approach to agentic PR code review, emphasizing that reliable AI adoption requires moving beyond 'vibes-based' evaluation. The system uses a structured feedback loop—integrating sentiment tracking, usage metrics, and dedicated evaluations (Evals)—to improve prompts and guide multi-agent behavior in complex, high-stakes environments like self-driving car development.
Key takeaways
-
Agent Reliability Requires Observability
2:08
Multi-agent systems are stochastic and difficult to predict. Kirk notes that observability is critical; if a single agent's behavior cannot be observed, building reliable, production-ready multi-agent workflows is extremely challenging.
-
The Pitfalls of Public Benchmarks
10:53
Public coding benchmarks are often untrustworthy because agents can learn to 'cheat' the tests. Performance gains may simply reflect improved cheating mechanisms rather than genuine capability improvements.
-
Structured Feedback Loops are Essential
22:30
Wayve implements a feedback loop by collecting data on code review outcomes, including sentiment (thumbs up/down) and usage tracking. This data is used to identify common mistakes in prompts and improve agent behavior iteratively.
-
The Value of Evals
23:25
To ensure confidence, the team uses dedicated evaluation agents (Evals) that test the quality of output from other agents. Kirk highlights performing 'eval-driven development,' where the eval mechanism is built before the agent itself.
Technical details
-
Agentic PR Review Pipeline
1032s
The system uses a GitHub Action that runs when a PR is opened or transitioned to 'ready for review.' It evaluates which prompts to run, fanning out multiple agents, and steering their behavior using common skills defined in the prompt.
-
Prompt Steering and Scope Control
1085s
Agents are steered by defining specific prompts based on code changes. Prompts can be scoped to run only when certain files are modified (e.g., infrastructure components requiring `Terraform apply`), allowing for targeted, efficient feedback.
-
Multi-Agent Coordination
310s
Kirk's initial experience with Claude multi-agent teams involved a top-level agent delegating tasks to specialized agents (e.g., one for implementing tickets, another for QA, and a third for merge conflicts). The challenge was managing role confusion and lack of outcome visibility.
-
Metrics Collection
1280s
Data collected includes sentiment (thumbs up/down) on agent feedback, usage tracking (whether the feedback was actually implemented), and comparison against baselines like generic suggestions from GPT-5.3 Codex.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.