AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile
Summary
The presentation analyzes the quality of AI-generated pull requests (PRs) compared to human-written code, using data from over a million PRs reviewed monthly at Greptile. Findings show that while AI agents (like Claude, Devin, and Codex) are highly capable, they exhibit distinct failure modes compared to humans. Specifically, Claude is 1.5x more likely than humans to introduce SQL injection, and N+1 queries are common from Cursor. The speaker argues that traditional manual code review is insufficient for modern enterprise scale (median user: <50 commits/month; 99th percentile: ~1,000 commits/month), proposing a validation framework that answers three questions: 1) Does the change violate the user contract? 2) Does it increase the propensity for future violations? 3) Does it fulfill the author's intent?
Key takeaways
-
AI PR Adoption Rate is Rapidly Increasing
5:33
A quarter (25%+) of PRs reviewed by Greptile in April were largely or entirely AI-generated, a significant increase from under 1% a year prior. This trend is accelerating as model performance improves.
-
AI PR Quality is Comparable to Human Code
7:40
When measured by revert rates, P0/P1/P2 bug counts, and review cycles to merge, agent-written PRs performed similarly to human-written ones. However, failure modes differ significantly.
-
AI Agents Show Distinct Failure Patterns
9:00
Specific agents have unique failure propensities: Claude is 1.5x more likely than humans to introduce SQL injection; Devin is half as likely to cause auth bypasses; and N+1 queries are noted as common from Cursor.
-
Code Validation Must Scale Beyond Manual Review
11:20
Given that the 99th percentile Greptile user writes nearly 1,000 PRs monthly, manual review is impossible. Validation must focus on answering if the change violates the user contract, increases future violation risk, and fulfills author intent.
Technical details
-
AI Agents and Code Review
110s
Greptile uses a swarm of agents that validate PRs by analyzing changed files, related files, spinning up code in a sandbox, installing dependencies, and running browser agents to attempt to break the application.
-
AI Coding Evolution
180s
The evolution moved from simple tab-complete features (e.g., Cursor, Copilot) to multi-file editing (Cursor, 2024) and finally to fully autonomous agents capable of creating entire PRs (2025).
-
PR Quality Metrics
460s
Metrics analyzed include revert rates, average PR size, counts of P0/P1/P2 bugs, and the number of review rounds required for merging.
-
Greptile Validation Framework
690s
The proposed validation framework requires answering three questions: 1) Does this change violate the user contract? 2) Does it increase the propensity of a future violation? 3) Does it do what the author intended?
Mentioned resources
- Greptile
- Codex
- Claude Code
- Devin
- Cursor
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.