# AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile

## Executive summary

The presentation analyzes the quality of AI-generated pull requests (PRs) compared to human-written code, using data from over a million PRs reviewed monthly at Greptile. Findings show that while AI agents (like Claude, Devin, and Codex) are highly capable, they exhibit distinct failure modes compared to humans. Specifically, Claude is 1.5x more likely than humans to introduce SQL injection, and N+1 queries are common from Cursor. The speaker argues that traditional manual code review is insufficient for modern enterprise scale (median user: <50 commits/month; 99th percentile: ~1,000 commits/month), proposing a validation framework that answers three questions: 1) Does the change violate the user contract? 2) Does it increase the propensity for future violations? 3) Does it fulfill the author's intent?

## Key takeaways

- AI PR Adoption Rate is Rapidly Increasing: A quarter (25%+) of PRs reviewed by Greptile in April were largely or entirely AI-generated, a significant increase from under 1% a year prior. This trend is accelerating as model performance improves.
- AI PR Quality is Comparable to Human Code: When measured by revert rates, P0/P1/P2 bug counts, and review cycles to merge, agent-written PRs performed similarly to human-written ones. However, failure modes differ significantly.
- AI Agents Show Distinct Failure Patterns: Specific agents have unique failure propensities: Claude is 1.5x more likely than humans to introduce SQL injection; Devin is half as likely to cause auth bypasses; and N+1 queries are noted as common from Cursor.
- Code Validation Must Scale Beyond Manual Review: Given that the 99th percentile Greptile user writes nearly 1,000 PRs monthly, manual review is impossible. Validation must focus on answering if the change violates the user contract, increases future violation risk, and fulfills author intent.

## Technical details

- AI Agents and Code Review: Greptile uses a swarm of agents that validate PRs by analyzing changed files, related files, spinning up code in a sandbox, installing dependencies, and running browser agents to attempt to break the application.
- AI Coding Evolution: The evolution moved from simple tab-complete features (e.g., Cursor, Copilot) to multi-file editing (Cursor, 2024) and finally to fully autonomous agents capable of creating entire PRs (2025).
- PR Quality Metrics: Metrics analyzed include revert rates, average PR size, counts of P0/P1/P2 bugs, and the number of review rounds required for merging.
- Greptile Validation Framework: The proposed validation framework requires answering three questions: 1) Does this change violate the user contract? 2) Does it increase the propensity of a future violation? 3) Does it do what the author intended?

## Practical implications

- Code review processes must adapt to handle massive volumes of AI-generated code.
- Validation should shift from manual line-by-line checking to systemic checks (e.g., user contract violation, blast-radius analysis).
- AI agents are viable contributors to enterprise codebases, but their failure modes require specialized detection methods.

## Topics

AI Coding, Code Review, Software Engineering, DevOps, Code Quality, Greptile, Codex, Claude Code, Devin, Cursor

Source: https://www.youtube.com/watch?v=474j-n1Ltxc
