# The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX

## Executive summary

This presentation analyzes the impact of AI on developer productivity using data from 200,000 engineers. While AI has increased deployment frequency and code maintainability, the data reveals critical tensions: Change Confidence has dropped 6%, and PR size has significantly increased (from ~44 to 72 lines). The core finding is that code generation was never the bottleneck; instead, non-AI factors like meeting overhead and context switching are the primary constraints on value generation. DX proposes a measurement framework focusing on Utilization, Impact, and Cost, and emphasizes that improving foundational Developer Experience (DX) metrics—such as reliable CI and modular code—is crucial for maximizing agent efficiency.

## Key takeaways

- Change Confidence vs. Maintainability: Code maintainability has risen by nearly 4%, but Change Confidence has dropped 6%. This suggests developers feel more capable of understanding AI-generated code but are more hesitant to trust it, indicating a psychological shift in risk perception.
- PR Size and Incremental Delivery: Average Pull Request (PR) size has increased from approximately 44 to 72 lines. This trend, coupled with a 10% drop in the perception of incremental delivery, suggests developers are consolidating changes, which increases the risk of bugs and makes code less portable.
- AI Efficiency by Role: While junior engineers use AI the most, staff+ engineers are achieving comparable time savings while consuming fewer tokens, suggesting that deep architectural understanding is key to efficient AI utilization.
- The True Bottleneck: The median increase in PR throughput was only about 7.7%, far from the expected 2x gain. The speaker asserts that time savings from AI are often outweighed by non-AI factors like meeting overhead and context switching, which are the true constraints on value generation.
- Agent Readiness Requires Good DX: The speaker argues that improving foundational developer experience (DX) metrics—such as clear documentation, modular code, and reliable, non-flaky test suites—is necessary to build effective AI agents.

## Technical details

- DORA Metrics: Deployment Frequency (DF) is shown to be steadily increasing, indicating improved speed of shipping work. The speaker notes that DF is a proxy metric and does not account for defect ratios or change failure rate.
- DX AI Measurement Framework: A proposed methodology for assessing AI impact based on three dimensions: Utilization (usage metrics), Impact (value generation metrics), and Cost (token/compute expenditure).
- Agent Experience Feedback: A method for assessing AI effectiveness by gathering qualitative feedback directly from the AI agents regarding issues encountered during human-agent collaboration and context provision.
- Legacy Code Agents: Morgan Stanley deployed a DevGen AI agent to interpret legacy code (e.g., COBOL/mainframe), creating PRs to eliminate manual reverse engineering steps, saving an estimated 300,000 hours annually.

## Practical implications

- Focus optimization efforts on non-code bottlenecks (e.g., CI/CD pipeline time, meeting overhead, context switching) rather than solely on code generation speed.
- Prioritize improving foundational DX elements (documentation, modularity, test coverage) as these directly enhance agent performance and reliability.
- Adopt a holistic measurement approach (Utilization, Impact, Cost) to prove the ROI of AI investments, moving beyond simple usage counts.
- Implement agent-based systems for repetitive, non-core tasks (e.g., administrative summaries, superficial code reviews) to maximize throughput and capacity per engineer.

## Topics

AI, Developer Experience (DevEx), DORA Metrics, CI/CD, Software Development Lifecycle (SDLC), Agentic Workflow, Productivity Measurement, DX AI Measurement Framework, Q2 Report, Developer Experience (DX) Platform

Source: https://www.youtube.com/watch?v=Se8jHLliLXE
