E
Emre Dundar
Guest
We can measure AI usage and software delivery. The hard part is understanding what happens between the two.
AI coding measurement has become surprisingly sophisticated in a very short time. We can measure active AI users. We can see which tools developers use. We can count prompts, accepted suggestions, tokens, AI-assisted commits, and increasingly even identify which parts of a codebase were created with AI assistance.
At the other end of the software delivery system, we already have mature engineering metrics:
- Pull request cycle time
- Review time
- Deployment frequency
- Change failure rate
- Rework
- Incidents
- Delivery lead time
Yet there is a large gap between these two worlds.
We know how much AI is being used.
We know how software is being delivered.
What we often don't know is what happened in between.
That missing middle is where I think most of the interesting AI productivity questions now live.
The First Generation of AI Metrics Was About Adoption
This was reasonable. When companies started rolling out GitHub Copilot, Cursor, Claude Code, and similar tools, engineering leaders first needed answers to basic questions:
- Who has access?
- Who is actually using it?
- Which tools are being used?
- How many suggestions are accepted?
- How much does it cost?
- How much code appears to be AI-assisted?
These are important operational metrics.
But they answer a rollout question: Are developers using AI?
They don't answer the much harder question: What changed in the engineering system because developers used AI?
The distinction becomes more important as adoption increases. Once most developers in an organization use an AI coding tool, increasing adoption from 70% to 80% tells an engineering leader relatively little about whether the investment is working. At that point, usage stops being the interesting variable. The outcomes become interesting.
Code Generation Is Only the Beginning of the Pipeline
Software doesn't become valuable when code is generated. It still has to survive:
Coding → Review → Testing → Merge → Deployment → Production → Maintenance
This sounds obvious, but many AI productivity dashboards effectively stop at the first step.
Imagine an AI assistant reduces coding time by 30%.
Great.
Now imagine that the resulting pull requests are larger, review takes 25% longer, and rework increases.
Did productivity improve?
Maybe.
Maybe not.
It depends on where the time went.
This is the first measurement principle I think engineering organizations need in the AI era:
A local productivity gain is not necessarily a system productivity gain.
Software engineering is a pipeline. Accelerating one stage can expose or create a bottleneck somewhere else.
Think About AI as an Intervention in a Queueing System
Consider a simplified engineering workflow:
Work Item
↓
Coding
↓
Pull Request
↓
Review
↓
Testing
↓
Deployment
↓
Production
Suppose AI dramatically increases the rate at which developers produce pull requests.
The arrival rate into code review increases. But reviewer capacity hasn't changed. You can end up with something like this:
| Metric | Change |
|---|---|
| Coding Time | -25% |
| PR Throughput | +30% |
| PR Pickup Time | +18% |
| Review Time | +27% |
Looking only at coding activity makes AI look highly successful.
Looking at the entire system tells a more complicated story.
The bottleneck moved.
This is not necessarily a failure of AI. In fact, it may mean the AI tool is doing exactly what it should.
The organization simply hasn't adapted the rest of its engineering system to the new throughput.
This Is Why "AI-Generated Code %" Is Both Useful and Dangerous
I actually think AI code attribution is becoming an important engineering signal. But not for the reason people sometimes assume.
If 60% of a team's changes are AI-assisted, that does not mean the team is 60% more productive. It doesn't even mean those developers saved 60% of their coding time. What it gives us is something much more useful:
An analytical dimension.
Now we can ask:
How do highly AI-assisted changes behave compared with less AI-assisted changes?
For example:
| Dimension | Question |
|---|---|
| Coding | Are AI-assisted changes completed faster? |
| PR Size | Are AI-assisted PRs larger? |
| Review | Do they require more review time? |
| Rework | How much code changes again shortly after merge? |
| Quality | Do quality issues change? |
| Delivery | Does lead time improve? |
| Stability | What happens to failed changes? |
AI contribution becomes useful when it is joined with engineering outcomes.
By itself, it is just another activity metric.
The Most Important Metric May Be Where the Work Moved
One of the mistakes we made with traditional developer productivity metrics was assuming visible activity corresponded closely to useful work: commits, lines of code, tickets closed, pull requests created.
AI makes that assumption even more dangerous because producing artifacts is becoming dramatically cheaper.
The question therefore changes from:
How much did developers produce?
to:
Where did engineering effort move?
This is closely related to a broader shift in software engineering measurement: activity metrics become far more useful when they are interpreted alongside flow, quality, delivery, and developer experience. I explored that broader measurement model in a separate guide to software engineering metrics, including why isolated activity counts can create misleading conclusions.
An AI assistant might reduce:
- Boilerplate coding
- Searching documentation
- Writing initial tests
- Creating first implementations
- Repetitive refactoring work
At the same time, it might increase:
- Verification
- Code review
- Debugging
- Architectural checking
- Security review
- Rework
If 40 minutes disappear from implementation but 25 minutes appear in verification, the productivity gain is not 40 minutes. And if that verification work falls on another developer, looking only at the original developer's metrics will completely miss it.
We Need an AI Engineering Funnel
Instead of a single AI productivity metric, I find it more useful to think of measurement as a funnel.
1. Exposure
Who can use AI?
Examples:
- Licensed users
- Eligible developers
- Available AI tools
2. Adoption
Who actually uses it?
Examples:
- Active AI users
- Weekly AI usage
- Tool adoption by team
- Model adoption
3. Contribution
Where does AI participate in engineering work?
Examples:
- AI-assisted changes
- AI-assisted commits
- AI-heavy pull requests
- AI-assisted development rate
4. Flow
What happens to development?
Examples:
- Coding Time
- PR Cycle Time
- PR Pickup Time
- Review Time
- Throughput
- Work Item Cycle Time
5. Quality
What happens after that code is created?
Examples:
- Rework
- Reverts
- Defects
- Maintainability issues
- Security findings
- Test failures
6. Delivery
Does the organization ship differently?
Examples:
- Change Lead Time
- Deployment Frequency
- Change Fail Rate
- Recovery Time
- Deployment Rework
7. Outcome
Did something economically meaningful change?
Examples:
- Engineering capacity
- Delivery predictability
- Customer outcomes
- Engineering cost
- AI cost
- Time to market
The further down this funnel you go, the closer you get to actual organizational impact.
The downside is that attribution becomes harder.
That is unavoidable.
There Probably Isn't One "AI Productivity Number"
Executives understandably like summary numbers. But I would be cautious about creating something like:
AI Productivity Score: 83
unless everyone understands exactly what went into it.
AI affects multiple dimensions that can move in opposite directions. For example:
| Metric | Change |
|---|---|
| Coding Time | -21% |
| PR Throughput | +17% |
| Review Time | +14% |
| Rework | +9% |
| Deployment Frequency | +8% |
| Change Fail Rate | +3% |
| Developer Satisfaction | +16% |
Is AI working?
That is a much more interesting engineering discussion than whether a score changed from 76 to 83.
It also exposes something important:
Productivity is multidimensional.
A productivity improvement may appear as:
- Faster delivery
- Higher quality
- Less cognitive load
- Better developer experience
- More capacity for previously neglected work
Trying to compress all of that into one number can destroy the information engineering leaders actually need.
Controlled Experiments and Production Systems Tell Different Stories
This is another reason the debate around AI productivity often becomes confusing.
Some controlled experiments have shown large improvements in task completion speed when developers use AI coding assistants.
Other real-world studies involving experienced developers working in mature repositories have found smaller gains, no gains, or even temporary slowdowns.
Those results sound contradictory.
They aren't necessarily.
They measure different environments.
A bounded programming task is different from changing a mature production system containing:
- Undocumented architectural decisions
- Historical tradeoffs
- Dependencies
- Internal conventions
- Operational constraints
- Domain knowledge
- Security requirements
- Legacy systems
This makes universal statements like:
"AI makes developers 30% faster"
almost meaningless without context.
The better question is:
Which developers, performing which tasks, in which codebases, using which AI tools, measured at which part of the delivery system?
Compare Cohorts Instead of Company-Wide Averages
Suppose an organization sees:
- AI adoption: 72%
- PR cycle time: -11%
- Deployment rate: +14%
It is tempting to connect those numbers. But many things may have changed simultaneously.
A better analysis starts segmenting.
By Repository
AI may be extremely effective in one codebase and much less useful in another.
By Type of Work
Feature development, maintenance, testing, refactoring, and incident fixes are different activities.
By Team
Different engineering practices can dramatically change AI outcomes.
By AI Contribution
Compare lower and higher AI-assisted work.
By PR Size
Otherwise, AI-heavy work may simply be larger or smaller.
By Time
Compare stable periods rather than only the week immediately before and after rollout.
The goal isn't to manufacture causality.
The goal is to eliminate obviously misleading comparisons.
Correlation Is Still Useful
Engineering analytics sometimes falls into another extreme.
If we cannot prove causality perfectly, some people conclude we shouldn't measure the relationship at all.
I disagree.
Observational engineering data can still reveal useful patterns.
Suppose high-AI pull requests repeatedly show:
- Coding Time ↓
- PR Size ↑
- Review Time ↑
- Rework ↑
That doesn't prove AI caused the pattern. But it gives an engineering leader a very useful hypothesis:
Maybe our AI-enabled development workflow needs smaller pull requests or stronger automated validation.
You can now change the system and observe what happens.
Measurement becomes an improvement loop rather than a performance judgment.
That's much more useful.
Measure AI at the Team Level Before the Individual Level
AI telemetry makes extremely granular measurement possible. That doesn't mean every possible metric should become a management metric.
I would be especially cautious about metrics such as:
- AI-generated lines per developer
- Prompt count per developer
- Commits per developer
- Acceptance-rate rankings
- AI productivity leaderboards
Once people realize a metric affects how they are evaluated, the metric stops behaving like neutral telemetry.
It becomes a target. And targets get optimized.
Instead, start with teams and workflows.
Ask:
Is this team's engineering system improving?
Then use deeper data diagnostically when the team needs to understand why.
The objective should be improving the engineering system, not building a more sophisticated surveillance system.
Every Speed Metric Needs a Counter-Metric
This is perhaps the simplest practical rule.
Whenever AI appears to improve one metric, place a balancing metric beside it.
| If you measure... | Also measure... |
|---|---|
| Coding Time | Rework |
| PR Throughput | Review Time |
| PR Size | Review Load |
| Deployment Frequency | Change Fail Rate |
| AI Contribution | Quality |
| Cycle Time | Developer Experience |
| AI Cost | Delivery Outcome |
Why?
Because optimization in engineering frequently transfers cost.
A team can increase deployment frequency by making smaller deployments.
Good.
Or by bypassing necessary controls.
Not good.
The number alone can't tell you which happened.
What I Would Put on an AI Engineering Dashboard
Not 40 AI metrics. I'd start with something much smaller.
Adoption:
Active AI Developers
Are people actually using the tools?
Contribution:
AI-Assisted Development Rate
Where is AI participating in actual engineering work?
Flow:
PR Cycle Time
Is work moving through development faster?
Review:
Review Time
Did the bottleneck move downstream?
Quality:
Rework Rate
Are we creating additional follow-up work?
Delivery:
Change Lead Time
Is local acceleration reaching production?
Stability:
Change Fail Rate / Deployment Rework
Are faster changes remaining reliable?
Experience:
Developer Perception
Do developers actually feel that the workflow improved?
Economics:
AI Cost
What are we paying to create those changes?
That's already enough to have a substantially better conversation.
The Question Has Changed
A few years ago, the interesting question was:
Can AI write useful production code?
Then it became:
Will developers adopt AI coding assistants?
Both questions are becoming less interesting.
The next question is harder:
What happens to a software engineering system when AI becomes a normal participant in development?
An AI tool's usage dashboard can't answer that question. And it cannot be answered by traditional engineering metrics alone.
You need both.
AI attribution without engineering outcomes tells you what AI did. Engineering outcomes without AI context tell you what changed. Connecting the two is where we begin to understand impact. And that missing middle may turn out to be the most important part of measuring software engineering in the AI era.