TLDRocket
Sign in

AI Is Making Activity-Based Engineering Metrics Obsolete

Bharat Sharma Covered by 7 sources

AI tools are skewing engineering metrics. Commit counts and PRs now hide more than they reveal.

Based on reporting by Bharat Sharma — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AI hasn’t made engineering metrics obsolete. It has made the wrong ones dangerous. The old habit of using commits, PR counts, and story points as a proxy for productivity was shaky before; with code generation in the loop, it starts to look actively misleading. The useful question is no longer “how much did engineers do?” but “what changed in the system?”

The evidence is messy, and that mess is the point. A 2026 Halkwinds survey of 758 engineering organizations found 76% had rolled out at least one AI coding assistant organization-wide, up from 41% in 2024. Yet only 34% could show a measurable, audited change in delivery metrics. Most teams have adopted the tool. Far fewer can tell whether it helped.

The studies on output split hard depending on context. METR’s 2025 randomized controlled trial put 16 experienced open-source developers on tasks in repositories they already knew, and AI made them 19% slower. By contrast, field experiments at Microsoft, Accenture, and a Fortune 100 company, covering 4,867 developers, found 26.08% more completed tasks. An observational study of 16,223 developers found up to 40.5% more PRs among heavy GitHub Copilot users, but that’s a self-selected group, so treat it as an upper bound, not a forecast.

That spread is not a contradiction. It says task type, experience, and codebase familiarity matter more than the marketing slide does. Small senior teams deep in legacy systems should expect friction. Larger teams doing more boilerplate-heavy work should expect gains in raw output. And if your scorecard can’t tell those apart, it’s not telling you anything useful.

The bigger problem is that AI pushes work downstream. GitClear’s analysis of 211 million lines of code from 2020 to 2024 found duplicate code overtaking moved code for the first time in its dataset. Veracode’s 2025 security report found 45% of generated samples failed security tests. Apiiro’s telemetry from Fortune 50 repositories found AI-assisted developers committing three to four times faster, with critical architectural flaws rising and privilege escalation paths up 322%. The machine writes faster. Humans then pay for it in review, validation, rework, and cleanup.

So the measure has to change. Keep delivery outcomes. Drop the vanity proxies. Track review queue time, short-horizon rework, validation failures, and post-deployment fixes against a pre-AI baseline. Keep it at team level, not individual level, unless the goal is to invent a leaderboard and watch it get gamed. AI is an amplifier, not a miracle. If the org was weak before, it will just get noisy faster.

My take — AI-written commentary, not fact-checked reporting

The board-pack obsession with commits and PR counts should probably be buried without ceremony. AI didn’t create that bad habit, it just made the lie easier to see. The smarter move is boring and slightly annoying: measure what survives review, not what was spat out in a hurry.

Read more about this at: Bharat Sharma

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.