Back to Blog
Engineering
By Anil Konur
July 31, 2026

Your Engineering Team Is Seeing 20% Gains from AI. The Teams at 3x Are Doing Something Different.

A synthesis of AI engineering productivity data across randomized controlled trials and company disclosures shows three tiers of outcome - and the gap between 20% and 3x is not the model. It is the operating model built around it.

Tomasz Tunguz published a synthesis this week that every SaaS engineering leader should read carefully.

He mapped the growing body of AI engineering productivity data - across randomized controlled trials, company disclosures, and published studies - and found that the outcomes are not random. They cluster into three distinct tiers.

The mean tier delivers 20-46% productivity gains. These are the teams that distributed an AI IDE and measured the result.

The frontier tier delivers 2.5-3x gains. These are the companies that built an operating layer around AI agents. NVIDIA reported a 3x increase in committed code across 30,000 developers with bug rates flat. Amplitude tripled weekly production commits, with an AI agent now a top-three contributor to the codebase. Anthropic measured a 2.5x increase in code written per engineer since adopting Claude Code internally.

The factory tier delivers 8x or more. These are the organizations where agents operate as first-class organizational units - Nubank with Devin is the documented example.

Tunguz's conclusion is the one that matters: the gap between 20% and 3x is not the model. It is the operating discipline built around it.

The Konur Consulting take: Most SaaS engineering leaders have deployed AI coding tools and are watching 20% gains. They are not behind on tooling. They are behind on operating model. The path from 20% to 3x is defined delegation policies, measurement frameworks, and the organizational structure to act on what agents produce.

What the frontier teams are actually doing differently

The companies reporting 3x gains did not get there by deploying a better tool. They got there by making a set of organizational decisions that most engineering teams have not made yet.

They defined what agents are authorized to do without human review. At NVIDIA, with 30,000 developers, the decision about which categories of engineering work can be delegated to agents - and which require a human checkpoint - is not being made informally, engineer by engineer. It is a policy. The policy is what makes 3x sustainable rather than a single-sprint experiment.

They built the measurement infrastructure before they scaled. Amplitude's metric - AI agent as a top-three contributor to the codebase - only means something if someone designed a way to track agent contribution separately from human contribution. Most engineering teams have not done this. They are measuring productivity the same way they did before AI, which means they cannot see what agents are contributing or where the bottlenecks are.

They restructured capacity planning around what agents can absorb. The teams at 3x are not simply adding AI tools to an unchanged team structure. They have re-examined which tasks humans do and which agents do, what the right team size is given agent capacity, and how sprint planning and backlog grooming change when agents can execute bounded tasks in parallel with human developers.

The operating model gap is where the 20% teams are stuck

If you deploy an AI IDE and measure the result, you will see something in the 20-46% range. That is the tool benefit - faster autocomplete, better suggestions, less time spent on boilerplate.

The tool benefit is real. It is also the floor, not the ceiling.

To get from 20% to 3x, you need to answer questions that are not engineering questions. They are operating model questions.

Which tasks are appropriate for agent execution without human review at each step? This is a delegation policy decision. It requires a definition of "bounded, testable, reviewable" that is specific enough for agents to execute and for humans to trust. Most teams have not written it down.

How do you measure agent contribution accurately? Velocity metrics built on human developer output - story points, PRs per engineer, lines reviewed - break when agents contribute. The denominator changes. If you are still measuring the same way you did before AI, you cannot see the 3x opportunity or understand why you are not capturing it.

How does offshore and nearshore team structure change when agents can absorb the high-volume, bounded tasks that previously required human developers? The teams using AI to augment offshore capacity - rather than simply replacing it - are getting compound leverage. The ones that have not rethought their offshore structure are paying for human capacity that agents could cover.

What is your governance model for agent-produced code? Agents make mistakes. When an agent-generated PR introduces a bug, what is the review protocol? Who is accountable? How does that differ from the review protocol for human-generated code? These are not hypothetical questions. They are questions that teams at 3x have answered, and teams at 20% have not.

What to do Monday

Run an honest audit of where your engineering team sits in the three tiers. If you distributed an AI IDE and are seeing 20-46% gains, that is the baseline - not the destination.

The next step is not a better tool. It is a delegation policy. Write down - in one page - which categories of engineering work agents can own without human review at each step, which require a checkpoint, and who has authority to expand or contract those boundaries as agents demonstrate reliability.

Then add agent contribution tracking to your output measurement before the next capacity planning cycle. Decide how you will attribute and measure what agents contribute alongside human developers. Without this, you cannot make a data-driven case for the operating model changes that get you from 20% to 3x.

Finally, if you run offshore or nearshore engineering capacity, schedule a structural review. The team that was right for pre-agent delivery is not necessarily right for agent-augmented delivery. The task intake process, the QA workflow, and the communication cadence all change when agents can execute bounded work in parallel.

FAQ

Is 3x realistic for a team that is not NVIDIA or Amplitude?

The 3x tier is documented at companies of various scales, not just large enterprises. The common thread is not company size - it is the presence of explicit delegation policies, measurement infrastructure, and restructured workflows. A 20-person engineering team can reach 3x if it makes the organizational decisions. A 500-person team will stay at 20% if it only deployed the tool.

What is the risk of moving too fast on agent delegation?

The primary risk is quality degradation that is not caught until it is in production. The mitigation is the same as for any delegation decision: define the scope tightly, build checkpoints at the boundaries, and expand scope as reliability is demonstrated. The teams at 3x did not get there by delegating everything at once. They got there by expanding the boundary systematically.

How long does it take to move from 20% to 3x?

The companies in the data set did not publish timelines. Based on the structural changes required - delegation policy, measurement redesign, capacity reallocation - a realistic horizon for a team making the full set of changes is 2-3 quarters. Teams that only address one or two of the levers will see incremental improvement but not the step-change.

The tool is deployed. The operating model is what is missing. That gap is where the difference between 20% and 3x lives.

Konur Consulting helps SaaS engineering organizations build the operating model infrastructure for agent-augmented delivery - from delegation policy to measurement frameworks to offshore team restructuring. Reach out at info@konurconsulting.com.


Source - AI engineering productivity synthesis: Tomasz Tunguz, "AI Engineering Productivity Is Anything But Normal," tomtunguz.com, July 27, 2026. tomtunguz.com/ai-engineering-productivity-anything-but-normal

Source - NVIDIA productivity data: Cursor, "How NVIDIA uses Cursor," February 2026. 3x committed code across 30,000 developers, bug rates flat. Cited in Tunguz.

Source - Amplitude data: Cursor, "Amplitude and Cursor cloud agents," April 2026. 3x weekly production commits, AI agent as top-three codebase contributor. Cited in Tunguz.

Source - Anthropic internal data: Boris Cherny, head of Claude Code, Big Technology podcast, July 2026. 2.5x code per engineer. Cited in Tunguz.

Source - Engineering productivity landscape: Faros AI Engineering Report 2026. Cited in Tunguz as corroborating source for the productivity tier framework.