Most enterprise engineering teams now have AI coding tools in some state of adoption. Copilot licences have been bought, a few developers are running Claude Code or Cursor, and there may be an official pilot. The question I get asked has shifted from "should we use these?" to "why are our results so uneven?"
They are uneven. I've seen teams where AI assistance visibly shortened delivery timelines. I've seen other teams at the same company, on the same licence, where the main result was a 40% jump in pull request volume and a review queue nobody could keep up with. Same tool. Everything around it was different.
What changes when the tools arrive
The naive model is "same process, faster typing." What really happens is that the bottleneck moves. Writing code gets cheaper, so more code gets written, and review, integration and understanding what you built all get relatively more expensive. If you don't adjust, the bottleneck doesn't just move. It backs up.
Three shifts are worth planning for.
First, review load goes up. There are more PRs, and they get bigger if you let them. The senior engineers who do most of the reviewing become the constraint on everything else.
Second, obviously wrong code gets replaced by plausible but wrong code. AI-generated code compiles, passes the happy-path test and reads cleanly. Its bugs are quieter than a rushed human's: wrong assumptions, edge cases just missed, confident use of an API that doesn't quite work that way.
Third, conventions drift faster. The model writes in the style of its training data unless you tell it otherwise. Left alone, every new file brings a slightly different idea of how you handle errors, logging and naming.
What the good rollouts do
Write down how your codebase works, because the tools will read it
The most useful hour in an AI tooling rollout is the one spent writing the conventions file. Call it CLAUDE.md, copilot-instructions or cursor rules, depending on the tool. Put in how you structure modules, how you handle errors, your test conventions, which patterns are deprecated and which internal libraries to use instead of reaching for npm.
Teams skip this because their conventions live in senior engineers' heads and in old review comments. That was fine when only humans wrote code, since humans learn from review. The model doesn't come to your retros. The conventions file is how it learns. Teams with a good one get code that looks like their codebase, and teams without one get code that looks like Stack Overflow.
Hold the line on PR size
AI tools make big changes effortless, and big changes are where review quality goes to die. The teams doing well kept their PR size discipline, and some tightened it, precisely because generating code got cheap. A reviewer can properly check a 200-line change. Nobody properly reviews 2,000 lines, whoever wrote them.
Strengthen CI, because review alone won't catch it
If plausible but wrong is the new failure mode, the safety net has to be mechanical. Teams that improved their test suites, linting and type coverage before scaling up got compounding returns, because the tools write better code when they can run the tests and fix what fails. A weak test suite combined with a lot of generated code is exactly where the horror stories come from.
Measure cycle time, not acceptance rate
Vendor dashboards love acceptance rates and lines generated. Neither tells you whether you're better off. The numbers that matter are the ones that always did: cycle time from first commit to production, change failure rate and time spent in review. If those are improving, the rollout is working. If PR volume is up and cycle time hasn't moved, you've automated the production of work in progress.
Don't let juniors skip the apprenticeship
This is the problem nobody has solved. Juniors with AI help produce senior-looking code, so the old signal that someone needs guidance (rough code) has vanished while the need is still there. Teams handling it on purpose run AI-free code reading sessions, ask juniors to explain every line of their PRs, and pair on debugging in particular, because debugging is where you learn what code really does. None of that makes experience free. The tools amplify judgement. They don't supply it.
How I'd sequence a rollout today
In the first two weeks, write the conventions files and tighten CI where it's weak. Pick one team to go first, ideally one that's keen rather than conscripted.
For the next six weeks or so, that team uses the tools every day. Capture what works in shared prompts and updated conventions. Keep a close eye on review load and adjust PR norms as you go.
After that, expand one team at a time and take the playbook with you. Each new team should inherit the conventions and the habits, not just a licence.
Giving everyone a licence on day one feels decisive. It also produces the uneven results that everyone then blames on the tool.
Where I've landed
I use these tools every day and wouldn't go back. But they're amplifiers, and an amplifier doesn't care what it amplifies. A team with strong conventions, solid tests and disciplined review gets a lot faster. A team without those gets the same chaos, sooner. Most of the rollout work is the dull part around the tool.
If your team's results have been patchy, I can help you work out which piece of the scaffolding is missing.

