What We Accept From AI That We Would Never Accept From an Offshore Vendor

A member of my team recently told me something he was genuinely excited about. He had been using Claude to write code, then asking Claude to review the code it had just written. The review found bugs. To him, that was evidence that the workflow was working. And to be clear, it was working. The bugs were found before production. The code improved. That is better than shipping defects and discovering them later.

But I asked him a different question: Would you be equally impressed if our offshore vendor wrote the code, introduced the bugs, and then proudly told us they had found those same bugs during their own review?

The answer is probably no. We would be glad the defects were caught, but we would not describe the process as exceptional. We would ask why the code contained those defects in the first place. We would want to understand whether the requirements had been misunderstood, whether the implementation had been rushed, whether the test strategy was weak, or whether the first review had simply become another stage of development.

Most importantly, we would count the full cost of the loop. That conversation stayed with me because it exposed a standard we have quietly changed. We are accepting behavior from AI coding tools that we would have treated as a delivery problem from an offshore vendor, consultant, systems integrator, or even another internal team.

This is not an argument that AI-generated code is bad. Human developers write bugs. Vendors write bugs. Internal teams write bugs. Every meaningful software process includes review, testing, correction, and learning. The problem is not that AI requires those things. The problem is that we often exclude them when describing AI’s productivity. We celebrate that the code appeared in four minutes. We do not always count the hour spent understanding it, the second prompt used to repair it, the tests added after the fact, the architectural cleanup, the extra review cycle, or the senior engineer who had to decide whether the result was safe to own.

If a vendor operated that way, all of those activities would count against the engagement. They would appear in delivery metrics, defect reports, invoices, retrospectives, and executive conversations. With AI, the same work is often treated as normal overhead surrounding an impressive generation event.

We have replaced time to trusted software with time to first draft. Those are not the same thing.

How We Used to Judge Engineering Partners

Engineering leaders have spent decades learning how to evaluate external delivery partners. We did not simply ask how quickly they could produce code. We asked whether they understood the business problem, whether they clarified ambiguity, whether their design fit the existing system, whether they handled failure paths, whether their tests challenged the implementation, and how much attention our own senior engineers had to provide before the work became production-ready.

A vendor could generate an enormous amount of code and still be a poor partner. In fact, excessive output was often a warning sign. More code meant more surface area to review, more places for defects to hide, and more software our internal teams would eventually need to maintain.

The metric that mattered was never raw output. It was the total cost required to reach an outcome we were willing to own. When a vendor misunderstood a requirement, we counted the rework. When a pull request took three review cycles, we counted the review effort. When weak tests allowed defects to escape, we counted the operational impact. When our architects had to redesign the solution after delivery, we counted the senior engineering time.

We understood that cheap construction could become expensive software. AI has not made that principle obsolete. But our language around AI often acts as though it has.

A New Kind of Supplier Without Accountability

There is one important difference between AI and an offshore vendor: AI is not actually a supplier. It cannot own a deadline, accept contractual accountability, absorb rework costs, support production, or explain itself in an escalation meeting.

That should make us more careful about how we measure it, not less. When a vendor delivers weak work, there is a party outside the organization that can be held accountable. When AI delivers weak work, the cost is absorbed almost entirely by the engineering team. Developers review it. Senior engineers interpret it. Architects correct it. Quality engineers validate it. Operations teams live with whatever survives.

AI did not eliminate the supplier’s work. It eliminated the supplier we could hold accountable. That creates a strange economic effect. Generation looks inexpensive because many of the downstream costs are paid in internal labor. The tool produces the draft. The organization pays the verification tax.

This is especially easy to miss because the verification work often feels productive. The developer is learning. The code is improving. The AI may even identify mistakes that a human reviewer would have missed. All of that can be valuable, but valuable work is still work. If an AI system writes a bug and then finds that bug, the second action does not erase the cost of the first. It may prevent a worse outcome, but it still represents another loop required to reach trusted software.

The Same Mistake, Evaluated Differently

Consider how differently we react to familiar delivery problems. AI misunderstands the requirement, and we say the prompt needed more detail. A vendor misunderstands the requirement, and we question whether the team understands the business. AI misses edge cases, and we run another review. A vendor misses edge cases, and we question the rigor of its testing. AI produces code that does not fit the architecture, and we add more repository context. A vendor does the same, and we ask why its architects were not aligned with ours. AI writes weak tests, and we ask it to regenerate them. A vendor writes weak tests, and we treat it as a quality issue. AI generates a large pull request that takes significant effort to understand, and we call the output impressive. A vendor sends the same pull request, and we ask why the change was not decomposed into something reviewable.

Replace the word AI with the name of your offshore partner and the excuses stop sounding innovative very quickly. The point is not that AI should never make mistakes. The point is that our response to those mistakes should be consistent. We should evaluate the complete workflow rather than isolating the most flattering part of it.

Code Is Not the Product

The fascination with generation speed reflects an old mistake in software engineering: confusing activity with value. We have measured lines of code, story points, velocity, ticket counts, and utilization. Each metric offered a convenient way to describe work while avoiding the harder question of whether the software actually improved the business.

AI code generation may be the newest version of that mistake. A thousand lines generated in minutes sounds impressive because the result is visible and immediate. But code is not the product. A working, maintainable, secure system that solves the intended problem is the product.

The distance between generated code and trusted software can be small. It can also be enormous. That distance is where review, testing, context, architecture, judgment, and accountability live.

Those are not secondary activities surrounding software development. They are software development. This is also why self-reported productivity can be misleading. Developers may genuinely feel faster because AI removes the discomfort of starting from a blank page. Editing a plausible solution can feel easier than constructing one. But easier is not always faster, and faster is not always cheaper once the full delivery loop is measured.

The Correct Standard Is Not Perfection

Holding AI to the same standard as an engineering partner does not mean demanding perfect code from the first prompt. We did not demand perfection from human teams either.

Good vendors asked questions. Good developers iterated. Good engineering organizations expected defects to be caught through layered controls. The standard was not perfection; it was disciplined delivery. The question was whether each stage added confidence rather than merely moving uncertainty downstream.

That is the standard AI-assisted development should meet. Did the tool help clarify the requirement, or did it simply implement an assumption faster? Did it generate tests that meaningfully challenged the solution, or tests designed to confirm its own interpretation? Did it produce code that fits the system, or code that happens to compile in isolation? Did it reduce the total effort required to reach production, or merely shift work from implementation into verification?

Those questions are more difficult than measuring generation time. They are also much closer to what engineering leaders actually care about.

Time to Trusted Software

The metric I keep coming back to is time to trusted software—not time to code completion, time to first pull request, tokens consumed, lines generated, or how quickly a demonstration can be produced. How long did it take to reach software that the team was willing to deploy, operate, explain, maintain, and change?

That measure includes generation, but it also includes everything generation creates around it: clarification, review, correction, testing, integration, security analysis, architecture alignment, and operational confidence.

It also creates a fairer comparison. AI may still win dramatically. In many situations, it almost certainly will. Boilerplate, transformations, repetitive integration work, test scaffolding, documentation, and well-defined implementation tasks can all move faster with AI.

But when AI wins under a complete measure, the productivity gain is real. We are no longer giving it credit for the first draft while hiding the rest of the work in the engineering team’s calendar. That distinction matters because organizations are making staffing, budgeting, and delivery commitments based on assumed productivity gains. If the measure is incomplete, the decisions built on top of it will be incomplete too.

Optimism Requires a Higher Standard

I am optimistic about AI in software development. My teams use it. I use it. I believe it will reshape how software is designed, built, tested, and operated.

But optimism should not require us to forget what we already know about managing engineering work. We know that output is not the same as value. We know that review effort is real effort. We know that rework is part of cost. We know that software is not complete when it compiles. We know that the person or team ultimately responsible for production must understand and trust what has been delivered.

AI does not change those principles. It makes them more important because it can create plausible output faster than organizations can evaluate it. My team member was right to be pleased that Claude found bugs before the code moved forward. That is a useful capability, and it may become an essential part of the development workflow.

But the better management question is still the one I asked him: Would we be equally impressed if our offshore vendor had done the same thing? If the answer is no, then the issue is not whether AI is good enough. The issue is why we changed the standard.

Continue Exploring

METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (2025). In a randomized study of experienced developers working in mature codebases, participants expected AI to make them faster but completed assigned tasks more slowly when AI tools were allowed.

Ruofan Gao, Amjed Tahir, Peng Liang, Teo Susnjak, and Foutse Khomh, “A Survey of Bugs in AI-Generated Code” (2025). A review of research on the categories, causes, and remediation of defects found in generated code.

Florian Tambon and colleagues, “Bugs in Large Language Models Generated Code: An Empirical Study” (2024). The authors identify recurring failure patterns including requirement misinterpretation, missing corner cases, hallucinated objects, incomplete generation, and prompt-biased code.

Suzhen Zhong, Shayan Noei, Ying Zou, and Bram Adams, “Human-AI Synergy in Agentic Code Review” (2026). An analysis of code-review conversations found that human reviewers contributed more contextual, testing, and knowledge-transfer feedback, while AI suggestions were adopted less often and required continued human oversight.

Anthropic, “Code Review for Claude Code” (2026). Anthropic’s review tooling reflects the emerging reality that faster code generation increases the need for scalable validation before changes can be trusted.

Previous
Previous

Stop Rolling Your Eyes at Ted Lasso — He’s Teaching Agile for Free

Next
Next

The Difference Between Finished and Done