Token Hell: Why AI Coding Tools Aren't Moving Your Engineering Margins

Murali Vivekanandan
Maheshwari Vigneswar

TL;DR

  • Using AI coding tools isn't what matters, instead tracking their usage is. Paying for AI licenses without measuring adoption or ROI is Token Hell.
  • The gap sits in the review layer. AI drafts faster, but requirements, review, and handoff still move at the old pace, staffed by the same number of people doing the same checking work.
  • Enterprise-scale AI-native delivery needs three things at once: a standardization and visibility platform, a redesigned Agile process, and training. Tool access alone substitutes for none of them.
  • The number that predicts margin movement is engineer-by-engineer AI impact: usage, output, and quality scored together, not usage alone.

Table of Content

Most engineering leaders run a version of the same experiment. License an AI coding assistant, roll it out to the team, and wait for velocity to show up in the numbers. Token spend climbs every quarter. Whether that spend is producing anything is a different question, and most organizations cannot answer it.

I have watched this pattern repeat from early-stage startups to Fortune 500 engineering orgs. You keep hearing Copilot this, Claude that, every day, but gross margin on the projects has not shifted one percent.

The reason is structural. The AI assistant produces a faster first draft, but the product manager still writes the story the old way, the developer still submits it into the same review process, and the same number of people do the same checking work. MIT NANDA's 2025 study on enterprise AI adoption found that 95 percent of generative AI pilots failed to produce measurable financial impact, and traced it back to workflows that were never redesigned around the tool. The pattern traces back to the same root cause every time: nobody set up a way to measure what the AI usage was actually returning before scaling it across the org. Access to the tool was never the constraint. Visibility into what the tool was doing was.

Who Gets Trapped in “Token Hell”? 

There are three groups in this market, and only one of them has a problem.

An organization that has not adopted AI coding tools at all is fine. There is no wasted spend, and whatever the delivery process's limits are, they are at least a known quantity. An organization running AI at true scale, with usage measured, benchmarked, and controlled, is also fine. That organization did the harder work and the numbers reflect it.

The group stuck in Token Hell is everyone in between. Licenses are issued. Token spend climbs. Nobody can say with precision which engineers are getting real output from AI and which are generating drafts that get rewritten anyway, what proportion of spend is going to a model that costs four times more than necessary for the task, or whether AI-assisted work is actually landing faster once review time is counted.

That inability to answer those questions, not the token spend itself, is the actual disease. An organization that cannot build an ROI story for its AI spend is negotiating every renewal on faith instead of evidence.

Camp Status Why
Not using AI coding tools at all Fine No wasted spend. The delivery process has known limits, but they're a known quantity.
Using AI at true enterprise scale, measured and controlled Fine The harder work was done up front. The numbers reflect it.
Licensed, adopted, not measuring Stuck in Token Hell Token spend climbs every quarter. Nobody can say with precision what it's buying.

What Actually Causes Token Hell?

Here are the reasons of what actually causes token hell.

AI Usage Is Measured instead of AI Output

Most orgs can tell you how many seats are licensed. Almost none can tell you which of those seats are producing work that ships clean versus work that comes back for two rounds of rework.

Review and Correction Absorb the Productivity Gain

The AI-generated draft still has to be checked. That checking time is real cost, and it's the cost most productivity claims quietly leave out.

Every Engineer Uses AI Differently

Without a shared standard, one engineer's AI-assisted output holds up in review and another's doesn't, and a usage percentage alone can't tell you which is which. Two engineers can show up identically in an adoption report and land in completely different places on actual delivered output.

Token Spend Grows Without a Clear ROI Model

No measurement means no ROI story. No ROI story means every renewal conversation about AI tooling spend happens on faith instead of evidence.

What Does an Enterprise AI-Native SDLC Actually Require? 

Individual developer productivity and organizational delivery velocity are different problems. Moving the second requires three things together, not any one alone.

Platform: Standardize and Measure AI Usage

Shared coding patterns, tools, and libraries, plus telemetry on what AI usage is actually producing engineer by engineer and dollar by dollar. Without it, every AI agent writes against a different convention and nobody can say what the spend is buying.

Process: Redesign Agile Around AI Participation

Structured work, specs, test cases, risk flags, goes to AI agents directly, with humans at defined checkpoints. Without it, the story still passes through a human before it reaches an engineer, and the old bottleneck stays with a tool bolted on top.

Training: Teach Teams How to Work With AI Agents

How to split work for an AI agent, and what belongs in config files versus a skill versus a manually invoked tool. Without it, every developer runs a different workflow, and output stays inconsistent even when platform and process are right.

Platform + Process + Training. None works independently. Platform without process still routes every draft through the old review queue. Process without training produces agents set up badly. Training without platform has no governed environment to run in.

Layer What It Does What Happens Without It
Platform Standardization and visibility: shared coding patterns, tools, and libraries, plus telemetry showing what AI usage produces engineer by engineer and dollar by dollar. Every AI agent writes against a different convention. Nobody can say what the spend is buying.
Process Agile redesigned so structured work-specs, test cases, and risk flags goes directly to AI agents, with humans at defined checkpoints. The story still passes through a human before it reaches an engineer. The old bottleneck stays, with a tool bolted on top.
Training How to split work for an AI agent, and what belongs in config files versus a skill versus a manually invoked tool. Every developer runs a different workflow. Output is inconsistent even when the platform and process are right.

Skip any one and the other two can't compensate. Platform without process still routes every draft through the old review queue. Process without training produces agents set up badly. Training without platform has no governed environment to run in.

How to Identify and Quantify Token Hell?

Getting out of Token Hell is not one capability. It is five separate questions an engineering org needs answered, and most tools on the market only answer one of them.

Engineering intelligence: This is the baseline question: where is your SDLC actually getting stuck, from the first prompt to production? Answering it means combining AI-specific telemetry, commits, pull requests, reviews, deploys, with established frameworks like DORA and SPACE into a single view, rather than reading AI adoption off of anecdote or a survey your team fills out once a quarter. Without this layer, "we're using AI" and "we know where AI is helping" remain two completely different claims.

Token intelligence: This is the direct answer to Token Hell. It means knowing where every token goes: which models are consuming the spend, what proportion is going to a frontier model for tasks a cheaper model would have handled just as well, and how your return on that spend compares to benchmarks across other engineering organizations. Most teams can tell you what they're spending on licenses. Almost none can tell you what they're getting per dollar of token spend, or whether that ratio is improving or quietly getting worse.

Prompt routing: Once you can see where tokens go, the next question is whether they're going to the right place. A structural edit does not need a frontier model. A prompt router classifies each request and sends it to the model with the best quality-per-token for that specific task, learning from feedback at the individual and organizational level over time. This is one of the few token-cost levers that reduces spend without asking a single engineer to change how they work, since the routing happens transparently underneath whatever client they're already using.

Code intelligence: Faster output that is worse code is not a win, and an org that only measures speed will not catch that trade-off until it shows up in production. Code intelligence closes the loop by scoring the output itself, not just the volume of it, tracking code quality and review outcomes alongside code output so that an increase in AI-assisted merges is a real signal and not just a warning sign in disguise.

Model observability: This is the layer that makes the other four trustworthy at an enterprise level: security, compliance, and audit-readiness designed in rather than bolted on. Without SOC 2, HIPAA, and GDPR compliance, SSO, and role-based access built into how the measurement itself works, none of the numbers above can be handed to a board or an auditor with confidence.

Token intelligence without code intelligence tells you spend is dropping without telling you whether quality dropped with it. Engineering intelligence without a prompt router tells you where the bottleneck is without giving you a lever to pull on it. All five have to operate on the same underlying data for the picture to be one a CFO or a board can act on.

The Maker-Checker Model: How Should AI Participate in the Software Development Lifecycle?

One AI agent writes the artifact, a spec, a test suite, a first code pass. A second agent verifies it. Only then does it reach a human product manager or tech lead, who still owns intent, prioritization, and the final go or no-go. What changes is the volume of first-draft work that arrives already checked once.

This is the same maker-checker logic banks have used for decades in financial controls. It doesn't remove human judgment, it relocates it, out of the checkpoints where a human was just re-doing work an AI could draft correctly, and into the checkpoints where judgment actually matters: plan review, edge cases, production readiness.

Here's the actual process:

  • AI Drafts the First Version - One AI agent writes the artifact, a spec, a test suite, a first code pass.
  • A Second AI Agent Verifies the Output - A second agent checks it before either one reaches a human. This is the same maker-checker logic banks have used for decades in financial controls, applied to software delivery.
  • Humans Own Intent, Judgment, and Final Approval - The human product manager or tech lead still decides intent, prioritization, and the final go or no-go. What changes is the volume of first-draft work that arrives already checked once.
  • Move Human Review to the Decisions That Actually Require Judgment - The model doesn't remove human judgment, it relocates it: out of the checkpoints where a human was just re-doing work an AI could draft correctly, and into the checkpoints where judgment actually matters, plan review, edge cases, production readiness.

How Do You Measure Whether AI Is Actually Improving Engineering Productivity?

Every AI coding tool vendor has a headline productivity number. The number that actually predicts margin movement looks different, and it has to be measured at the individual level before it means anything in aggregate.

Practical lift = the productivity gain from AI − the review and correction overhead it creates.

A team can look substantially faster on raw output. Once the hours spent correcting, re-prompting, and validating are subtracted, the net gain that reaches a sprint is smaller, sometimes dramatically. That gap is exactly where token costs stop translating into anything a CFO can see: the org pays for the draft, then pays again in engineering hours to make it production-ready.

None of this is calculable without measuring first. Most orgs in Token Hell skip straight to optimizing, a new tool, a tighter prompt library, without ever establishing what current usage returns. That sequence runs backward.

An engineer using AI on 72 percent of their work and producing output 31 percent above baseline, at a quality score in the low 90s, is in a completely different position than an engineer using AI on 31 percent of their work with an output gain of 9 percent. Both engineers show up in an org-wide "AI adoption" number as adopters. Only one of them is actually moving the team's delivery numbers. Averaging those two together into a single company-wide productivity claim is exactly how organizations end up unable to explain, a year later, why the AI spend never showed up in the P&L.

This is also why token consumption on its own is a misleading metric. Tokens show consumption, not value. Two engineers can burn the same number of tokens in a month and produce completely different amounts of usable, review-ready work. Scoring cost, efficiency, and quality together, rather than reporting spend as a single number, is what turns "we're using a lot of AI" into a claim you can actually defend.

Why Raw AI Output Is a Misleading Productivity Metric

A team can look substantially faster on raw output. Once the hours spent correcting, re-prompting, and validating are subtracted, the net gain that reaches a sprint is smaller, sometimes dramatically. That gap is exactly where token costs stop translating into anything a CFO can see.

Measure AI Usage at the Engineer and Model Level

An org-wide adoption number hides more than it reveals. Usage has to be measured per engineer and per model before it means anything in aggregate.

Measure Quality and Review Time Alongside Output

Output speed without a quality and review-time check is a trap. None of this is calculable without measuring first, and most orgs in Token Hell skip straight to optimizing, a new tool, a tighter prompt library, without ever establishing what current usage returns. That sequence runs backward.

Signs you're in Token Hell:

  • [ ] You know licenses are issued, but not which engineers are getting real output from AI
  • [ ] You can't say what proportion of token spend is going to a model more expensive than the task needs
  • [ ] You can't say whether AI-assisted work is landing faster once review time is counted
  • [ ] Your only AI metric is a usage or license count, not an output or quality score
  • [ ] "We already have Copilot" is a claim you can make; "we know what Copilot is doing for us" is not

Why a DORA Dashboard or a Survey Doesn't Answer This

Most orgs already have some measurement, a DORA dashboard, a SPACE survey, a quarterly AI adoption pulse check. None of it separates what a human wrote from what an AI generated. A thousand AI-generated lines in ten seconds and a thousand lines a senior engineer wrote over two days look identical in a lines-shipped count, so the volume spike that should raise a flag reads as a productivity win instead.

Building this internally is possible in theory. In practice it means standing up human-versus-AI attribution at the pull-request level and maintaining it as new models ship every few months, on top of the engineering work it's supposed to measure. That's the project that stays half-built while the token bill keeps climbing. It's why I build this in as phase one of an engagement, not something a client stands up before we start.

How Ideas2IT Build This Into an Engagement

This runs through our Forward-Deployed Engineering model: our engineers embedded inside the client's delivery process, working alongside their team on real sprint work, not a training environment disconnected from production.

  • The measurement layer goes in first, before process redesign or training. It answers five questions most orgs have never had answered together.
  • Build the Prompt and Pattern Library Around the Client's Stack: The library gets built from the client's own tech stack during the engagement, not handed over as a generic template.
  • Train Teams Through Real Sprint Work: Our AI-Native Developer Academy sits on top of the measurement layer, built from our own internal rollout across more than 700 engineers: safe usage, prompt structure, review discipline, and the judgment calls behind what gets automated.
  • Leave the Client With an Operating Model They Can Run Independently: By the end of the engagement, the client owns the operating model, the prompt and pattern library, and a trained cohort that runs the process without us in the room.

How Do You Build an AI-Native SDLC That Can Survive Security Review?

A measurement layer asking to see every commit, pull request, and prompt across an org is asking for real access. That access has to be governed the same way any other sensitive system is. Role-based access and audit trails on the measurement layer itself, not just on the codebase it's reporting on.

A measurement layer asking to see every commit, pull request, and prompt across an org is asking for real access. An evaluation that skips the compliance question stalls the moment it reaches security review.

  • [x] SOC 2 Type II
  • [x] ISO 27001
  • [x] HIPAA-compliant delivery
  • [x] AWS Healthcare Competency certified

How Ideas2IT Helps Engineering Teams Move Beyond AI Tool Adoption

Here's how to proceed from tool adoption.

Measure Where AI Is Actually Creating Value: 

The engagement starts by establishing your practical lift baseline, engineer by engineer and model by model, before anything else changes.

Redesign the SDLC Around AI Participation:

Agile gets rebuilt around the maker-checker model, with a decision register that reflects your actual risk posture, not a generic policy.

Train Engineers Through an AI-Native Developer Academy:

Your team goes through the same platform, prompt, and review discipline training built from our internal rollout across more than 700 engineers.

Scale the Model From Pilot Squads to the Engineering Organization:

What's validated in one or two pilot squads becomes the roadmap for the rest of the org, with a trained cohort that can run it without us.

What Does an AI-Native SDLC Operating Model Look Like in Practice?

The measurement layer goes in first, before process redesign or training begins. It answers five questions most orgs have never had answered together.

Question What It Reveals
Where does the SDLC actually get stuck? Telemetry from first prompt to production, so the bottleneck shows up in data rather than a quarterly guess.
Where does every token go? Which models are consuming spend, and whether work is running on a model that costs more than the task requires.
Is the model choice wasting money? Structural, low-complexity work is routed to cheaper models, while frontier models are reserved for tasks that actually need them.
Is the output good, not just fast? Code quality and review outcomes are tracked alongside output, so a spike in AI-assisted merges becomes a signal worth investigating.
Can the numbers survive scrutiny? A compliance foundation is built in from the start, so metrics presented to a board or auditor can withstand scrutiny.

Start With an AI Coding Cost and Productivity Assessment

Here's how the assessment engagement actually looks like.

Measure What You're Spending

The entry point is a token cost-of-ownership assessment: what you're actually paying for AI coding tools, seat by seat and model by model, against what you're getting back once review time is counted.

Identify Where Practical Lift Is Being Lost

Most orgs in Token Hell find the same thing once the numbers are in: spend climbing with no matching lift, and the gap sitting in review and correction overhead rather than in the tools themselves.

Build the Roadmap From the Baseline

If the assessment confirms that gap, the engagement moves into the full redesign: maker-checker process, decision register, training that gets your team running it without us in the room.

If your organization is paying for AI coding tools and cannot say which engineers are getting real value from them, where the token spend is actually going, or whether AI-assisted work is landing faster once review is counted, that visibility gap is Token Hell.

Book a demo to see engineering intelligence, token intelligence, and code quality scored together for your own team.