Measuring productivity purely by AI token usage misses the point. There are far better metrics to monitor.

Measuring productivity purely by AI token usage misses the point. There are far better metrics to monitor.

(©Karola G – canva.com)

The conversation about AI and enterprise productivity has drifted into a peculiar place. Organizations that spent years refining outcome-based performance frameworks are now, in some cases, reverting to curious metrics as a measure of an employee’s “productivity.” Not how much code they write or how many customer resolutions they manage – purely AI usage time and token consumption. The logic, presumably, is that more AI use equals more value delivered.

This is simply untrue, and two recent examples illustrate the problem well. Starbucks has structured a quarter of tech workers’ bonuses around department-wide AI adoption goals, defined as using an AI assistant multiple times a week. Certainly, the intent is to incentivize adoption – but the effect, if history is any guide, is to incentivize the appearance of adoption.

Separately, Uber’s engineering teams burned through their entire 2026 AI budget by April after ranking engineers on internal leaderboards based on tool usage. This happened because engineers responded to the incentive structure they were provided by using the tools. A lot. Whether the outcomes justified the spend is a different question.

These are not isolated mistakes. They reflect a genuine measurement problem that most enterprises haven’t solved yet – how do you quantify the value of AI-assisted work when the output of that work is often qualitative, contextual, and difficult to compare to a pre-AI baseline?

But many organizations are choosing the wrong approach – measuring inputs instead of results.

Why activity metrics fail for AI

In the early days of knowledge work automation, organizations measured productivity by the number of reports generated, emails sent, or documents processed. These were measurable, and logic dictated that they become the measured metrics du jour. These also had limited correlation with actual business outcomes, which is why most serious organizations ultimately moved away from them.

AI usage metrics occupy the same category of error. Token consumption, time spent in AI tools, and percentage of code written by AI are measures of activity, not productivity. An engineer who uses an AI coding assistant to generate 500 lines of mediocre, buggy code has consumed more tokens than one who used it sparingly to solve a specific architectural problem. Under a token-volume incentive structure, the first engineer looks more productive—and might even skate by a few rounds of layoffs.

The Uber case is instructive. The article states that engineers spend between $500 and $2,000 per person per month on AI tools, with 95% of engineers using those tools every month, meaning the company had achieved near-universal adoption. It had also exhausted its annual budget in four months. Praveen Neppalli Naga, Uber’s CTO, acknowledges the company was heading back to the drawing board, as the tools proved too successful to afford at scale.

The Starbucks approach occupies a different failure niche but shares the same root cause. When 25% of a bonus is tied to whether developers used an AI tool a minimum number of times per week, employees optimize for frequency of use rather than the quality of the work those tools help produce. The metric creates a perverse incentive – using AI inefficiently, and often, is better than using it sparingly and well, because only the former registers on the dashboard. It is worth noting that Starbucks also recently rolled back an AI initiative it had deployed for inventory counting at stores—a reminder that adoption metrics and outcomes are not the same thing.

One CEO recently described to me a pattern that captures the absurdity precisely. His team was producing multi-page reports drafted with AI assistance. He was then using a separate AI tool to compress those reports into single-page summaries. Two AI tools, two sets of tokens, net productivity negative. If you squint hard enough, every step of that workflow resembles AI adoption. None of it would have shown up as a problem in a usage-based measurement system.

The alternative isn’t obvious, but it exists

This is not to say we should abandon all hope of measuring productivity – enterprise leaders are right to want accountability for AI investment. The question is what to measure, and the answer depends on a distinction that most AI governance frameworks haven’t made explicit yet – the difference between AI as a process accelerator and AI as a reasoning partner.

Process acceleration is the easier case. If AI is used to compress a task that previously took four hours into one hour, productivity is straightforward to measure – time saved, cost reduced, throughput increased. This is where AI has delivered the most consistent and verifiable gains, and it’s where output-based metrics work well.

Reasoning partnership is harder to quantify, but it’s where the higher-value applications live. When AI is used to model scenarios, stress-test assumptions, synthesize complex inputs, or generate options that a human then evaluates and decides on, the value isn’t in the AI’s output alone, but rather in the quality of the decisions that result. Token consumption reveals nothing useful in this context.

Organizations that are getting this right tend to share a few characteristics:

  • They define AI success at the business-unit level before deploying tools, identifying specific outcomes they want to improve and how they’ll know if they have.
  • They treat AI usage data as a diagnostic rather than a performance indicator — high usage in a role where outcomes haven’t improved is a flag, not a badge.
  • They distinguish between roles where efficiency gains are the primary value and roles where judgment quality is the primary value, applying different measurement frameworks accordingly.

Cultures of experimentation vs. quota-driven adoption

There’s a meaningful difference between organizations that are building cultures of AI experimentation and those that are driving adoption through quotas and incentives. Experimentation cultures give employees permission to try AI tools on real work, fail productively, and share what they learn. The measurement in these environments can be the speed at which they learn and the amount of knowledge that gets propagated. This builds a culture of durable adoption because employees develop genuine intuition about where AI adds value.

Quota-driven adoption cultures, the kind that emerge from token leaderboards and usage-time bonuses, produce a strikingly different outcome. Employees learn to use the tools in ways that satisfy the metric, not necessarily to improve their work. Worse, the measurement system masks this distinction. From the dashboard, quota-driven adoption looks identical to value-creating adoption. The difference only surfaces when you try to find the business outcomes.

Major corporations are measuring usage frequency as their primary compliance metric, and the fact that they’ve landed here suggests this is not a failure of sophistication, but rather of frameworks. The industry simply doesn’t have well-established playbooks for measuring AI productivity in knowledge work, and organizations are filling that gap with the metrics that are easiest to produce.

What enterprise leaders should do now

Separate adoption metrics from productivity metrics – rack usage to understand deployment coverage. Do not use usage data as a proxy for value delivered. These are related but distinct, and conflating them produces bad incentives.

Define outcome metrics before deploying AI in each function – for each use case, be it customer support, engineering, financial analysis, or others, identify what a meaningful improvement looks like before the tools go live. If you can’t define success in advance, you’re not ready to measure productivity – you’re only ready to measure activity.

Audit roles for the type of value AI is expected to deliver – process acceleration and reasoning support require different measurement frameworks. Applying time-saved metrics to roles where judgment quality is the primary output will systematically undervalue AI in those roles and, more importantly, give you no signal when AI is degrading output quality.

Build AI spend governance that connects cost to outcomes – the Uber situation is a governance failure as much as it’s a measurement failure. AI infrastructure costs are real and growing. Organizations need accountability structures that connect expenditure to demonstrable business value, not just to usage volume.

Treat leaderboard-style usage metrics as a risk signal, not a success signal – any measurement system that can be gamed will be gamed. Token leaderboards and usage-time bonuses are gameable. The organizations that avoid this trap are those that make outcome accountability visible at the same level as activity visibility, ensuring high usage paired with flat outcomes registers as a problem.

Measurement needs to catch up

None of this is an argument against AI adoption or against measuring its impact. The productivity gains from well-deployed AI are real, and organizations that figure out how to capture them will have structural advantages. The question is whether the measurement is oriented toward the thing that matters.

Right now, many enterprises are measuring what’s easy to measure. Token consumption is countable. Usage time is trackable. These metrics exist, they’re real-time, and they satisfy the organizational appetite for accountability. The problem is that they measure compliance with AI adoption, not the value of AI-assisted work.

The organizations worth watching are the ones building measurement frameworks that start from business outcomes and work backward to AI activity, not the ones starting from AI activity and hoping the outcomes follow.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *