Meta built an internal dashboard that ranked roughly 85,000 employees by how many AI tokens they burned. The top user hit about 281 billion tokens in a single month. Top-rankers earned titles — “Token Legend,” “Cache Wizard.” The leaderboard was dead within weeks. Not because it was controversial, but because it measured nothing anyone actually cared about. That’s the whole story of AI-usage dashboards in one company: the moment you put usage on a scoreboard, you stop measuring work and start manufacturing activity.

Same mistake, new bill. Lines of code — a vanity metric the industry abandoned. Hours saved — the one we just retired. Tokens burned — the one we're making right now. Each one measured activity, not outcomes, and each one got gamed.

Should I track AI usage to prove it’s working?

No. The instant usage becomes a number people are judged on, they optimize the number, not the work — and the two come apart fast.

This is just Goodhart’s law with a GPU bill attached: when a measure becomes a target, it stops being a good measure. Meta’s “Claudenomics” leaderboard is the cleanest proof I’ve seen. Rank 85,000 people by token consumption and you don’t get 85,000 people doing better work — you get people finding ways to consume more tokens. Employees were caught automating unnecessary tasks purely to inflate their numbers (CIO). The dashboard didn’t reward productivity. It rewarded burn. Meta shut it down in April 2026 once that was obvious (HR Executive).

A mock AI-usage leaderboard: rank, name, tokens burned — "Token Legend" at 281B, "Cache Wizard" behind it — with a stamp across it reading "Measures burn, not value."

“Tokenmaxxing” got a name in 2026 for exactly this reason: treating token consumption as a productivity badge — leaderboards, gamified titles, budgets that reward burn rate. It peaked early in the year alongside agentic coding tools, then ran straight into weak ROI and obvious metric-gaming (CIO).

Why does an AI-usage dashboard get gamed so fast?

Because it can’t tell useful from useless. A worker who spends 10 million tokens automating a real workflow and a worker who spends 10 million tokens running queries they didn’t need look identical on the dashboard. Same bar height, same rank.

So the number carries no signal about value — only about volume. And the second you attach a reward or a ranking to it, the cheapest way to move it up is to manufacture activity: more prompts, more runs, more automated busywork that no one needed. You’re not measuring output. You’re measuring effort you can fake. That’s why usage metrics don’t just fail to help — they actively pull behavior in the wrong direction, toward whoever is best at looking busy in the tool.

Isn’t this just the lines-of-code mistake again?

Yes. Literally the same one. We spent years learning that ranking developers by lines of code produced bloated, worse software — because more code is not more value, it’s often less. We abandoned it. Tokens burned is that vanity metric reborn, just with a compute invoice instead of a keystroke count. Critics made the comparison directly: it’s counting lines of code, “reborn with a GPU bill attached” (HR Executive).

Even the vendors are backing away from usage as a proxy. Salesforce’s public line is that token usage isn’t how you prove productivity — what you do with the tokens is (Axios). When the companies selling you the tokens are telling you to stop counting them, the metric is over.

So what should I measure instead?

Outcomes, not activity. The one question that survives a budget review is “what shipped, and what did it cost?” — not “how much AI did we use getting there.”

Most organizations still report AI success with open rates, click-throughs, or isolated productivity gains — activity metrics that rarely survive a budget review — even as 90%+ of executives plan to increase AI spend in 2026 (UC Today). When someone has to defend the AI line item, “our Token Legend burned 281 billion tokens” is not an answer. “This workflow costs 40% less per completed case than it did before” is.

Pick one workflow. Baseline it before AI — cost per successful outcome, not hours logged or prompts sent. Then measure the same number after. That’s the whole method, and it’s the same argument I made in how to actually measure AI ROI: stop counting activity, count outcomes that shipped. (It’s also why the AI-jobs headlines got the story wrong — companies measured the wrong thing and cut the wrong roles.)

What to do with this

Open your AI dashboard this week and kill one activity metric — tokens, prompts, adoption %, whatever’s there — and replace it with one shipped-outcome number for one workflow. Cost per successful case is a good default.

The test is simple: if a metric can be topped by burning more tokens, it’s theater. Meta found that out with 85,000 employees and a Token Legend. You can find it out with a five-minute look at your own dashboard.