Your CFO is about to Discover your AI Bill

discover-your-ai-bills-2

Somewhere in your company right now, someone is spending money you don’t know about.

A support lead has wired an AI summarizer into every inbound ticket. An engineer left an agent running overnight because retries are cheap, at the moment. A marketing manager generates forty draft variants to pick one. Each decision is small and each one is defensible. None of them usually appears in the budget you had.

Chamath Palihapitiya predicted on CNBC that the discovery moment for most companies will be an earnings miss traced backward to AI usage nobody governed.

The discovery moment

A product VP I’ll call Ash [a composite drawn from several conversations] sat in a quarterly business review last spring. The CFO had one slide and one question: AI-related spend was $1.9 million for the quarter, up 60 percent, what is the ROI?

Ash has three options. Guess, and be wrong on the record. Stall, and look like he wasn’t managing it. Or promise an audit. He took the logical choice and promised the audit.

The audit took two analysts six weeks. It found that more than half the spend couldn’t be traced to any unit of work: not to a feature or an outcome or an experiment with a name. The CFO responded the way CFOs respond to untraceable spend, with an approval gate on every new AI expense. Ash’s two open headcount requests was denied in the same meeting cycle. By the way this is the reason we have had less hiring and higher layoffs due to AI. Not that AI is effective but AI is costing more and there is a need to manage costs.

For the next two quarters, every AI proposal and business case from such organizations was reviewed. The team with the strongest AI roadmap in the company spent six months paying down a credibility debt.

Why this happens: Tokenmaxxing

Software spend used to be predictable. A seat license costs the same in month one and month nine. Your finance team built its entire model of software economics on that assumption, and for twenty years the assumption held.

Marginal cost is back. Every AI feature you ship and every workflow your teams automate has a per-use cost, and an entire generation of product leaders built instincts on software that cost whenever you onboard a new user. Those instincts now fail, and they fail silently, because this spend is variable: it scales with usage, and usage scales with enthusiasm, which is exactly the thing you’ve been trying to increase. This has given rise to TokenMaxxing a weird trend that started in 2026.

Garglely Orez wrote a great article explaining this trend. Gergely Orosz’s piece digs into “tokenmaxxing,” where engineers at companies like Meta, Microsoft, and Salesforce burn AI tokens deliberately to look productive on internal leaderboards. Meta’s leaderboard, nicknamed Claudeonomics,” tracked 85,000 employees and racked up 60.2 trillion tokens in 30 days (roughly $900M at list price), before getting shut down after The Information reported on it.

Microsoft and Salesforce built similar tracking tools, and engineers describe padding their usage with throwaway prototypes and unnecessary AI queries just to avoid looking “insufficiently AI-native.” Shopify offers the counterexample: it renamed its leaderboard to a “usage dashboard,” added circuit breakers for runaway agents, and had leadership actually check in on why people were spending heavily, rather than treating raw volume as the goal. Orosz’s takeaway is that token count is becoming the new lines-of-code metric. An output easy to game, and a company that rewards it will get gamed.


How to Audit: The Token Ledger: five columns and an honest afternoon

Here’s a tool I call it the Token Ledger. One simple table with five columns.

Cost center | Monthly spend | Unit of work | Cost per unit | Owner and alert threshold

Budget an afternoon for this work. Send a couple of emails one to finance to get 90 days of expense items tagged software and the second to Platform engineering to give you API usage broken out by key

1. Pull what you own. Your three biggest AI vendors, first pass. Model provider invoices. Cloud AI services. The tool subscriptions that meter usage under the seat price. Skip the ones you’re not sure about.

2. Split invoices into cost centers. Each shipped feature should be a line. Each internal workflow also is a line. Each experiment in flight is a line. This can be hard, because one invoice usually hides three or four cost centers behind a single shared key.

Separate keys per team, going forward. A tag field on requests, where the platform supports it. For everything already spent, a percentage split negotiated in one meeting with the two engineers who actually know the traffic. You’re building a number good enough. Crude beats absent, and you refine it monthly.

3. Attach a unit of work to your top five lines. This is the step that turns a bill into information. On an invoice, healthy growth and waste look identical. Per unit, they don’t.

Worked example:

Support runs AI summarization and draft replies on 40,000 tickets a month.

Invoice: $4,400. Ledger line: $0.11 per ticket resolved. A resolved ticket saves about eight minutes of agent time.

At a loaded cost of $40 an hour, that’s $5.30 of labor bought for eleven cents of compute. Call it 45 to 1. You want that line growing.

Next month the invoice reads $9,100. Ticket volume hasn’t moved. Per-unit cost just doubled. This time the culprit was a retry loop re-summarizing on every escalation. What the ledger did was point at the right haystack

4. Assign one owner and one threshold per line. Not who could notice. Who gets the alert, and what number trips it. Rule of thumb: flag any line where per-unit cost moves more than 20% in a month. Flag any line whose total doubles in a quarter.

5. Review it monthly. Fifteen minutes. Review it monthly and it’s governance. One question per line: is this per-unit drift, an efficiency problem, or volume growth at flat unit cost, which is adoption, which is what you wanted? You should answer that for every line in under a minute.

Do this and the “what did we spend and what did it return” meeting can’t happen to you. It stops being an audit you owe instead its something you already walk into the meeting with.

Article content

The counterargument, taken seriously

A reasonable person will object. our AI spend is a rounding error next to payroll, and building governance for it now is busywork. For many companies that’s true today, and if the spend were static I’d agree.

But this spend compounds with every feature launch and every quiet automated workflow, and cost that doubles silently is exactly the kind that stays too small to govern until it’s too large to explain. The ledger costs you an afternoon. The audit costs six weeks and two headcount, and the difference between them is only visible in hindsight.

Where this goes

The ledger gives you visibility, and visibility is where the real questions start. Should the summarization feature move to a cheaper model? Should the workflow that’s scaling leave the API for self-hosting? Is the line finance question a cost problem or a pricing problem? Those are decisions rather than measurements, and the decision layer is part of my upcoming udemy course and book: Product Manager to Product Leader in an AI world teaches. What AI actually costs, when to build, buy, or partner, and how to walk into the CFO conversation with numbers that survive scrutiny.

If you are interested in the course “Crafting a business case and budget for AI” please join my waitlist at https://forms.gle/XRkcmWjimju3gipc9

If you like to be the first to review my book “Product Manager to Product Leader in an AI Era” please sign up at https://forms.gle/XE6S9VhaxiADahZ28

Next week: I am going to share build vs buy: why self-hosting AI is a forecasting bet

Comments are closed.