I kept a screenshot for about a year. It was a chart of one AI company's annualized revenue run-rate climbing past a number that, at the time, sounded like a verdict: the boom was real, the doubters were wrong, this thing prints money. I treated that number the way a lot of people treated it — as evidence. Then I started reading the footnotes under the footnotes, and the number stopped meaning what I thought it meant.
If you are trying to decide whether the current AI boom is sustainable, the single most useful thing you can learn is this: AI pricing pressure and the valuation are not two separate stories, they are the same story told at two different times. The valuation prices the revenue. The pricing pressure prices the durability of that revenue. A run-rate tells you the company sold something. It does not tell you whether it can keep selling it at a price that covers the cost of delivering it.
So before you read the next headline about a record valuation, walk the mechanism that actually produces — and erodes — that number. In order. Because the order is the whole point.
What happens first: the revenue arrives, and it looks durable
A vendor signs enterprise customers at a per-token or per-seat price. The contracts are real, the logos are recognizable — a bank, a retailer, a logistics company you've heard of. Annualized, the bookings produce a run-rate that gets reported as a headline, and the valuation gets marked up to some multiple of it.
Here is the trap. The multiple is not paying for the revenue that exists today. It is paying for the assumption that this revenue compounds and persists — that next year's number is bigger and the customers don't leave. A 30x revenue multiple is a bet that the seat you just won is a seat you keep, at roughly the price you charged. Everything that follows is about whether that bet holds.
What happens next: the customer tests a substitute
This is the step the headline skips. The enterprise that signed your contract is not loyal to you. It runs a procurement review every budget cycle, and the question on the table is not "is this model good," it's "is this model worth the line item versus the one that costs forty percent less."
For a lot of workloads, the honest answer is that a cheaper model is close enough. Summarizing a ticket, drafting an email, classifying a document — these do not require the frontier. The customer doesn't need the best model; they need an adequate one at a price their CFO will renew. And switching, for a well-built integration, is often a config change pointed at a different endpoint, not a six-month migration.
So the customer either moves, or — more commonly — uses the threat of moving. Which brings the next step.
What happens after that: the price cut you didn't choose
Once one credible competitor will do the job for less, your price is no longer set by you. It's set by the cheapest tolerable substitute. You cut to keep the seat, or you hold and watch the seat walk.
This is what people mean, loosely, when they talk about a price war, but "war" makes it sound like a choice. It isn't. It's gravity. When several well-funded vendors sell a product that customers experience as interchangeable, the price drifts toward the cost of providing it. That's not a failure of strategy. That's what a commodity does.
And the run-rate? It can still be growing while this happens. More usage at a lower price can look like a healthy chart and feel like a slow bleed at the same time.
What happens last: the compute bill doesn't move, and the multiple resets
Here is the part that turns margin pressure into a valuation problem. The revenue side compresses — you're charging less per unit. But the cost side does not give you the same relief. The expensive part of serving a model is inference compute, and that is rented from a handful of cloud providers who are not in a hurry to discount.
So you get squeezed from both ends at once: falling price per token, sticky cost per token. Gross margin — the real number, the one that says whether each dollar of revenue is worth having — gets thin. And a thin-margin business does not deserve a software multiple. Software multiples assume that revenue is mostly profit at scale. If serving the next customer costs nearly what you charge them, you are not a software company; you are something closer to a reseller of someone else's GPUs.
When the market notices that — and it notices late, all at once — the multiple doesn't drift down. It resets. The same run-rate that justified the headline valuation suddenly supports a much smaller one, because the assumption underneath it (durable, high-margin revenue) turned out to be wrong.
How I'd actually read a valuation
When a number lands in front of you, here's what I'd ask before I treated it as proof of anything:
- Revenue quality, not revenue size. Is it from contracted enterprise seats with switching costs, or from usage that can re-route to a competitor's endpoint by Friday? Two companies with identical run-rates can be worth wildly different amounts.
- Switching cost. What actually keeps the customer? Proprietary data, deep workflow integration, and fine-tuned models are sticky. A raw API call is not. Stickiness is the only thing standing between you and the price floor.
- Gross margin, disclosed or estimated. Find out — or guess hard — what it costs to serve a dollar of revenue. If that number isn't disclosed, assume there's a reason.
- Who pays the compute bill. A vendor renting capacity at hyperscaler prices has a structurally worse cost position than one that owns its silicon or has a sweetheart compute deal. The pressure cascades: squeeze the vendors and you eventually squeeze whoever sold them the GPUs.
Who should care, and who shouldn't
If you're trading these names, allocating into the sector, or deciding whether your employer's AI bet is sound, this mechanism is the thing to track — not the press-release number. If you're a builder choosing a model to ship on, you mostly benefit from the price war; cheaper inference is a gift to you and a problem for your vendor.
And if you only want the headline to be true because you're already long, this is exactly the piece you should sit with.
That run-rate I screenshotted still looks impressive. What I understand now is that it was never the answer. It was the question — can this revenue survive its own pricing pressure — printed in a font large enough to be mistaken for a conclusion.