Tokenmaxxing
Why it doesn't matter if frontier token prices go up drastically.
You and your team should be tokenmaxxing right now.
Not because I know frontier tokens are about to get expensive. I suspect they are (and I’ll get to why), but the argument I actually want to make does not depend on it. There is a real fight happening over whether frontier tokens already cost too much and whether open source has closed the gap, and I am not going to settle it here. My point is that founders and operators are not being asked to settle it either. You are being asked something simpler: what is speed worth right now, and what happens to the price of it. Tokens are the input to execution speed, and execution speed is the only advantage in this market that expires.
Run both branches.
(A) You spend hard now and intelligence gets more expensive: you bought speed and quality at a price your competitors will never see again.
(B) You spend hard now and it never does: you still have the progress, you just pulled it forward and overpaid in operating expense.
One outcome is a durable lead. The other is a line item you will not remember in two years. Decisions under uncertainty get made on the shape of the payoff, not on the confidence of the prediction, and this payoff is lopsided enough that being wrong about prices costs you almost nothing.
Here is the part that argues against me. Over the last seven months my token volume went up 276x while my bill went up 40x. My blended cost per million tokens fell about seven times over that stretch. Tokens got dramatically cheaper for me, in real numbers, over exactly the period I was becoming convinced that intelligence is about to get rationed. I have spent harder every month anyway.
That deflation is the strongest case against everything I am about to say, and it is not a fluke of my account. Cost per token has collapsed every year of this industry, by orders of magnitude, and betting against that trend has been a reliable way to look foolish. So the burden is on me to explain why my own bill does not refute me.
The answer is that the thing getting cheap and the thing I want are not the same thing. What collapses in price is a fixed capability. Last year’s model at this year’s prices is astonishingly cheap and will keep getting cheaper, because the industry gets better at serving a known workload every quarter. What may not get cheaper is access to the frontier, because the frontier is definitionally whatever runs on the scarcest compute available. Yesterday’s intelligence gets commoditized on schedule. Today’s intelligence is what gets rationed, and it gets rationed precisely when capacity is tight.
That scarcity is not hypothetical, which is why I lean toward the expensive branch. Dwarkesh Patel has made the case that compute is the binding constraint on this industry, that memory supply is the constraint sitting behind compute, and that when demand outruns supply the cost has to land somewhere. You can watch that pressure arrive in the hardware. Supply chain reporting has NVIDIA halving the SOCAMM capacity on Vera Rubin modules, cutting system memory per Vera slot from 192GB to 96GB, because LPDDR5X supply is expected to stay tight through 2027. The HBM configuration on Rubin Ultra is under the same kind of review, with reported evaluations well below the original spec. Read across the platform and the pattern repeats: less memory per accelerator so more accelerators can ship. Morgan Stanley’s teardown has memory going from roughly 9.4% of the bill of materials on Grace Blackwell to roughly 25.7% on Vera Rubin, a 435% increase in memory cost while total system cost about doubled. Nobody ships a chip with less memory than they designed for because things are going well.
If that read is right, the cheap thing and the good thing keep separating, which is also why I keep circling distillation and rejecting it. Taking frontier output and distilling it onto a small cheap model is a real and often correct engineering move. It is also, structurally, a decision to freeze your capability at today’s level in exchange for lower cost. If the models were finished improving, that would be an obviously good trade. They are not finished improving, and the delta between generations has been the single most valuable input available to me as an operator. Optimizing for cost per token means optimizing against the improvement curve. Fine for a mature workload. Bad for anything still being figured out.
Now the part that matters most to a founder, and the part I initially argued my way out of. My instinct was to say that tokens spent on shipped features are not a moat, since a competitor builds the same feature later at whatever price exists then and arrives at the same place. That is true about the feature and false about the business. Features are copyable. Distribution is not. If you ship six months earlier, you are not six months ahead on a roadmap. You are six months into acquiring users while your competitor is still in a planning doc, and every week of that head start compounds into retention, data, referrals, and word of mouth they have to pay to overcome. Anyone can build a product. Without distribution it dies. Speed to distribution is the highest-return use of an expensive token there is, and it is the one with a hard clock on it.
Six months, not eighteen. Eighteen months was the right unit of advantage two years ago. Teams with a decent harness and real context now ship in weeks what used to take quarters, which means the window in which being early is decisive is compressing at the same time the input is getting scarcer. That is the actual squeeze. Not that AI is expensive, but that the period in which speed converts into a defensible position is shorter than it has ever been, and the fuel for that speed may be about to get rationed.
So what deserves the spend. Two things compound and one does not. Speed to distribution compounds externally, for the reasons above. Context compounds internally: the knowledge you encoded, the process you pulled out of somebody’s head into a form a machine can use, the judgment your team built about where the model can be trusted and where it cannot. What does not compound is running a solved workload more times. If a job works reliably and produces the same shape of output every day, that is a candidate for the cheapest model that clears the bar, not for frontier spend. Confusing those two is how a strategy turns into an expensive habit.
Look again at the shape of my usage and you can see which one I was buying. Input tokens grew 346x over those seven months. Output tokens grew 9x. My input to output ratio went from about 4:1 in January to about 150:1 in July. Almost none of that increase went into generating more text. It went into feeding models context: repositories, prior decisions, retrieved history, the accumulated state of the products I am building. The bill is not a record of how much output I produced. It is a record of how much context I built and how often I made a model read it.
It is worth saying plainly that the absolute numbers are small. The monthly bill has never crossed two thousand dollars, which is less than a week of contract engineering, and the output of that spend has been production software and a materially higher personal ceiling. The reason to think hard about this line item is not that it is large. It is that the ratio of what you get to what you pay may never be this good again.
Which leaves one way this actually goes wrong, and it is not branch B. It is spending aggressively on repeat inference over solved work, so that neither branch pays: prices never move and you bought nothing that lasted. That failure is entirely inside your control, which is what separates it from the price question. The discipline was never about how much you spend. It is about what you point it at.
The frontier has never been cheaper than it is today, and it has never been more likely that this sentence stops being true. Spend it on the things that outlive the price.

