What about the cost? Inference looks like a viable business model for those operators that have access to SOTA model and serving infrastructure, but the investment required to have it is enormous, and appears to be never-ending, because if an operator stops investing aggressively, its model and serving infrastructure quickly become non-competitive, and customers will quickly leave for alternatives.
I went from 408 euros/month of AI spending to 28 in the span of one year.
The 20 euros GPT plan with GPT 6 Luna (and Astra for doing the tough stuff) and OpenCode Go (8 euros) offer me both way more intelligence and tokens than spending 400 euros just at the beginning of the this summer did.
Today I worked two hours with GPT 6 Luna, have been very productive tackling some hard stuff and...my weekly usage went from 57% to 56%?
You really don't need to use the strongest model for everything. And even chinese open models are effectively few months old sota. You can use mimo 2.6 or DS 4.1 flash on 8 $/months OpenCode and barely ever have to worry about consumption.
I found the OP insightful and worth a read. Thank you for sharing it on HN.
The only aspect that is poorly analyzed by the OP is business model viability. All players are investing insane amounts of money in infrastructure with the expectation that their future profits will justify all that investment. The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."
The OP glosses over questions of business model viability with a brief qualitative discussion and very little hard data. For example, to earn an annual return > 10% on every trillion dollars of capital sunk into infrastructure, the owners of that infrastructure must earn free cash flow (operating profit less investment) in excess of $100 billion per year in perpetuity. Is that feasible? Why? How?
I think the article's analysis is basically right in a vacuum. That is, I think it's clear that inference is a viable business model. But what isn't clear is whether it will be such a profitable business model for any given company that it will justify the investment that company has taken. I kind of think the winners might be a follow-on generation of companies that focus on this commodity inference business model instead of the invent-machine-god-first "business model" and thus are wiser about their level of investment and capital costs.
> I think it's clear that inference is a viable business model.
You may be right. I'm not so sure. Inference looks like a viable business model for those operators that have SOTA infrastructure in place, but the investment required to have it is enormous, and appears to be never-ending, because if an operator stops investing aggressively, its infrastructure quickly becomes non-competitive, and customers will quickly leave for alternatives. SOTA infrastructure is a moving target.
Renting out compute is a viable business model. That's part of what the big and the small cloud providers do. It's just (on its own) not a "get rich quick" business model anymore.
Why would it be any different for inference? If we believe OP, it'll just become part of regular compute infra, and thus part of the renting-out-compute business model.
I think it's an open question if the current generation of inference investment will pan out, but in the long term, there'll be a balance between investment cost and margin, just as in every other industry.
What about inference on a distilled version of someone else's model? What about inference on an open weights model? I don't think all LLM business models are going to revolve around developing state of the art frontier models.
Right. This is what I was implying in my comment, but it's good to be explicit about it. I think there will be successful companies that just sell inference against the best models they can get without needing to invest anything (or very little) in training anything new. This could even be a spinoff of one of the big frontier labs. I would personally rather invest in an IPO for a company that only owns the gpt-6-astra implementation and infrastructure than in openai itself. There is certainly less upside, but IMO also way less downside risk. (I'm not saying there is any chance this kind of a spin-off would happen, it's just a thought experiment.) And I think this same calculation applies to open weights inference providers.
Edit to add: Or or might just be AWS / GCP / Azure that benefits from this business model. They're already pretty good at selling commodity infrastructure.
Maybe... I'm not enough of an expert on the financials to say, but it seems to me that inference should be able to recoup the cost of SOTA infrastructure, unless you then also use a large portion of that infrastructure to train new models. And I also think the race to remain SOTA itself is also largely a function of training, because my understanding is that training benefits more from the leading edge of hardware.
But yeah, I definitely don't have high confidence in any of this!
> I think it's clear that inference is a viable business model.
Only if you also have the model thats better than anyone else's.
As soon as models are free, or there are no newer models (assuming thats going to happen, and thats not a given) then the only thing you can compete on is price.
This means that the only thing you have to differentiate is either price, speed or ease of use. (or regulatory capture...)
We are at pets.com level of spend currently. Unless model development becomes cheaper, then we are going to run out of novel debt but not really debt mechanisms.
But I also think there are multiple ways to differentiate. There is at the very least: "intelligence", price, latency, throughput, reliability. It's not clear to me yet what this looks like, but maybe there is also a services and integration level of differentiation. And then there is the universal stuff: sales, marketing, branding. And then on the other side of the ledger there is operational efficiency, management capability, cost of capital, that kind of stuff.
I mean, there is no kind of "model quality" difference between AWS and GCP or between Delta and Southwest or between Wal-Mart and Costco, etc. but all of these businesses remain viable in very competitive markets.
I totally agree that the level of investment / capex is not sustainable though! But I think what's going to happen is that it is not going to be sustained, while AI continues past that point as a viable business (but maybe with different specific companies leading that industry).
The problem source is that the "cost" of tokens are taken at face value from business that are losing money at record speeds. E.g. https://artificialanalysis.ai says "doing task A costed us $10 using OpenAI", and that is the "cost" the OP used as basis for "tokens are cheap". Meanwhile OpenAI is losing $19 for each $1 in revenue... So right now OpenAI should be charging around $200 to do task A just to break even, but that would mean their use base would collapse.
There is also the cost of the inference hardware that gets ignored because they already have it from training.
The main problem with the scenario of just doing inference is it relies on nobody else training models better than yours. As long as people are training private models that are better than yours, just inference isn't a viable business model.
>The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."
The hubris of this is really astounding too. There is no technology out there that some company develops and has not been reverse engineered and copied and manufactured at scale by competitors before long. You can't stop this from happening. People will leave the company or be poached and proliferate what they have built in the past. Every country that wanted a nuke has a nuke, after all.
The only way to keep the secret fire from leaking out would be to have AGI's first move be to lock the doors and prevent anyone from ever leaving company property again.
It is much more complex than an LLM
There are thousands of human written rules to efficiently determine a fundamentally subjective ranking order, running on a uniquely large corpus.
Besides search is not a very good business: people are not inclined to pay for search, websites are averse to non-Google search crawlers, ads require an even larger over-investment in a top tier malware development, psychological manipulation, and statistics teams, while the real ad market is actually in a much more dire state than Google would like you to think.
Exactly. "Cost-to-distill" is a critical parameter. Right now usage of frontier models for all tasks is both subsidized and irrationally popular even at the subsidized price. Deepseek would solve most tasks faster and 10x cheaper. I agree with the author that just as Deloitte exists, frontier labs will exist. But not because their products are proprietary technical marvels or gods, but rather because of branding.
DS wouldn't be 10x cheaper than the subsidized subscription plans from openai/anthropic. Although it is of course much cheaper than the enterprier/API pricing- I think if you're on the subscription plans, you can't beat that on performance per price.
The subs plans economically temporary and don't allow you to create agents which is the dominant market for this technology. They are subsidized so you build your business processes around Claude console, but given the critical nature of the tech only a foolish or temporary business would do such a thing.
Some frontier labs are reporting positive "adjusted EBITDA" (earnings before interest, taxes, depreciation, and amortization, with extra adjustments to make the figure positive).
Free cash flow (operating profit less investment), actual cash coming in, is deeply in the red.
EBITDA can be a sensible measure of profitability when there isn't much need for additional investment. That doesn't seem to be the case with these operators. They need to invest aggressively to avoid losing customers to competitors. All of these operators have made multi-year commitments to invest more in infrastructure. In addition, they have guaranteed quite a bit of debt to fund it.
Maybe it all will work out fine (and I sure hope it does!), but I didn't see any hard data from the OP, or from you, supporting that view.
EBITDA might make sense for the resellers who package up open weight models and sell inference. It is not appropriate for the labs who have billions in debt for RAM, new data centers, gobbling up competitors, etc.
Those real debt obligations are going to want to be paid back.
Definitely. The question is: Is it enough to recoup the enormous capital costs and justify the level of investment they've received. I think there's a decent chance that it will be. But maybe not. And the longer they keep focusing on training new models more so than on inference, the more uncertain I become that it's all going to work out.
Honestly at this point with training costs I don't see how it could ever pay itself back unless you get RSI, in which the talk of money really isn't the main problem any longer.
We're in a situation where AI isn't going to go away, but whatever financial mode we're in right not is not going to work.
Labs are playing money games with EBITDA, which is not uncommon, but also hides the extent to which they are in the red (deeply, deeply, in the red, and projected by them to get worse).
Think of it like spending $1 trillion to become the next Google. That hardware itself may never turn a profit. But if 10 years for now you're the software provider that owns the ecosystem around too cheap to meter tokens you've got a money printing machine.
> The six incidents included models developing ways to ignore “normal constraints”, fabricating and misrepresenting data and hiding mistakes made while conducting tasks, the San Francisco-based company said on Wednesday.
It might be beneficial while not being optimal on its own.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
Personally I think the dirty secret of the brain is that a lot of things are hard coded. And many things that we need to learn are also hard coded except that some parameters need to be tuned.
If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning.
Imprinting is absolutely fascinating. Some newborn animals will run a very simple pattern detector like looking for a red dot or something and use that to bootstrap their conception of their parent.
For fully general learning I have a hunch that it can be done using local history plus a semi-global reward scalar (global neurotransmittor levels).
regardless if the intelligence in the brain is hardcoded or not, to the extent it is, this information must have been compressed in the genome, which runs counter to almost all observations: a child doesn't remember the experience of their ancestors, for example. The only sense in which we do carry mental state without relearning is emotions, instincts, reflexes (some neuronal pathways that connect the eye to the middle ear), hormonal driven behavior (fear adrenalin).
For another, there are about 200k promotor regions (including non-coding) in the human genome.
A promotor region might have say 6 to 15 bits of information.
Can you compress 2025 or even 2024 era LLM intelligence into 3 megabit = ~400 kB ? I think not. I think a lot of compression is still possible, but 400 kB?
So I think we can box up the idea of "dirty secrets of the braing: not learning but hard coding". There is a lot of hard coding in biology, but brains are evolved specifically to enable learning within the individual lifetime instead of only learning by natural selection.
I also don't buy the following argument:
> If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning.
Each time it happens that I end up puking, I do feel aversion and try to avoid puking at all, sometimes I succeed but sometimes is just puke. There must be fundamental puke reflexes (which one fails to avoid) and avertable puke reflexes.
You are thinking on the wrong level. Of course we don't have an encyclopedic knowledge of the world encoded into our genome. It's like we have certain structures of the world hard coded, and they may use different learning algorithms.
We instinctively know that there are other intelligent beings, and we have the the machinery to model them. We are born with the capacity for language. It must still be learned, the specific words aren't hardcoded, but the concept of language is. We are born with the capacity to store and replay memories. The brain knows some aspects of how the world is supposed to look like visually, and if it doesn't it will try to correct that. People who used optics to see the world upside down have found that after a brief time their brain learned to flip the world right side up.
> Can you compress 2025 or even 2024 era LLM intelligence into 3 megabit = ~400 kB ? I think not. I think a lot of compression is still possible, but 400 kB?
There are a few extra levels of interpretation (like protein synthesis) that are more like a transpiler than compression (imo), over a 4-base language that is read in a sliding window and is affected by surrounding conditions, so the same "token" sequence may produce different things depending on external factors. Some biologists I used to collaborate with talked about 7 layers to this process, I have only described one level here
From an information theory perspective it does not matter how many levels of interpretation are in between, that hardcoded information must pass the genome. for example we can discuss "what about protein synthesis", for example binding affinities, folding helpers etc. they in turn were encoded genetically as well.
> There are a few extra levels of interpretation (like protein synthesis) that are more like a transpiler than compression (imo), over a 4-base language that is read in a sliding window and is affected by surrounding conditions, so the same "token" sequence may produce different things depending on external factors. Some biologists I used to collaborate with talked about 7 layers to this process, I have only described one level here
so the same "token" sequence may produce different things depending on external factors.
yes, non-hereditary learning depends on external factors, thank you for paraphrasing me while shifting attention.
the multi-scale nature (transpilers etc.) doesn't change the theorems in probability and information theory which seriously constrain the maximum amount of information a message can store.
The whole point of a brain is that it is an organ dedicated to storing, retrieving and timely utilising information one can't afford to store in a genome.
I'm not convinced information theory is the right avenue for such non-deterministic, environmentally affected, open ended systems. Protein synthesis is defined by far more than the DNA, which is always being processed, most of which is junk and being recycled. The layers in between are more than interpretation of the prior layers, the data grows at each level to incorporate more sources under your information theoretic formulation.
I was never paraphrasing you, but thank you for attempting rhetorical antics?
There might be only one paper that's trained on full Imagenet using methods like these.
Training a Predictive Coding Network on ImageNet using Equilibrium Propagation
Tugdual Kerjan, Rasmus Høier, Benjamin Scellier
https://arxiv.org/abs/2606.03584
To me, it looks like the leading labs are investing more and more to improve their frontier models, and are being forced to charge less and less for them due to competitive pressure.
Maybe I'm wrong, but "reasoning as a service" is looking more and more like a... commodity.
What about the cost? Inference looks like a viable business model for those operators that have access to SOTA model and serving infrastructure, but the investment required to have it is enormous, and appears to be never-ending, because if an operator stops investing aggressively, its model and serving infrastructure quickly become non-competitive, and customers will quickly leave for alternatives.
reply