More interesting take on Nvidia's position than I've come across before. One thing to be noted is 1) Nvidia is already making moves in robotics so even if their position in AI (moreso llms) diminished, they certainly have another big avenue arguably harder to just get into (although I'm not sure what efforts Google is doing for the tpu in robotics). Another point is Nvidia is still the main player in the west, that is, China certainly can and will create their own full stack without reliance on US companies. That puts Europe and other countries in an interesting, do you buy Nvidia because it's the only option or for security. That's to say I believe Nvidia's position relied on many different things being true at the same time, and we're moving towards an environment where those things are certainly being contested at (roughly) the same time.
Even in the west, Nvidia's dominance is bound to weaken. There is a notable uptick of articles on HN about people running large models on AMD hardware. And while I don't know official sales figures, I know we have trouble getting our AMD system delivered
AMD's software story is still a lot worse than Nvidia's. But patching up vllm to run one or two models you care about on AMD hardware is a much easier proposition than using them in most other fields of AI.
I think this is a great example of the disconnect people have in these types of conversations.
You can both become a company that supplies 80% of the world with your type of product, and then still have your stock go down in value.
All it takes is over evaluation by the stock market. Then a course correction from unsustained growth on growth (second order). So even if you continually replace YoY 80% of the world's hardware on a rotating business, but you don't increase market share or increase demand (aka growth)... Your business looks stagnant to the stock market, and there isn't really anything you can do about it. The best you can do is track inflation +/- 2%.
And that's why a lot of older established companies were dividend stocks. You don't expect to growth anymore, but that's not where the value is anymore... The value is in the reliable sales that will happen after infinitum because your company controls a majority share of the business... And that's ok! Unfortunately, silicon valley has created a philosophy of 'you gotta expand into new fields or your on the decline' - aka neo-monopolization
Nobody is saying they’re outright failing, but that they’re not going to be printing money the way they have been recently. Think about Intel circa 2010: most of their competitors like POWER or MIPS were marginalized, they owned the desktop and server markets with a bit of competition from AMD well contained, and their biggest desktop competitor (Apple) had just switched. A lot of analyst predictions … did not match what happened next. The same was true of Cisco a decade earlier. Both companies are still there but they don’t set the terms in their market segments.
I’m not predicting Nvidia will MBA themselves to death in the near future but I think there’s a tendency to overstate how profitable companies will stay. The more money Nvidia makes, the more motivated their competitors will be to get a piece of that market and the more customers will be looking for alternatives like the push into TPUs which the article discussed.
The current administration is definitely corrupt enough that you could imagine an anti-competitive deal of some sort but I don’t think there’s a way for even that to change matters because key competitors are well-connected American companies willing to play that game, too.
> The more money Nvidia makes, the more motivated their competitors will be to get a piece of that market and the more customers will be looking for alternatives like the push into TPUs which the article discussed.
But people have been saying this about CUDA for twenty years, and we are not any closer to a replacement GPGPU paradigm today.
The root comment in this thread was about Nvidia hedging their bet on lost AI market share. They recognize that a reduced pace in training and inference competition will undercut their business, but CUDA isn't a one-trick pony for LLMs alone. TPUs are - you can't even reuse the same architecture for training and inference, they're separate ASICs unlike CUDA cores/ALUs. Veterans of crypto mining will tell you that the ASICs lost in the end, as Nvidia was evolving their hardware faster than the ASIC manufacturers could iterate. When the crypto acceleration landscape diversified away from ETH/BTC into altcoins, Nvidia was still there making money hand-over-fist from mining hardware.
I guess you could argue that robotics, world models or computer vision won't be a trillion-dollar market. But Nvidia is positioned to be the first mover in all of these markets, and none of their competitors are even coming close to the integrated stack that they sell consumers.
I stand corrected, only Google's post-Ironwood TPUs have the split as well.
Nonetheless, TPU architectures are still a systolic array, and have their own limitations for scalability and flexibility. CUDA is no silver bullet, but it satisfies the demands of the edge and research customers very well.
> But people have been saying this about CUDA for twenty years, and we are not any closer to a replacement GPGPU paradigm today.
How much money was in it for the first decade or so? I think AMD was asleep at the switch but e.g. Apple just did their own thing for the parts which they prioritized.
My understanding is also that Anthropic and OpenAI have also worked to decouple themselves so I think it’s likely that the CUDA moat is going to be less of a barrier than it used to be from the perspective of guaranteeing Nvidia profits.
Money wasn't really the problem. Apple pulled OpenCL together pro-bono, and worked with Khronos to find willing industry stakeholders that would oppose Nvidia. OpenCL needed hardware standardization though, and nobody wanted to design or implement on a scalable GPGPU architecture like CUDA had. AMD and Apple both bet big on raster efficiency, which turned out to be a terrible play when Nvidia was already putting dedicated ray tracing and tensor hardware into their GPUs. They both bet the farm against each other, and only Nvidia won.
Once Apple fully left Khronos, AMD played the smartest card they had; they architecturally split RDNA and CDNA into separate product lines, so they could optimize them independently. This staunched the bleeding, and gave AMD a datacenter presence that Apple Silicon could only dream of. Still not a scalable architecture, but better than nothing.
Ever free newsletter and talking head spouts narratives like this free. If you want something that quantifies and gives actionable information, you must do it yourself or pay for it. What are your below $2000/month sources for good analysis?
The website from this article. It's not free. Ben Thompson releases 1 free article a week but the other 3 articles published each week requires a $15/month subscription. Ben Thompson is also very influential in Silicon Valley and the overall tech/media industry.
I know him, I subscribed for a while. Even his paid content is lacking. I'm looking more Valens Research kind of analysis. SemiAnalysis is also good in the higher tier.
ps. Being influential in Silicon Valley just means you are influential, it does not mean substantial. Leopold is still influential and gets money thrown at him at $100s of million despite having no substance.
you might be looking for SemiAnalysis? I only read the free portions of articles but they have various paid options, mostly targeting investors with information and tools.
They are the marketing wing of the AI ecosystem. Their recent article on how SpaceX would drive 500B in data-center revenue was ludicrous-mode. Lets revisit this in a few years.
Phew. I thought we were just about to leave it, albeit not quite getting to space, just dissolving the financial system so our AI overlords can make some more harvester drones.
In many investment theses - like Nvidia's bet that demand for compute will keep growing - the first order assumption is usually correct. Yes, demand for more compute, chips, infrastructure is huge and each year some additional data centers will be built. Where such investment bets usually fail is in the second-order assumptions: Ie. the expectation of the growth of demand. This is where there's a high chance that the current expectations are likely exaggerated. So: demand is likely to persist for the foreseeable future but not increase every year. And that can upend the whole investment story. That can be enough to make these bonds a huge burden for Nvidia in the end. Not because people stopped buying more compute but because they stopped buying more every year.
What makes this insanely hard to predict is that the compute needed for the same quality output has roughly gone down 90% every 18 months for ~5 years.
1) We don't know how long that trend will continue, but you do know where to look for when it may end (if smaller sized models continue to compress the knowledge effectively of larger models).
2) We don't know when the appetite for higher cost models might go down and by how much if smaller models get "good enough" and price becomes far more important.
It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less.
It is also entirely possible that at some size - LLMs pick up some emergent capability that doesn't scale well to smaller sizes - and that there's an incredible boost to demand to get that capability.
I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not going to run out of economically useful things to do with it anytime soon on the demand side.
The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have to switch to something like burning the model weights into silicon to continue to make gains, then the current math on general purpose accelerators might be upside down.
What's really interesting is that if you scale it to higher densities (eg 3nm and stacked die) with ComputeInMemory for fp8 you can reasonably start to fit 30B-70B models. With MoE and multiple stacked die, just like HBM, you could fit an open weight near frontier 1T model (like GLM5.2) at similarly much lower power <10kW and high token rates >2ktps. For running a bunch of agents where fill rate and speed/latency are important it may not matter that you're 6-18mo behind on weights. The process for the chip design could be largely automated, and new silicon pumped out as new weights are available (with a 3-6mo delay).
It could be quite a while before we reach that point though. 5+ years easily.
I've been keenly interested in the ability to run local models, but the hardware is just not there. Consumer RAM speeds and capacity will have to significantly increase before local models will be able to perform as well as even the lowest end GPT-5.6 Luna model.
This is on the backdrop of RAM becoming prohibitively expensive. And without the speed and quantity of RAM, it becomes impossible to generate tokens at interactive speeds, regardless of model. There is a fundamental dependency between calculating all of the active params with the given RAM speed.
Even with a model that has been quantized all the way down to Q4, the DGX/RTX Spark chip with 128GB of RAM can only generate ~18 tokens/sec for a MoE model with only 30B active parameters. There haven't been any broadly useful models below 30B active parameters. And that is for a $5000+ piece of hardware that will be one of the best for running on-device models.
I really want to buy instead of rent my AI, but the economics are truly terrible.
Very few consumers are going to spend multiple thousands of dollars to save $10 per month. Companies absolutely will to save hundreds per month per employee, but that's not consumer hardware.
If you can integrate AI accelerators into consumer cards (you can), you can have local AI for "reasonably" cheap. This is Nvidia's long term goal if you listen to what Jensen has to say.
The limitation is entirely on memory right now. Just a few years ago we could of been strapping 80-100GB to cards for under $200 (BoM).
Well if the trend that the comment further up in this thread claimed continues and compute requirements keep dropping exponentially then perhaps in a few years you can have today’s frontier performance on the normal laptop you already have on your desk anyway.
> efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute
I buy that. Jevon's Paradox, sure.
> and we are not going to run out of economically useful things to do with it anytime soon on the demand side
This I don't buy. Not fully, at least. Whether or not there's demand for LLMs in some particular field is one thing, whether or not there is a sustainable business model to be built out of that demand is another thing entirely.
There is a staggering amount of money pouring into startups looking for novel use cases for LLM-based agents. As usual, 99% of them will fail, but those other 1% are going to have to look harder and harder to find a novel use case that can actually be served profitably.
First of all, there's only so many places where a chatbot is going to sell. But, that also seems to be the only interface anyone can come up with that allows a user to steer an agent through a long-running task reliably. I'd love to be proven wrong here.
Also, if current trends plateau and large datacenters are still needed for complex tasks, that would stimy growth of LLM usage across entire industries.
But, if present trends continue, then local inference will become feasible for most tasks. That would lower the barrier to entry across tons of heavily-regulated and/or cost-sensitive industries. But, widespread local inference will almost certainly come with a painful market correction centered around hyperscalers, which would itself dry up the pool for ventures into new markets.
Replace "chat-bot" with "Human Being" because the models I've been using are significantly better than 90% of the human-chat-bots that I must talk to on the phone while scheduling and coordinating my internet installation for example.
Now for every human replacement, that is 1 unit less of communication and bureaucratic burden (HR, middle management etc) that the org requires.
> the models I've been using are significantly better than 90% of the human-chat-bots that I must talk to on the phone while scheduling and coordinating my internet installation for example.
I guess I don't know what to say except that my experience is the polar opposite of yours.
I moved to a new state at the beginning of the year. Needed a new doctor, needed to schedule apartment tours, needed to talk to my employer about insurance and relocation stuff, etc etc. Lots of chatbots, a handful of humans. Humans consistently did what I needed them to do, the chatbots just didn't. I could list examples but I'd be typing all night.
And yknow what, my one call with Comcast to get my internet set up was downright pleasant. The rep was knowledgeable and a good conversationalist.
Nvidia's great superpower is flexibility. You can easily run models of very different types on the same card; and their hardware is great for R&D.
However, at some point AI may be good enough for most people and then it makes sense to make an ASIC for the model (or group of models); and at that point you don't need Nvidia.
I suppose this scenario will happen in various moments at different levels.
It's also hard to predict how much money will be burned going down wrong avenues. The internet was the future, but it took a lot of failed companies to eventually land on a sustainable model that brought us the giants we have today.
Railways were also the future, but that didn't stop a rush to build out (often subsidized) lines that were ultimately uneconomical (either because they were corrupt or the planned settlements never arrived).
If AI is similar, then there's going to be a long slowdown on compute spend until the surplus is worked through. A good historical analogy could be the fiber optic buildouts of the late 1990s. The demand for data never really went down much, but the industry eventually commodified and took down some large companies (Nortel, especially)
I rhink there's a difference with AI because -- it brings true value because you pay for the tokens, you only pay for what you use. That is true value.
Compare to just paying for an internet connection, you have bandwidth but not sure what you can do with it that is valuable.
Let's say you use AI to produce software. There;s no limit as to how high the quality you want your software to have. And how fast you want your project to be complete. There's plenty of room for higher quality, and more performant AI. As AI becomes chepaer people will use more of it, they're not going to say "We have enough AI".
Compare to railroads. Yes you pay for the distance travelled but there's a limit to how much people wwill want to travel, how it will benefit them.
It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less.
I doubt it. If LLM inference efficiency is 50x better than today, then there could be 1000x increase in inference volume due and we'll end up needing even more chips.
Jevons paradox should win out for a long time for AI.
When internet connections got faster than 56k modems, we didn't use the same amount of bandwidth but faster. We used more bandwidth doing things like 4k streaming. I see the same in AI inference. If AI inference is that much more efficient, it will just enable more use cases for AI.
Plus on top of that 1st and 2nd order can be correct, but then the price is too high, meaning people lose money even if correct about the future, but over pay for it.
They also have to be feeling the heat of the ASIC vendors. AMD just acquired Taalas and they work with Cerebras all the time on special projects. ASICs outgun nVidia's chips by an order of magnitude.
Each of the hyperscalers has put like 250B each in the last year for infra. That means that they need to be writing AI profits to the tune of 20B per year just to keep up with the cost of the cash they burned.
We are not there. But they better figure it out soon. The cash flows dried up, and everyone is taking debt to support the capex. Google for the first time in its public history is cash flow negative. Amazon too.
There’s a massive difference between “going up forever” and “point B is higher than point A.”
The current setup can’t sustain a downturn, even if yes 20 years from now point B is likely to be higher than present.
That’s the danger. Those that are going to get wiped out by the AI bubble burst aren’t wrong about AI being huge long term, they just put themselves in a position to not survive the storms that happen between points A and B.
IMO focusing on the hyperscalers is kind of misleading.
Yes, for programmers and tech companies AI is kinda boring now, but AI integration in general is still kind of uncharted territory.
There are so many small companies and individuals just getting started with AI today and I believe a large the customer base (and revenue) is still untapped. Hell, I’m discovering new use cases regularly still and the average mismanaged 30 people whatever SaaS vendor probably didn’t even get started yet.
This is what so many people on HN and the market are constantly missing.
Jason Kottke almost didn't found his blog in 1998, famously quoted as saying: "I thought I was too late, that no one would be interested." Needless to say, the internet was a tiny joke in 1998 compared to what it is now.
We are just barely scratching the surface of what's possible with AI, both in terms of the leading edge and in the 'torso' of the economy (the portion you're describing).
Folks from Silicon Valley working in AI-forward companies have a skewed perception of how many people have adopted this technology so far. Codex recently celebrated hitting 10 million users. This is a great milestone and all, but to put it in context, Microsoft office has a billion users. Sure, many people use Claude Code and or some other harness and the growth is staggering, but the overall scale is tiny compared to software as a whole. Costs of serving and usage are still very high, prohibitively so for many, so we aren't even close to market saturation.
And even at the leading edge, people who do work in those AI-forward companies; models are still slow, require hand holding, and produce suboptimal outcomes sometimes. Imagine the value when instead of needing to prompt it once per 30 mins, you prompt it once per day. Then once per week. Then once per month. Imagine all this running not on 3 trillion parameter models, not on 10 trillion, but 100 trillion. What kind of computer infra will be needed then? Certainly more than we have today.
I've always marveled at how one can pick any year since the internet went mainstream and in that year people thought, "Oh my goodness, the internet is amazing!"
Then, move five years forward from that year and look back. In every case, people think, "Hah! The internet was so simple then!"
So much of the dotcom bust was essentially: "anything internet will work", but even after it was pretty easy to find different ways to publish web pages or migrate to new forms of social media, video, short-video, etc.
It's orders of magnitude harder to automate work, which is the business value proposition of AI (whether you are displacing worker or adding breakthrough capacity), and it's not clear the LLM hyperscalers will get the value-add from that.
On the consumer side, it might be a race to the bottom: same ad profits, but now you need AI to produce it.
Ben is wrong; demand for compute, aka revenue backlogs, is mythical and will collapse, simply because of two reasons :
1. Circular investment/spending.
2. Too much capital in the system, so returns cannot be hit regardless because the barrier is too high. (Evidence being every capital cycle in history)
I think more interesting take here would be WHEN this will happen. I don't think Ben, or anyone else, thinks we wont have some sort of correction or stabilization in supply/demand (he has said as much) But when will that occur? 2 months? 2 years? 20 years?
well its typically when it becomes clear that the private equity firms taking the risk decide they can not get the returns they need, forcing the backstoppers such as Nvidia to take that burden, and the whole ecosystem collapses.
I guess my point is that until that collapse, returns are very very very good. This can (and probably will) go on for a decent amount of time more, regardless of the inevitable things you point out.
My guess is at least another 2 years, as most people don't use AI yet, or maybe more precisely, AI is not used in the underlying workflows (which are invisible to the consumer) that make up most people's jobs.
Who knows if I am right. My original post was just pointing out that what you say is about timing, not whether it's true or not, because of course its true.
Yes, agreed, it is timing. I think a sign we are getting closer is the new equity issuances, which kind of are leveraging the current environment and the retail excitement.
161 comments
AMD's software story is still a lot worse than Nvidia's. But patching up vllm to run one or two models you care about on AMD hardware is a much easier proposition than using them in most other fields of AI.
Trust me guys, it's over!
You can both become a company that supplies 80% of the world with your type of product, and then still have your stock go down in value.
All it takes is over evaluation by the stock market. Then a course correction from unsustained growth on growth (second order). So even if you continually replace YoY 80% of the world's hardware on a rotating business, but you don't increase market share or increase demand (aka growth)... Your business looks stagnant to the stock market, and there isn't really anything you can do about it. The best you can do is track inflation +/- 2%.
And that's why a lot of older established companies were dividend stocks. You don't expect to growth anymore, but that's not where the value is anymore... The value is in the reliable sales that will happen after infinitum because your company controls a majority share of the business... And that's ok! Unfortunately, silicon valley has created a philosophy of 'you gotta expand into new fields or your on the decline' - aka neo-monopolization
I’m not predicting Nvidia will MBA themselves to death in the near future but I think there’s a tendency to overstate how profitable companies will stay. The more money Nvidia makes, the more motivated their competitors will be to get a piece of that market and the more customers will be looking for alternatives like the push into TPUs which the article discussed.
The current administration is definitely corrupt enough that you could imagine an anti-competitive deal of some sort but I don’t think there’s a way for even that to change matters because key competitors are well-connected American companies willing to play that game, too.
Your margin is my opportunity - Jeff Bezos
The root comment in this thread was about Nvidia hedging their bet on lost AI market share. They recognize that a reduced pace in training and inference competition will undercut their business, but CUDA isn't a one-trick pony for LLMs alone. TPUs are - you can't even reuse the same architecture for training and inference, they're separate ASICs unlike CUDA cores/ALUs. Veterans of crypto mining will tell you that the ASICs lost in the end, as Nvidia was evolving their hardware faster than the ASIC manufacturers could iterate. When the crypto acceleration landscape diversified away from ETH/BTC into altcoins, Nvidia was still there making money hand-over-fist from mining hardware.
I guess you could argue that robotics, world models or computer vision won't be a trillion-dollar market. But Nvidia is positioned to be the first mover in all of these markets, and none of their competitors are even coming close to the integrated stack that they sell consumers.
AWS begs to differ. They originally split between `Trainium` and `Inferentia` but now support both with `Trainium`
Nonetheless, TPU architectures are still a systolic array, and have their own limitations for scalability and flexibility. CUDA is no silver bullet, but it satisfies the demands of the edge and research customers very well.
How much money was in it for the first decade or so? I think AMD was asleep at the switch but e.g. Apple just did their own thing for the parts which they prioritized.
My understanding is also that Anthropic and OpenAI have also worked to decouple themselves so I think it’s likely that the CUDA moat is going to be less of a barrier than it used to be from the perspective of guaranteeing Nvidia profits.
Once Apple fully left Khronos, AMD played the smartest card they had; they architecturally split RDNA and CDNA into separate product lines, so they could optimize them independently. This staunched the bleeding, and gave AMD a datacenter presence that Apple Silicon could only dream of. Still not a scalable architecture, but better than nothing.
https://deepmind.google/models/gemini-robotics/
Google is mostly the party behind the whole VLA principle.
ps. Being influential in Silicon Valley just means you are influential, it does not mean substantial. Leopold is still influential and gets money thrown at him at $100s of million despite having no substance.
https://semianalysis.com/
They are the marketing wing of the AI ecosystem. Their recent article on how SpaceX would drive 500B in data-center revenue was ludicrous-mode. Lets revisit this in a few years.
1) We don't know how long that trend will continue, but you do know where to look for when it may end (if smaller sized models continue to compress the knowledge effectively of larger models).
2) We don't know when the appetite for higher cost models might go down and by how much if smaller models get "good enough" and price becomes far more important.
It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less.
It is also entirely possible that at some size - LLMs pick up some emergent capability that doesn't scale well to smaller sizes - and that there's an incredible boost to demand to get that capability.
It's just very hard to predict.
The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have to switch to something like burning the model weights into silicon to continue to make gains, then the current math on general purpose accelerators might be upside down.
> If we have to switch to something like burning the model weights into silicon to continue to make gains
I think that's already being considered semi-seriously [0][1]
[0] https://taalas.com/products/
[1] https://ir.amd.com/news-events/press-releases/detail/1296/am...
I've been keenly interested in the ability to run local models, but the hardware is just not there. Consumer RAM speeds and capacity will have to significantly increase before local models will be able to perform as well as even the lowest end GPT-5.6 Luna model.
This is on the backdrop of RAM becoming prohibitively expensive. And without the speed and quantity of RAM, it becomes impossible to generate tokens at interactive speeds, regardless of model. There is a fundamental dependency between calculating all of the active params with the given RAM speed.
Even with a model that has been quantized all the way down to Q4, the DGX/RTX Spark chip with 128GB of RAM can only generate ~18 tokens/sec for a MoE model with only 30B active parameters. There haven't been any broadly useful models below 30B active parameters. And that is for a $5000+ piece of hardware that will be one of the best for running on-device models.
I really want to buy instead of rent my AI, but the economics are truly terrible.
If you can integrate AI accelerators into consumer cards (you can), you can have local AI for "reasonably" cheap. This is Nvidia's long term goal if you listen to what Jensen has to say.
The limitation is entirely on memory right now. Just a few years ago we could of been strapping 80-100GB to cards for under $200 (BoM).
You can't save yourself rich.
I buy that. Jevon's Paradox, sure.
> and we are not going to run out of economically useful things to do with it anytime soon on the demand side
This I don't buy. Not fully, at least. Whether or not there's demand for LLMs in some particular field is one thing, whether or not there is a sustainable business model to be built out of that demand is another thing entirely.
There is a staggering amount of money pouring into startups looking for novel use cases for LLM-based agents. As usual, 99% of them will fail, but those other 1% are going to have to look harder and harder to find a novel use case that can actually be served profitably.
First of all, there's only so many places where a chatbot is going to sell. But, that also seems to be the only interface anyone can come up with that allows a user to steer an agent through a long-running task reliably. I'd love to be proven wrong here.
Also, if current trends plateau and large datacenters are still needed for complex tasks, that would stimy growth of LLM usage across entire industries.
But, if present trends continue, then local inference will become feasible for most tasks. That would lower the barrier to entry across tons of heavily-regulated and/or cost-sensitive industries. But, widespread local inference will almost certainly come with a painful market correction centered around hyperscalers, which would itself dry up the pool for ventures into new markets.
Now for every human replacement, that is 1 unit less of communication and bureaucratic burden (HR, middle management etc) that the org requires.
I guess I don't know what to say except that my experience is the polar opposite of yours.
I moved to a new state at the beginning of the year. Needed a new doctor, needed to schedule apartment tours, needed to talk to my employer about insurance and relocation stuff, etc etc. Lots of chatbots, a handful of humans. Humans consistently did what I needed them to do, the chatbots just didn't. I could list examples but I'd be typing all night.
And yknow what, my one call with Comcast to get my internet set up was downright pleasant. The rep was knowledgeable and a good conversationalist.
However, at some point AI may be good enough for most people and then it makes sense to make an ASIC for the model (or group of models); and at that point you don't need Nvidia.
I suppose this scenario will happen in various moments at different levels.
Railways were also the future, but that didn't stop a rush to build out (often subsidized) lines that were ultimately uneconomical (either because they were corrupt or the planned settlements never arrived).
If AI is similar, then there's going to be a long slowdown on compute spend until the surplus is worked through. A good historical analogy could be the fiber optic buildouts of the late 1990s. The demand for data never really went down much, but the industry eventually commodified and took down some large companies (Nortel, especially)
Compare to just paying for an internet connection, you have bandwidth but not sure what you can do with it that is valuable.
Let's say you use AI to produce software. There;s no limit as to how high the quality you want your software to have. And how fast you want your project to be complete. There's plenty of room for higher quality, and more performant AI. As AI becomes chepaer people will use more of it, they're not going to say "We have enough AI".
Compare to railroads. Yes you pay for the distance travelled but there's a limit to how much people wwill want to travel, how it will benefit them.
Jevons paradox should win out for a long time for AI.
When internet connections got faster than 56k modems, we didn't use the same amount of bandwidth but faster. We used more bandwidth doing things like 4k streaming. I see the same in AI inference. If AI inference is that much more efficient, it will just enable more use cases for AI.
See for example, internet traffic over time: https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcS779VS...
Even after so many years, internet traffic continues to grow at an increasing rate.
We are not there. But they better figure it out soon. The cash flows dried up, and everyone is taking debt to support the capex. Google for the first time in its public history is cash flow negative. Amazon too.
Building a business model on the belief that “this time is different” always finds storms on the horizon.
The current setup can’t sustain a downturn, even if yes 20 years from now point B is likely to be higher than present.
That’s the danger. Those that are going to get wiped out by the AI bubble burst aren’t wrong about AI being huge long term, they just put themselves in a position to not survive the storms that happen between points A and B.
Yes, for programmers and tech companies AI is kinda boring now, but AI integration in general is still kind of uncharted territory.
There are so many small companies and individuals just getting started with AI today and I believe a large the customer base (and revenue) is still untapped. Hell, I’m discovering new use cases regularly still and the average mismanaged 30 people whatever SaaS vendor probably didn’t even get started yet.
Jason Kottke almost didn't found his blog in 1998, famously quoted as saying: "I thought I was too late, that no one would be interested." Needless to say, the internet was a tiny joke in 1998 compared to what it is now.
We are just barely scratching the surface of what's possible with AI, both in terms of the leading edge and in the 'torso' of the economy (the portion you're describing).
Folks from Silicon Valley working in AI-forward companies have a skewed perception of how many people have adopted this technology so far. Codex recently celebrated hitting 10 million users. This is a great milestone and all, but to put it in context, Microsoft office has a billion users. Sure, many people use Claude Code and or some other harness and the growth is staggering, but the overall scale is tiny compared to software as a whole. Costs of serving and usage are still very high, prohibitively so for many, so we aren't even close to market saturation.
And even at the leading edge, people who do work in those AI-forward companies; models are still slow, require hand holding, and produce suboptimal outcomes sometimes. Imagine the value when instead of needing to prompt it once per 30 mins, you prompt it once per day. Then once per week. Then once per month. Imagine all this running not on 3 trillion parameter models, not on 10 trillion, but 100 trillion. What kind of computer infra will be needed then? Certainly more than we have today.
Then, move five years forward from that year and look back. In every case, people think, "Hah! The internet was so simple then!"
In 2031, I suspect we'll say the same about 2026.
So much of the dotcom bust was essentially: "anything internet will work", but even after it was pretty easy to find different ways to publish web pages or migrate to new forms of social media, video, short-video, etc.
It's orders of magnitude harder to automate work, which is the business value proposition of AI (whether you are displacing worker or adding breakthrough capacity), and it's not clear the LLM hyperscalers will get the value-add from that.
On the consumer side, it might be a race to the bottom: same ad profits, but now you need AI to produce it.
Disappointed by the lack of Tom Cruise.
1. Circular investment/spending.
2. Too much capital in the system, so returns cannot be hit regardless because the barrier is too high. (Evidence being every capital cycle in history)
My guess is at least another 2 years, as most people don't use AI yet, or maybe more precisely, AI is not used in the underlying workflows (which are invisible to the consumer) that make up most people's jobs.
Who knows if I am right. My original post was just pointing out that what you say is about timing, not whether it's true or not, because of course its true.