Are We in an AI Bubble?

September 14, 2026

Our house view on the AI buildout: what's real, what's overbuilt, and where the returns will actually show up.

Defining a bubble


When people ask whether AI is a bubble, they’re asking one of two different questions, and it helps to separate them.

The first is an infrastructure bubble. Will we build more capacity than there turns out to be demand for? The parallel people draw is the fiber built during the dot-com boom. Companies laid enormous amounts of fiber-optic cable on the bet that internet traffic would need it. Too much got built, demand didn’t show up on schedule, and much of that cable sat unused for years. The question for AI is the same. Will the data centers and compute we’re building now find enough real use to justify the spend?

The second is a financial bubble. Has the money flowing into the asset class detached from any reasonable estimate of what these assets will earn? The dot-com crash worked the same way on the equity side. Companies with no earnings (and no clear path to any) traded at valuations that assumed the internet would pay off immediately and enormously. It did pay off, eventually, but over a much longer time frame than the market had priced in. That’s the financial question for AI. Has price run past the fundamentals?

This section is primarily about the first question. Will the money spent on all this capacity generate enough real economic value to justify itself? I’ll circle back to the financial bubble piece at the end.

Is the demand real?

The revenue trajectories we’re witnessing in AI are truly unprecedented. Anthropic is at $65B of annual recurring revenue, with reported net dollar retention of 500% and ~7x YoY growth in accounts spending over $100,000. OpenAI counts more than 1M business customers, over 9M paying business seats, and 92% of the Fortune 500, with its largest API customers each spending tens of millions of dollars a year.

The breadth matters as much as the big two: at least 10 AI products now generate over $1B in annual recurring revenue and 50 generate over $100M, per Menlo Ventures. Overall, the AI model and application companies have reached a $175B run-rate, growing at ~260% YoY.

Notably, this growth is showing no signs of slowing down anytime soon. In fact, it’s accelerating. In 2023, the AI industry needed 180 days to add $1 billion in revenue. It now needs less than two days (Exponential View).

These vendors were mostly unprepared for this massive inbound demand, and the best evidence is in how they’re behaving: confronted with a capacity crunch, the labs are rationing supply. Anthropic had been pulling Claude Opus from lower subscription tiers until its SpaceX Colossus 1 deal closed on May 6, bringing over 300 MW online within the month at a large premium to market prices. Immediately, Anthropic doubled Claude Code usage limits, dropped peak-hour throttling for Pro and Max subscribers, and lifted API rate limits on its best model. Elsewhere, GitHub paused new Copilot signups in April citing compute demand and OpenAI made the difficult decision to discontinue its Sora product given limited resources. Turning away willing buyers and pulling performing products from paying customers is only rational when you physically cannot serve the load.

The standard counterargument is that excess supply is already appearing: SpaceX renting out Colossus capacity, Meta planning a cloud business. On inspection these are the opposite of glut behavior. Take Meta, the case the bears lean on hardest. Meta contracted over 5 GW of cloud and colocation capacity in the first half of 2026 alone, and SemiAnalysis (independent semiconductor and data center research firm that tracks construction site by site) expects its procurement to keep accelerating. Every option Meta has for this compute is high-margin: frontier training at Meta Superintelligence Labs, scaling up ad models, premium rentals (like the recent SpaceX deals), and tokens-as-a-service. That optionality is why it can keep contracting aggressively. Renting compute at record prices while buying more of it from everyone else is what shortage looks like.

Is the demand durable, or just experimentation?

Durability primarily comes down to return on investment, and the strongest ROI evidence comes from coding.

Source: Company reporting

The output shows up in shipping speed. Anthropic released 74 products and features in one 52-day stretch. At Ramp, 84% of employees use coding agents weekly, and 800+ builders shipped 1,500+ internal apps in six weeks. These companies are demonstrating product velocity that was previously unheard of.

It goes beyond coding. In security, Anthropic's Mythos Preview identified thousands of zero-day vulnerabilities (flaws nobody had previously found) across major operating systems and browsers; partners have confirmed more than 10K high- or critical-severity flaws, and, in a separate open-source scanning program, six independent security research firms assessed 1,752 high and critical findings and validated 90.6% as true positives. On broad knowledge work, Anthropic analyzed 100K real conversation logs, comparing how long tasks took the model against how long they take professionals. The median task saw an 80% reduction in completion time, and the study estimates current models could lift US labor productivity growth by 1.8 points a year, roughly double the recent rate.

Now the honest caveat. That productivity lift is not yet reflected in the macro data. Real results exist, but broad, indisputable evidence of delivered ROI is still elusive, and many claims are self-reported or task-specific. The usage data explains the gap: adoption is concentrated in a small set of power users, and most of the world has barely started.

Alex Sacerdote, founder of the hedge fund Whale Rock Capital, estimates that ~10% of the world touches AI at all. On Sundar Pichai's estimate, only about 10 basis points of knowledge workers use it in a genuinely agentic way; most usage is tinkering. That sentiment is echoed by the spend and survey data: the top 1% of firms on Ramp spend $7,500 per employee per month on AI products, while the median spends just $11.

There are two ways to read the concentration. The bearish read is revealed preference: after three and a half years of free trials, the median firm spends just $11 a month per worker and OpenAI monetizes only about 5% of its weekly active users. Most people have sampled the products and found modest value.

The bullish read is that we are simply early, and the frontier's results will diffuse. The trend data settles it for me. On Ramp's index, the median and the top percentile are both rising steadily, and measured enterprise adoption keeps climbing well above the census figures the market prices off of.

Meanwhile, there is a second force underneath the diffusion. For a large enterprise, adopting AI is becoming less a matter of whether the math pencils out on any single deployment and more a matter of survival. If a competitor uses these tools to ship faster, serve customers cheaper, or undercut on price, standing still means falling behind. The early proof is in the data; early AI adopters are outgrowing their peers:

That threat pulls in the boardroom too: shareholders will push management to keep pace. So every firm is forced to push further, which forces its competitors to push further still, and the cycle feeds on itself. Adoption becomes a race nobody can afford to sit out, and that dynamic diffuses capability far faster than a sober ROI test would predict—and renders the “experimentation vs. production” question effectively moot.

Ultimately, every technology that ends up everywhere starts here: with a sliver of obsessives wringing real value out of something the average person has barely touched. The $11 median says almost nobody is paying for AI yet, which is a statement about adoption rather than value. Penetration is around 2.5%, maybe less. Every market that started there eventually grew, and willingness to pay grew with it.

Who pays for the buildout?

The AI infrastructure buildout forecasts are enormous. Morgan Stanley expects $2.9T over the next three years, while SemiAnalysis expects more than $8T. Someone has to fund it.

Today the hyperscalers fund most of it–but why put all those trillions of dollars at risk?

  1. They build against contracted demand. Cloud revenue growth and backlogs are accelerating
  1. The unit economics work. Inference produces cloud-like gross margins on hardware with a four-to-six-year useful life. As long as pricing, utilization, and asset lives hold, the return math on a rack clears.

The SpaceX GPU rental deals with Anthropic and Google are reported to be particularly lucrative. Public research from Epoch AI estimates that building 1 GW of AI data center capacity requires roughly $38 billion in CapEx. Meanwhile, based on publicly reported contract sizes, hyperscalers like Anthropic and Google are paying SpaceX annualized rates equating to $50 billion and $55 billion per GW, respectively. This means the infrastructure generates more revenue in its first 12 months than it costs to build, resulting in massive operating margins and a cash payback period of under a year.

  1. They consume the compute themselves–and it pays. Meta's generative ad-ranking model is 4x more efficient per unit of data and compute than its predecessor, and Meta's AI ad tools now run at a $75B annualized rate. Google's $250B+ search revenue business has accelerated from 10% to 17% year over year even as AI answers cannibalize some queries, because AI Mode queries run 3x longer and monetize questions that were previously hard to monetize.
  2. The compute is buying operating leverage. Across all four hyperscalers, revenue growth is accelerating while headcount stays flat (or declines) due to internal AI adoption
  1. The risk is asymmetric. Overbuild and you eat a few years of depreciation. Underbuild while AGI-scale demand materializes (AGI, meaning Artificial General Intelligence; AI that can do most economically valuable work) and you cede the platform layer of the next computing era permanently. Power, grid connections, and land are multi-year bottlenecks, so locking up GWs now is imperative and serves as an option on future demand that cannot be bought later at any price.

That fifth point cuts both ways. An industry that fears underbuilding more than overbuilding will overbuild eventually. The question is when, and the answer starts with funding capacity.

To date, much of the buildout has been funded with hyperscaler cash flows. However, capex spend has been eating at hyperscaler free cash flows for the past few years, and analysts are closely watching whether these totals dip negative in the coming quarters–and where they go from there. Bulls will point to consensus forecasts, which show capex growth decelerating while AI revenue keeps compounding–leading free cash flow to snap back.

Bears point to capex guidance inevitably creeping up again (as it has every cycle so far), rental rates compressing and depreciation expenses compounding faster than revenues.[3]

If the bear case wins out, as it has historically, will hyperscalers slow their spending to satisfy investors, or keep building for the long term regardless of the market? History says they keep building. These CEOs watched their stocks fall double digits on capex raises and raised again anyway. This time, though, they no longer have the free cash flow to fund it alone; the buildout increasingly needs outside capital, and that is already happening.

Paying more per GW without adding revenue-generating capacity erodes free cash flow directly, and it sharpens shareholder and creditor pressure to slow down. We are already seeing this play out, with industry capex funding diversifying beyond the hyperscalers.

So if the hyperscalers moderate, who picks up the bill? Morgan Stanley sizes the hole at about $1.5T beyond hyperscaler cash flow through 2028. Four groups line up:

  1. Customers and neoclouds, with creditors behind them. The standard structure is a take-or-pay contract: a lab or hyperscaler commits to fixed capacity at a fixed rate for around five years, whether or not it uses it, and the neocloud borrows against that contracted cash flow to fund chips and buildings. Some neoclouds now require prepayment of a vast majority of contract value, enough to fund an entire cluster upfront.
  2. Nvidia. It now backstops GPU rental contracts for neoclouds, guaranteeing minimum revenue in exchange for upside sharing, which substitutes Nvidia's balance sheet for a hyperscaler's. It has ~$100B of liquidity and upwards of ~$85B of annualized free cash flow. It also set up a $500B partnership with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to stand up "independent compute financing platforms." These moves would make it a potentially significant player in an AI debt market Semianalysis expects to exceed $7T outstanding by 2029, which would make it the second-largest asset-backed credit market in the US after mortgages (about $13T).
  3. The labs. Anthropic is now profitable, but profits are marginal relative to capex needs, so lab-funded buildout depends on outside equity. The swing factor is how much credit the capital markets give to AGI-scale payoffs (e.g. massive breakthroughs in drug discovery and materials science) and what that does to the labs' cost of capital.
  4. The government. The most underappreciated backstop. The US is racing China to AGI. China has the power, the construction capacity, the talent, and the capital. It lacks chips. Building as much compute in the US as physically possible is therefore a national security matter. Washington's toolkit spans grants, loans, equity (a rumored 5% stake in OpenAI, per the Financial Times in July), and regulatory relief. There is a second motive: AI infrastructure is driving more than half of US GDP growth in the first half of 2026, and David Sacks expects that share to reach 75% by year end on consensus capex. The buildout is becoming too big to fail. With the economy and national security posture vs. China ranking as two of the top issues in increasingly contentious election cycles, politicians have a strong incentive to subsidize the AI buildout and remove any red tape or regulation that could slow it down.

The government point, however, is nuanced.

Washington and the states are pulling in opposite directions. The federal government treats the AI buildout as an existential national security race, while state and local politics increasingly treat data centers as an industrial villain. Over a dozen states have considered moratorium bills, alongside over 100 municipal-level pauses. Historically, most state bills stalled, were vetoed (like Maine's 20MW pause), or targeted jurisdictions without a real project in the pipeline. That made early moratoriums largely empty gestures aimed at currying favor with voters worried about skyrocketing ratepayer bills, AI-driven job losses, and resource depletion.

With that said, the dynamic shifted over the summer. New York broke the pattern first with Governor Hochul’s targeted executive order pausing environmental permits for 50-megawatt-plus hyperscale facilities. Shortly after, Texas Governor Abbott instituted a statewide freeze in the nation’s second-largest market with roughly 250 planned projects.

The near-term effect is that capacity is forced into a dwindling list of friendly jurisdictions. Over time, however, the core complaint should fade on its own as new projects are shifting to power generated on site, never touching the grid or the ratepayer’s bill.

That said, none of this is predictable. Public perception of AI can shift fast, and there is real contagion risk that the backlash spreads to states with actual pipelines. Ultimately, public sentiment on AI, and how government and society respond to its impact, are the biggest wild cards in this entire analysis. In my view, however, the federal incentives toward national security and economic growth are too existential to make secondary to a fickle popularity contest among the general public.

Two conclusions follow with reasonable confidence.

First, I expect everything that can physically be built to be funded by someone, with the government as the final backstop. Second, hyperscalers and the government are both incentivized to overshoot eventually, but neither will spend meaningfully beyond what can realistically be brought online.

That reframes the bubble question as a sequencing question: does supply arrive before the demand to absorb it? Two questions follow, in order. How fast can funded capex become energized compute? And once energized, can demand generate the revenue that compute has to earn in order to justify the investment?

How fast can funded capex become energized compute?

The forecasts diverge widely. The street sits around 20-25% annual growth in energized GWs through 2030. SemiAnalysis sits above 50%.

Every input that the buildout requires is short right now, and you can see it in prices and lead times.

The question becomes: which one of these factors is the biggest bottleneck?

Today, it’s power. The operators have essentially settled this question. Microsoft's CEO describes chips sitting in inventory with nowhere to plug in, and its CFO says the company has been short of space and power for many quarters. Jensen Huang, Nvidia’s CEO with a clear incentive to emphasize a scarcity narrative in his own product, says power is the biggest bottleneck as it takes years to source while everything else can be secured in months. The power problem has three parts, but each part has an escape hatch.

  1. Interconnection: permission to connect to the grid. Queues run 7-10 years in major metros. Ranked by length of delay, this is the most pressing problem in power. It is also the only bottleneck in the entire cascade made of rules rather than atoms, and rules can move fast when the government wants something. FERC ordered faster interconnection for large loads in December 2025, and by June PJM had an expedited track cutting the study process to roughly ten months.
  2. Generation: making the watts. Grid-connected generation is lagging badly: ~15 GW a year of new grid capacity over the next four years against ~60 GW a year of need, and that capacity must also serve ordinary load growth from communities already protesting data-center connections over electricity prices. What makes generation tractable is the on-site layer: gas gensets, aeroderivative turbines, and fuel cells that are small, containerized, and factory-built, with no queue. In other words, operators can solve generation behind the meter (power generated on site that never touches public grids).
  1. Grid equipment: the hardware of connection. Transformer and switchgear lead times have lengthened every year, from about 140 weeks in 2023 to over 160 now (Wood Mackenzie), with demand headed from ~1.5K large units a year to over 9K by 2030. This is the smallest delay of the three, at three to four years, but the hardest to compress: the other two have escape hatches while equipment lead times are still stretching. Two mitigants. Behind-the-meter builds need far less grid-scale equipment, so this constraint pushes even more of the buildout on site. And supply is responding: new domestic factories arrive around 2028, Korean and European makers are expanding, Chinese manufacturers are quietly taking orders, and a market for refurbished and rewound transformers is growing.

Put together: a government that wants AGI can shrink interconnection queues with signatures, and behind-the-meter power mitigates all three parts at once. Hyperscaler take-or-pay commitments have also given generation, equipment, and cooling suppliers the confidence to expand capacity. Funded demand is pulling supply into existence.

Then the constraint flips back to chips. Once power is routable, silicon inherits the bottleneck, and it is the hardest input to expand. Labs want more compute than the supply chain is preparing to deliver because most suppliers have been through several boom-bust cycles over the years and thus refuse to underwrite AGI-scale forecasts, so everyone plans one notch too low. Gavin Baker calls TSMC's deliberate conservatism "the bottleneck that will prevent a bubble." Even Musk, the loudest voice on power, agrees chips become the terminal constraint once power is solved, which is the logic behind TeraFab.

The wild card is orbital data centers. A solar panel in the right orbit produces up to 8x the power of the same panel on Earth and, importantly, produces it continuously rather than relying on daytime hours on Earth. A data center in space needs no grid connection, no permits, no land, no cooling or water, and no zoning fight. It skips every bottleneck in this section at once.

This has moved past the napkin stage. SpaceX filed with the FCC for a constellation of up to one million data center satellites. Nvidia announced a space-rated chip platform at its March conference. Sundar Pichai has said that within a decade this will look like a normal way to build data centers. Gavin Baker thinks it could be viable within a few years—and JP Morgan agrees:

While upgrades and repair processes and unit economics still need to be ironed out, and scaling ultimately depends on the success of SpaceX’s unproven Starship rocket, the flip to chips comes sooner and harder if orbital data centers arrive anywhere near Musk’s 2028 target.

The overlooked constraint is labor, and capital cannot buy it directly. McKinsey projects a US gap of 130K electricians and 240K construction workers by 2030. Certifying an electrician takes four years, so contractors poach crews from each other instead of growing the pool. In Jefferies' model it is labor that caps additions over the next two years, at 10.4 GW in 2026 and 12.1 GW in 2027, well below other forecasts, because electrician headcount has historically grown high single digits a year.

That forecast underestimates the workaround: turning data-center construction into manufacturing. Prefab builders assemble electrical skids, cooling modules, and full power rooms in factories and ship them to site, which Schneider Electric and Vertiv case studies show can reduce on-site labor by 20-40%. Factory lines raise output per worker, hold quality, ignore weather, and assemble entire IT halls in weeks, with cost per megawatt running roughly half a traditional build in some configurations (DCD). While prefab cannot remove the burden entirely–its factories need their own skilled workers, and commissioning, substations, and high-voltage utility work cannot be put in a shipping container–it should ease the binding constraint on the 2026-27 ramp.

Whose forecast do we lean on? The street has been consistently conservative on forward capex (see: “Who pays for the buildout?”). SemiAnalysis has been more accurate historically, with bottoms-up project tracking and real industry depth, but its business is levered to the AI trade, so it carries its own bias. The truth likely sits between the two, but closer to SemiAnalysis, because its approach more accurately models how workarounds dynamically change the path.

The question then shifts from supply chain capacity to execution. The bears point to headlines about data-center delays, but on further inspection those are unfounded. SemiAnalysis's proprietary tracking shows its year-end 2026 hyperscaler self-build forecast moved about 1% over six months. The claim that half of 2026 capacity slipped uses a denominator inflated by announcement-stage press releases and non-binding letters of intent. The cancellations sit in the early-stage layer that never ordered equipment, while projects with site control, ordered equipment, and signed interconnection agreements are largely on schedule.

Bottom line: Real physical constraints will govern the ramp, but the street underestimates the solutions. Energized GWs land above street forecasts, and capex lands well above, because every workaround costs more per watt than the bottleneck it bypasses. More money flows into the supply chain than the market prices in. And more supply raises the bar for the revenue that supply must earn.

Can demand earn what the compute costs?

Shanu Mathew, an energy and infrastructure PM at Lazard, provides the machinery to size the bar. Working bottom-up from facility economics, a frontier GW costs about ~$51B to stand up. At a 10% after-tax return, a 50% full-stack operating margin, and realistic asset lives, each GW must generate about $22B of revenue per year. That converts any buildout forecast into a revenue requirement: GWs added times $22B[4]. The requirement by 2030 is as follows:

Before diving into demand, one of the key swing factors in these forecasts is GPU useful life. The $22B per GW bar assumes chips depreciate over roughly six years, but market pricing says they earn far longer: an 8-year-old T4 still yields 42% of its original list price annually at 50% utilization, well above the 17% annual charge a six-year schedule implies. If GPUs keep earning for 8 to 9 years instead of 6, the same CapEx buys more compute-years and the revenue bar each GW must clear falls with it.

On demand, the size of the ultimate destination is not the issue. SpaceX pegs its addressable AI market at $24T ($760B consumer subscriptions, $600B advertising, $22.7T enterprise). The issue is speed of penetration, defined by two key forces: falling costs and rising capability.

The baseline: cheaper tokens, more usage.

The base growth rate runs on a simple rule. As tokens get cheaper, people use more of them, and total spending rises even as the price per token falls. This has a name from energy economics: the Jevons paradox, first observed when more efficient steam engines made coal cheaper to use and consumption exploded. The AI evidence is direct. Token prices fell about 10x over the past two years while enterprise AI spend rose 22x (Apollo, Menlo Ventures). This only happens when demand is highly elastic, meaning each price cut is more than offset by higher volume.

What's actually driving it. The pattern holds because falling prices do more than deepen existing usage–they unlock work that was uneconomic before. Faster internet did not just load the same websites faster; it created streaming, cloud software, and video calls. AI has reached true liftoff in only one place, coding, which is 3-6% of US knowledge labor. Consumers mostly use free chatbots, which is why OpenAI monetizes only 5% of its weekly active users. Personal assistants, health coaches, agentic shopping and travel, and personalized media all sit further down the cost curve, waiting for the next price decline to cross their threshold.

Moreover, the capability we already have is barely adopted. The top 1% of users are proving what today's models can do, automating real code and compounding their own productivity. The other 99% haven't caught up. The binding constraint there is not the product. It's human and organizational inertia: people and companies change how they work far more slowly than a token-vs.-labor cost comparison says they should. There is a large pool of demand hiding in plain sight and, as that gap closes, spend rises on capability that already exists.

Will it hold? Believing this continues means believing token costs keep falling. Experts believe they will. Two independent engines drive the decline and neither is near its limit. Hardware gets about 40% more energy-efficient per year, and the compute needed for a fixed capability falls about 3x a year from better methods alone (Epoch).

Measured price elasticity by Google estimates demand elasticity around 1.7. Extrapolate the historical price declines forward and volume rises about 2,500x while revenue rises about 25x through the end of the decade That is roughly 125% annual growth, taking today’s ~$175B market to about $4.4T by 2030.

Is $4.4T reasonable? If the current tasks exposed to AI today were fully automated, AI spending would absorb $3.4T of annual US labor spend: $1.7T goes to tasks AI can already do end to end, and another $1.7T sits in copilot territory, where AI can power the work without fully replacing the worker. So while this doesn’t quite reach the $4.4T mark (and full automation of this entire labor pool is extremely unlikely), this notably doesn’t include the international opportunity (roughly 5x the US), not to mention the new surfaces that will be unlocked over the next several years. This makes the $4.4T more achievable.

The engine: rising capability.

Everything so far assumes the models stay roughly as capable as they are today and simply get cheaper and more widely used. That alone plausibly carries demand to the low end of what the buildout requires. But there is a second engine: capability. This could be the driver that separates a healthy market from a runaway one, as the evidence shows that incremental gains in AI capabilities have an outsized and compounding impact on value creation:

Part of what’s driving this is that, as models get smarter, each task consumes far more tokens. A reasoning model, one that works through a problem step by step before answering, burns about 10x the output tokens of a standard model on the same task. An agentic workflow, a model that plans, uses tools, checks its own results, and retries in a loop, can burn several thousand times the number of tokens per completed task compared with a single call. The cost per interaction reflects it: a typical 2023 workflow ran about $0.04, while a 2026 agentic one runs about $1.20, ~30x more.

So can capability keep compounding?

The best framework here is Leopold Aschenbrenner's (author of the seminal Situational Awareness paper and manager of the hedge fund of the same name, which grew from $200M to ~$24B AUM at peak in two years[6]). He shows that capability advances on three multiplying levers: physical compute (bigger training runs), efficiency (more capability per unit of compute, in hardware and in algorithms), and "unhobbling" (post-training techniques such as reinforcement learning, reasoning, agents, and tools that convert raw models into useful systems). The question is how much headroom each lever has. In order from most derisked to most contested:

Compute: on trend, while hardware efficiency covers the outer-years gap.

We are executing almost exactly on the maximum-progress trendline Aschenbrenner drew in 2024, with the 2026 milestone landing on schedule: xAI's Colossus 2, Anthropic and Amazon's New Carlisle site, Meta's Prometheus, OpenAI's Stargate Abilene.

Per Epoch AI, a prominent AI research institute, the largest sites now in construction reach roughly half of the forecast 2028 scale on a single campus (Meta's Hyperion at about 3.7M H100e and up to 5 GW; Microsoft's Fairwater at about 5M H100e and 3.3 GW), and multi-site training closes the rest of the gap. Google proved multi-site training years ago and has extended it to distant, low-bandwidth sites.

The 2030 milestone is the contested one: about 100M H100e per run against a total installed base CSIS projects at about 159M, and a headline 100 GW of power, ~40% of even the SemiAnalysis industry-wide bull case.

Those milestone figures, however, hold chip efficiency constant at H100 levels. As discussed, hardware gets about 40% more energy-efficient per year. Moreover, even more efficiency gains are expected to come from co-designing chips and models together: models trained in the number formats the silicon natively speaks and hardware specialized for the workload. Compounding those gains, Epoch and EPRI estimate the theoretical 2030 flagship training run would draw 4-16 GW, with about 6 GW as the central case; just 6% of the headline.

Algorithmic efficiency: derisked. Also mentioned above, pre-training efficiency improves about 3x per year, meaning the same capability costs a third the compute every year purely from better methods.

Data: the one real wall. Pre-training's fuel is human text, and we are running out. The indexed web holds about 500T tokens of unique text, and researchers expect the high-quality portion to be fully used between 2027 and 2029 (Epoch; Sutskever). Epoch's bull case traverses the wall: multimodal data (image, video, audio) plausibly triples the effective stock, and synthetic data could remove the ceiling entirely.

The people running the labs are more wary. Ilya Sutskever (OpenAI co-founder and Founder/CEO of Safe Superintelligence) claimed in 2024 that "pre-training as we know it will unquestionably end," comparing data to fossil fuel. Google CEO Pichai says the low-hanging fruit is gone. As a result, the returns to pure pre-training scale are softening.

Unhobbling is why capability did not stall, and it is where the variance lives. Sutskever named the paths forward himself: agents, synthetic data, inference-time compute. All of them lead to reinforcement learning (RL). If pre-training is reading everything ever written, RL is practice: the model attempts real tasks, gets graded on outcomes, and reinforces what worked. It converts raw compute into capability without new human text, a partial answer to the data wall, and it trains exactly what pre-training could not: judgment and multi-step reliability. The industry has moved its marginal compute dollar from pre-training to RL, just as Aschenbrenner predicted before the first reasoning models shipped.

So is it working? The metric that fits the new paradigm replaces test scores with task length: how long a task can a model complete autonomously and reliably? METR measures this horizon doubling about every seven months from minutes-long tasks in 2023 to hours-long today. If the trend holds, tasks that take a person days come into range this year, and week-long tasks in 2027.

Stacked S-curves frame what is at stake. If intelligence follows the S-curve of every prior technology (slow start, steep ramp, plateau), frontier progress stalls, and once today’s capability set diffuses through the economy, adoption stalls with it. We watched the first curve flatten in 2024, when returns to pre-training scale diminished against the data wall. RL started a second curve, and we are on its steep part now. So the capability question is really two questions. How much more can we squeeze out of the RL curve? And if its gains diminish too, does a new training method or architecture start a third curve before this one flattens? The answer is the difference between the red line below and the blue and green ones.

To evaluate this more granularly, we need two things: which work falls, in what order, on what timeline; and what has to be built to keep pushing the frontier outward.

The task frontier: work falls in the order it can be graded

Capability does not arrive as one thing on one date. It arrives in tiers, and the tiers are ordered by a single question: how hard is it to tell whether the machine did the job right? That ordering explains almost everything about the sequence we have already seen. Coding came first because code runs or it doesn't. Judgment work comes last because two experts disagree about it.

Tier 1: Checkable work. Falling now.

  • What it is: tasks a program can grade objectively. Code passes the test suite or it does not. A proof checks or it does not. This covers essentially all software, math, and much of quantitative science. RL with verifiable rewards (RLVR, where an automated checker provides the grade) has already produced expert-level or superhuman performance across this tier, and plausibly carries coding to full automation on its own.
  • Timeline and impact: now. ~10% of knowledge labor spend (~$600B-1.2T of the US knowledge-labor pool; $3.5-5T of the global pool)

Tier 2: Answer-key work. Falls 2027-28.

  • What it is: a right answer exists, but it arrives as free-form language no program can grade mechanically. "Write a function that deduplicates this list" is graded by a test suite in milliseconds. "A 55-year-old presents with joint pain, fatigue, and a facial rash; what is the diagnosis?" has one correct answer wrapped in a paragraph of reasoning, and something must judge that paragraph against the key. Much of medicine, chemistry, and applied economics lives here.
  • The gate: Requires reliable model judges plus large curated sets of expert reference answers that mostly do not exist yet.
  • The fix: this is engineering; no major breakthroughs required
  • Timeline and impact: the International AI Safety Report 2026 forecasts major progress in reasoning-based problem solving by 2027-28. Another ~20% ($1.2-2.4T US; $7-10T globally).

Tier 3: Rubric work. 2028-29 onward, domain by domain.

  • What it is: no answer key exists, but experts agree well enough on what good looks like to write it down as a rubric (i.e. diligence memos, client emails, design reviews). This is most of what most knowledge workers do every day.
  • The science already works. Rubric-based RL went from a proposal in mid-2025 to matching expert-written rubrics on health tasks and powering near-frontier research agents within about a year. Anthropic took open-ended task success from 25% to 76% in nine months. And perhaps the most surprising: models drafting their own rubrics matched physician-written ones, which suggests the cost of encoding "good" falls over time.
  • The gates.

1.    Encoding judgment at scale: someone must decompose "good" into rubrics across hundreds of domains. Requires thousands of experts doing slow work.

2.    Reward hacking: AI is a ruthless loophole-finder, and fuzzy graders are easier to exploit than test suites. Environments must be patched faster than models learn to game them.

3.    The compute bill: RL needs millions of practice runs per skill, which is fine when a run is a math problem but expensive when it is a 40-minute session in a simulated enterprise software stack.

4.    Calibration: knowing when to stop and ask a human. Benchmarks reward attempting not abstaining, so models learn overconfidence. The risk is calibration stays good enough only for supervised use, which would gate deployed hours and revenue even as benchmark scores climb.

  • The fix: data collection and engineering throughput. Expert pipelines, industrial-scale environment construction, and brute-forcing RL's sample inefficiency with volume. The uncertainty is schedule and cost, not science.
  • Timeline and impact: end of the decade for leading domains is reasonable, with a long tail behind. The METR trendline cross-checks the date: doubling every 7 months mechanically reaches day-long tasks around 2027 and week-long tasks around 2028-29, and long tasks are built from exactly this kind of self-judgment. The falsifier to watch: the METR curve bending outside code. The largest slice at ~50% of knowledge labor spend (~$3-6T US; $17.5-25T globally).

Tier 4: Taste work. Early 2030s for automated AI research; mid-2030s for meaningful coverage elsewhere.

  • What it is: senior experts genuinely disagree about what good looks like, and feedback arrives in years. Strategy, organizational design, research direction, negotiation, etc.
  • The precedent: AlphaFold. Predicting protein structures resisted fifty years of human effort and was filed under scientific intuition. DeepMind cracked it well enough to win the 2024 Nobel Prize in Chemistry. Machine judgment can conquer domains we call “taste”.
  • Where the analogy breaks: AlphaFold had a grader where every prediction was scored against experimental ground truth. Business strategy has no grader, so you train in a simulation of the world, and the model gets good at beating the simulation. When a strategic call takes years to score and the sample size is a handful, you cannot tell whether the simulator taught the right lessons.
  • The gates: the only genuinely unsolved problems in the stack. Cross-session memory (no proven mechanism yet for a model to accumulate months of context the way an employee does), continual learning (unsolved as research and as a business model), and simulators whose lessons survive contact with reality.

Timeline: the split inside the tier comes down to whether you can cheaply check that practice-world skill transfers to the real world. Automated AI research can, because experiments produce checkable results in days (labs are attacking this first). Strategy and social judgment cannot, which is why domain experts consistently put them last, usually 5-15 years behind technical work. This tier is a seniority slice across occupations: only a small % of labor hours, but the highest-paid hours in the economy: ~20% of spend (~$1.2-2.4T US; ~$7-10T globally).

The takeaway is the sequence. We are through the first tier, in the middle of the second, and the third is where the money in this decade sits. The fourth is the long tail of remaining work left to be automated that will be the focus of next decade, but as the next section shows, isn’t load-bearing for the broader argument.

The buildout pays for itself if AI gets good at work you can check (Tiers 1 and 2). Take today's market, compound by demand expansion from falling token prices ($4.4T). Then add US checkable work, cut in half for adoption speed, and you land at $5.3-6.2T by 2030. The street bar is $1.6T: the price-decline baseline clears it nearly three times over on its own, before any new capability. The aggressive bar is $5.8T: the full stack reaches it at the midpoint and clears it across the upper half of the range, with only the easy problems solved and only the US counted. The two pools ignored so far, international checkable work (~5x the US pool) and the first rubric-tier use cases opening around 2028 (that Anthropic's models are already making strong progress toward), carry the aggressive bar the rest of the way. A modest slice of either secures it even at the bottom of the range.


Part I Conclusion

I do not think this is a bubble on a three to five year view. This was the infrastructure question: Will we build more than we can use? I don’t think we will. While there are no certainties in this market, I believe that everything that can be built will be funded, and what can be built is much more than the market is pricing in. Demand is here, it is real, and it expands on two legs: falling prices deepen usage and unlock new markets, and rising capability pushes work into higher-value, more compute-hungry reasoning and agentic tasks on the path to full knowledge-work automation. We will execute reasonably well on capability advancement (at least for the foreseeable future), so the capacity we’re building now gets absorbed over the near-to-mid term.

The financial question is harder, and I won’t pretend to settle it here. History is blunt on this point: most of the money going into these companies today will be destroyed. In the dot-com era, for every Amazon or Google that grew into many times its peak, tens or hundreds of others never came back (eToys.com and Webvan went from billions of market cap to zero). But here’s the other side: we are very early in the adoption curve, which means the eventual winners have enormous room left to run. That’s the whole point of doing this work. Build the world view now, pick the winners early, and ride them up the adoption curve. That’s where the returns are.

 

Appendix: What has to be built to progress along this curve?

I bring up this question not for completeness of the core argument, but because it is important to understand where value might accrue in the software stack—and a clearer criteria set for the categories in Part III. Every capability gate above reduces to four inputs:

  1. Verification supply: data, rubrics, environments. The most valuable remaining knowledge has never been written down. It lives as undocumented process inside organizations and judgment inside heads. Frontier training data increasingly has to be manufactured: extract the judgment (expert demonstrations, workflow recordings, rubrics) and build the practice worlds to run it in (simulated codebases, browsers, and enterprise stacks with gradable tasks).
  2. Algorithms: manufacturing intermediate feedback so long tasks stop being long. A week-long task emits one pass-fail signal at the end, and credit has to reach the right decision among ten thousand. Solved in principle; per-domain application and the compute cost remain. Models are wildly sample-inefficient compared with humans, but volume substitutes for efficiency. The genuinely unsolved piece is continual learning, and it is unsolved twice: as research (models forget, and learning from live data invites poisoning) and as a business model (whose data can be used, what consents do vendors need, who captures the value). Cracking it would turn every customer deployment into free training and change the economics of the industry.
  3. Harnesses: the software wrapped around the model. How tasks get decomposed, how the model checks its work, saves progress, and recovers from errors. A meaningful share of METR's measured horizon gains came from harness improvements. Stanford's Meta-Harness studythe Stanford and MIT Meta-Harness paper found a 6x performance gap from harness choice alone on a fixed model, and showed harness design itself can be automated. The weak spots: recovery from novel failures in messy environments and calibration, which is now a stated lab priority because enterprise agent deployments fail on exactly this.
  4. Memory and context. Within a task: largely solved at the harness level. Agents compact their history, write notes to consult later, and delegate bounded chunks to fresh sub-agents. Across sessions: open. Every session still starts from scratch. This is the gate Tier 4 waits on.

The next question is who is positioned to deliver these advances, and who captures the value when they do.

Notes & Disclaimers:

[1] This material is distributed by VNTR Securities, LLC, member FINRA/SIPC, a subsidiary of Manhattan Venture Holdings, LLC (d/b/a Manhattan Venture Partners, “MVP”). MVP Manager, LLC, an affiliate, is an investment adviser registered with the Securities and Exchange Commission; registration does not imply any level of skill or training. This material is for informational purposes only, reflects the author’s views as of the date shown, and is not investment, legal, tax or accounting advice. It does not take into account the objectives, financial situation or needs of any recipient, and no statement here is a recommendation to any particular person to buy or sell any security. This material is distributed by VNTR Securities, LLC only, through its registered representatives.

[2] This analysis contains forward-looking statements, estimates and projections, including market-size estimates, adoption timelines and growth rates. These are opinions based on assumptions that may prove incorrect; they are not projections of the performance of any MVP fund, account or investment, and no reader should rely on them as an indication of any investment result. Actual outcomes may differ materially, and nothing here is a promise, guarantee or prediction. Certain information, including charts and statistics, was obtained from third-party sources identified in the text; MVP has not independently verified it, does not represent that it is accurate or complete, and presents it as of the dates indicated.

[3] Component inflation is an accelerant: Cost per GW has doubled to tripled since the start of the year, per Gavin Baker. Microsoft pinned about $25B of its raised 2026 capex guidance on component prices alone, and Meta about $10B.

[4] One refinement: a growing share of builds are lower-cost custom-silicon campuses at $30-35B per GW, which need only $13-15B per GW per year. At a realistic mix the bars below run 10-15% lower.

Ready to partner with MVP?