a16z
For Limited Partners
June 2026

The Long Play
Portfolio Intelligence for Institutional Allocators
LPs looking to hire 3 roles

A small set of open LP roles shared with a16z. If you're hiring and would like to add to the list, please click the button below.

Organization Role Location Posting
Texas Permanent School Fund Head of VC Austin Link
Goldman Sachs Associate Various locations Link
Sammons Financial Group Head of Alts Geography Agnostic Link

Sends the request to the a16z team.

The Demand Curve

The Infinite Demand For Intelligence

"This is the biggest technological revolution of my lifetime. This is clearly bigger than the Internet. The comps on this are the microprocessor, the steam engine, and electricity. Or maybe the wheel." - Marc Andreessen
a16z
a16z Global Partnerships team · June 2026
Listen to this article
Speed

We are in the earliest phase of one of the most consequential infrastructure buildouts in modern history, and the reason is demand. Most discussion of AI demand stays in financial terms, revenue and market size, which makes it sound like a market cycle. It is better understood in physical terms, because AI demand has a unit.

That unit is the token: the small pieces of language a model reads and writes, one at a time, to do anything at all. Every answer, every document reviewed, every line of code, every step an agent takes is served as tokens. They are to AI what kilowatt-hours are to the grid, the unit the entire system exists to deliver, and what is happening to token consumption right now reframes everything else about this moment.

The Demand You Can Already Measure

Start with what’s measurable. AI company annualized revenue reached approximately $80 billion as of May 2026, up roughly five times year over year. Revenue is the financial expression of demand. Behind that number sits the raw expression, the tokens themselves, and the token figures are where the story stops sounding like a market cycle and starts sounding like physics.

Google's models processed 9.7 trillion tokens per month in May 2024. By May 2026, the figure was 3.2 quadrillion, an increase of more than 300x in two years.

To put this into perspective: that is the equivalent of the entire English Wikipedia roughly every six seconds.1

Tokens Processed Per Month · Google
May 2024 vs. May 2026
330×
in two years
9.7T
May 2024
3.2Q
May 2026
T = trillion · Q = quadrillion. On a linear scale, May 2024 is barely a sliver next to May 2026.
Tokens at that rate since you opened this page
0
0 English Wikipedias
Extrapolated from May 2026's volume of 3.2 quadrillion tokens per month — more than a billion every second, around the clock, for a single company's models.
For illustrative purposes only. Based on size of Wikipedia ~5 billion words, assuming that each word utilizes approximately 1.3–1.5 tokens.

That is happening today, before agents have fully arrived in force.

Agents Change The Shape of the Curve

A chatbot answers a question and stops. An agent plans, executes, checks its work, calls tools, and loops, and each of those steps is served in tokens. A single agentic task can consume ten to fifty times the tokens of the chatbot request it replaces. In coding, where agents arrived first, an April 2026 study found agents consuming up to 1,000 times the tokens of conversational use, with identical tasks varying in cost by up to 30x.

A knowledge worker equipped with agents will consume vastly more tokens than one issuing chat prompts today. While consumer agents are forming behind them: flight bookings, inbox cleanups, the always-on background tasks already appearing on smartphones, projected to push daily queries toward 11 billion.

We Are Measuring the Leading Edge

The most important feature about this demand curve is that most of today's measurable AI compute demand comes from a tiny sliver of what may be the eventual user base.

Today's demand is concentrated among developers, AI-native companies, and early enterprise agent deployments. There are roughly 30 million developers worldwide, each consuming on the order of 10 percent of a GPU in tokens. Our research suggests that the entire global developer footprint sums to about 3 million GPUs and roughly 4.5 gigawatts of power.2

Penetration of agentic AI across the broader workforce is still under 5 percent, and AI workloads do not scale with the developer population. They scale with the human workforce. The next phase is AI moving into the white-collar base: legal, finance, accounting, operations, roughly 1.5 billion knowledge workers worldwide, about 50 times today's developer base. Consumers, the billions who shop, bank, and travel, come behind them, with always-on agents projected to multiply consumer token consumption 12x by 2030.

Who Generates the Demand
Developers vs. knowledge workers · areas to scale (developers : knowledge workers = 1 : 50)
1.5B 30M
Circles are area-proportional (developers : knowledge workers = 1 : 50).
(1) a16z survey of 10 startups and enterprises engineering leaders regarding AI coding usage. Additional information provided upon request.
(2) World Bank, 2026.

This is why the honest answer to "how big does demand get?" is that demand for intelligence is limitless.

Cheaper Intelligence Means More Demand

There is an intuitive objection to all of this. For a given level of capability, the price of intelligence is collapsing. The cost of an LLM at equivalent performance has been falling roughly 10x per year, while pricing of frontier models has been stable. With the cost of a given capability plummeting, surely cheaper intelligence relieves pressure on the physical layer?

The economic term for this objection is the Jevons Paradox: when a resource becomes dramatically cheaper, total consumption rises, because cheap access opens categories of use that were not viable before. Cheaper transistors expanded chip demand. Cheaper bandwidth expanded internet traffic.

Cheaper intelligence can expand compute demand the same way, and the mechanism is already visible. Enterprises that could not justify $2,000 per engineer per month can spin up full deployments the moment the same intelligence costs a fraction of that. Workflows that were too expensive to automate become a workload. The addressable market for viable AI work will likely expand faster than unit costs fall.

This is the engine connecting the two curves of the moment. The price of a given unit of intelligence falls by an order of magnitude a year. Token consumption is projected to grow 24x across four years and the second follows from the first. Cost compression at the model layer intensifies pressure on the physical layer.

Demand Is Meeting The Physical World

Demand outpacing supply is only one side of the problem. There's also the physical problem. SaaS is compute bound too, but it hits diminishing returns on compute orders of magnitude earlier than AI does. For AI, more compute still means a better product, so it's electricity bound, bound by the physical infrastructure that enables it.

One example being power delivery itself. NVIDIA’s next-generation racks draw on the order of a megawatt each. At that density, the conventional AC electrical architecture that has run data centers for decades becomes increasingly inefficient. The fix is to switch from AC to DC distribution inside the facility to deliver power efficiently at that scale. This requires a new way to wire data centers that few electricians are currently trained to do.

Power delivery is only the start. Rack power is moving to 100 to 250 kilowatts to a MW, compute density is rising roughly 70x, and bandwidth requirements roughly 20x. Cooling is moving from air to liquid, campuses are moving from tens of megawatts to hundreds, and in some cases to gigawatt scale, with power sourced from behind-the-meter and captive generation as well as the grid. New data centers are already projected to need 44 gigawatts of additional power by 2028 against only 25 gigawatts expected to come online, a 19 gigawatt gap.

What This Means

Demand of this magnitude and shape was never what today’s infrastructure was designed to handle.

Every prior platform shift demanded a rebuild of its underlying compute infrastructure. The web rebuilt computing for traffic, the cloud rebuilt it for software delivery, and AI is demanding a larger and faster version of that shift, because the unit of demand is model inference itself.

The gap between compute demand and compute supply is the defining feature of this era.

1.Based on size of Wikipedia ~5 billion words, assuming that each word utilized approximately 1.3-1.5 tokens
2.Source: a16z survey of 10 startups and enterprise engineering leaders on AI coding usage. Additional information available upon request.
This newsletter is provided for informational purposes only, and should not be relied upon as legal, business, investment, or tax advice. Furthermore, this content is not investment advice, nor is it intended for use by any investors or prospective investors in any a16z funds. This newsletter may link to other websites or contain other information obtained from third-party sources - a16z has not independently verified nor makes any representations about the current or enduring accuracy of such information. If this content includes third-party advertisements, a16z has not reviewed such advertisements and does not endorse any advertising content or related companies contained therein. Any investments or portfolio companies mentioned, referred to, or described are not representative of all investments in vehicles managed by a16z; visit https://a16z.com/investment-list/ for a full list of investments. Other important information can be found at a16z.com/disclosures. You’re receiving this newsletter since you opted in earlier; if you would like to opt out of future newsletters you may unsubscribe immediately.