A small set of open LP roles shared with a16z. If you're hiring and would like to add to the list, please click the button below.
| Organization | Role | Location | Posting |
|---|---|---|---|
| Texas Permanent School Fund | Head of VC | Austin | Link |
| Goldman Sachs | Associate | Various locations | Link |
| Sammons Financial Group | Head of Alts | Geography Agnostic | Link |
We are in the earliest phase of one of the most consequential infrastructure buildouts in modern history, and the reason is demand. Most discussion of AI demand stays in financial terms, revenue and market size, which makes it sound like a market cycle. It is better understood in physical terms, because AI demand has a unit.
That unit is the token: the small pieces of language a model reads and writes, one at a time, to do anything at all. Every answer, every document reviewed, every line of code, every step an agent takes is served as tokens. They are to AI what kilowatt-hours are to the grid, the unit the entire system exists to deliver, and what is happening to token consumption right now reframes everything else about this moment.
Start with what’s measurable. AI company annualized revenue reached approximately $80 billion as of May 2026, up roughly five times year over year. Revenue is the financial expression of demand. Behind that number sits the raw expression, the tokens themselves, and the token figures are where the story stops sounding like a market cycle and starts sounding like physics.
Google's models processed 9.7 trillion tokens per month in May 2024. By May 2026, the figure was 3.2 quadrillion, an increase of more than 300x in two years.
To put this into perspective: that is the equivalent of the entire English Wikipedia roughly every six seconds.1
That is happening today, before agents have fully arrived in force.
A chatbot answers a question and stops. An agent plans, executes, checks its work, calls tools, and loops, and each of those steps is served in tokens. A single agentic task can consume ten to fifty times the tokens of the chatbot request it replaces. In coding, where agents arrived first, an April 2026 study found agents consuming up to 1,000 times the tokens of conversational use, with identical tasks varying in cost by up to 30x.
A knowledge worker equipped with agents will consume vastly more tokens than one issuing chat prompts today. While consumer agents are forming behind them: flight bookings, inbox cleanups, the always-on background tasks already appearing on smartphones, projected to push daily queries toward 11 billion.
The most important feature about this demand curve is that most of today's measurable AI compute demand comes from a tiny sliver of what may be the eventual user base.
Today's demand is concentrated among developers, AI-native companies, and early enterprise agent deployments. There are roughly 30 million developers worldwide, each consuming on the order of 10 percent of a GPU in tokens. Our research suggests that the entire global developer footprint sums to about 3 million GPUs and roughly 4.5 gigawatts of power.2
Penetration of agentic AI across the broader workforce is still under 5 percent, and AI workloads do not scale with the developer population. They scale with the human workforce. The next phase is AI moving into the white-collar base: legal, finance, accounting, operations, roughly 1.5 billion knowledge workers worldwide, about 50 times today's developer base. Consumers, the billions who shop, bank, and travel, come behind them, with always-on agents projected to multiply consumer token consumption 12x by 2030.
This is why the honest answer to "how big does demand get?" is that demand for intelligence is limitless.
There is an intuitive objection to all of this. For a given level of capability, the price of intelligence is collapsing. The cost of an LLM at equivalent performance has been falling roughly 10x per year, while pricing of frontier models has been stable. With the cost of a given capability plummeting, surely cheaper intelligence relieves pressure on the physical layer?
The economic term for this objection is the Jevons Paradox: when a resource becomes dramatically cheaper, total consumption rises, because cheap access opens categories of use that were not viable before. Cheaper transistors expanded chip demand. Cheaper bandwidth expanded internet traffic.
Cheaper intelligence can expand compute demand the same way, and the mechanism is already visible. Enterprises that could not justify $2,000 per engineer per month can spin up full deployments the moment the same intelligence costs a fraction of that. Workflows that were too expensive to automate become a workload. The addressable market for viable AI work will likely expand faster than unit costs fall.
This is the engine connecting the two curves of the moment. The price of a given unit of intelligence falls by an order of magnitude a year. Token consumption is projected to grow 24x across four years and the second follows from the first. Cost compression at the model layer intensifies pressure on the physical layer.
Demand outpacing supply is only one side of the problem. There's also the physical problem. SaaS is compute bound too, but it hits diminishing returns on compute orders of magnitude earlier than AI does. For AI, more compute still means a better product, so it's electricity bound, bound by the physical infrastructure that enables it.
One example being power delivery itself. NVIDIA’s next-generation racks draw on the order of a megawatt each. At that density, the conventional AC electrical architecture that has run data centers for decades becomes increasingly inefficient. The fix is to switch from AC to DC distribution inside the facility to deliver power efficiently at that scale. This requires a new way to wire data centers that few electricians are currently trained to do.
Power delivery is only the start. Rack power is moving to 100 to 250 kilowatts to a MW, compute density is rising roughly 70x, and bandwidth requirements roughly 20x. Cooling is moving from air to liquid, campuses are moving from tens of megawatts to hundreds, and in some cases to gigawatt scale, with power sourced from behind-the-meter and captive generation as well as the grid. New data centers are already projected to need 44 gigawatts of additional power by 2028 against only 25 gigawatts expected to come online, a 19 gigawatt gap.
Demand of this magnitude and shape was never what today’s infrastructure was designed to handle.
Every prior platform shift demanded a rebuild of its underlying compute infrastructure. The web rebuilt computing for traffic, the cloud rebuilt it for software delivery, and AI is demanding a larger and faster version of that shift, because the unit of demand is model inference itself.
The gap between compute demand and compute supply is the defining feature of this era.