Huawei filed that sentence under a networking trend, as one bullet among ten. Read it as a market size and it is unverifiable. Read it as a ratio and it is checkable, load-bearing, and already behind schedule.
It is a single sentence in Huawei's Intelligent World 2035 report, sitting in Trend 9, the "agentic Internet," between a section about storage and a section about energy.
"By 2035, the global population will reach 9 billion and be served by 900 billion AI agents."
Read that as a market size and it sounds like every other vendor forecast: enormous, unfalsifiable, and safe to ignore. Read it as a ratio and it stops being a forecast at all. It becomes a staffing plan.
The whole forecast, as a ratio
Same axis, same unit, one scale. The human bar is the thin black one.
Huawei is a networking company. Trend 9 is about networks, and in the next paragraph it explains what a network needs to do about it: connect every person on Earth, "as well as a staggering 900 billion agents," through something Huawei names the agentic Internet, in order to absorb "the 100-fold increase in data traffic in the next 10 years."
That is the part worth stopping on. A vendor with a decade of optical and wireless roadmaps just told the market that its growth case is agents, and that the volume is not 10% or 30% above human traffic. It is 100 times the current shape of the internet.
The same week, Reuters reported the forecast's other half: that agents will account for more than 90% of global AI token traffic by 2035, that the report "makes no call to slow development," and that Huawei lists security and privacy protection for autonomous agents as one of ten technologies required to support it.
One company's forecast is not evidence of anything. So here is the test we ran instead: take Huawei's ratio, put it beside the only independent measurements that already exist, and see whether the world is on schedule. It is not on schedule. It is ahead.
In February 2026, agent traffic on OpenRouter passed human traffic. It has since grown roughly 14x. That means the crossover everyone treats as a 2035 question already has a date, and the date is behind us.
The milestone, dated
No modelled curve here. Three reported points on one line.
This is where most coverage will go wrong. A forecast of "90% of traffic by 2035" will be filed as a prediction, compared against other predictions, and forgotten by Thursday, because nobody can be held to a number ten years out. But the variable that decides the headline is not on a ten-year clock. It is on a six-month clock, and it has already moved.
There is a second reason the number is smaller than it looks. Agent traffic and human traffic are not the same kind of unit, and traffic share is not population share. Goldman Sachs modeled the structure in a 41-page note on May 5: an average LLM chatbot session is roughly 1,000 tokens, and the 2025 average query runs about 1,715 tokens. An agent does something categorically different. It runs a loop.
Why one agent is not one user
Both bars to the same scale: one session for a person, a full day for a single agent.
Parse intent, plan, call tools, read results, validate, retry, update state, repeat. Goldman's simulated email-monitor agent runs 24 hourly scans, a 50-loop triage pass, a 12-loop reply pass, and a scheduling sub-loop. That is a single agent doing nothing you would recognize as work, and it is burning a hundred times a chat session before it has made a decision.
So when Huawei says 90% of token traffic, it is not saying 90% of activity. It is saying that the parts of the economy that talk to machines will be outnumbered, unattended, and continuous. Half your traffic being agents and half being people is entirely compatible with most people still using AI the way they do now.
Goldman built the forecast from the bottom up, simulating travel booking, email monitoring, coding, data entry and call centers in pseudo-code, counting tokens per task and tokens per day, then layering adoption curves on top. Knowledge-worker adoption lands near 12% by 2030 and 37% by 2040. The output of that exercise is the number the infrastructure industry is actually building against.
Global token consumption, on one scale
Quadrillion tokens per month. Linear, so the 2026 bar is honestly a sliver.
144 quadrillion. That is the 2030 gap between consumer agents (~60q) and enterprise agents (~56q) summed, and it is the same order of magnitude the entire human internet produces in a year. Enterprise workloads pass 70% of the total by 2040. Source: Goldman Sachs via Asymmetric Curiosity, grade B.
Divide 120 quadrillion tokens a month by the number of seconds in a month and the industry's 2030 target comes out at roughly 46 billion tokens per second, sustained, forever, across everything. That is not a spike. That is a utility.
Two things follow from that table, and they point in opposite directions.
The first is that the demand case for the build-out is no longer speculative. It is the arithmetic that lets $725 billion in annual hyperscaler capex, up roughly 77% year over year, look like a plan instead of a bet. Whoever signs off on a data center needs exactly this number, and now they have it, with a bank's name attached.
The second is that the same numbers demolish the pricing model. Goldman notes compute cost per token from leading silicon is falling 60% to 70% per year. Compounded across the nine years between now and Huawei's 2035, even the conservative end of that band is a collapse of roughly 1 in 3,800. At 65% a year, it is 1 in 12,688. And the money being spent assumes tokens go from cheap to nearly free while volume goes from large to infinite.
The price of a token, logged
Y axis is logarithmic. Today's price is 1. The band is the 60-70% annual decline range.
Read the last label carefully. A person reads about five words a second and then sleeps. No amount of human attention supplies a volume curve that survives a price collapse of this size. The customer has to be software. Source: 60-70% silicon decline, grade B in, and the compounding is our arithmetic. Band edges are labelled; the solid line is 65% per year.
This is the sentence that reframes the whole story: 90% agent traffic is not a prediction of the future. It is a precondition of the present. Diverging prices and rising volumes only both work if the customer is software. The forecast and the capex are the same claim. Huawei is not forecasting agent dominance so much as describing, in a networking company's vocabulary, the only world in which anyone's current spending closes.
Now do what Huawei did and take the number literally, but follow it one step further than the report does. If a hundred agents per human actually arrive, every one of them needs the things a thing needs before it can act on your behalf. An identity. A permission set. Somewhere to run continuously. A record of what it did. A way to be stopped.
Huawei's own report concedes this is unsolved. Security and privacy protection for autonomous agents is one of the ten technologies the report says are required. Item one of ten, from the company selling the network. That is not a talking point. That is a requirements list with an empty box at the top.
Put a number on the box. Take 900 billion agents, assume each makes a modest 100 tool calls a day, and assume each call writes a single kilobyte of log.
How the audit trail gets to 92 petabytes
Three multiplications. Bar lengths are not to scale; the arithmetic is.
And the alternative to storing it is worse. An agent economy that keeps no log is not a rogue-agent story, it is an unattributable one. When an agent moves money, rewrites a record, or sends something it should not have, the only three questions that matter are who it acted for, what it was permitted to do, and what it did. Without a per-call record, none of them has an answer.
That is why the safety debate keeps failing to land. It is argued in the register of catastrophe, and catastrophe is easy to wave away. Storage and retention are boring, measurable, and impossible to wave away. The honest version of the autonomy problem is a capacity problem: who keeps the record, who deletes it, who may read it, and who is accountable when it is gone.
On September 15 the House moved the Ratepayer Protection Act under suspension of the rules. It concerns electricity for large data centers. That is the visible half of the cost of running a hundred agents per human, and it is now a live legislative question with sponsors and a floor vote.
The invisible half does not have a sponsor. When an agent runs on the machine in front of you instead of in a facility, the compute bill arrives as depreciation. We have the receipts on one desk, running a handful of agents for three weeks:
Nothing here is exotic. It is what the architecture does. The upstream trackers have the same story in public, filed as bug reports rather than as cost: memory issues 11315, 67433, 79815, 84960 and 86984 on anthropics/claude-code; disk-write issues 30612, 32496, 35401 and 35482 on openai/codex.
Scale that by Huawei's number and the arithmetic is not ambiguous. Ninety-two petabytes a day cannot be written to a laptop. It can only be written to a facility, by software that lives there. The forecast is not just a demand signal for networks and power. It is a demand signal for somewhere for the agent to live.
Goldman's model separates agents into two species, and the gap between them is the most useful number in the note.
| Agent type | Token load | What it implies |
|---|---|---|
| On-demand | >10,000 / session | You fire it, it finishes, it stops. A tool. |
| Embedded copilot | >5,000 / active day | It lives inside an app someone else runs. |
| Always-on | >100,000 / day / user | It runs continuously and acts when needed. This is the multiplier Goldman flags as larger. |
The 24x and 55x growth figures in the chart in section 02 are produced almost entirely by the bottom row. On-demand agents do not get you to 120 quadrillion tokens a month. Persistent ones do.
So the honest question raised by Huawei's 900 billion is not which model wins, or how good agents get. It is whether the thing running your agent is something you switch on, or something that stays on. A model is a component. An environment is the product. The distinction sounds like positioning until you try to run something continuously and watch it compete with your own laptop for memory, disk and thermals.
Which is the whole reason the ratio matters. A hundred agents per human is not a fleet anyone operates by hand, opening windows. It is closer to a population, and populations need somewhere to live: state, addressability, a record, and a way to reach the person they work for. That is the layer this decade will actually be decided on, and it is a layer the current debate is barely discussing.
One. Does the always-on number hold up? Goldman's >100,000 tokens per day per user is the multiplier behind the whole curve. Watch for the note itself, and for anyone reproducing the figure without citing the note. If the number softens, the 24x softens with it.
Two. Does 100:1 survive contact with identity? A hundred agents per human requires a hundred sets of credentials, permissions and revocations. Watch for the first serious standard, incident disclosure, or regulator that treats an agent as a principal with its own audit trail. An empty box at the top of Huawei's list is an opportunity, and someone is going to take it.
Three. Who publishes the first number that would prove them wrong? Every participant here is a seller of compute, networks, models or environments, including us. The industry's forecasts will improve the day one participant publishes a falsifier with the same font size as the headline. Watch for that, and expect to wait.
| Claim in this piece | Source | Grade |
|---|---|---|
| "By 2035, the global population will reach 9 billion and be served by 900 billion AI agents" · "agentic Internet" · "100-fold increase in data traffic" · ten megatrends · 200+ workshops, 100+ experts over two years · going physical as the essential path to AGI | Huawei, Intelligent World 2035, huawei.com/en/giv and the Huawei newsroom release | A |
| 90% of global AI token traffic by 2035 · no call to slow development · agent security and privacy as one of ten required technologies · April security-ops platform · Guo Ping on Huawei compute as "China's Nvidia" | Reuters, "China's Huawei forecasts billions of agents will dominate AI traffic by 2035," Eduardo Baptista, Sept 16 2026, reporting Huawei's Intelligent World 2035: Turning Vision into Action | B for the reporting; A for the underlying report, which we have not read end to end |
| 24x tokens by 2030 to ~120 quadrillion/month · 55x by 2040 to ~278 quadrillion · chatbot session ~1,000 tokens · 2025 average query ~1,715 tokens · always-on agents >100,000 tokens/day/user · on-demand >10,000/session · embedded copilot >5,000/active day · 12% adoption 2030, 37% 2040 · consumer queries ~5B/day to ~23B/day · search share 68% to 36% · cost per query ~$0.075 to ~$0.018 · 60-70% annual silicon cost decline | Goldman Sachs, 41-page note published May 5 2026, as detailed by Asymmetric Curiosity | B · needs the note itself for A |
| Agentic token usage on OpenRouter surpassed human usage in February 2026 and has grown ~14x since | Two independent secondary reports of OpenRouter data (stack-archive.com, ai-brainer.com), both dated Aug 23 2026 | B · verify against OpenRouter directly |
| Inference at roughly two-thirds of global AI compute in 2026, versus one-third in 2023 and half in 2025 | Deloitte TMT Predictions 2026, cited second-hand | C · cited only as context, never load-bearing here |
| Roughly $725B combined hyperscaler 2026 capex, up ~77% year over year | Aggregated company guidance as summarized in the same secondary report | C · used qualitatively, no precise figure relied upon |
| 37 TB written in 21 days · ~65 GB/day · ~911 GB/hour RAM growth peak · 98°C CPU under load | Measured on our own hardware, disclosed as ours | A for the measurement; conflicted interest disclosed |
Public memory and disk-write issue numbers on anthropics/claude-code and openai/codex | The respective public issue trackers | B · cited as symptom reports, not as verified root causes |
| Ratepayer Protection Act moved to the floor under suspension of the rules, Sept 15 2026 | Our earlier reporting on the same week's legislation | A · primary bill text and floor record |
| Photographs: Huawei Shenzhen campus (x2); Ameca humanoid; NERSC data center rack (x2); technician with server rack; laptop on a desk; server rack | Wikimedia Commons, licences as credited in each caption (CC BY-SA 4.0, CC BY 2.0, CC0) | A · credit and licence reproduced under every image |
Grades. A primary document, filing, or first-hand measurement. B reported by a major outlet or a detailed single-source account. C secondary aggregation we could not trace to a primary document, used only where it is not carrying weight.
Derivations are ours and labelled. The 100:1 ratio, the 46 billion tokens per second, the 1 in 12,688 price collapse, the 92 petabytes per day and the 2035 extrapolation are all our arithmetic on the inputs in the table above. Assumptions are stated where they appear. None of them is a vendor projection and none should be quoted as one.
What would prove this wrong. If per-token cost decline flattens to single digits per year, the pricing argument in section 02 collapses. If always-on agent token load is an order of magnitude below Goldman's figure, the 24x and 55x curves collapse with it and Huawei's 90% becomes a networking sales pitch rather than a structural claim. If agent traffic growth on the major gateways decelerates through 2027, section 01 inverts. Write them down now so the scorekeeping is honest later.