The US Army Burned a Year of AI Tokens in a Month - a Warning for Every Enterprise
The US Army rolled out near-unlimited AI via Ask Sage - then blew through its 100-million-token annual pool in about a month and had to reinstate limits. Why it happened, and the AI cost-planning lesson for every organization.

Table of contents
In May, the Pentagon was celebrating: roughly half of its 3.5 million employees were reportedly using AI, and the US Army had rolled out effectively unlimited access to large language models through Ask Sage, a platform that routes requests to models from Google, Meta and OpenAI. By mid-June, that "unlimited" access was gone. The Army had blown through its annual token allotment in about a month, and staff started receiving emails asking them to cut back.
It's a small story with a very large lesson for anyone deploying AI at scale.
What actually happened
The Army's subscription included an annual "enterprise pack" of about 100 million tokens — the units of text an LLM reads and writes. Usage climbed far faster than anyone planned: reporting (via Wired) describes the Army's CIO reinstating limits on token use after the pool drained, on a rollout valued at around $49 million. Rather than immediately buy more capacity, the response was to ration — nudge the idle, throttle the eager.
One internal example shows how quickly agentic and bulk workloads burn resources: an AI review of position descriptions reportedly tied up multiple GPUs at a cloud provider for eight hours. Token and compute cost, in other words, started dictating which use cases the Army could actually scale.
Why this matters beyond the military
Strip away the uniform and this is the story of every enterprise that treats AI usage as free. A year of budget lasting thirty days isn't an adoption triumph; it's a capacity-planning failure. The Army pushed usage as proof of transformation — leaderboards, top-ups, nudge emails — without modeling what "everybody uses AI" actually costs when each query, and especially each autonomous multi-step agent, meters tokens.
That's the same warning we've made about AI agents becoming infrastructure that needs cost controls: the moment agents run long, unattended tasks, spend stops being a rounding error. It also reframes vendor choice — comparing cloud AI providers for business use has to include realistic token-burn projections, not just headline per-token prices.
The takeaways for anyone rolling out AI
- Model the cost before the mandate. "Everyone should use AI" is a budget decision, not just a culture one.
- Watch agents especially. Long-horizon, multi-step agents can consume tokens (and GPU time) at a rate single prompts never approach.
- Instrument usage from day one. Per-team metering, caps and alerts turn a surprise outage into a managed cost.
- Plan for the demand curve, not the pilot. Enthusiastic early adoption is exactly what breaks an unmetered plan.
Bottom line
The Army said AI would transform national defense. What it actually demonstrated first was that unmetered AI transforms your budget forecast — into fiction. For any organization about to declare its own "unlimited" AI program, the cheapest lesson available is someone else's: the meter is always running, and enthusiasm is the fastest way to drain it.
Reporting: Wired (via Techmeme), AI Weekly. Figures are as reported and may be updated as more detail emerges.


