← BlogAugust 30, 20267 min read

Nobody can quote you what an AI agent will cost next month

You can put a number on nearly everything else. The truck payment, the insurance renewal, the software the office runs on. They arrive on the same day for the same amount, and somebody writes to you before that changes. AI agents do not work that way, and this month three of the largest vendors made it official.

The seat price is a floor, not a price

Look at how the three biggest vendors sell agents right now and the same structure shows up in all three.

Anthropic sells Claude Team seats at $25 a month billed monthly, or $20 billed annually, with a premium seat at $125 monthly carrying five times the usage. Underneath that, once a person hits the usage included with their seat, an owner can switch on usage credits, which the company says are "billed at standard API rates." Its pricing page describes the Enterprise arrangement plainly as seat price plus usage at API rates. Microsoft sells Copilot Studio as tenant-wide packs of 25,000 Copilot Credits at $200 per pack per month, or on a pay-as-you-go meter with no commitment. Google moved the same way on August 26, adding a consumption edition of Gemini Enterprise with no upfront commitment and no base subscription fee.

So the seat price is still there and it is still predictable. It is just no longer the bill. The second half is decided by how much your people use the thing, and by how much work the thing decides to do on their behalf, and neither of those is a number you know in August about September.

What the meter is actually counting

Microsoft publishes a rate card worth reading even if you never touch Copilot Studio, because it is the clearest public description of how agent billing works anywhere.

A canned answer a person wrote in advance costs 1 credit. A generated answer costs 2. An agent action, meaning a trigger or a step of reasoning or a handoff between topics, costs 5. Grounding a response in your company's own indexed data costs 10.

The part owners miss is that one exchange can hit several of those at once. Microsoft's own example: an agent grounded in tenant data spends 12 credits answering a single complex question, 10 for the grounding and 2 for the generated answer. If it uses a reasoning model it bills twice, once at the feature rate and again for the reasoning at 10 credits per thousand tokens. And on the newer of Microsoft's two agent harnesses, billing begins the moment you start building rather than when you publish. Previewing, testing, generating evaluations, all of it spends money before anything is live.

Now run Microsoft's own worked example forward. It describes a support agent on a website handling 900 people a day, four canned answers and two generated answers each, which comes to 7,200 credits a day. Thirty days of that is 216,000 credits, or nine packs, or $1,800 a month. That arithmetic is mine, but every number in it is theirs. Nobody quoted you $1,800. You quoted yourself, by writing an agent that answers questions well and putting it somewhere popular.

The internal example in the same document lands very differently. An agent that watches for new orders, checks stock, confirms a ship date, approves, and emails the customer costs 20 credits per order, which at a hundred orders a day is a rounding error. That is the real lesson. The rate card is knowable and the volume is not.

Three vendors, one answer, and the answer is a shutoff

Here is the part that actually changed this month. All three companies have now shipped cost controls, and all three shipped the same one.

Google's August 26 release puts firm monthly spend caps at the project level in the Cloud billing console. Cross the line and agent API calls temporarily pause. You get automated emails at 50, 80 and 100 percent of budget, plus anomaly detection that flags a deviation and names the top three line items driving it, and you can optionally allow overage if you would rather keep running and pay for it.

Anthropic lets an owner set a monthly spend limit across the organization or per person, and checks it before every single request. Its documentation is direct about what happens next: exceed the limit and "you won't be able to use Claude, Cowork, or Claude Code again until the next billing period, or until your limits are adjusted."

Microsoft's enforcement kicks in at 125 percent of prepaid capacity, at which point custom agents are disabled. The current conversation finishes, and after that anyone who tries to use the agent gets one of two messages: "There is a billing issue" or "This agent is currently unavailable. It has reached its usage limit." Unused credits do not carry into the next month.

None of this is a criticism. When cost is driven by usage, a hard stop is the only control that actually works, and shipping one is more honest than not. But it changes what you are buying. The safety feature on offer is not a discount or a warning. It is an off switch, and you decide where to set it.

The failure mode is not the invoice

Owners worry about the surprise bill, and the surprise bill is the part these releases mostly solved. Alerts at 50 and 80 percent and a firm cap at 100 mean a runaway invoice now takes effort to produce. What is not solved is that the agent stops, mid-month, and the damage depends entirely on where you put it.

Put it in front of customers and the failure is invisible to you. Somebody who has never dealt with your company asks a question at four in the afternoon on the 22nd and is told there is a billing issue. They do not call to complain. They just go somewhere else, and the only trace is a number that did not happen.

Put it inside the business and the failure is quieter but slower to catch. The thing that reads inbound requests and files them into your system stops filing them. Two weeks later somebody notices the backlog, and by then the person who used to do it by hand has stopped checking whether it needs doing.

The monthly reset makes both of these worse in a way that is easy to miss. Capacity resets on a billing cycle, not on your business cycle. If your heaviest week is the last week of the month, that is precisely the week the meter runs dry.

There is a second trap in the shared pool. In Microsoft's model credits are pooled across the whole company unless somebody deliberately carves them up, which is why their documentation spends so long on environment-level allocation. A 60-person company has no dedicated ops person, so whoever is experimenting with a new agent on a Tuesday afternoon and the agent answering your customers are drawing from the same tank, and nobody set that up on purpose.

What to do next

None of this costs anything. Most of it is an hour and a decision.

  • Before a pilot, write down how many times a week you expect the thing to run and multiply it by the vendor's published per-action rate. If the vendor publishes no rate card, you have learned something useful about how predictable the bill will be.
  • Budget the pilot itself, not just the live version. On at least one major platform, building and testing an agent spends real money before a customer ever touches it.
  • Set the cap on day one, at a number you would be irritated but not alarmed to pay. Set it while the tool is still boring, because once it is useful nobody wants to be the person who capped it.
  • Decide in writing what happens when the cap hits. For anything a customer touches, "stop" is almost never the right answer, and "hand it to a person" is a path you have to build yourself. The vendor's version of that path is an error message.
  • Separate the budgets. Whatever runs in front of a customer should not draw from the same pool as somebody's experiment.
  • Put one person's name on the meter and have them read the consumption report once a month. It does not have to be the owner. It has to be a name.
  • Check the numbers again at 90 days. The first month of any tool is not representative, and usage of the things people like goes up.

The useful question here was never what an agent costs. It is what happens on the day it stops, and whether anyone in your building would notice.

Sources

Want this looked at for your company?

An AI Blueprint is a short scoping session and a written plan: which of your processes are worth automating, in what order, and what each one would take. Quoted in full before anything starts.

Get your AI Blueprint