Hard budget caps stop usage when spending reaches a limit you set. Alerts only send an email. Since July 2026, Google Cloud, AWS, OpenAI and Anthropic have all offered some kind of real cap. Each works differently, each has gaps, and only some are on by default. This guide covers what each one does when you hit the limit, where the gaps are, and how to set them up before a coding agent or a buggy loop spends money while you sleep.
On October 3, 2026, Simon Willison made the case in a post titled "We're going to need default hard budget caps on pretty much everything", which reached the top of Hacker News. He argues that caps should be the default and that removing one should require an explicit opt-in. He describes the familiar nightmare where a rogue service "consumed several hundred (or several thousand) more dollars" overnight. He also expects coding agents to make this more common, because agents now provision infrastructure and call paid APIs without asking. Defaults are a provider decision. Turning on the caps that exist today is up to you.
Alerts vs. hard caps: what changed in 2026
For most of cloud computing's history, a "budget" was a reporting tool. It told you that you'd overspent. It didn't stop anything. Three things changed this year:
- July 29, 2026: Google Cloud announced Spend Caps in public preview. They let you set "a monthly financial cap on specific services within a project."
- September 16, 2026: AWS launched a new builder experience. It includes per-project monthly spend limits: "If a project's usage reaches its spend limit, your project is paused for that month."
- LLM APIs: OpenAI documents hard spend limits at both organization and project level. Anthropic's Claude Console supports monthly spend limits per workspace.
Every one of these caps has to be configured, and the rules differ by provider. The table below compares them.
Provider-by-provider comparison
| Provider | Scope | What happens at the limit | Known gaps |
|---|---|---|---|
| AWS (new AWS Settings) | Per project, up to 10 projects | Project paused, all resources stopped, data preserved | Limited rollout; Paid Plan required; data deleted after 90 days if not reactivated |
| Google Cloud Spend Caps | One project + one service | Cost-incurring usage of that service restricted within minutes | Preview; only Gemini API, Agent Platform, Cloud Run, Cloud Run Functions |
| OpenAI API | Organization and/or project, monthly | Requests return HTTP 429 (organization_spend_limit_exceeded / project_spend_limit_exceeded) |
Not instantaneous; spend can slightly exceed the limit |
| Anthropic Claude API | Per workspace, monthly, below the org limit | Hard monthly cap (per Anthropic docs) | Can't set limits on the Default Workspace |
| Vercel Spend Management | Whole team, per billing cycle | Optional: pause production deployments for all projects (503) | Checks every few minutes; AI Gateway and v0 usage not stopped by pausing |
AWS: project-level limits with graduated controls
The AWS spend limit documentation is the most detailed of the group. Points that matter in practice:
- Availability is limited. The page warns that the new experience is being released "to a limited number of customers," so you may not have it yet.
- There's a minimum. The lowest limit you can set is the greater of $20 or a "conservative estimate" of your likely spend, based on month-to-date activity, running resources and the previous month. If many resources are running, you'll have to stop some before you can set a low cap.
- Credits are ignored. Limits apply to pre-tax charges and exclude credits.
- Optional early controls. You can opt into three escalating actions. About 7 days before you'd hit the limit, new resource launches are blocked through an AWS-managed service control policy. About 5 days out, idle EC2, RDS and SageMaker resources are paused. About 4 days out, your top cost drivers among EC2, RDS, Lambda, Bedrock and SageMaker are paused. AWS says the last one "might prevent you from incurring costs if you have a runaway Lambda or unexpected Bedrock spike."
- Hitting the limit is disruptive. The project pauses and all resources stop. To reactivate, you raise the limit, and some resources may need a manual restart. If you take no action within 90 days, AWS permanently deletes the project data.
Be careful with auto-scaling. AWS notes that once new launches are blocked, scale-out events won't start new instances.
Google Cloud: per-service caps, starting with AI and serverless
Google's Spend Caps are narrower but well targeted. They cover the services most likely to run away in an AI app: the Gemini API, Agent Platform, Cloud Run and Cloud Run Functions. Google calls enforcement "near real-time," triggering within minutes. It's non-destructive: data and resources aren't deleted, and services outside the cap keep running. You lift a cap with one click in the Budgets UI. Committed Use Discounts and Provisioned Throughput keep billing at their contracted rates. During preview, each cap covers one project and one service for a fixed monthly period.
The same announcement added early anomaly detection. It monitors early cost signals per service and can alert before costs reach your bill, and it names the top three SKUs behind a spike.
OpenAI: 429s instead of surprise invoices
OpenAI's spend limits guide separates alerts ("API traffic continues") from hard limits, where affected requests return a 429. Limits can be set at the organization level, which covers all projects, and at the project level. Both can apply to the same request. They reset monthly. The guide warns that "enforcement is not instantaneous," so recorded spend "can slightly exceed the configured amount," and that hard limits "can interrupt production traffic."
Your code needs to handle this. A 429 usually means "retry with backoff." If your client retries project_spend_limit_exceeded the same way it retries rate limits, it will loop pointlessly until the next billing cycle. Check the error code and fail fast.
Anthropic: workspace spend limits
In the Claude Console, each workspace has a Spend limits tab for capping monthly spend and setting alerts, and a Rate limits tab for requests and tokens per minute. Workspace limits must be lower than the organization's limits, and organization-wide limits still apply even if workspace limits add up to more. The detail people miss: you can't set limits on the Default Workspace. If all your keys live there, you have no workspace-level cap. Create dedicated workspaces such as dev, staging, production and agents, and scope keys to them.
The Claude Code workspace, created automatically when members sign in to Claude Code with Console accounts, is the only workspace that supports per-user monthly spend limits. That's useful if your team runs coding agents on API billing.
Vercel: a pause switch you have to turn on
Vercel's Spend Management is available on Pro and on Enterprise with Flexible Commitment. The docs are clear that setting a spend amount "does not stop usage on its own." You also have to enable Pause Production Deployments, which pauses production for every project on the team. Visitors then get a 503 DEPLOYMENT_PAUSED error. Spend is checked every few minutes, so projects can keep accruing usage after you cross the line. Pausing doesn't stop AI Gateway or v0 usage. Raising the limit doesn't unpause projects either; you have to resume each one manually. A webhook can handle softer responses, such as switching to a static fallback.
A setup checklist for builders
This order covers most solo founders and small teams:
- List everything with a metered bill. That means cloud accounts, LLM APIs, hosting, databases, vector stores, email and SMS. Include any account that a coding agent can access with your credentials.
- Isolate before you cap. Caps work at the project or workspace level, so split experiments, agents and production into separate AWS projects, GCP projects, OpenAI projects and Claude workspaces. Each experiment then has its own small cap.
- Use tight caps for sandboxes and agents, looser ones for production. AWS says its limits are designed for "experimentation, learning, and sandbox workloads," and suit production only if a brief pause is acceptable. For production, set the cap well above normal usage and rely on alerts plus anomaly detection to catch problems early.
- Set caps below your true maximum. Every provider here enforces with a delay of minutes. Give yourself room.
- Handle the failure mode in code. Detect spend-limit errors (OpenAI's
*_spend_limit_exceededcodes, a paused deployment, a stopped service), stop retrying, and show users a clear degraded state. - Route alerts to somewhere you'll actually see them. Email at 50%, 75% and 90% does nothing if it lands in a folder you don't read. Vercel supports SMS at 100% and webhooks. Send them to Slack or a pager.
- Write down how to recover. Note who can raise each cap, whether resources restart on their own (on AWS and Vercel, some don't), and AWS's 90-day deletion deadline.
To decide what a cap should be, estimate a normal month first. The LLM API Cost Calculator turns request volume and token counts into monthly spend across GPT, Claude and Gemini. The Token Counter shows what a single large prompt actually costs.
Caps for your own product, too
The same logic applies to what you sell. If your SaaS wraps an LLM, one customer's agent loop can burn through your margins before any provider cap kicks in, because the provider cap protects your account as a whole, not individual customers. Track spend per customer and per agent run: count tokens per request, keep a running total, and refuse work over a per-account ceiling. Willison's proposal, where removing the cap requires an explicit "I will be responsible for subsequent charges" checkbox, works for usage-based products too. Customers who know they can't be surprised by a bill are easier to sign up. That matters more as AI tools compete on trust; browse comparable products in the AI coding assistants directory to see how others present usage limits.
What to watch next
- General availability and wider coverage. Google's caps are in preview for four services. AWS's limits are on a limited rollout and act only on certain services before the full pause. Watch for GA announcements and expanded service lists.
- Defaults. No provider here caps new accounts by default, which is the change Willison is asking for. Vercel's own changelog has headlines like "Spend Management now enabled by default on Pro," so check the details of each provider's defaults rather than assuming.
- Agent tooling. AWS's new onboarding installs an Agent Toolkit by default. Expect agent frameworks to start reading cap settings and warning when they're about to provision into an uncapped account.
For now: create a separate project or workspace for anything an agent touches, put a low hard cap on it, and make sure your code fails fast instead of retrying when it hits the cap.
Frequently asked questions
Does AWS have a hard spending limit now?
Yes, but only in its new AWS Settings experience, which AWS says is rolling out to a limited number of customers. You need a Paid Plan, limits apply per project (up to 10 projects), and when a project hits its limit AWS pauses it and stops all its resources. Data is kept, but AWS deletes it permanently if you take no action within 90 days.
What's the difference between a budget alert and a hard spend cap?
An alert sends you a notification and lets usage keep running. A hard cap blocks or pauses usage once spend reaches the limit. OpenAI's docs put it this way: spend alerts let API traffic continue, while hard limits make requests fail with a 429 error. Classic cloud budgets were alert-only, so they couldn't stop a runaway bill on their own.
Can my bill go over a hard spend cap?
Yes, by a little. OpenAI says enforcement isn't instantaneous and recorded spend can slightly exceed the limit. Vercel checks spend every few minutes, and Google says its caps take effect within minutes. Set the cap below the most you're actually willing to pay.
How do I cap Gemini API spending?
Google Cloud's Spend Caps, in public preview since July 29, 2026, cover the Gemini API, Agent Platform, Cloud Run and Cloud Run Functions. Create one on the Budgets & Alerts page by choosing the Spend Cap budget type for a project and service. In preview, each cap covers a single project and service for a fixed monthly period.
Can I set a spend limit on Anthropic's default workspace?
No. Anthropic's docs say you can't set limits on the Default Workspace. Create a separate workspace, set its monthly spend limit on the Spend limits tab, and issue API keys scoped to that workspace.
Sources
- Simon Willison: We're going to need default hard budget caps on pretty much everything (Oct 3, 2026)— simonwillison.net
- AWS What's New: New AWS experience helps builders get started and ship faster (Sep 16, 2026)— aws.amazon.com
- AWS docs: Create a spend limit in AWS Settings— docs.aws.amazon.com
- Google Cloud blog: New early anomalies and spend caps on Google Cloud budgets (Jul 29, 2026)— cloud.google.com
- OpenAI API docs: Spend limits— developers.openai.com
- Claude Platform docs: Workspaces (spend and rate limits)— platform.claude.com
- Vercel docs: Spend Management— vercel.com
