Claude Code cost and usage: /cost, /usage, rate limits and what actually saves money

Claude Code Published:

How Claude Code billing works on a subscription versus API usage, checking consumption with /cost and /usage, how the rate limits behave, and the habits that cut token spend.

Verified on Sep 7, 2026 These tools change quickly. Please also check the latest official documentation.
Contents
  1. The two ways to pay
  2. Checking your usage
    1. /cost (for API billing)
    2. /usage (for subscriptions)
    3. The Console dashboard (API)
  3. How the rate limits behave
  4. What actually saves money
    1. 1. Match the model to the task
    2. 2. Keep the context short
    3. 3. Narrow what gets read
    4. 4. Cap non-interactive runs
    5. 5. Work with the prompt cache
  5. Managing cost across a team
  6. Summary

What Claude Code costs depends on which of two paths you're on. A subscription is a fixed amount with usage limits; the API has no ceiling but bills what you use. Either way, the habits that reduce waste are the same.

This article covers the two models, the commands for checking consumption, how the rate limits behave, and the savings that actually move the number. Prices and limits get revised, so treat the official pricing page as current.

KEY POINT

What you will learn

  • Subscription versus API billing, and how to pick
  • /cost, /usage, and the Console's usage view
  • Savings from model choice, context hygiene and non-interactive limits

The two ways to pay

ModelWho it's forBillingLimits
SubscriptionPro / Max for individuals, Team / Enterprise for organizationsFixed monthlyA limit per multi-hour window and a weekly limit
API usageAn Anthropic Console API keyPer input and output tokenNo ceiling; budget alerts available in the Console

Claude Code works out which one applies from how you signed in. /status tells you which is active.

Choosing between them:

  • Several hours a day favors a subscription, Max in particular: hitting the limit never produces an extra bill
  • A few short sessions a week is usually within Pro
  • CI, automation, or wanting per-team usage visibility favors the API

用語解説

Max tiers: Max comes in tiers with several times Pro's limits. If you lean on Opus-class models all day, the lower tiers reach their limit quickly.

Checking your usage

/cost (for API billing)

Shows what the current session has spent in USD, the input and output token counts, and elapsed time. On a subscription the dollar figure is indicative only.

/cost

/usage (for subscriptions)

Shows how much of your current limit you've consumed and how long until it resets. This is the one for "how much have I got left".

/usage

The Console dashboard (API)

The Usage view in the Anthropic Console breaks consumption down by day, by model and by API key. For team use, separate keys per purpose make the breakdown meaningful. Budget caps and alerts are configured there too.

How the rate limits behave

A subscription has a limit for concentrated use within a multi-hour window, plus a weekly limit. Reach it and Claude Code is unavailable until the reset.

What burns through a limit fastest:

  • Running an Opus-class model for everything
  • Re-reading huge files or logs repeatedly
  • Long sessions never cleared, since the history is resent every turn
  • Many subagents or parallel sessions at once

API billing has its own per-minute request and token limits. If you run things in parallel in CI, check your tier's limits in the Console first.

What actually saves money

1. Match the model to the task

/model switches mid-session. Routine fixes, adding tests and documentation updates are usually fine on a Sonnet-class model; reserving Opus-class for design decisions and hard debugging cuts consumption substantially.

/model claude-sonnet-5

Subagents can pin their own model in their definition file — see Claude Code subagents.

2. Keep the context short

Conversation history is resent on every request. Using /clear at task boundaries and /compact when a session gets long reduces the tokens per request on its own — see Managing context in Claude Code. A bloated CLAUDE.md costs you for the same reason.

3. Narrow what gets read

Ask for "20 lines either side of the error" rather than "read the log". Writing an instruction in CLAUDE.md to exclude generated trees such as node_modules and dist from searches helps too.

4. Cap non-interactive runs

For automation, limit the loop with --max-turns and record total_cost_usd from --output-format json — see Running Claude Code in CI.

5. Work with the prompt cache

Claude Code uses prompt caching internally so the repeated parts of a request, such as the system prompt and CLAUDE.md, cost less. Rewriting CLAUDE.md constantly undercuts that, so batch its edits rather than tweaking it mid-session.

Managing cost across a team

  • Separate API keys by purpose and read the breakdown in the Console
  • Set budget alerts so an unexpected climb surfaces early
  • Set a Sonnet-class default model in .claude/settings.json and let individuals switch with /model
  • Compare monthly: per-seat subscriptions often come out cheaper than API usage for the same work

Summary

  • Two models: a subscription with limits and no overage, or API billing with no ceiling
  • /cost for money, /usage for how much of the limit is left, the Console for day and key breakdowns
  • The basics of saving are model choice, /clear and /compact, and narrowing what gets read
  • In automation, --max-turns plus the JSON cost field stops a run from getting away from you

FAQ

What do I need to use Claude Code?
Either a subscription such as Claude Pro or Max, or an API key from the Anthropic Console billed per token. The free plan doesn't include it.
Why does a subscription stop working mid-session?
Subscriptions have usage limits per window and per week. They reset after a period; /usage shows how much is left and when.
On API billing, what does one session cost?
It varies widely by task and model. /cost shows the figure for the current session, so work from your own numbers rather than an estimate.

Primary sources

This article was drafted by AI from official documentation and reviewed by the site operator before publishing. Found a mistake? Let us know via the contact page.