How to Calculate Claude Code API Cost from Token Usage

A Claude Code session cannot be priced exactly from elapsed time and one combined token number. API billing separates input tokens, output tokens, prompt-cache writes, and prompt-cache reads. Each category has a different rate, and the final invoice is denominated in US dollars before tax and currency conversion.

The session duration—such as 16 minutes—does not multiply the model-token price. Time may matter for separately billed tools or infrastructure, but ordinary model API usage is token-based.

The General Formula

For a model whose prices are quoted per million tokens:

model cost USD =
  input_tokens       / 1,000,000 × input_rate
+ cache_write_tokens / 1,000,000 × cache_write_rate
+ cache_read_tokens  / 1,000,000 × cache_read_rate
+ output_tokens      / 1,000,000 × output_rate

Add any separately itemised web-search, code-execution, batch, regional, or fast-mode charges. Then convert the USD total using the card provider’s or accounting system’s actual exchange rate and add applicable VAT or fees.

Current Sonnet Rates

Anthropic’s current pricing page lists Claude Sonnet 4.6 at:

Token categoryPrice per million tokens
Standard input$3.00
Output$15.00
5-minute prompt-cache write$3.75
Prompt-cache read$0.30

The page lists the same four rates for legacy Sonnet 4.5. Prices and model availability can change, so preserve the model ID and pricing date used for any estimate.

Why “25.1k Tokens” Is Not Enough

Suppose a session summary shows 25,100 tokens but does not say which category they belong to.

If all 25,100 were standard input:

25,100 / 1,000,000 × $3 = $0.0753

If all were output:

25,100 / 1,000,000 × $15 = $0.3765

The exact model cost lies somewhere else if the number mixes input and output, and caching can make the relationship less intuitive. Therefore, quoting one precise euro amount from that display would create false accuracy.

As a simple mixed example, assume 20,000 ordinary input tokens and 5,100 output tokens:

input  = 20,000 / 1,000,000 × $3  = $0.0600
output =  5,100 / 1,000,000 × $15 = $0.0765
total  = $0.1365

At an illustrative exchange rate of €0.92 per US dollar, that would be about €0.126 before tax and card fees. Use the actual rate on the transaction date; €0.92 is an example, not a live financial quote.

Prompt Caching Changes the Calculation

Claude Code often reuses repository context and conversation prefixes. Prompt caching charges more when a reusable prefix is first written, then much less when that cached material is read again.

For Sonnet 4.6, 100,000 cache-read tokens cost:

100,000 / 1,000,000 × $0.30 = $0.03

The same quantity as ordinary input would cost $0.30. However, creating a new 5-minute cache entry for 100,000 tokens costs $0.375. A long session can therefore show a large volume of processed context while costing less than a naive “all tokens × input price” calculation—or more during repeated cache creation with few hits.

Subscription Usage Is Not the Same as API Billing

Claude Pro, Max, Team, or Enterprise access may include Claude Code usage under plan limits and rolling windows. Direct API usage through an Anthropic Console organisation is metered to that organisation. A third-party platform can apply its own credits or markup.

Determine which authentication path the CLI used:

  • Claude account plan;
  • Anthropic API key and billing organisation;
  • cloud provider integration;
  • third-party gateway or reseller.

Do not assume that seeing a dollar estimate means a separate card charge, or that paying for a chat subscription covers every API key.

Capture the Data Needed for an Exact Number

For each session or billing period, record:

  • exact model ID;
  • standard input tokens;
  • output tokens;
  • cache creation by duration where reported;
  • cache-read tokens;
  • service tier, regional or fast-mode multiplier;
  • separately billed tool calls;
  • credits and discounts;
  • invoice exchange rate, VAT, and fees.

Use the provider’s usage and billing console as the authoritative total. A local CLI summary is useful for estimation but can omit credits, delayed usage, or provider-side adjustments.

Reduce Cost Without Reducing Correctness

Keep repository instructions concise, exclude generated files and large irrelevant directories, start a fresh conversation when old context no longer helps, and use a smaller model for routine transformations. Preserve caching by avoiding needless changes to a large shared prompt prefix. Use the strongest model where reasoning prevents expensive mistakes rather than for every trivial edit.

The reliable calculation is category by category. Token direction, cache status, model version, and billing path matter far more than the elapsed session clock.

Related Guides