Skip to content

Understanding AWS Bedrock Pricing Before It Understands Your Wallet

AWS Bedrock pricing is straightforward in theory and surprising in practice. You pay per token — input tokens and output tokens, priced separately, varying by model. Simple enough until you are running five models across three environments and nobody knows which team is burning what.

The pricing model

Every Bedrock model charges differently. Here is a simplified view:

Model Input (per 1K tokens) Output (per 1K tokens)
Claude Sonnet $0.003 $0.015
Claude Haiku $0.00025 $0.00125
Model Input (per 1K tokens) Output (per 1K tokens)
Titan Text Express $0.0002 $0.0006
Titan Text Lite $0.00015 $0.0002
Model Input (per 1K tokens) Output (per 1K tokens)
Llama 3 70B $0.00265 $0.0035
Llama 3 8B $0.0003 $0.0006

The spread between the cheapest and most expensive option is 75x on output tokens. That is not a rounding error — it is the difference between a $50 monthly bill and a $3,750 one for the same traffic pattern.

Where costs hide

The bill is always bigger than you expect

The obvious cost is the per-call token charge. The hidden costs below are what actually blow budgets.

  1. Retries — a failed call still costs tokens for the input
  2. Verbose prompts — system prompts that never change but get billed every call
  3. Model overselection — using Sonnet for tasks Haiku handles fine
  4. Missing caching — prompt caching can cut input costs by 90% for repeated prefixes
import boto3

bedrock = boto3.client("bedrock-runtime") # (1)!

response = bedrock.invoke_model(
    modelId="anthropic.claude-sonnet-4-20250514",
    body='{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 100}',
    contentType="application/json",
)

usage = response["ResponseMetadata"]["HTTPHeaders"] # (2)!
print(f"Input tokens: {usage.get('x-amzn-bedrock-input-token-count', 'N/A')}")
print(f"Output tokens: {usage.get('x-amzn-bedrock-output-token-count', 'N/A')}")
  1. Creates a Bedrock runtime client using your default AWS credentials and region.
  2. Token counts are returned in HTTP response headers, not the body — easy to miss.

What I built to solve this

I got tired of guessing, so I built two tools:

Bedrock Cost Tracker

A decorator-based library that wraps every Bedrock call and logs token usage plus USD cost to a local SQLite database. Zero config — just add the decorator:

from bedrock_cost_tracker import track_cost

@track_cost
def summarize(text: str) -> str:
    response = bedrock.invoke_model(...)
    return response["body"].read()

Then run bedrock-costs report to see a breakdown by model, by day, by function.

Bedrock Model Compare

A CLI tool that runs one prompt across multiple models and generates a cheapest-first Markdown report with latency, token counts, and cost per call. Useful for deciding which model to route to.

The takeaway

Bedrock costs are predictable if you measure them. The problem is that most teams do not measure until the bill arrives. Instrument early, route intelligently, and cache aggressively.