Understanding AWS Bedrock Pricing Before It Understands Your Wallet¶
AWS Bedrock pricing is straightforward in theory and surprising in practice. You pay per token — input tokens and output tokens, priced separately, varying by model. Simple enough until you are running five models across three environments and nobody knows which team is burning what.
The pricing model¶
Every Bedrock model charges differently. Here is a simplified view:
| Model | Input (per 1K tokens) | Output (per 1K tokens) |
|---|---|---|
| Claude Sonnet | $0.003 | $0.015 |
| Claude Haiku | $0.00025 | $0.00125 |
| Model | Input (per 1K tokens) | Output (per 1K tokens) |
|---|---|---|
| Titan Text Express | $0.0002 | $0.0006 |
| Titan Text Lite | $0.00015 | $0.0002 |
| Model | Input (per 1K tokens) | Output (per 1K tokens) |
|---|---|---|
| Llama 3 70B | $0.00265 | $0.0035 |
| Llama 3 8B | $0.0003 | $0.0006 |
The spread between the cheapest and most expensive option is 75x on output tokens. That is not a rounding error — it is the difference between a $50 monthly bill and a $3,750 one for the same traffic pattern.
Where costs hide¶
The bill is always bigger than you expect
The obvious cost is the per-call token charge. The hidden costs below are what actually blow budgets.
- Retries — a failed call still costs tokens for the input
- Verbose prompts — system prompts that never change but get billed every call
- Model overselection — using Sonnet for tasks Haiku handles fine
- Missing caching — prompt caching can cut input costs by 90% for repeated prefixes
import boto3
bedrock = boto3.client("bedrock-runtime") # (1)!
response = bedrock.invoke_model(
modelId="anthropic.claude-sonnet-4-20250514",
body='{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 100}',
contentType="application/json",
)
usage = response["ResponseMetadata"]["HTTPHeaders"] # (2)!
print(f"Input tokens: {usage.get('x-amzn-bedrock-input-token-count', 'N/A')}")
print(f"Output tokens: {usage.get('x-amzn-bedrock-output-token-count', 'N/A')}")
- Creates a Bedrock runtime client using your default AWS credentials and region.
- Token counts are returned in HTTP response headers, not the body — easy to miss.
What I built to solve this¶
I got tired of guessing, so I built two tools:
Bedrock Cost Tracker¶
A decorator-based library that wraps every Bedrock call and logs token usage plus USD cost to a local SQLite database. Zero config — just add the decorator:
from bedrock_cost_tracker import track_cost
@track_cost
def summarize(text: str) -> str:
response = bedrock.invoke_model(...)
return response["body"].read()
Then run bedrock-costs report to see a breakdown by model, by day, by function.
Bedrock Model Compare¶
A CLI tool that runs one prompt across multiple models and generates a cheapest-first Markdown report with latency, token counts, and cost per call. Useful for deciding which model to route to.
The takeaway¶
Bedrock costs are predictable if you measure them. The problem is that most teams do not measure until the bill arrives. Instrument early, route intelligently, and cache aggressively.