Split input correctly
Cached tokens are carved out of total input. In Standard mode, only that subset receives a recorded cache rate; Batch uses its own published input rate.
Turn a token workload into per-request, monthly, and annual estimates. Compare official prices and see which model context windows can actually hold the request.
Private, browser-only estimate
No API key is needed. Your workload inputs stay on this device and are not saved or sent to a provider.
Workload
Cached input is a subset of input, never an additional token count.
Selected estimate
gpt-5.6-terra
Per request
$0.0688
22K total tokens
Per month
$1,512.50
22K requests
Per year
$18,150.00
12 equal months
22K combined input + requested output
Context: 1.1M
Max output: 128K
1M context tokens remain before the published window limit.
Standard short-context processing. Requests above 272K input tokens receive long-context multipliers.
Ranked comparison
Sorted by estimated monthly cost. Select a row to inspect its rates and constraints above.
| Model | Per request | Per month | Context fit | Source |
|---|---|---|---|---|
| $0.002350 | $51.70 | 2.1% | Official ↗ | |
| $0.004200 | $92.40 | 8.6% | Official ↗ | |
| $0.009650 | $212.30 | 2.1% | Official ↗ | |
| $0.009650 | $212.30 | 2.1% | Official ↗ | |
| $0.0130 | $286.00 | Does not fit | Official ↗ | |
| $0.0255 | $561.00 | 11% | Official ↗ | |
| $0.0267 | $586.85 | 11% | Official ↗ | |
| $0.0275 | $605.00 | 2.1% | Official ↗ | |
| $0.0281 | $617.10 | 11% | Official ↗ | |
| $0.0383 | $841.50 | 2.1% | Official ↗ | |
| $0.0394 | $866.25 | 2.1% | Official ↗ | |
| $0.0413 | $907.50 | 2.1% | Official ↗ | |
| $0.0510 | $1,122.00 | 2.2% | Official ↗ | |
| $0.0550 | $1,210.00 | 2.1% | Official ↗ | |
| $0.0688 | $1,512.50 | 2.1% | Official ↗ | |
| $0.1275 | $2,805.00 | 2.2% | Official ↗ | |
| $0.1375 | $3,025.00 | 2.1% | Official ↗ | |
| $0.2550 | $5,610.00 | 2.2% | Official ↗ | |
| $0.2600 | $5,720.00 | 17% | Official ↗ | |
| $0.3825 | $8,415.00 | 11% | Official ↗ | |
| $0.3825 | $8,415.00 | 11% | Official ↗ | |
| $0.7200 | $15,840.00 | Does not fit | Official ↗ |
Transparent methodology
Cached tokens are carved out of total input. In Standard mode, only that subset receives a recorded cache rate; Batch uses its own published input rate.
Input cost and output cost form a per-request estimate. Requests per day multiplied by working days produces the monthly volume; twelve equal months produces the annual figure.
Input plus requested output is compared with the context window. Output is also checked independently when the provider publishes a separate maximum.
Common questions
The calculator multiplies the selected model's published per-million-token input and output rates by your tokens per request, then multiplies the request estimate by requests per working day and working days per month. The yearly figure uses twelve equal months.
No. Cached input is treated as a subset of the total input. The cached subset receives a published cached-input rate when one is recorded; otherwise it receives the ordinary input rate.
No. Batch mode appears only when the model registry contains official Batch input and output rates. The calculator does not assume a percentage discount, combine Batch with a cached-input discount, or infer missing prices.
No. It covers the recorded text-token prices only. Taxes, regional pricing, minimum commitments, long-context tiers, storage, searches, grounding, tools, fine-tuning, media, partner platforms, and other feature charges may change the actual bill.