API migration intelligence

AI model deprecations, without the deadline panic

See what is retiring, how long you have, and which model each provider recommends next. Search 27 curated API and open-weight records with release history, context limits, standard text pricing, and links back to official notices.

5 providersOfficial first-party sourcesReviewed July 23, 2026Calendar export

Model lifecycle overview

Next 90 days

4

Documented retirement or earliest shutdown milestones.

Deprecated

10

Models that need an active migration plan.

Active

16

Current first-party API and open-weight records.

Open weights

6

Checkpoints whose hosting limits and cost can vary.

Migration queue

Upcoming lifecycle deadlines

Explore

Find a model or migration

Provider

Showing 27 of 27 curated records

AnthropicDeprecated

Claude Opus 4.1

claude-opus-4-1-20250805

Aliases: claude-opus-4-1

Retirement

Aug 5, 2026

13d 0h remaining

Announced Jun 5, 2026

Recommended replacement

Claude Opus 4.8claude-opus-4-8

Audit every dated ID and alias, then test prompting, tool calls, effort settings, output length, and tokenization on Opus 4.8.

Context
200K
Max output
32K
Input / output
$15 / $75

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Caveat: This lifecycle date covers Anthropic-operated platforms. Amazon Bedrock and Google Cloud publish their own schedules.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 2.5 Pro

gemini-2.5-pro

Earliest shutdown

Oct 16, 2026

85d 0h remaining

Recommended replacement

Gemini 3.1 Pro Previewgemini-3.1-pro-preview

Google currently recommends a preview replacement. Test thought signatures, tools, structured output, and long prompts before migrating.

Context
1.05M
Max output
65.5K
Input / output
$1.25 / $10

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: Google labels October 16 as the earliest possible shutdown, not a guaranteed exact shutdown day.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 2.5 Flash

gemini-2.5-flash

Earliest shutdown

Oct 16, 2026

85d 0h remaining

Recommended replacement

Gemini 3.6 Flashgemini-3.6-flash

Compare thinking behavior, latency, tool calls, and multimodal prompts on the stable 3.6 endpoint before switching.

Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: Google labels October 16 as the earliest possible shutdown, not a guaranteed exact shutdown day.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite

Earliest shutdown

Oct 16, 2026

85d 0h remaining

Recommended replacement

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite

Follow Google's documented migration target, then separately evaluate the newer Gemini 3.5 Flash-Lite for current workloads.

This replacement is documented by the provider but is not a separate record in this curated catalog.

Context
1.05M
Max output
65.5K
Input / output
$0.10 / $0.40

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: Google labels October 16 as the earliest possible shutdown, not a guaranteed exact shutdown day.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4

gpt-4

Aliases: gpt-4-0613, gpt-4-0613-completions, gpt-4-completions

Retirement

Oct 23, 2026

3 months remaining

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-evaluate prompts and tool schemas on GPT-5.6 Sol, then replace every GPT-4 snapshot and Completions alias before shutdown.

Context
8.2K
Max output
8.2K
Input / output
$30 / $60

Capabilities

function callingstreaming

Input: text

Output: text

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4 Turbo

gpt-4-turbo

Aliases: gpt-4-turbo-2024-04-09, gpt-4-turbo-completions

Retirement

Oct 23, 2026

3 months remaining

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Benchmark long prompts and multimodal inputs on GPT-5.6 Sol, then move both the rolling alias and dated snapshot.

Context
128K
Max output
4.1K
Input / output
$10 / $30

Capabilities

function callingstructured outputsstreaming

Input: text, image

Output: text

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-3.5 Turbo

gpt-3.5-turbo

Aliases: gpt-3.5-turbo-0125, gpt-3.5-turbo-completions

Retirement

Oct 23, 2026

3 months remaining

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Run representative low-latency and high-volume evaluations on Terra and update legacy Chat Completions assumptions.

Context
16.4K
Max output
4.1K
Input / output
$0.50 / $1.50

Capabilities

function callingstreaming

Input: text

Output: text

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o1

o1

Aliases: o1-2024-12-17

Retirement

Oct 23, 2026

3 months remaining

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-test reasoning effort, token budgets, and tool behavior on Sol before changing the production model ID.

Context
200K
Max output
100K
Input / output
$15 / $60

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o3-mini

o3-mini

Aliases: o3-mini-2025-01-31

Retirement

Oct 23, 2026

3 months remaining

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Validate reasoning-intensive workloads on Sol and account for its different price and latency profile.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text

Output: text

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o4-mini

o4-mini

Aliases: o4-mini-2025-04-16

Retirement

Oct 23, 2026

3 months remaining

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Test Terra with the same tools and reasoning workloads; check quality, latency, and token use before cutover.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GooglePreview

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Aliases: gemini-pro-latest

Context
1.05M
Max output
65.5K
Input / output
$2 / $12

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Fable 5

claude-fable-5

Availability commitment

No retirement sooner than Jun 9, 2027

Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: Anthropic temporarily suspended Fable 5 in June 2026 and restored global access on July 1. Availability and data-retention eligibility can vary.

Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Haiku 4.5

claude-haiku-4-5-20251001

Aliases: claude-haiku-4-5

Availability commitment

No retirement sooner than Oct 15, 2026

Context
200K
Max output
64K
Input / output
$1 / $5

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: The dated API ID is pinned. The shorter claude-haiku-4-5 value is a convenience alias.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.8

claude-opus-4-8

Availability commitment

No retirement sooner than May 28, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 5

claude-sonnet-5

Availability commitment

No retirement sooner than Jun 30, 2027

Context
1M
Max output
128K
Input / output
$2 / $10

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Pricing note: Introductory pricing through August 31, 2026; Anthropic lists $3 input / $15 output afterward.

Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.5 Flash

gemini-3.5-flash

Aliases: gemini-flash-latest

Context
1.05M
Max output
65.5K
Input / output
$1.50 / $9

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate.

Caveat: gemini-flash-latest is a rolling alias and can change target. Use the stable model ID for production pinning.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite
Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.6 Flash

gemini-3.6-flash
Context
1.05M
Max output
65.5K
Input / output
$1.50 / $7.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate.

Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActiveAPI + open weights

Gemma 4 31B

gemma-4-31b-it
Context
256K
Max output
Not published
Input / output
Varies

Capabilities

reasoningfunction callingcodingmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 30.7B dense

License: Apache 2.0

Caveat: Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.

Official sources (2)
OpenAIActive

GPT-5.6 Sol

gpt-5.6-sol

Aliases: gpt-5.6

Context
1.05M
Max output
128K
Input / output
$5 / $30

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Standard short-context processing. Requests above 272K input tokens receive long-context multipliers.

Caveat: The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.

Official sources (3)
OpenAIActiveOpen weights

gpt-oss-120b

gpt-oss-120b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 117B total / 5.1B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Official sources (3)
OpenAIActiveOpen weights

gpt-oss-20b

gpt-oss-20b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 20.9B total / 3.6B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Official sources (3)
MetaActiveOpen weights

Llama 4 Maverick

llama-4-maverick
Context
1M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 400B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.

Official sources (2)
MetaActiveOpen weights

Llama 4 Scout

llama-4-scout
Context
10M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 109B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.

Official sources (2)
MistralActiveAPI + open weights

Mistral Small 4

mistral-small-2603+1

Aliases: mistral-small-4, mistral-small-latest

Context
256K
Max output
Not published
Input / output
$0.15 / $0.60

Capabilities

reasoningfunction callingstructured outputscodingself-hosting

Input: text, image

Output: text

Parameters: 119B total / 6.5B active

License: Apache 2.0

Pricing note: First-party Mistral API price; self-hosting cost is deployment-specific.

Caveat: mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

Official sources (2)

Methodology

How to read this tracker

Deprecated means a provider has published a replacement and shutdown milestone. Previewmeans the endpoint is usable but may change on a shorter lifecycle. An availability commitment such as Anthropic's “not sooner than” date is shown separately and is not treated as a retirement announcement.

Prices are standard paid text rates in US dollars per one million tokens. Cached reads and Batch API prices appear in the underlying registry when published, while the cards emphasize the comparable standard input/output pair. Long-context tiers and tool charges are called out but not blended into that number.

Model facts are versioned in the site source so price changes, alias changes, and lifecycle updates can be reviewed rather than silently rewriting history. Every record carries its own verification date and first-party links.

Coverage boundaries

What this does not imply

  • This is a curated set of major general-purpose models, not every image, audio, embedding, or experimental endpoint.
  • API lifecycle dates do not automatically apply to ChatGPT, claude.ai, Gemini apps, or third-party cloud catalogs.
  • “Open-weight” describes downloadable weights; it does not guarantee an open-source license or identical hosted limits.
  • A recommended replacement still needs workload-specific evaluation for quality, safety, latency, tools, and cost.

Source index

Provider pricing and lifecycle notices

These are the core documents used across multiple records. Each model card also exposes its specific model card and release sources.

FAQ

Model lifecycle questions

Are all model retirement dates exact?+

No. OpenAI and Anthropic publish confirmed retirement dates for deprecated API models. Google describes dates in its Gemini deprecation table as the earliest possible shutdown dates and says it will communicate the exact date later. The tracker labels that distinction on every affected record.

Does an API model retirement affect ChatGPT, Claude, or Gemini apps?+

Not necessarily. This tracker follows developer model IDs and first-party API lifecycle notices. Consumer apps can switch their underlying models on a separate schedule, and partner platforms such as Amazon Bedrock or Google Cloud can publish different retirement dates.

What happens when an open-weight model is retired?+

Downloaded weights do not disappear. A hosting provider can retire its endpoint, alias, or managed deployment, but self-hosted checkpoints remain available subject to their license. This is why open-weight records show model facts separately from provider-specific hosting cost.

Can I compare the prices as a complete cost estimate?+

The displayed figures are sourced standard text input and output prices per one million tokens. They are a baseline, not a quote: long-context uplifts, tools, regional processing, priority tiers, cache writes, images, audio, and partner margins can change the total.

Planning an API migration?

Export the lifecycle dates above, inventory every hard-coded model ID and alias, and test the replacement against representative production inputs well before the deadline. For related background, browse large language models, OpenAI, Anthropic, and Google DeepMind.