Skip to main content
Google price change (July 2026): Gemini 2.5 Flash moved from 0.15/0.15/0.60 to 0.30/0.30/2.50 per 1M tokens (2–4× increase). Cheaper alternatives: google/gemini-2.5-flash-lite (0.10/0.10/0.40) or the new google/gemini-3.6-flash.

Legacy (backward compat)

Still operational for compatibility but not recommended for new integrations: Retired in July 2026: openai/gpt-4.1, openai/o4-mini, deepseek/deepseek-chat, deepseek/deepseek-reasoner, moonshot/kimi-k2, moonshot/moonshot-v1-128k, xai/grok-3, xai/grok-3-mini, xai/grok-4.

When to use each

Critical tasks

Claude Fable 5, Claude Opus 5 or GPT-5.6 Sol. Top of class in reasoning.

Production at scale

Claude Sonnet 5. Price/quality balance — pick it blind if you don’t know the rest.

High volume / low cost

Gemini 2.5 Flash Lite, GPT-5.6 Luna or DeepSeek V4 Flash. Sub-dollar per 1M tokens.

Reasoning (chain of thought)

DeepSeek V4 Pro, o3-mini or Grok 4.20 Reasoning. Designed for structured reasoning.

Very long context (1M tokens)

Gemini 3.6 Flash / 2.5 Pro / 2.5 Flash, Grok 4.3 or Claude Fable/Opus/Sonnet 5. Process whole documents.

Code

Kimi K2.7 Code, Claude Sonnet 5 or Grok Build. Strong code performance.

Natural failover

Because all models share the same endpoint and SDK, failover across providers is trivial: