Effective public catalog
Model rankings
Deterministic views of RouteShift's effective public model catalog. These tables rank catalog facts, not adoption or traffic estimates.
Method. Prices sort ascending; context and sourced capability indices sort descending. Ties use provider, then model ID.
Source. RouteShift curated registry plus the generated LiteLLM supplement. Generated .
Cheapest input
Positive catalog input prices, lowest cost per million tokens first. Zero or unavailable prices are omitted.
| Rank | Model | Provider | Input price | Routing |
|---|---|---|---|---|
| 1 | text-embedding-3-small | OpenAI | $0.02 / 1M | Explicit only |
| 2 | text-embedding-004 | $0.025 / 1M | Explicit only | |
| 3 | gemma-7b-it | Groq | $0.05 / 1M | Explicit only |
| 4 | llama-3.1-8b-instant | Groq | $0.05 / 1M | Explicit only |
| 5 | gpt-5-nano | OpenAI | $0.05 / 1M | Explicit only |
| 6 | gpt-5-nano-2025-08-07 | OpenAI | $0.05 / 1M | Explicit only |
| 7 | qwen-turbo | Qwen | $0.05 / 1M | Explicit only |
| 8 | qwen-turbo-2024-11-01 | Qwen | $0.05 / 1M | Explicit only |
| 9 | qwen-turbo-2025-04-28 | Qwen | $0.05 / 1M | Explicit only |
| 10 | qwen-turbo-latest | Qwen | $0.05 / 1M | Explicit only |
Cheapest output
Positive catalog output prices, lowest cost per million tokens first. Embeddings, zero prices, and unavailable prices are omitted.
| Rank | Model | Provider | Output price | Routing |
|---|---|---|---|---|
| 1 | gemma-7b-it | Groq | $0.08 / 1M | Explicit only |
| 2 | llama-3.1-8b-instant | Groq | $0.08 / 1M | Explicit only |
| 3 | glm-4-32b-0414-128k | Z.ai | $0.10 / 1M | Explicit only |
| 4 | qwen-turbo | Qwen | $0.20 / 1M | Explicit only |
| 5 | qwen-turbo-2024-11-01 | Qwen | $0.20 / 1M | Explicit only |
| 6 | qwen-turbo-2025-04-28 | Qwen | $0.20 / 1M | Explicit only |
| 7 | qwen-turbo-latest | Qwen | $0.20 / 1M | Explicit only |
| 8 | glm-4.5-air | Z.ai | $0.20 / 1M | Auto or explicit |
| 9 | gpt-oss-20b | OpenAI | $0.25 / 1M | Auto or explicit |
| 10 | together-ai-8.1b-21b | Together | $0.30 / 1M | Explicit only |
Largest context
Published catalog context windows, highest token count first.
| Rank | Model | Provider | Context window | Routing |
|---|---|---|---|---|
| 1 | gpt-5.4-2026-03-05 | OpenAI | 1,050,000 tokens | Explicit only |
| 2 | gpt-5.5-2026-04-23 | OpenAI | 1,050,000 tokens | Explicit only |
| 3 | gpt-5.6 | OpenAI | 1,050,000 tokens | Explicit only |
| 4 | gpt-5.6-luna | OpenAI | 1,050,000 tokens | Explicit only |
| 5 | gpt-5.6-terra | OpenAI | 1,050,000 tokens | Explicit only |
| 6 | @cf/zai-org/glm-5.3-flash | Cloudflare Workers AI | 1,048,576 tokens | Auto or explicit |
| 7 | gemini-2.0-flash-001 | 1,048,576 tokens | Explicit only | |
| 8 | gemini-2.5-flash-lite | 1,048,576 tokens | Explicit only | |
| 9 | gemini-3-pro-preview | 1,048,576 tokens | Explicit only | |
| 10 | gemini-3.5-flash-lite | 1,048,576 tokens | Explicit only |