AI21 Models
Usage
6.2K
10
2 of 29
- siliconflow-deepseek-v3.2
- glm-image
Which models that traffic went to
- Siliconflow Deepseek V3.2100.0%6.2K
- Siliconflow Deepseek V3.290.0%9
- GLM Image10.0%1
All 29 AI21 Models
Open in model list| Modalities | ||||||||
|---|---|---|---|---|---|---|---|---|
| gemini-3.1-flash-lite-preview | Takes text, vision, audio, video, PDF. Output modality not published. | 1.05M | 66K | $2$2/M | — | — | — | — |
| glm-5.2 | Takes text. Output modality not published. | 1M | 131K | $1.1268$3.9438/M | $0.2817/M | — | 107 tok/s | 0.75 s |
| doubao-seed-2-0-lite-260215 | Takes text, vision, video. Output modality not published. | 256K | 128K | $2$2/M | $0.4/M | — | — | — |
| minimax-m2.5 | Takes text. Output modality not published. | 205K | — | $0.288$1.152/M | — | — | — | — |
| deepseek-v3.2-think | Takes text. Output modality not published. | 164K | — | $0.274$0.411/M | — | — | 15 tok/s | 0.36 s |
| baidu-deepseek-v3.2-think | — | — | $0.274$0.411/M | — | — | — | — | |
| siliconflow-deepseek-v3.2 | — | — | $0.274$0.411/M | — | — | 4 tok/s | 3.03 s | |
| sophnet-deepseek-v3.2 | — | — | $0.274$0.411/M | — | — | 8 tok/s | 1.73 s | |
| mm-minimax-m2.5 | — | — | $0.288$1.152/M | — | — | — | — | |
| sophnet-minimax-m2.5 | — | — | $0.338$1.352/M | — | — | — | — | |
| baidu-kimi-k2.5 | — | — | $0.548$2.877/M | — | — | — | — | |
| s-kimi-k2-thinking | — | — | $0.548$2.192/M | — | — | — | — | |
| sf-kimi-k2-thinking | — | — | $0.548$2.192/M | — | — | — | — | |
| sophnet-kimi-k2.5 | — | — | $0.548$2.877/M | — | — | — | — | |
| azure-deepseek-v3.2 | — | — | $0.58$1.68/M | — | — | — | — | |
| azure-deepseek-v3.2-speciale | — | — | $0.58$1.68/M | — | — | — | — | |
| azure-kimi-k2-thinking | — | — | $0.6$2.5002/M | — | — | — | — | |
| azure-kimi-k2.5 | — | — | $0.6$3/M | — | — | — | — | |
| deepinfra-glm-5 | — | — | $0.88$2.816/M | $0.176/M | — | — | — | |
| alicloud-deepseek-v3.2 | — | — | $2$2/M | $0.4/M | $2.5/M | — | — | |
| alicloud-glm-5 | — | — | $2$2/M | $0.4/M | — | — | — | |
| baidu-glm-5 | — | — | $2$2/M | — | — | — | — | |
| glm-image | Takes text, returns vision. | — | — | $2$2/M | — | — | — | — |
| sophnet-glm-5 | — | — | $2$2/M | — | — | — | — | |
| xiaomi-mimo-v2-omni | — | — | $2$2/M | — | — | — | — | |
| xiaomi-mimo-v2-pro | — | — | $2$2/M | — | — | — | — | |
| aihubmix-Jamba-1-5-Large | — | — | $2.2$8.8/M | — | — | — | — | |
| o1-global | — | — | $15$60/M | $7.5/M | — | — | — | |
| jimeng-3.0-pro | — | — | $20$20/M | — | — | — | — |
AI21 on AIHubMix
Which AI21 model should I start with?
baidu-deepseek-v3.2-think at $0.274/M input — the cheapest entry here that declares a token price. Move up to jimeng-3.0-pro when answer quality matters more than cost.
Why are there several entries for the same model?
Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), and some differ only in capitalisation, kept so older integrations keep working.
The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.
How is cached input billed?
The Cache read column is the rate for input tokens served from the prompt cache — for example deepinfra-glm-5 bills cache hits at 20% of the input rate and glm-5.2 bills cache hits at 25% of the input rate. Cache write is the surcharge for putting a prompt into the cache in the first place, and only a few upstreams bill it separately. A dash in either column means the catalog carries no cache rate for that model, so plan on paying the full input rate.
Do I need a separate AI21 account?
No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.
Start calling AI21 in one line
One key, one endpoint, 750 models across 29 model authors.