Moonshot AI Models
Usage
0
4
1 of 20
All 20 Moonshot AI Models
Open in model list| Modalities | ||||||
|---|---|---|---|---|---|---|
| kimi-k3 | Takes text, vision, video. Output modality not published. | 1.05M | $3$15/M | $0.3/M | — | — |
| kimi-k2-thinking | Takes text. Output modality not published. | 262K | $0.548$2.192/M | — | 42 tok/s | 1.63 s |
| kimi-k2-turbo-preview | Takes text, returns text. | 262K | $1.2$4.8/M | $0.3/M | 20 tok/s | 3.06 s |
| Kimi-K2-0905 | Takes text, returns text. | 256K | $0.548$2.192/M | — | 49 tok/s | 0.10 s |
| moonshotai/kimi-k2-instruct | Takes text, returns text. | 256K | $0.62$2.48/M | — | — | — |
| kimi-k2-0711 | Takes text, returns text. | 128K | $0.54$2.16/M | — | 50 tok/s | 0.10 s |
| moonshotai/Moonlight-16B-A3B-Instruct | — | $0.2$0.2/M | — | — | — | |
| moonshotai/Kimi-Dev-72B | — | $0.32$1.28/M | — | — | — | |
| alicloud-kimi-k2.5 | — | $0.548$2.877/M | $0.0959/M | — | — | |
| kimi-k2.5 | Takes text, vision, video. Output modality not published. | — | $0.548$2.877/M | $0.0959/M | — | — |
| cc-Kimi-K2-Instruct | Takes text, returns text. | — | $1.1$3.3/M | — | — | — |
| moonshot-v1-8k | — | $2$2/M | — | — | — | |
| moonshot-v1-8k-vision-preview | — | $2$2/M | — | — | — | |
| moonshot-kimi-k3 | — | $3$15/M | $0.3/M | — | — | |
| kimi-latest | — | $4$4/M | — | — | — | |
| moonshot-v1-32k | — | $4$4/M | — | — | — | |
| moonshot-v1-32k-vision-preview | — | $4$4/M | — | — | — | |
| moonshot-v1-128k | — | $10$10/M | — | — | — | |
| moonshot-v1-128k-vision-preview | — | $10$10/M | — | — | — | |
| kimi-thinking-preview | — | $30$30/M | — | — | — |
Moonshot AI on AIHubMix
Which Moonshot AI model should I start with?
kimi-k2-0711 at $0.54/M input — the cheapest entry here that declares tool calling, and it carries a 128K context. Move up to kimi-thinking-preview when answer quality matters more than cost.
Why are there several entries for the same model?
Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), some are the open-weight repository form (moonshotai/…), and some differ only in capitalisation, kept so older integrations keep working.
The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.
How is cached input billed?
The Cache read column is the rate for input tokens served from the prompt cache — for example kimi-k3 bills cache hits at 10% of the input rate and moonshot-kimi-k3 bills cache hits at 10% of the input rate. Cache write is the surcharge for putting a prompt into the cache in the first place, and only a few upstreams bill it separately. A dash in either column means the catalog carries no cache rate for that model, so plan on paying the full input rate.
Do I need a separate Moonshot AI account?
No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.
Start calling Moonshot AI in one line
One key, one endpoint, 750 models across 29 model authors.