Pricing that follows the work
Your bill is computed from three quantities: the tokens you train, the tokens you serve, and the storage you keep.
Training tokens
Fine-tuning jobs and training-loop steps are metered per million tokens processed, whichever way you train.
Inference tokens
Calls to your private endpoints are metered per million input and output tokens, through the OpenAI- and Anthropic-compatible APIs.
Storage
Datasets, checkpoints and exported weights are metered per GB-hour, so you only pay while you keep them.
Base model prices
Base models are priced per training token and per served token. Input and output tokens cost the same, so there is no prefill and sampling matrix to decode.
| Model | Prices |
|---|---|
| openai/gpt-oss-20bMoE · 21B params · 3.6B active | $0.36 train$0.25 infer128K |
| nvidia/Nemotron-3-Nano-30B-A3BMoE · 30B params · 3B active | $0.40 train$0.27 infer64K |
| Qwen/Qwen3-8BDense · 8B params | $0.40 train$0.30 infer32K |
| Qwen/Qwen3.6-35B-A3BMoE · 35B params · 3B active | $0.48 train$0.34 infer256K |
| openai/gpt-oss-120bMoE · 117B params · 5.1B active | $0.68 train$0.45 infer128K |
| nvidia/Nemotron-3-Super-120B-A12BMoE · 120B params · 12B active | $1.15 train$0.72 infer64K |
| Qwen/Qwen3-32BDense · 32B params | $1.35 train$0.90 infer32K |
| deepseek-ai/DeepSeek-V4-FlashMoE · 284B params · 13B active | $1.70 train$1.10 infer1M |
| meta-llama/Llama-3.3-70BDense · 70B params | $2.75 train$1.90 infer128K |
| deepseek-ai/DeepSeek-V3.1MoE · 671B params · 37B active | $3.40 train$2.30 infer128K |
| zai-org/GLM-5.2MoE · 753B params · 40B active | $3.95 train$2.55 infer1M |
| moonshotai/Kimi-K2.6MoE · 1T params · 32B active | $4.40 train$2.75 infer128K |
| nvidia/Nemotron-3-Ultra-550B-A55BMoE · 550B params · 55B active | $5.00 train$3.10 infer64K |
| deepseek-ai/DeepSeek-V4-ProMoE · 1.6T params · 49B active | $5.60 train$3.50 infer1M |
| moonshotai/Kimi-K3MoE · 2.8T params · 104B active | Not trainable$6.50 infer1M |
| Model | Prices |
|---|---|
| google/gemma-4-E4BDense · 8B params · 4.5B effective | $0.35 train$0.24 infer128K |
| Qwen/Qwen3-VL-30B-A3B-InstructMoE · 30B params · 3B active | $0.45 train$0.32 infer128K |
| Qwen/Qwen3-VL-8B-InstructDense · 8B params | $0.48 train$0.35 infer128K |
| Qwen/Qwen3.5-9BDense · 9B params | $1.30 train$0.90 infer64K |
| google/gemma-4-31BDense · 31B params | $1.45 train$0.95 infer256K |
| Qwen/Qwen3-VL-235B-A22B-InstructMoE · 235B params · 22B active | $2.40 train$1.60 infer128K |
| Qwen/Qwen3.5-397B-A17BMoE · 397B params · 17B active | $6.00 train$3.75 infer64K |
| Model | Prices |
|---|---|
| black-forest-labs/FLUX.2-kleinFlow transformer · 4B params | $0.50 train$2.50 inferUp to 2MP |
| stabilityai/stable-diffusion-3.5-largeMMDiT · 8B params | $0.90 train$4.50 inferUp to 2MP |
| Qwen/Qwen-ImageMMDiT · 20B params | $1.30 train$7.50 inferUp to 4MP |
| black-forest-labs/FLUX.2-devFlow transformer · 32B params | $1.40 train$8.00 inferUp to 4MP |
| Model | Prices |
|---|---|
| tencent/HunyuanVideo-1.5DiT · 8.3B params | $1.40 train$5.00 infer720p · 10s |
| Lightricks/LTX-2DiT · native audio + video | $1.60 train$5.50 inferUp to 4K · 10s |
| Wan-AI/Wan2.2-T2V-A14BMoE DiT · 27B params · 14B active | $1.80 train$6.50 infer720p · 5s |
| Model | Prices |
|---|---|
| nvidia/parakeet-tdt-1.1bASR · 1.1B params | $0.30 train$0.20 inferStreaming |
| openai/whisper-large-v3ASR · 1.5B params | $0.60 train$0.40 infer30s windows |
| mistralai/Voxtral-4B-TTS-2603TTS · 4B params | $0.90 train$4.50 infer9 languages |
Prices are in USD per million tokens and apply to LoRA fine-tuning and serving through the platform. Image, video and audio models meter the tokens of the encoded media: a megapixel image is roughly 4K tokens, a five-second 720p clip roughly 70K, and a minute of audio roughly 3K. Storage is $0.10 per GB per month, metered by the GB-hour. Mixture-of-experts models are priced by their active parameters. A model marked not trainable is served from the base weights only; you cannot fine-tune it yet.
Pay as you go
Usage-based from the first token, with the whole API included.
- Managed fine-tuning jobs
- The low-level training loop
- Evaluations
- OpenAI + Anthropic compatible serving
- Weight export, any checkpoint, any time
Enterprise
For organisations running fine-tuning at scale or under specific compliance requirements.
- Volume pricing
- Tailored onboarding and migration support
- SLA-backed dedicated support
- PDPA-aligned data handling
- Custom fine-tuned models built with our team
Every token accounted for
The usage API reports training tokens, inference tokens and storage for any date range, grouped by day, project, or base model, so you can forecast spend before the bill arrives and reconcile it afterwards.
Explore the usage APIFrequently asked questions
Can't find the answer to your question? Contact our team and we will get back to you.