Pricing that follows the work

Your bill is computed from three quantities: the tokens you train, the tokens you serve, and the storage you keep.

Training tokens

Fine-tuning jobs and training-loop steps are metered per million tokens processed, whichever way you train.

Inference tokens

Calls to your private endpoints are metered per million input and output tokens, through the OpenAI- and Anthropic-compatible APIs.

Storage

Datasets, checkpoints and exported weights are metered per GB-hour, so you only pay while you keep them.

Base model prices

Base models are priced per training token and per served token. Input and output tokens cost the same, so there is no prefill and sampling matrix to decode.

Language
ModelPrices
openai/gpt-oss-20bMoE · 21B params · 3.6B active$0.36 train$0.25 infer128K
nvidia/Nemotron-3-Nano-30B-A3BMoE · 30B params · 3B active$0.40 train$0.27 infer64K
Qwen/Qwen3-8BDense · 8B params$0.40 train$0.30 infer32K
Qwen/Qwen3.6-35B-A3BMoE · 35B params · 3B active$0.48 train$0.34 infer256K
openai/gpt-oss-120bMoE · 117B params · 5.1B active$0.68 train$0.45 infer128K
nvidia/Nemotron-3-Super-120B-A12BMoE · 120B params · 12B active$1.15 train$0.72 infer64K
Qwen/Qwen3-32BDense · 32B params$1.35 train$0.90 infer32K
deepseek-ai/DeepSeek-V4-FlashMoE · 284B params · 13B active$1.70 train$1.10 infer1M
meta-llama/Llama-3.3-70BDense · 70B params$2.75 train$1.90 infer128K
deepseek-ai/DeepSeek-V3.1MoE · 671B params · 37B active$3.40 train$2.30 infer128K
zai-org/GLM-5.2MoE · 753B params · 40B active$3.95 train$2.55 infer1M
moonshotai/Kimi-K2.6MoE · 1T params · 32B active$4.40 train$2.75 infer128K
nvidia/Nemotron-3-Ultra-550B-A55BMoE · 550B params · 55B active$5.00 train$3.10 infer64K
deepseek-ai/DeepSeek-V4-ProMoE · 1.6T params · 49B active$5.60 train$3.50 infer1M
moonshotai/Kimi-K3MoE · 2.8T params · 104B activeNot trainable$6.50 infer1M
Vision
ModelPrices
google/gemma-4-E4BDense · 8B params · 4.5B effective$0.35 train$0.24 infer128K
Qwen/Qwen3-VL-30B-A3B-InstructMoE · 30B params · 3B active$0.45 train$0.32 infer128K
Qwen/Qwen3-VL-8B-InstructDense · 8B params$0.48 train$0.35 infer128K
Qwen/Qwen3.5-9BDense · 9B params$1.30 train$0.90 infer64K
google/gemma-4-31BDense · 31B params$1.45 train$0.95 infer256K
Qwen/Qwen3-VL-235B-A22B-InstructMoE · 235B params · 22B active$2.40 train$1.60 infer128K
Qwen/Qwen3.5-397B-A17BMoE · 397B params · 17B active$6.00 train$3.75 infer64K
Image generation
ModelPrices
black-forest-labs/FLUX.2-kleinFlow transformer · 4B params$0.50 train$2.50 inferUp to 2MP
stabilityai/stable-diffusion-3.5-largeMMDiT · 8B params$0.90 train$4.50 inferUp to 2MP
Qwen/Qwen-ImageMMDiT · 20B params$1.30 train$7.50 inferUp to 4MP
black-forest-labs/FLUX.2-devFlow transformer · 32B params$1.40 train$8.00 inferUp to 4MP
Video generation
ModelPrices
tencent/HunyuanVideo-1.5DiT · 8.3B params$1.40 train$5.00 infer720p · 10s
Lightricks/LTX-2DiT · native audio + video$1.60 train$5.50 inferUp to 4K · 10s
Wan-AI/Wan2.2-T2V-A14BMoE DiT · 27B params · 14B active$1.80 train$6.50 infer720p · 5s
Audio
ModelPrices
nvidia/parakeet-tdt-1.1bASR · 1.1B params$0.30 train$0.20 inferStreaming
openai/whisper-large-v3ASR · 1.5B params$0.60 train$0.40 infer30s windows
mistralai/Voxtral-4B-TTS-2603TTS · 4B params$0.90 train$4.50 infer9 languages

Prices are in USD per million tokens and apply to LoRA fine-tuning and serving through the platform. Image, video and audio models meter the tokens of the encoded media: a megapixel image is roughly 4K tokens, a five-second 720p clip roughly 70K, and a minute of audio roughly 3K. Storage is $0.10 per GB per month, metered by the GB-hour. Mixture-of-experts models are priced by their active parameters. A model marked not trainable is served from the base weights only; you cannot fine-tune it yet.

Pay as you go

Usage-based from the first token, with the whole API included.

  • Managed fine-tuning jobs
  • The low-level training loop
  • Evaluations
  • OpenAI + Anthropic compatible serving
  • Weight export, any checkpoint, any time
Talk to us

Enterprise

For organisations running fine-tuning at scale or under specific compliance requirements.

  • Volume pricing
  • Tailored onboarding and migration support
  • SLA-backed dedicated support
  • PDPA-aligned data handling
  • Custom fine-tuned models built with our team
Talk to sales

Every token accounted for

The usage API reports training tokens, inference tokens and storage for any date range, grouped by day, project, or base model, so you can forecast spend before the bill arrives and reconcile it afterwards.

Explore the usage API

Frequently asked questions

Can't find the answer to your question? Contact our team and we will get back to you.

Your tuned and private LLM. At a fraction of the cost.