Fireworks AI logo

Fireworks AI

Model Aggregator

Fireworks AI is a high-performance serverless inference platform for open and open-weight frontier models (DeepSeek, GLM, Qwen, Kimi, MiniMax, GPT-OSS), offering Standard and Priority throughput tiers with cached-input discounts.

Models
29
Prompt Caching
Batch Discount
Tool Calling
Model29
DeepSeek V4 FlashDeepSeek$0.14$0.28
DeepSeek V4 Flash (0731)DeepSeek$0.22$0.66
DeepSeek V4 Flash Vision ExpDeepSeek$0.22$0.66
DeepSeek V4 ProDeepSeek$1.74$3.48
DeepSeek V4 Pro (0813)DeepSeek$1.32$3.96
GLM 5.1GLM$1.4$4.4
GLM 5.1 FastGLM$2.8$8.8
GLM 5.2GLM$1.4$4.4
GLM 5.2 FastGLM$2.1$6.6
GLM 5.2 Fast USGLM$2.1$6.6
GLM 5.3GLM$1.4$4.4
GLM 5.3 FastGLM$2.1$6.6
GLM 5.3 FlashGLM$0.15$0.5
Kimi K2.6Kimi$0.95$4
Kimi K2.6 FastKimi$2$8
Kimi K2.7 CodeKimi$0.95$4
Kimi K2.7 Code FastKimi$1.9$8
Kimi K3Kimi$3$15
Kimi K3 FastKimi$4.5$22.5
Kimi K3 USKimi$3.3$16.5
MiniMax M2.7MiniMax$0.3$1.2
MiniMax M3MiniMax$0.3$1.2
Muse Glimmer 30B$0.35$1.5
NVIDIA Nemotron 3 Ultra (Preview)Nemotron$0.6$2.4
NVIDIA Nemotron 3.5 Lightning 30B A3BNemotron$0.05$0.2
OpenAI GPT OSS 120BGPT-OSS$0.15$0.6
OpenAI GPT OSS 20BGPT-OSS$0.07$0.3
Qwen 3.7 PlusQwen$0.4$1.6
Qwen 3.8 MaxQwen$2$6