Access models from the top open-source labs through a single, unified Inference API. Pay only for the tokens you use.
Showing GLM-5.3, item 1 of 4.
Model 01Open weights
GLM-5.3
prism/glm-5.3
Model 02Open weights
Kimi K3
prism/kimi-k3
Model 03Open weights
DeepSeek-V4-Flash
prism/deepseek-v4-flash
Model 04Open weights
Qwen3.6 35B-A3B
prism/qwen
Model profile
Z.AI’s strongest GLM-5 agentic coder — the same base as GLM-5.2, with every gain coming from post-training.
Input
$1.40
/ M tokens
Cache read
$0.26
/ M tokens
Output
$4.40
/ M tokens
Context
1M
tokens
Max output
128K
tokens
Serving
Serverless
Pay per token
02 / Production workloads
Built for production agent workloads
Prism serves open-weight models through one reliable inference API built for long tool traces, large context windows, and production-scale agent workloads.
Low latency
Speculative decoding and batching shaped around code generation, not general chat.
High throughput
Batched serving across a GPU fleet so a customer agent can stay in a tight loop.
Reliable uptime
Redundant serving paths keep inference available when individual workers or GPUs go offline.