Cheap, fast inference for models worth building with.

GLM 5.3

z-ai/glm-5.3
Latency
4.18s
ThroughputP50 · rolling 5 min
30.0 tps
Uptime
100.00%
TTFT
2.36s
Recent uptimeUpdated programmatically.
100.00%

GLM 5.3 Flash

z-ai/glm-5.3-flash
Latency
8.82s
ThroughputP50 · rolling 5 min
14.3 tps
Uptime
100.00%
TTFT
5.48s
Recent uptimeUpdated programmatically.
100.00%

DeepSeek V4 Flash

deepseek/deepseek-v4-flash
Latency
7.77s
ThroughputP50 · rolling 5 min
28.7 tps
Uptime
100.00%
TTFT
2.36s
Recent uptimeUpdated programmatically.
100.00%

Built with experience and support from