All systems operational
Live status and speed of the Umans Code gateway and its models, refreshed every 30 seconds. For each model we show its output speed and median time to first token.
Models in production
tok/s = output tokens per second · TTFT = time to first token · p50 = median over the last 5 minutes. On each gauge the midpoint is that model's target; the marker sits further right when it's beating target (faster TTFT, higher throughput).
Umans DeepSeek V4 Flash Recommended Operational107.6tok/sthroughput · p50 · last 5 min876msTTFT · p50 · last 5 min100.00%uptime · 24h
DeepSeek V4 Flash: DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release on a 1M-token context. The cheapest production model in the lineup for real agentic work. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Served on our own GPU infrastructure with high availability.
90-day speed trends & events →Operational61.4tok/sthroughput · p50 · last 5 min2.10sTTFT · p50 · last 5 min100.00%uptime · 24h
Kimi K3: Moonshot's most capable model and the first open 3T-class release - 2.8T parameters, a 1M-token context window, and native vision, built for repository-scale understanding and long agentic runs at closed-frontier quality (92.4% vs Claude Fable 5's 92.6% across ~1,030 agentic tasks in Fireworks' independent study). It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed.
90-day speed trends & events →Operational67.9tok/sthroughput · p50 · last 5 min2.10sTTFT · p50 · last 5 min99.90%uptime · 24h
DeepSeek V4 Pro, served from the official 0813 release: DeepSeek's flagship coding and reasoning MoE, built for long-horizon agentic coding and demanding tool-heavy workloads on a 1M-token context window. The 0813 release supersedes the April preview with substantially stronger agentic performance. Reasoning has three modes: non-think (none), think high (high, the default) and think max (max). Billed per token ($1.32 / $3.96 / $0.044 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability.
90-day speed trends & events →Umans Flash Fastest also served as umans-qwen3.6-35b-a3bOperational398.5tok/sthroughput · p50 · last 5 min330msTTFT · p50 · last 5 min99.34%uptime · 24h
Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.
90-day speed trends & events →Umans GLM 5.2 Deprecateddeprecated · serving until Aug 23, 2026 · successor umans-deepseek-v4-pro-0813Operational104.7tok/sthroughput · p50 · last 5 min3.94sTTFT · p50 · last 5 min100.00%uptime · 24h
GLM 5.2 is our best model for coding right now, with a 400K context window for large codebases. Vision is available on the Anthropic Messages API (`/v1/messages`) only, through a server-side handoff (GLM 5.2 generates the text, Kimi preprocesses the image); that handoff will be retired soon in favour of more efficient client-side image handling. Deprecated: use `umans-deepseek-v4-pro-0813` instead (sunset 2026-08-23).
90-day speed trends & events →Umans Kimi K2.7 Code Deprecateddeprecated · serving until Aug 10, 2026 · successor umans-kimi-k3successor to Kimi K2.6also served as umans-coderOperational167.2tok/sthroughput · p50 · last 5 min244msTTFT · p50 · last 5 min100.00%uptime · 24h
Kimi K2.7-Code via Umans Code - Moonshot's strongest coding model and the successor to Kimi K2.6. Built for complex, tool-heavy agentic coding; it reasons more efficiently than K2.6, so agent sessions run faster at the same depth. Deprecated: use `umans-kimi-k3` instead (sunset 2026-08-10).
90-day speed trends & events →