umans/status
Live · updated just now

All systems operational

Live status and speed of the Umans Code gateway and its models, refreshed every 30 seconds. For each model we show its output speed and median time to first token.

6 / 6production models operational
98.89%gateway uptime · 90 days
0models in testing
0active issues
Production

Models in production

6 listed

tok/s = output tokens per second · TTFT = time to first token · p50 = median over the last 5 minutes. On each gauge the midpoint is that model's target; the marker sits further right when it's beating target (faster TTFT, higher throughput).

umans-deepseek-v4-flash-0731 · DeepSeek-V4-Flash · DeepSeek
Operational
107.6tok/s
throughput · p50 · last 5 min
876ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h

DeepSeek V4 Flash: DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release on a 1M-token context. The cheapest production model in the lineup for real agentic work. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Served on our own GPU infrastructure with high availability.

90-day speed trends & events →
90 days agoin production since Aug 3, 2026today
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
umans-kimi-k3 · Kimi K3 · Moonshot
Operational
61.4tok/s
throughput · p50 · last 5 min
2.10s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

Kimi K3: Moonshot's most capable model and the first open 3T-class release - 2.8T parameters, a 1M-token context window, and native vision, built for repository-scale understanding and long agentic runs at closed-frontier quality (92.4% vs Claude Fable 5's 92.6% across ~1,030 agentic tasks in Fireworks' independent study). It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed.

90-day speed trends & events →
90 days agoin production since Jul 31, 2026today
Context
1049K
Max output
131K
Recommended
131K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/high/max
umans-deepseek-v4-pro-0813 · DeepSeek-V4-Pro-0813 · DeepSeek
Operational
67.9tok/s
throughput · p50 · last 5 min
2.10s
TTFT · p50 · last 5 min
99.90%
uptime · 24h

DeepSeek V4 Pro, served from the official 0813 release: DeepSeek's flagship coding and reasoning MoE, built for long-horizon agentic coding and demanding tool-heavy workloads on a 1M-token context window. The 0813 release supersedes the April preview with substantially stronger agentic performance. Reasoning has three modes: non-think (none), think high (high, the default) and think max (max). Billed per token ($1.32 / $3.96 / $0.044 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability.

90-day speed trends & events →
90 days agoin production since Aug 15, 2026today
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/high/max
Umans Flash Fastest
umans-flash · Qwen3.6-35B-A3B · Qwen
also served as umans-qwen3.6-35b-a3b
Operational
398.5tok/s
throughput · p50 · last 5 min
330ms
TTFT · p50 · last 5 min
99.34%
uptime · 24h

Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.

90-day speed trends & events →
90 days agoin production since May 3, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/medium/high
Weights
Umans GLM 5.2 Deprecated
umans-glm-5.2 · GLM
deprecated · serving until Aug 23, 2026 · successor umans-deepseek-v4-pro-0813
Operational
104.7tok/s
throughput · p50 · last 5 min
3.94s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

GLM 5.2 is our best model for coding right now, with a 400K context window for large codebases. Vision is available on the Anthropic Messages API (`/v1/messages`) only, through a server-side handoff (GLM 5.2 generates the text, Kimi preprocesses the image); that handoff will be retired soon in favour of more efficient client-side image handling. Deprecated: use `umans-deepseek-v4-pro-0813` instead (sunset 2026-08-23).

90-day speed trends & events →
90 days agoin production since Jun 21, 2026today
Context
406K
Max output
131K
Recommended
131K
Vision
Via handoff
Tools
Yes
Reasoning
Toggle · none/high/max
Weights
umans-kimi-k2.7 · Kimi K2.7-Code · Moonshot
deprecated · serving until Aug 10, 2026 · successor umans-kimi-k3
successor to Kimi K2.6
also served as umans-coder
Operational
167.2tok/s
throughput · p50 · last 5 min
244ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h

Kimi K2.7-Code via Umans Code - Moonshot's strongest coding model and the successor to Kimi K2.6. Built for complex, tool-heavy agentic coding; it reasons more efficiently than K2.6, so agent sessions run faster at the same depth. Deprecated: use `umans-kimi-k3` instead (sunset 2026-08-10).

90-day speed trends & events →
90 days agoin production since Jun 12, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Always on
History

Gateway uptime

90-day window · daily worst status
90 days ago98.89% operationaltoday
Changelog

Recent events & model lifecycle

releases, retirements, and playground changes
Aug 182026
New Labs experiment: Umans Qwen3.8 27B Testing
A new lab opened on umans-qwen3.8-27b-lab: Qwen3.8 27B served from Qwen's official FP8 checkpoint, a compact dense model with native image and video understanding and a 256K context window, thinking at xhigh effort by default (dial to low or medium, or turn thinking off). Free and seat-gated while the experiment runs, served at limited capacity and low availability: it is a lab, so expect it to be flaky and to go down under load. Crash it, give it a moment, and try again. The window closes August 19, 2026.
Aug 152026
Released pay-per-token: Umans DeepSeek V4 Pro Released
umans-deepseek-v4-pro-0813 joins the lineup as the long-context coding flagship: DeepSeek's official 0813 release of V4 Pro, the checkpoint that served here as the seat-gated pre-release lab since August 13, with a 1M context window and thinking at high effort by default (dial to max when a task deserves more). Billed per token: $1.32 / $3.96 / $0.044 per 1M (input / output / cache read). It succeeds GLM 5.2, which is deprecated and sunsets on August 23, 2026. The pre-release window's metrics stay on the model's status page as its pre-release period (before Aug 15). Served on our own GPU infrastructure with high availability.
Aug 152026
Playground closed: Umans DeepSeek V4 Pro 0813 (pre-release) Testing
The V4 Pro 0813 pre-release window closed at the pay-per-token release: the production id now bills per token, so the seat-gated experiment on umans-deepseek-v4-pro-0813-lab ended rather than charge anyone by surprise. Thanks to everyone who pushed it, including through the August 14 outage, and shared findings; both the fixes and the metrics carried straight into the release. The model keeps serving as umans-deepseek-v4-pro-0813: the model stays.
Aug 132026
New Labs experiment: Umans DeepSeek V4 Pro 0813 Testing
A new lab opened on umans-deepseek-v4-pro-0813-lab: DeepSeek V4 Pro served from DeepSeek's own pre-release 0813 checkpoint - the flagship coding and reasoning MoE, ahead of its official release. Free and seat-gated while the experiment runs, served at limited capacity and low availability: it is a lab, so expect it to be flaky and to go down under load - crash it, give it a moment, and try again.
Aug 132026
Playground closed: Umans DeepSeek V4 Flash Vision Testing
The vision V4 Flash test window closed after two days. Thanks to everyone who pushed images through it and shared findings. The umans-deepseek-v4-flash-0731-vision-lab id is unlisted and stops serving; the text V4 Flash keeps serving as the pay-per-token umans-deepseek-v4-flash-0731 regardless: the model stays.
Aug 112026
New Labs experiment: Umans DeepSeek V4 Flash Vision Testing
The text V4 Flash lab window ended and a new vision lab opened on umans-deepseek-v4-flash-0731-vision-lab: our own V4 Flash with vision added, the same performance and speed as the text V4 Flash plus the ability to read images. Free and seat-gated while the experiment runs, served at limited capacity and low availability: it is a lab, so expect it to be flaky and to go down under load - crash it, give it a moment, and try again. (The vision lab closed on 2026-08-13 - see the entry above.)
Aug 32026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup as the cheapest way we serve real agentic work: $0.14 / $0.28 / $0.028 per 1M (input / output / cache read), a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Founding users pay the 10x cheaper cache rate until Monday, August 10, 2026 (see /pricing). Served on our own GPU infrastructure with high availability.
Aug 32026
The V4 Flash lab continues as umans-deepseek-v4-flash-0731-lab Testing
The V4 Flash pilot closed at the pay-per-token release: the production id now bills per token, so the seat-gated pilot on it ended rather than charge anyone by surprise. The lab reopened on the new umans-deepseek-v4-flash-0731-lab id with a smaller cohort - free, seat-gated, same experimental capacity as before. The model keeps serving as umans-deepseek-v4-flash-0731 regardless: the model stays. (The text lab closed on 2026-08-11 when the vision lab opened - see the next entry.)
Aug 12026
Generally available: Umans Kimi K3 Released
The Kimi K3 Labs pilot closed and umans-kimi-k3 is generally available: no Labs seat required anymore - plan and pay-per-token keys alike get the 1M context window, native vision, and max-effort reasoning at $3.00 / $15.00 / $0.30 per 1M (input / output / cache read).
Jul 312026
Released pay-per-token: Umans Kimi K3 Released
umans-kimi-k3 opened to pay-per-token (wallet and service-account keys): Moonshot's largest open model, a 1M context window, native vision, and max reasoning effort by default, billed at $3.00 / $15.00 / $0.30 per 1M (input / output / cache read).
Older events 10
Jul 162026
Playground closed: Umans DeepSeek V4 Pro DSpark Testing
The DSpark test window closed after two days. What we took from it: DSpark speculative decoding lets us serve more users at once while keeping each session fast enough, the DeepSeek V4 architecture is now mature enough to serve at scale, and the model itself is solid. V4 Pro is not joining the lineup though: it is still a preview build, and the issues testers hit (DSML leaks, language bleed, long-context artifacts) are model-side. DeepSeek confirmed a better version is coming this month, so we would rather roll the learnings into that. Thanks to everyone who tested.
Jul 142026
Playground opened: Umans DeepSeek V4 Pro DSpark Testing
umans-deepseek-v4-pro-dspark entered the playground for a short, seat-gated test window: DeepSeek V4 Pro served from the original weights with DSpark speculative decoding. Experimental and temporary; not for production.
Jul 22026
Playground closed: Umans GLM 5.2 NVFP4 Testing
The short NVFP4 test window ended after four days. Thanks to everyone who pushed it and shared findings.
Jun 292026
Playground opened: Umans GLM 5.2 NVFP4 Testing
umans-glm-5.2-nvfp4 entered the playground for a short, low-capacity test window. Experimental and temporary; not for production.
Jun 242026
Retired: Umans GLM 5.1 Retired
umans-glm-5.1 was retired in favour of GLM 5.2. Requests to the old id now return a clear deprecation error pointing to umans-glm-5.2.
Jun 212026
Released to production: Umans GLM 5.2 Released
umans-glm-5.2 was released as the long-context model, with a 405K context window, after a pre-release period that started Jun 16.
Jun 182026
Retired: Umans Kimi K2.6 Code Retired
umans-kimi-k2.6 was retired and superseded by K2.7. Requests to the old id now return a clear deprecation error pointing to umans-kimi-k2.7.
Jun 122026
Released to production: Umans Kimi K2.7 Code Released
umans-kimi-k2.7 was released as the recommended coding model (also served as umans-coder).
May 132026
Retired: Umans Kimi K2.5 Retired
umans-kimi-k2.5 was retired and superseded by Kimi K2.6.
May 92026
Retired: Umans MiniMax M2.5 Retired
umans-minimax-m2.5 was retired without a direct replacement.
Past models 9 retired
umans-deepseek-v4-pro-0813-lab · DeepSeek
retired Aug 15, 2026
umans-deepseek-v4-pro-0813
umans-deepseek-v4-flash-0731-vision-lab · DeepSeek
retired Aug 13, 2026
umans-deepseek-v4-flash-0731-lab · DeepSeek
retired Aug 11, 2026
umans-deepseek-v4-pro-dspark · DeepSeek
retired Jul 16, 2026
umans-glm-5.2-nvfp4 · GLM
retired Jul 2, 2026
umans-glm-5.2
umans-glm-5.1 · GLM
retired Jun 24, 2026
umans-glm-5.2
umans-kimi-k2.6 · Moonshot
retired Jun 18, 2026
umans-kimi-k2.7
umans-kimi-k2.5 · Moonshot
retired May 13, 2026
umans-kimi-k2.6
umans-minimax-m2.5 · MiniMax
retired May 9, 2026