umans/status/umans-qwen3.8-flash-next-lab
Live · refreshes every 30s
← all models
Umans Qwen3.8 Flash Next (lab) Experimental
umans-qwen3.8-flash-next-lab · Qwen3.8-Flash-Next · Qwen
In testing
168.3tok/s
throughput · p50 · last 5 min
4.82s
TTFT · p50 · last 5 min
96.88%
uptime · 24h

Qwen3.8 Flash-Next as a Labs experiment, open for a short test window: temporary, not a permanent id. Served from Qwen's official FP8 checkpoint: a 125B mixture-of-experts model with 6B active parameters per token, the first open release of the next-generation architecture that powers the upcoming Qwen4 family, with native image and video understanding, built for fast coding and agentic tasks on a 256K context window. It thinks by default at xhigh effort; reasoning can be tuned down (low, medium) or turned off (none). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and low availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. For production work we recommend umans-coder or umans-deepseek-v4-pro-0813.

90 days agotoday
Context
262K
Max output
131K
Recommended
131K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/medium/xhigh
Weights
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
now 163.4 tok/s
90 days agotoday
TTFT p50 · time to first token, lower is better
now 2.71s
90 days agotoday
Changelog

Events for Umans Qwen3.8 Flash Next (lab)

incl. gateway-wide announcements
Aug 262026
New Labs experiment: Umans Qwen3.8 Flash Next Testing
A new lab opened on umans-qwen3.8-flash-next-lab: Qwen3.8 Flash-Next served from Qwen's official FP8 checkpoint, a 125B mixture-of-experts model (6B active per token) and the first open release of the next-generation architecture that powers the upcoming Qwen4 family, with native image understanding and a 256K context window, thinking at xhigh effort by default (dial to low or medium, or turn thinking off). Free and seat-gated while the experiment runs, served at limited capacity and low availability: it is a lab, so expect it to be flaky and to go down under load. Crash it, give it a moment, and try again. The window closes August 27, 2026.