Qwen3.8 Flash-Next
A sparse model with extra embedding memory for long-context agent workloads.
- Published size
- 125B+
- Active parameters
- 6B
- Context window
- 256K native
- Architecture
- MoE · vision
- License
- Qwen Community 1.0
Deployment considerations
Include the large n-gram embedding tables and MTP head when sizing memory. The 6B activated figure is not the total weight footprint.
Official scope: 125B backbone with 6B activated, plus 51B n-gram embeddings and 4B MTP. Native context is 262,144 tokens; extension to 1M is separately configured.
GPU memory estimates and minimum / recommended configurations are pending. This profile does not contain measured deployment results.
GPU requirements
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Memory estimates and tested configurations will appear here after checkpoint review and deployment testing.
No invented leaderboards.
Latency, throughput, and cost per token will appear here after a reproducible run. Until then, this page helps you understand the model—not predict its performance.
Qwen3.8 Flash-Next deployment FAQ
How much GPU memory does Qwen3.8 Flash-Next need?
The full checkpoint weight footprint is pending review. Minimum and recommended GPU configurations will be added after testing; the model name or active parameter count alone is not a memory requirement.
Has BenchGrid benchmarked Qwen3.8 Flash-Next?
Not yet. This profile contains publisher specifications and calculated weight-memory estimates. We do not currently publish measured latency, throughput, or cost per token for this model.
Where do these specifications come from?
The specifications are based on the official Alibaba model card linked on this page. Memory estimates use the stated total parameter count, including inactive experts for MoE models. Nominal model sizes are labeled with ~.
Explore your compute options.
Check available hardware, quotas, and current pricing with the provider. These links are not verified deployments or performance recommendations.