MultimodalAlibaba / Dense · vision
A mid-sized vision-language model for reasoning, coding, and agent workloads.
- Parameters
- 27B
- Context
- 256K native
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
DistilledXiaomi / Dense · vision
A compact Qwen-based MiMo distillation for a smaller deployment footprint.
- Parameters
- ~9B
- Context
- Under review
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
MultimodalGoogle / Dense · multimodal
Unified text, image, and audio understanding in a compact Gemma 4 model.
- Parameters
- 11.95B
- Context
- 256K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
AgenticAlibaba / MoE · vision
A sparse model with extra embedding memory for long-context agent workloads.
- Parameters
- 125B+ / 6B active
- Context
- 256K native
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
AgenticXiaomi / MoE · multimodal
MiMo's multimodal Flash model with sparse experts and a 1M context window.
- Parameters
- 309B+ / 15B active
- Context
- 1M
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
MultimodalZ.ai / MoE · vision
A multimodal GLM model combining sparse experts with hybrid attention.
- Parameters
- 320B / 18B active
- Context
- Under review
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
MultimodalAlibaba / Dense · vision
A smaller Qwen vision-language model for comparing compact deployments.
- Parameters
- ~9B
- Context
- Under review
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
AgenticAlibaba / MoE · vision
A smaller sparse Qwen model for coding agents and vision-language tasks.
- Parameters
- 35B+ / 3B active
- Context
- 256K native
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
MultimodalGoogle / MoE · vision
Gemma's sparse vision-language model, a useful counterpart to the dense 31B.
- Parameters
- ~26B / 3.8B active
- Context
- 256K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
MultimodalGoogle / Dense · vision
A dense Gemma 4 model for text and image understanding at a larger scale.
- Parameters
- ~31B
- Context
- 256K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
On-deviceGoogle / Dense · multimodal
An edge-focused Gemma model with text, image, and audio inputs.
- Parameters
- 8B+
- Context
- 128K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
AgenticNVIDIA / Hybrid MoE
A hybrid NVIDIA model with an official NVFP4 checkpoint for agent inference.
- Parameters
- 30B / 3B active
- Context
- 1M
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
CodingAlibaba / Hybrid MoE
A coding-focused sparse model for long-running agents and local development.
- Parameters
- 80B / 3B active
- Context
- 256K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
ReasoningMistral AI / MoE · vision
A multimodal MoE unifying instruction following, reasoning, and coding.
- Parameters
- 119B / 6.5B active
- Context
- 256K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
ReasoningDeepSeek / MoE · vision
A multimodal model using encoder-decoder attention and KV cache compression.
- Parameters
- 552B+
- Context
- 1M
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
AgenticXiaomi / MoE · multimodal
MiMo's large multimodal model for long-context, tool-using agent workloads.
- Parameters
- 1.02T+ / 42B active
- Context
- 1M
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
MultimodalMiniMax / MoE · vision
A native multimodal model with sparse attention for million-token contexts.
- Parameters
- ~428B / 23B active
- Context
- 1M
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
AgenticMoonshot AI / MoE · vision
A large-scale vision-language MoE for coding and long-horizon agent tasks.
- Parameters
- 2.8T / 104B active
- Context
- 1M
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
CodingZ.ai / MoE
A large GLM checkpoint for coding, with a separate license from Flash.
- Parameters
- ~753B*
- Context
- Under review
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
ReasoningAlibaba / MoE
A large-scale Qwen model for distributed text-generation deployments.
- Parameters
- ~2.4T / 95B active
- Context
- Under review
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
Weight memory estimatePending
BaselineAlibaba / Dense
A compact starting point for reasoning, chat, and your first self-hosted deployment.
- Parameters
- 8.2B
- Context
- 32K native
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
16-bit weight estimate16.4 GB
BaselineXiaomi / MoE
Sparse compute, substantial memory. A closer look at the economics of a large MoE.
- Parameters
- 309B / 15B active
- Context
- 256K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
16-bit weight estimate618 GB
BaselineMeta / Dense
An established instruction-tuned baseline for a practical deployment comparison.
- Parameters
- ~8B
- Context
- 128K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
16-bit weight estimate16 GB
BaselineGoogle / Dense · vision
Text and image understanding, with a memory budget that deserves a closer look.
- Parameters
- ~12B
- Context
- 128K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
16-bit weight estimate24 GB
BaselineAlibaba / Dense
A larger dense reasoning model for exploring precision and memory trade-offs.
- Parameters
- 32.8B
- Context
- 32K native
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
16-bit weight estimate65.6 GB
BaselineGoogle / Dense · vision
A smaller multimodal model to explore when your memory budget comes first.
- Parameters
- ~4B
- Context
- 128K
- Minimum
- —GPU / VRAM · pending
- Recommended
- —GPU / VRAM · pending
16-bit weight estimate8 GB
Specifications from official model cards. Memory figures are calculated estimates, not measured VRAM. How to read the data