Model rankings & shortlists
Find candidates for your next deployment. Compare the specifications we can verify, then narrow the test to your workload.
Specifications first. Calculated rankings and editorial shortlists are labeled separately. Performance benchmarks are not available yet.
Our methodologyExplore the lists 6
Sources and ordering explained on every pageSmallest weight footprints
How much space do the weights take?
Long-context model tiers
Which models advertise room for longer inputs?
Compact models to evaluate
Where should a smaller deployment experiment start?
Coding & agent model shortlist
Which candidates belong in a coding-agent test?
Multimodal deployment shortlist
Which models should we test with more than text?
MoE deployment shortlist
What does sparse compute mean for deployment?
The measured rankings come next
These need reproducible GPU runs. No scores or winners have been assigned.
24 / 48 / 80 GB GPU fit
Working configurations at a stated precision, context, and concurrency.
Inference speed
Successful throughput and tail latency on the same hardware and workload.
Deployment cost
Measured capacity paired with dated instance pricing and utilization assumptions.
How to use these lists
Choose a question, inspect the inclusion rules, then open a model profile to check the original card. The lists cover a curated directory, not the whole open-model ecosystem. Position in an alphabetical shortlist does not indicate quality.
To compare individual candidates side by side, use the model comparison tool. For sizing context and runtime overhead, start with our GPU memory field note.