Workload shortlist

Coding & agent model shortlist

Compare deployment candidates for coding and tool-using agents, with model-card sources, runtime questions, and an application-level evaluation plan.

What is included
Selected profiles whose publisher descriptions cover coding or agent workloads. Model sizes vary substantially; inclusion does not mean equivalent hardware needs or verified tool compatibility.
How this list is ordered
Alphabetical by model name. No benchmark score or performance order is assigned.

Publisher sources are linked for every model. No independent performance scores have been measured. On smaller screens, scroll the table horizontally.

Coding & agent model shortlist. Alphabetical by model name. No benchmark score or performance order is assigned.
ModelArchitectureContextDeployment note
XiaomiSource
MoE · multimodal1MA large sparse checkpoint for testing multi-step tool workflows.Compare model
Mistral AISource
MoE · vision256KReasoning and coding candidate; Small is a product name, not a memory guarantee.Compare model
NVIDIASource
Hybrid MoE1MAn official quantized agent candidate; validate the NVFP4 deployment path.Compare model
AlibabaSource
Hybrid MoE256KCoding-focused; evaluate the actual tool parser and agent scaffold.Compare model
AlibabaSource
Dense · vision256K nativeA dense candidate for a different hardware budget from the large sparse entries.Compare model

Rank completed tasks when measurements arrive

For a coding agent, generating tokens is only one part of the job. Record whether the change passes its tests, whether tool calls are valid, and how many attempts are required. Keep the repository snapshot, agent instructions, tool access, and task budget fixed so a future ranking measures comparable work.

The serving configuration includes the scaffold

Chat templates, tool-call parsing, reasoning settings, and maximum output length affect the behavior an agent sees. Validate one complete tool cycle before a longer evaluation. A successful plain-text chat request does not prove that a checkpoint works with a particular coding harness.

Separate quality from capacity

First reject configurations that cannot complete the target tasks reliably. Then compare end-to-end time and completed tasks within a fixed resource budget. No independent coding scores or deployment winners have been measured by BenchGrid yet; these entries are candidates for that experiment.

Before you choose hardware

A good benchmark shows its working. explains the next planning step. For the distinction between publisher specifications, calculations, and measurements, read the BenchGrid methodology.