Coverage — planned vs landed
Every planned model × effort × arm, matched against landed data. Batches land incrementally — this board is the honest map of what exists and what is still a gap.
as of Jul 20, 2026, 4:07 PM157/ 264
59%landed · 59% of 264 planned
157
landed
1
partial
103
gap
3
paused
264
planned
742 valid slots collected
Chinese models: 27 official-API seats still unrun, 14 families with zero data so far.
This is the gap we most want closed. Every figure here is computed from the plan against landed data — never a hand-written count. Click any gap cell below to run it yourself with your own key.
17families planned27 / 36api seats open14zero-data families42 / 86cells with data
landed— every variant has ≥1 valid slotpartial— some valid slots, but not every variantattempted— config exists, zero valid slotsmissing— no config yetpaused— planned but disabled in the matrix
claude-fable-5claude
claude-haiku-4-5claude
claude-opus-4-8claude
claude-sonnet-5claude
deepseek-v4-flashdeepseek
highmax
deepseek-v4-prodeepseek
highmax
doubao-seed-2-1-pro-260628doubao
disabledenabled
ernie-5.1ernie
gemini-3.1-progemini
gemini-3.5-flashgemini
minimallowmediumhigh
google/gemma-4-31b-itgemma
glm-5.2glm
gpt-5.2-codexgpt
lowmediumhighxhigh
gpt-5.4gpt
nonelowmediumhighxhigh
gpt-5.4-minigpt
nonelowmediumhigh
gpt-5.5gpt
high
gpt-5.6-lunagpt
gpt-5.6-solgpt
gpt-5.6-terragpt
grok-4.3grok
lowmediumhighdefault
grok-4.5grok
lowmediumhighdefault
hy3hunyuan
intern-s1-prointern
kwaipilot/kat-coder-pro-v2.5kat
kimi-k2.6kimi
offon
kimi-k2.7-codekimi
kimi-k3kimi
meta-llama/llama-4-maverickllama
LongCat-2.0longcat
mai-thinking-1mai
mercury-2mercury
mimo-v2.5-promimo
MiniMax-M3minimax
mistral-medium-2604mistral
nvidia/nemotron-3-ultra-550b-a55bnemotron
amazon/nova-2-lite-v1nova
allenai/olmo-3.1-32b-thinkolmo
openpangu-2.0-flashpangu
default
qwen3.7-maxqwen
inclusionai/ring-2.6-1tring
spark-xspark
step-3.7-flashstep
claude-fable-5claude
lowmediumhighxhighmax
claude-haiku-4-5claude
lowmediumhighxhighmax
claude-opus-4-8claude
lowmediumhighxhighmax
claude-sonnet-5claude
lowmediumhighxhighmax
gpt-5.6-solgpt
gpt-5.6-terragpt
gpt-5.4gpt
lowmediumhighxhigh
gpt-5.4-minigpt
lowmediumhigh
gpt-5.5gpt
lowmediumhighxhigh
gpt-5.6-lunagpt
lowmediumhighxhigh
gpt-5.6-solgpt
lowmediumhighxhigh
gpt-5.6-terragpt
lowmediumhighxhigh
glm-5.2glm
deepseek-v4-flashmulti
deepseek-v4-promulti
glm-5multi
glm-5.1multi
glm-5.2multi
grok-4.5multi
hy3-previewmulti
kimi-k2.5multi
kimi-k2.6multi
kimi-k2.7-codemulti
kimi-k3multi
mimo-v2-omnimulti
mimo-v2-promulti
mimo-v2.5multi
mimo-v2.5-promulti
minimax-m2.5multi
minimax-m2.7multi
minimax-m3multi
qwen3.5-plusmulti
qwen3.6-plusmulti
qwen3.7-maxmulti
qwen3.7-plusmulti
grok-4.5grok
lowmediumhigh
kimi-k3kimi
claude-haiku-4-5multi
claude-opus-4-5multi
claude-opus-4-6multi
lowmediumhighmax
claude-opus-4-7multi
lowmediumhighxhighmax
claude-opus-4-8multi
lowmediumhighxhighmax
claude-sonnet-4multi
claude-sonnet-4-5multi
claude-sonnet-4-6multi
lowmediumhighmax
claude-sonnet-5multi
lowmediumhighxhighmax
deepseek-v3.2multi
glm-5multi
gpt-5.6-lunamulti
lowmediumhighxhighmax
gpt-5.6-solmulti
lowmediumhighxhighmax
gpt-5.6-terramulti
lowmediumhighxhighmax
minimax-m2.1multi
minimax-m2.5multi
qwen3-coder-nextmulti
deepseek-v4-flashmulti
deepseek-v4-promulti
glm-5multi
glm-5.1multi
glm-5.2multi
grok-4.5multi
hy3-previewmulti
kimi-k2.5multi
kimi-k2.6multi
kimi-k2.7-codemulti
kimi-k3multi
mimo-v2-omnimulti
mimo-v2-promulti
mimo-v2.5multi
mimo-v2.5-promulti
minimax-m2.5multi
minimax-m2.7multi
minimax-m3multi
qwen3.5-plusmulti
qwen3.6-plusmulti
qwen3.7-maxmulti
qwen3.7-plusmulti
qwen3.7-maxqwen
Cost referencePer-call USD ranges observed while building this dataset. A config is N valid slots × 2 prompt variants.
| family | arm | per-call USD | basis | note |
|---|---|---|---|---|
| gpt | api | $0.05 – $2 | observed | Official /responses, background mode. Averaged ~$0.45/call across effort tiers on gpt-5.4 / 5.4-mini / 5.2-codex; xhigh on full-size models reaches ~$2/call. gpt-5.6 family not yet measured, expected higher. |
| grok | api | $0.02 – $0.12 | observed | Official x.ai API: full grok-4.3 + grok-4.5 sweep (64 calls) cost ~$3.1 total. |
| claude | CC | $0.01 – $0.6 | plan | Claude Code harness on a flat-rate plan; the metered-equivalent of our 4-model x 5-effort sweep was ~$64 (fable-5 alone ~$49). Zero marginal cost on plan. |
| claude | api | $0.05 – $1 | estimated | Official Anthropic API not yet run; range projected from Claude Code metered-equivalents. |
| gpt | codex-oauth | plan / free | plan | codex CLI signed into a ChatGPT plan; zero marginal cost. |
| grok | grok-cli | plan / free | plan | grok CLI OAuth subscription; zero marginal cost. |
| multi | opencode | plan / free | plan | opencode go subscription (22 models, CLI harness + gateway api); zero marginal cost. |
| multi | go | plan / free | plan | opencode go gateway api line, same subscription. |
| deepseek | api | $0.001 – $0.03 | estimated | Official first-party API list price ballpark; not yet run. |
| qwen | api | $0.002 – $0.05 | estimated | DashScope first-party list price ballpark; not yet run. |
| glm | api | $0.002 – $0.05 | estimated | Zhipu first-party list price ballpark; not yet run. |
| kimi | api | $0.002 – $0.05 | estimated | Moonshot first-party list price ballpark; not yet run. |
| gemini | api | $0.01 – $0.3 | estimated | Not yet run (deliberately deferred). |
Observed ranges — they vary strongly with effort tier; treat as planning numbers, not quotes.
Click any missing or attempted cell to add it to a contribute batch.