cloud venue · open weights & multi-model mixes
How far do open-weight coding agents get?
The same squeezer agent and rubrics as the local 4070 benchmark, pointed at open-weight models - and multi-model mixes of them rented on OpenRouter. For each arm it reports score, cost per task, latency, and trial count, so you can see which setup extracts the most real coding ability per dollar.
What should you run?
the best open-weight coding agent for your hardware vs the cloud, by score on this benchmark. quality is venue-independent, so we measure it once in the cloud.Best fit for your VRAM12 GB
No measured model fits this VRAM yet.
Best rented open-weight model
No cloud data yet.
Local quality shown is the cloud (higher-precision) score; a local q4 build runs ~3-5 points lower in our tests. VRAM is an approximate q4 estimate.
Best local (4070)
-
no local runs
Best cloud arm
-
no cloud runs
Best value (score / $)
-
-
Cloud spend so far
$0.00
0 cloud runs · 0 arms
Looking for the ranking? There's now one leaderboard covering every contender - local GPUs and cloud arms together, with $/task - on the main dashboard. Below is the cloud-only per-eval breakdown and the separate hard tier.