Which open-weight coding agents
actually finish the work?
lemon-squeezer tests open-weight LLM coding agents across model, harness, config, and venue (a single local GPU, or open weights rented in the cloud), then scores each run by executing the code the agent produced against a deterministic rubric. No human grading, no vibes.
Leaderboard
0 contendersEvery contender, every venue ranked together by capability. Mean score per (harness × model × eval), averaged over trials; the rank column is coverage-adjusted, so a 1/0 row of 100% can't outrank a 0/0 row of 95%. $/task is mean cloud cost (blank for local). The separate hard tier is on the cloud page. Click any cell for its breakdown.
harness = the agent scaffold around the model (single-pass, plan-then-build, write-tests-and-self-correct, ...). It moves scores as much as the model does - hover any harness chip for what it does. The venue glyph is cloud, local GPU.
| # | host | harness | model | avg | rank | cov | $/task |
|---|
The harness gap
0 harnessesMean score per harness on each eval (averaged across models, tags, and trials). The spread on a single eval is how much the harness choice alone is worth.