open research · a reproducible benchmark for open-weight coding agents
0scored runs

Which open-weight coding agents actually finish the work?

lemon-squeezer tests open-weight LLM coding agents across model, harness, config, and venue (a single local GPU, or open weights rented in the cloud), then scores each run by executing the code the agent produced against a deterministic rubric. No human grading, no vibes.

0 scored runs40 executable coding evalsmodel × harness × config × venuescore = code that runscost + latency tracked
Filters

Leaderboard

0 contenders

Every contender, every venue ranked together by capability. Mean score per (harness × model × eval), averaged over trials; the rank column is coverage-adjusted, so a 1/0 row of 100% can't outrank a 0/0 row of 95%. $/task is mean cloud cost (blank for local). The separate hard tier is on the cloud page. Click any cell for its breakdown.

harness = the agent scaffold around the model (single-pass, plan-then-build, write-tests-and-self-correct, ...). It moves scores as much as the model does - hover any harness chip for what it does. The venue glyph is cloud, local GPU.

#hostharnessmodelavgrankcov$/task

The harness gap

0 harnesses

Mean score per harness on each eval (averaged across models, tags, and trials). The spread on a single eval is how much the harness choice alone is worth.