Text2WetLab Leaderboard

Agents write Opentrons OT-2 protocols for 11 wet-lab tasks: 7 easy tasks that give the steps, 4 hard tasks that give only a goal and the source paper (paper-only). Each protocol is simulated, checked against a ground truth, and scored by a three-vote rubric judge. One attempt per task, same Claude Code agent for every model. Run 2026-10-07-openrouter.

Overall

Refused tasks are not scored: a refusal is the provider's safety policy, not a protocol. Models are ranked on the 8 tasks every model answered; refusals are listed beside the score.

Per task

Hover a score for the judge votes and any failed check; † marks a provisional score. Every failure with its evidence is in the full report.

TaskClaude Sonnet 5.5Claude Opus 5.5Claude Fable 5.1GPT-6.1 Sol
Easy · steps given
a1-a12-100ul1.001.001.001.00
ampure-bead-cleanup1.001.001.001.00
colony-pcr-screening1.001.001.001.00
ecoli-heat-shock-transformation1.001.001.001.00
golden-gate-assembly1.00refusedrefused1.00
opentrons-rna-extraction0.301.000.751.00
split-200ul-two-wells1.001.001.001.00
Paper-only (hard) · goal and paper
colony-pcr-screening-hard0.75†1.001.000.50†
ecoli-heat-shock-transformation-hard0.75†0.75refused0.75
golden-gate-assembly-hard1.00refusedrefused0.75
opentrons-rna-extraction-hard0.69†1.001.000.69†

How to read it

RewardRubric score: three core items at 25% each (robot practice, tips and contamination, fidelity to the task or paper) and task items sharing 25%. A failed critical check caps it at 0.30.
VerifierLint gate, 10 reward-hacking traps, the Opentrons 7.5 simulator and deterministic checks against the ground truth come first. All 11 reference solutions pass; 152 broken or cheating protocols all fail.
Provisional (†)5 scores were judged before the paper-only audit removed two requirements the papers do not support, or with fewer than three judge votes. Hover a † for the reason; they are re-judged in the next run.