Text2WetLab is the first end-to-end benchmark that converts published wet-lab papers into executable Opentrons OT-2 protocols and grades correctness in a physics-accurate MuJoCo simulation.
We chain a structured extraction model, an intermediate representation (IR), an Opentrons code generator, and a simulation-based verifier to turn any methods-section PDF into a graded, reproducible robot protocol.
Models invent pipetting actions not present in the source protocol— mixing competent cells, spurious washes, or phantom reagent additions that would ruin a transformation.
Off-by-20µL elution volumes in RNA extraction risk bead carryover into downstream qPCR—a silent failure invisible to the operator but lethal to a clinical assay.
Because the OT-2 simulator tracks every µL with perfect fidelity, we can assign deterministic pass/fail rewards without a wet lab— enabling large-scale, reproducible AI evaluation.
Real published wet-lab protocols, graded at full simulation fidelity. Each exposes a distinct failure mode — or a surprising success.
This four-step protocol (add DNA → heat shock → SOC recovery → incubate)
is conceptually simple — but both Opus and Fable
added an unrequested mix() call immediately after dispensing
the plasmid DNA into competent cells. Vigorous mixing at that stage
disrupts osmotic conditions and dramatically reduces transformation efficiency.
Sonnet followed the spec exactly.
Assembling four four-fragment chromoprotein expression plasmids with BsaI-HFv2 and T4 ligase demands precise volume fractions, correct fragment ordering, DpnI digestion timing, and a thermocycler programme across 33 protocol steps. All three models — Opus, Sonnet, and Fable — achieved a perfect 1.00 reward. This is the benchmark at its best: a complex protocol faithfully rendered by every contender.
A magnetic bead RNA extraction over 48 samples — 16 simultaneous checks cover bead binding order, wash volumes, magnet timing, air-dry duration, and elution recovery. The critical failure: all models recovered the full 100µL elution volume instead of the protocol-specified ~80µL. The remaining 20µL near the pellet carries magnetic beads. In a clinical RT-qPCR workflow, those beads inhibit polymerase — a silent assay failure that looks like RNA degradation.
Mean reward across all 7 tasks (pass@1, single attempt per task). Grading combines deterministic end-state checking and LLM-judged rubrics.
| Task | Grader | Oracle | Opus | Sonnet | Fable |
|---|---|---|---|---|---|
| a1-a12-100ul | end-state | 1.000 | 1.000 | 0.917 | 0.833 |
| ampure-bead-cleanup | end-state | 1.000 | 0.944 | 0.889 | 1.000 |
| colony-pcr-screening | end-state | 1.000 | 1.000 | 0.875 | 0.875 |
| ecoli-heat-shock-transformation | end-state | 1.000 | 0.714 | 0.857 | 0.714 |
| golden-gate-assembly | end-state | 1.000 | 1.000 | 1.000 | 1.000 |
| opentrons-rna-extraction | LLM judge + 16 checks | 0.944 | 0.889 | 0.833 | 0.833 |
| split-200ul-two-wells | end-state | — | 1.000 | 1.000 | 0.917 |
| Mean | 0.991 | 0.935 | 0.910 | 0.882 | |
Text2WetLab is open source. Add a new model, contribute a protocol task, or replicate our results with a single Docker command. Full dataset on HuggingFace.