Benchmark · 2026

Millions of lab protocols will be executed by AI.
How do we know they're right?

Text2WetLab is the first end-to-end benchmark that converts published wet-lab papers into executable Opentrons OT-2 protocols and grades correctness in a physics-accurate MuJoCo simulation.

7 Protocol tasks
3 Models benchmarked
0.935 Top score (Opus)
✓ sim_pass
reward: 1.000
12 commands
Section 01 — The Pipeline

From PDF to Pipette in Five Steps

We chain a structured extraction model, an intermediate representation (IR), an Opentrons code generator, and a simulation-based verifier to turn any methods-section PDF into a graded, reproducible robot protocol.

📄
Paper PDF
Methods section
→
💬
NL Protocol
paper2protocol
→
🗂
IR (JSON)
Structured steps
→
🤖
Python Protocol
OT-2 API v2
→
🎮
MuJoCo Sim
Physics reward
⚠️

Hallucinated Steps

Models invent pipetting actions not present in the source protocol— mixing competent cells, spurious washes, or phantom reagent additions that would ruin a transformation.

🧮

Volume Errors

Off-by-20µL elution volumes in RNA extraction risk bead carryover into downstream qPCR—a silent failure invisible to the operator but lethal to a clinical assay.

🔁

Simulation as Ground Truth

Because the OT-2 simulator tracks every µL with perfect fidelity, we can assign deterministic pass/fail rewards without a wet lab— enabling large-scale, reproducible AI evaluation.

Section 02 — Task Showcases

Three Protocols That Tell the Story

Real published wet-lab protocols, graded at full simulation fidelity. Each exposes a distinct failure mode — or a surprising success.

Task 01 / E. coli Heat Shock Transformation

The Phantom Mix

"The robot mixed the cells it was never asked to touch."

This four-step protocol (add DNA → heat shock → SOC recovery → incubate) is conceptually simple — but both Opus and Fable added an unrequested mix() call immediately after dispensing the plasmid DNA into competent cells. Vigorous mixing at that stage disrupts osmotic conditions and dramatically reduces transformation efficiency. Sonnet followed the spec exactly.

# Correct (Sonnet) + p20.transfer(2, plasmid['A1'], cells['A1']) # Hallucination (Opus / Fable) - p20.transfer(2, plasmid['A1'], cells['A1'], - mix_after=(3, 10)) # ← phantom step
✓ Sonnet — 0.857 ⚠ Opus — 0.714 ✗ Fable — 0.714
Task 02 / Golden Gate Assembly

Perfect at 33 Steps

"613 simulation commands. Every check green."

Assembling four four-fragment chromoprotein expression plasmids with BsaI-HFv2 and T4 ligase demands precise volume fractions, correct fragment ordering, DpnI digestion timing, and a thermocycler programme across 33 protocol steps. All three models — Opus, Sonnet, and Fable — achieved a perfect 1.00 reward. This is the benchmark at its best: a complex protocol faithfully rendered by every contender.

✓ Opus — 1.000 ✓ Sonnet — 1.000 ✓ Fable — 1.000
# 9 rubric categories, all scored 1.0 ✓ deck_and_hardware ✓ pcr_setup ✓ dpni_and_cleanup ✓ assembly_mix ✓ thermocycler_prog ✓ transformation ✓ tips_contamination ✓ robot_practice ✓ fidelity_to_paper
2D IR visualization of RNA extraction elution step
Task 03 / Opentrons RNA Extraction

The 20µL That Matters

"Recovering 100µL instead of 80µL risks bead carryover."

A magnetic bead RNA extraction over 48 samples — 16 simultaneous checks cover bead binding order, wash volumes, magnet timing, air-dry duration, and elution recovery. The critical failure: all models recovered the full 100µL elution volume instead of the protocol-specified ~80µL. The remaining 20µL near the pellet carries magnetic beads. In a clinical RT-qPCR workflow, those beads inhibit polymerase — a silent assay failure that looks like RNA degradation.

⚠ Opus — 0.889 ⚠ Sonnet — 0.833 ⚠ Fable — 0.833
# All models wrote: - p1000.transfer(ELUTION_VOL, elution_plate, # 100 µL ← wrong # Correct approach: + p1000.transfer(80, elution_plate, # ~80 µL ✓
Section 03 — Results

Model Leaderboard

Mean reward across all 7 tasks (pass@1, single attempt per task). Grading combines deterministic end-state checking and LLM-judged rubrics.

Model
Mean reward (7 tasks)
Score
#1
claude-opus-5-5
0.935
#2
claude-sonnet-5-5
0.910
#3
claude-fable-5-1
0.882
Task Grader Oracle Opus Sonnet Fable
a1-a12-100ul end-state 1.000 1.000 0.917 0.833
ampure-bead-cleanup end-state 1.000 0.944 0.889 1.000
colony-pcr-screening end-state 1.000 1.000 0.875 0.875
ecoli-heat-shock-transformation end-state 1.000 0.714 0.857 0.714
golden-gate-assembly end-state 1.000 1.000 1.000 1.000
opentrons-rna-extraction LLM judge + 16 checks 0.944 0.889 0.833 0.833
split-200ul-two-wells end-state — 1.000 1.000 0.917
Mean 0.991 0.935 0.910 0.882
Section 04 — Get Involved

Run the Benchmark on Your Model

Text2WetLab is open source. Add a new model, contribute a protocol task, or replicate our results with a single Docker command. Full dataset on HuggingFace.