LatchBio built a 100-question benchmark called TxBench-Antibody Discovery to test whether AI agents can turn real experimental data into defensible antibody discovery decisions, and the best-performing configuration, Claude Opus 5 running inside Claude Code, passed only 53.0 percent of evaluable attempts. Ten competencies are covered, running from judging a target and designing an assay through cellular pharmacology to de-risking a candidate. Every task comes out of a published experimental study and is graded against a fixed answer key. No system cleared 55 percent.

The gap at the top was thin. Grok 4.6 paired with the PI harness reached 51.7 percent, Grok 4.6 with Grok Build hit 51.0 percent, Opus 5 with the PI harness reached 51.4 percent, and Gemini 3.7 Flash with PI reached 50.0 percent. Four configurations from three model families landed within three percentage points of the leader, according to LatchBio’s paper. Every task in the set of 100 was cracked by some configuration, and 89 were cracked by models from two or more families, so nothing in the set depends on knowledge only one system carries.

OpenAI’s configurations trailed the field by a wide margin. GPT-5.6 Sol passed 33.3 to 33.8 percent depending on harness, GPT-5.5 passed 31.5 to 33.6 percent, GPT-5.6 Terra passed 29.7 to 31.0 percent, and GPT-5.6 Luna passed just 18.2 to 19.1 percent, the lowest of the 20 configurations LatchBio tested. Harness choice mattered independently of the underlying model: Gemini 3.7 Flash dropped from 50.0 percent with PI to 45.3 percent with Mini-SWE-Agent, a nearly five-point swing on identical questions.

Spending more did not buy accuracy. Opus 5 with Claude Code, the top performer, averaged roughly $1.70 and 1.12 million tokens per run. Grok 4.6 with PI matched it within two points at $0.59 and 0.50 million tokens, less than half the cost. Gemini 3.7 Flash with PI ran even cheaper, at $0.46 per evaluation, while using more tokens and tool calls than either. At the other extreme, Opus 4.8, an earlier model, cost $1.99 to $2.38 per run and still passed only 42.0 to 45.6 percent, and GPT-5.6 Luna was the cheapest configuration tested at $0.06 to $0.07 per run and the least accurate.

LatchBio’s failure analysis is the most consequential finding for anyone deploying these systems on real discovery work. The typical failure was not bad arithmetic. Models ran coherent calculations aimed at a nearby question instead of the one the experimental evidence was actually posing. One task on translating receptor findings across species illustrates it: Opus 5 in Claude Code went three for three on about 0.29 million tokens, while Gemini 3.7 Flash with PI went zero for three while spending over twice the tokens and three times the tool calls. Each model noticed the same inconsistency in the data; Gemini simply carried a wrong experimental-arm assignment through to its answer. All results are LatchBio’s own on a benchmark it designed and has not had independently replicated.

For teams piloting AI agents on antibody discovery workflows, a roughly 50 percent pass rate on grounded decisions means these systems are better suited to accelerating a scientist’s first pass than to standing in for one. Any deployment plan should budget for a human review step specifically targeted at scientific framing errors, since LatchBio’s data shows that is where most failures originate, not in the arithmetic.

According to LatchBio’s TxBench-Antibody Discovery paper, published September 2, 2026.