Quick answer
On September 26, 2026 Anthropic published an account of running roughly 950 Claude agents in a coordinated swarm on a biology research task involving reverse transcriptase enzymes, the class of proteins that copy RNA into DNA. The agents divided literature review, hypothesis generation, and analysis among themselves and aggregated results. It is a demonstration of scale for multi-agent research, not a claimed discovery, and the methods and results are Anthropic's own report pending peer review.
Multi-agent systems have mostly been demos with five or ten agents. Nine hundred and fifty is a different kind of experiment, and it raises a different kind of question: what happens to scientific work when reading everything is cheap.
What the swarm did
- Divided a large body of literature on reverse transcriptases among agents for reading and extraction
- Generated and ranked hypotheses about enzyme variants and their properties
- Ran analyses and cross-checked each other's claims against sources
- Aggregated findings into a summary for human researchers to evaluate
Why reverse transcriptases
They are central to biotechnology: essential in RNA sequencing, diagnostics, and gene-editing tools, with a large literature and a real engineering problem (finding variants with better properties). That makes them a good test: enough prior work for agents to read, and useful outputs a lab could check.
What to be sceptical about
- This is a vendor account of a vendor experiment; no independent replication yet
- Reading and proposing is not the same as discovering; the value depends on wet-lab validation
- Aggregating 950 agents multiplies errors as well as coverage; the cross-checking design matters and is hard to assess from outside
- Cost: nothing was said about what the run cost, which matters for whether labs can repeat it
Why it matters anyway
Even with every caveat, the experiment shows that literature-scale reading and hypothesis generation is now a thing a lab can buy. That changes the bottleneck in science from reading to validating, which was already the expensive part. Tools that help researchers check AI-generated hypotheses, from scite's citation analysis to lab automation, become more valuable, not less.
Related reading
Bottom line
A swarm that reads everything is a real capability. Whether it produces real science is a question for the labs that validate its output, and that is where the interesting results will come from.

