Key takeaways
- Anthropic announced the result on September 23, 2026. Its preprint says Claude Mythos 5 agents searched 1.94 billion protein clusters and identified a previously uncharacterized family the authors call array-associated reverse transcriptases, or ART.
- The reported campaign used 119 tasks, 949 agent sessions, 215.6 million tokens and 21.5 hours of wall time. It produced 19 reports, but only three of 17 candidate partner families survived as previously unreported associations.
- The strongest evidence is narrower than “AI discovered a CRISPR.” The authors report a distinct RT family, repeat arrays and abundant array-derived RNA during phage infection; they have not shown that the RT is active, that those RNAs are its substrates or what ART does.
- Discovery was not reliably repeatable end to end: ten additional campaigns all missed the defining upstream DNA array. Fixed-input tests performed much better when capable models actually read contiguous DNA into context.
- Treat this as evidence for supervised anomaly discovery over primary scientific data, not autonomous biological truth. Pilot only with preserved traces, deterministic filters, negative-run reporting and qualified experimental validation.
Anthropic reported the ART finding on September 23
Anthropic announced on September 23, 2026 that it had formed a life-sciences research group and laboratory and that Claude agents had surfaced a previously uncharacterized enzyme system in bacteriophages. The linked preprint names it array-associated reverse transcriptase, or ART, because a reverse-transcriptase gene sits beside an array of non-coding DNA repeats and a dedicated partner gene.
This is a research release, not a new model or API launch. The campaign ran on Claude Mythos 5 through Claude Code and an internal multi-agent harness; follow-up work used interactive Claude Science sessions. Claude Science is publicly described as a beta product, but the reviewed materials provide no reproducible campaign package, public endpoint, exact configuration or token price for this run.
The campaign searched 1.94 billion protein clusters
The authors gave the system a research brief asking it to find novel reverse-transcriptase systems through new partner-gene associations. Worker agents planned and executed analyses; supervisor agents reviewed results and opened follow-up tasks; curators maintained shared records; editors produced reports. The agents built sequence-profile models, recovered 198,290 RT clusters, classified them into nine classes, sampled 10,983 loci and scored 3,564 recurring neighboring protein families.
The complete campaign comprised 119 tasks and 949 agent sessions, including 98 follow-up tasks. The preprint reports 77 agent-hours, 21.5 hours of wall-clock time and 215.6 million tokens without human intervention during the computational campaign. Those are vendor-reported resource counts, not a bill: Anthropic does not disclose an internal or customer-equivalent dollar cost for the run.
The useful signal emerged after an initial idea failed
A worker first investigated an apparent association between an RT and a phage RNA-polymerase subunit, then rejected that association as spurious. Its supervisor opened a follow-up on the RT. A later worker loaded upstream DNA directly into context and noticed repeated sequences that resembled a tandem array, even though the original brief had focused on protein partners rather than non-coding DNA.
That path is important because it includes refutation rather than a straight success narrative. Of 17 candidate partner families promoted for deeper investigation, the authors retained only three as previously unreported RT associations; 14 were annotation artifacts, parts of known systems or ordinary neighbors. The campaign also reported three new RT lineages. A useful discovery workflow needs the rejected majority and the reasons for rejection, not only the winning trace.
The preprint supports a new family, not a known function
Follow-up profile searches identified 95 distinct ART RT clusters at 90% identity in cultured jumbo phages and predicted viral contigs. Twenty-eight had a detectable upstream repeat array. Those arrays spanned 0.3 to 4.1 kilobases, with three to 21 copies of short repeats separated by longer unique spacers. The authors distinguish this architecture from CRISPR arrays and report no nearby cas genes.
Analysis of public RNA-sequencing data from Staphylococcus phage SA1 showed the array was highly expressed during infection and appeared as discrete RNA units. The paper proposes that ART might use a repertoire of RNAs with its RT and partner protein. That is a hypothesis. The authors explicitly say they have not shown that the RT is active, that the unit RNAs are its substrates, that the RT and partner interact or what the system does for the phage.
Ten full reruns missed the defining array
Anthropic reran the same campaign ten times. Nearly every run that completed the candidate census sampled ART loci, and two opened follow-up investigations, but none read the upstream DNA and none rediscovered the repeat array. The authors attribute this to the broad search space and non-deterministic harness behavior.
That failure is operationally more informative than a single impressive transcript. It shows that having the capable model and relevant data in the workflow did not make discovery dependable. A research team should report yield across full runs, the cost of misses and the human effort needed to identify a promising report—not treat one selected trajectory as a repeatability estimate.
Directly reading the data mattered more than adding tools
The authors then built fixed-input tests around ART sequences. They report that Opus 5.5, Mythos 5.1, Mythos 5 and Opus 5 outperformed Opus 4.6, Opus 4.8 and Sonnet 5 on a model-judged ten-feature rubric. When only the loci were placed directly in context, the stronger models described the array in at least 90% of attempts; in a tool-rich condition, performance fell as low as 32% for Opus 5.
Transcript inspection supplied a plausible mechanism: 39% of file-based attempts never read a contiguous stretch of at least 200 nucleotides, so they did not observe more than roughly one repeat unit. Recognition improved by 16 to 32 percentage points when models read at least 200 contiguous nucleotides. These are author-designed, author-run and model-judged benchmarks, not independent evaluations, but they expose a practical failure mode: access to a file is not evidence that an agent inspected the relevant data.
The evidence is substantial but institutionally concentrated
The release is unusually detailed for an AI-for-science claim: it includes a 40-page preprint, resource counts, a candidate funnel, selected traces, negative reruns, fixed-input tests, public-data identifiers and explicit unknowns. The wet-lab and computational follow-up also goes beyond accepting the agent’s prose at face value.
The evidence is nevertheless concentrated inside Anthropic. Every listed preprint author is affiliated with Anthropic; Anthropic supplied the model, harness, internal signals, database and laboratory; and the document is a preprint. AccessAllGPT found no peer-reviewed paper or independent replication in the reviewed materials. The current record supports “Anthropic reports a promising AI-originated lead with follow-up evidence,” not “Claude independently proved a new programmable biotechnology.”
What research teams should do next
A sensible pilot starts with a bounded discovery problem over data a qualified team already understands. Freeze the model and harness, preserve every task and tool trace, define deterministic data-coverage checks, record all candidates and rejections, and reserve validation data that the agent cannot use during exploration. Measure accepted leads per model token, compute hour and expert-review hour—not the number of polished reports.
Route every candidate through domain review and an experimental plan before biological action or publication. Require at least one rerun or independent path that could falsify the claim, and publish failure rates beside the best trace. Wait when the system cannot prove which primary data it inspected. Reject autonomous scientific claims when the only evidence is model-generated narrative, a selected run or a vendor-controlled judge.
Copy-ready AI discovery campaign record
Complete this record for one model, harness, dataset and scientific target before accepting an agent-generated lead.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Pre-specified target, known baseline, what a positive lead means, what it does not establish and accountable domain scientist.
Model and version, harness revision, prompts, worker and supervisor roles, tools, permissions, context limits, judge and configuration hashes.
Dataset versions, licenses, size, filters, expected units, units actually opened, contiguous-data checks, missingness and contamination controls.
Tasks, sessions, tokens, CPU/GPU hours, wall time, storage, failed runs, review labor, laboratory cost and cost per accepted lead.
Initial candidates, promotions, rejections, rejection reasons, duplicates, known-system checks, selected reports and all negative outcomes.
Number of full reruns, rediscovery rate, fixed-input tests, stochastic settings, data-read rate and sensitivity to model or tooling changes.
Validator and conflicts, held-out computational checks, experimental assay, controls, result, artifacts, peer-review state and unresolved discrepancies.
Permitted analyses and experiments, biosafety and ethics review, prohibited autonomous actions, approval owners and stop conditions.
Explore, validate, publish, wait or reject; approved scope; evidence gaps; budget; expiry; correction owner; triggers for a fresh review.
Primary sources
Browse the publication-wide evidence index →
- Claude discovers a novel enzyme system with CRISPR-like repeatsAnthropic · Reviewed: September 23, 2026 announcement; life-sciences lab launch; autonomous search description; array-associated reverse transcriptase finding; stated human involvement; early experimental evidence; open questions; linked preprint · Retrieved · Supports: Anthropic announced its new life-sciences research group and laboratory on September 23, 2026 and reported that Claude agents identified array-associated reverse transcriptases. The post gives the approximate search scale and says Anthropic still does not know the system’s primary function.
- Autonomous AI agents discover reverse transcriptases with tandem repeat arraysAnthropic preprint (Peter H. Yoon et al.) · Reviewed: Abstract; introduction; agentic genome-mining methods; candidate funnel; ART discovery trace; family characterization; phage-infection RNA evidence; reproducibility reruns; fixed-input model benchmarks; internal-signal analysis; discussion; methods and supplementary figures · Retrieved · Supports: The 40-page preprint reports the harness, model, dataset, task and token counts, candidate funnel, sequence and expression analyses, ten campaign reruns, benchmark results and unresolved biological questions. All listed authors are affiliated with Anthropic; the document is a preprint rather than peer-reviewed independent replication.
- Claude Science (beta)Anthropic · Reviewed: Beta status; product positioning; analysis and database-search claims; artifact provenance; managed compute; download and contact-sales access paths · Retrieved · Supports: Anthropic currently labels Claude Science as beta and describes an application that runs analyses, searches databases and retains code and artifact history. The reviewed page offers download and contact-sales paths but does not publish a model-token price for reproducing the reported campaign.
Limitations
AccessAllGPT reviewed only Anthropic-controlled sources. We did not run Claude Mythos 5, Claude Code or Claude Science; inspect the proprietary harness, database or internal model signals; reproduce 119 tasks or 215.6 million tokens; verify invoices; validate the 1.94-billion-cluster search; audit the 198,290-cluster catalog; rerun sequence profiles, phylogenies, structure predictions or RNA analysis; inspect every trace; culture phage; assay RT activity; test substrates or partners; establish ART function; or interview independent biologists. The paper is a preprint whose listed authors are all Anthropic affiliates. The fixed-input benchmark uses an Anthropic model as judge, and full campaign reruns did not rediscover the defining array. Claude Science access, model availability, names, pricing and capabilities may change. This is not biological, medical, biosafety, ethics or investment advice.
Disclosures
AccessAllGPT did not receive Anthropic access, credits, data, traces, preprint files, laboratory material, a briefing, review or compensation for this article. Anthropic and the paper authors did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Anthropic or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Claude Computes a Nine-Loop Physics Result—With Validation and a $1,000–$2,000 Cost Estimate
- From Paper to Production Without Losing the Plot
- Design an Agent Benchmark That Predicts Production
- Where Human Approval Belongs in AI Automation
- Choose a Model Without Chasing the Leaderboard
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.