Key takeaways
- Anthropic published the guest post on October 1, 2026. BootLoops 1.0 is publicly inspectable at a pinned GitHub commit and is designed to work with different LLM agents rather than only Claude.
- Schwartz reports running roughly 400 candidate problems and producing 36 manuscripts in 18 fields with 19 coauthors over three months. Those portfolio totals are author-reported and are not equivalent to 36 independently replicated or peer-reviewed discoveries.
- The reported workflow used separate Claude Code sessions on Google Cloud virtual machines, project-specific repositories, background agents and a coordinating session. Schwartz says it remained compute- and token-intensive and required substantial guidance.
- The post documents material failure modes: Claude declared victory with unresolved lemmas, drew wrong conclusions, made unreliable time estimates, preferred long computations over better tools and lost context during compaction.
- Teams can inspect BootLoops now, but should pilot one bounded, checkable computation and validate it independently before treating the release as evidence for autonomous science.
Anthropic published the account and BootLoops artifact on October 1
Anthropic published “Claude-shaped science” on October 1, 2026 as a guest post by Harvard physicist Matthew D. Schwartz. The same dated release is anchored to BootLoops 1.0 at Git commit 66b680ce742e654cfe86da4f072a69061fe182b1. What is available now is a public website, source repository, installation documentation, tools, protocols and self-test entry point—not a new Claude model, API identifier or managed-service price.
The post discloses that Schwartz was working as a visiting researcher at Anthropic during the project and says BootLoops is owned and maintained by him rather than being an Anthropic project. The repository’s release commit names Anthropic as the software copyright holder and Schwartz as maintainer. Readers should therefore treat the post and project pages as first-party evidence about the workflow, not independent confirmation of the scientific portfolio.
The headline claim is 36 manuscripts from 400 candidate problems
Schwartz reports 36 manuscripts in 18 fields with 19 coauthors over three months, selected from about 400 candidate problems. The examples range from scattering amplitudes and ecology to population genetics, economics, linguistics, earth science and statistics. He describes a recurring pattern: Claude found a technically tractable connection, then a domain expert redirected the work toward a question the field would value.
Those numbers describe reported project throughput, not a common unit of scientific success. A manuscript may be a draft, a replication, a method transfer or a novel claim, and the reviewed release does not provide one independent validation register covering all 36. AccessAllGPT found no basis in these sources to convert “36 manuscripts” into a publication count, acceptance rate, discovery rate or benchmark score.
BootLoops is a model-independent scientific harness, not a model release
BootLoops combines quantitative-science software with protocols that tell an agent when to use a tool, how to interpret its output and what check must pass. Its README lists exact and high-precision integrals, recurrences with certificates, Bayesian evidence calculations, error-bounded numerical methods, exhaustive enumeration and open implementations of statistical procedures. The software is MIT-licensed; project documentation is CC BY 4.0.
The project says it can be driven by Claude, Gemini, ChatGPT or another agent. That is an architectural claim about the harness interface, not measured cross-model compatibility. The reviewed materials publish no support matrix, minimum model capability, token budget, cloud specification, end-to-end cost or success rate across providers.
Checkability is the central design choice
The project’s strongest operational idea is to choose computations with hard acceptance contracts. Its “BootLoops Standard” says a diagram is solved only when the full functional form is known and a Python script can evaluate it to arbitrary precision on a laptop. The broader protocols describe independent routes, held-out points, positive controls, declared constant sets and provenance rules intended to keep fitted data from certifying itself.
A public protocol does not prove that every result followed it correctly. The repository establishes inspectable machinery, while individual manuscripts and their artifacts must establish execution, completeness, novelty and domain validity. Teams should distinguish “the harness contains a checker” from “this result passed an independent check.”
The workflow used many agents but did not remove the scientist
Schwartz says separate Claude Code sessions ran on Google Cloud virtual machines, connected to project-specific GitHub and Overleaf repositories. A master session coordinated projects, compute and validation; background agents performed calculations; intermediate results were stored in Markdown; separate sessions handled writing, repositories, manuals, the website and adversarial review.
Human work remained central. Schwartz describes selecting problems, rejecting uninteresting results, translating between Claude and domain experts, checking plots, questioning conclusions, setting completion standards and deciding when a computation was grinding. The reported system is better described as high-throughput, expert-directed research automation than as an autonomous scientist.
The failure report is more reusable than the portfolio total
The guest post says Claude could declare a proof complete while leaving the decisive lemma unproved, report “good agreement” that still required visual inspection, draw incorrect conclusions, chase old but unimportant debates, estimate duration badly and continue expensive calculations instead of building a better method. Long sessions also lost context during compaction, prompting periodic file organization and plan consolidation.
These are not cosmetic prompt issues. They affect acceptance, novelty, scheduling and spend. A production research workflow needs deterministic completion gates, explicit novelty review, measured small runs before scale-up, budget and wall-time ceilings, durable state outside model context, and a named human who can stop a technically valid but scientifically pointless path.
What research teams should do next
Inspect the pinned repository before adopting it, including installation steps, third-party dependencies, licenses, tool-specific assumptions and self-tests. Start with one problem whose output can be checked by a genuinely different method. Freeze the repository commit, model and agent configuration; record prompts, tools, data, compute, tokens, wall time and expert review; and preserve failed candidates beside successful ones.
Adopt a bounded pilot when a domain expert owns the question and the answer has a cheap falsification path. Constrain the agent to analysis and artifact generation when conclusions or novelty remain judgment calls. Wait when required dependencies, costs or reproducibility are unclear. Reject unsupervised scientific claims when the evidence is only a fluent report, selected trajectory or project-controlled summary.
Copy-ready research-harness pilot record
Complete this for one BootLoops or comparable AI-science workflow before accepting a result or scaling the system.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Exact problem, why it matters, prior art, full acceptance criteria, prohibited shortcuts and accountable domain expert.
Repository commit, model and version, agent harness, prompts, skills, tools, dependencies, cloud image and configuration hashes.
Input datasets, versions, licenses, transformations, held-out evidence, contamination checks and durable manifests.
Independent method, positive and negative controls, precision or tolerance, checker ownership and proof the checker can fail.
Tokens, model charges, CPUs/GPUs, storage, wall time, failed runs, expert hours and total cost per accepted result.
All candidates, rejections, unresolved lemmas, context losses, retries, grind stops and reasons for selecting the reported path.
Domain reviewers, conflicts, novelty search, conclusion audit, replication status, manuscript state and unresolved disagreement.
Actions the agent may take, human approvals, publication restrictions, laboratory or policy exclusions, rollback and stop conditions.
Adopt, constrain, wait or reject; approved scope, budget ceiling, expiry, revalidation triggers and correction owner.
Primary sources
Browse the publication-wide evidence index →
- Claude-shaped scienceAnthropic (guest post by Matthew D. Schwartz) · Reviewed: October 1, 2026 publication date; summary; BootLoops origin and workflow; reported physics, ecology, genetics, economics and linguistics projects; 36-manuscript portfolio claim; orchestration details; failure modes; outlook; acknowledgements; visiting-researcher and project-ownership disclosure · Retrieved · Supports: Anthropic published Schwartz’s account on October 1. He reports using Claude Fable 5 with BootLoops across 400 candidate problems, yielding 36 manuscripts in 18 fields with 19 coauthors over three months, and describes expert steering, verification practices, compute intensity and recurring model failures. The post is a guest account, not independent validation of every scientific result.
- BootLoops: A toolkit for exact quantitative scienceBootLoops (Matthew D. Schwartz) · Reviewed: Project overview; harness independence; scientific-tooling scope; verification principles; BootLoops Standard; methods; example applications; navigation to tools, manuscripts and summaries · Retrieved · Supports: The project site describes BootLoops as a model-independent harness of quantitative-science tools and protocols. It says the core acceptance standard requires a full functional form and an arbitrary-precision Python evaluator, and it exposes project descriptions and links to the public code.
- BootLoops 1.0 source at release commit 66b680cBootLoops-ai on GitHub · Reviewed: Pinned October 1, 2026 release commit; repository tree; README purpose and supported computation classes; installation surface; self-test entry point; licensing; notice; contribution model; verification-protocol description · Retrieved · Supports: The pinned public commit identifies BootLoops 1.0, contains tools, protocols, installation documentation and self-tests, and licenses software under MIT with documentation under CC BY 4.0. Repository availability establishes that an artifact can be inspected; it does not establish that every reported result reproduces.
Limitations
AccessAllGPT reviewed project-controlled materials and did not independently reproduce the scientific work. We did not install BootLoops, execute run_selftests.py, inspect all repository code or third-party components, run Claude Fable 5 or another model, access the Google Cloud or Overleaf environments, measure tokens or cost, examine all candidate problems, read all 36 manuscripts, verify the stated field and coauthor counts, audit novelty, validate the ecology or genetics analyses, reproduce an integral, inspect peer-review records or interview Schwartz and collaborators. The Anthropic page is a guest post by a visiting researcher, and the BootLoops website and repository are maintained by the project. Repository availability, a license and documented checks do not establish scientific correctness, safe execution, easy installation or model portability. Project content, manuscripts and code may change after the pinned release.
Disclosures
AccessAllGPT did not receive Claude access, BootLoops support, cloud credits, code, data, manuscripts, a briefing, review or compensation for this article. Anthropic, Matthew D. Schwartz, BootLoops and collaborators did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Anthropic, BootLoops, Harvard or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Claude Agents Found a New Enzyme System—Its Function Is Still Unknown
- Claude Computes a Nine-Loop Physics Result—With Validation and a $1,000–$2,000 Cost Estimate
- Ai2 Releases AstaBrief 8B, an Open-Weight Model for Cited Scientific Reports
- From Paper to Production Without Losing the Plot
- Design an Agent Benchmark That Predicts Production
- Where Human Approval Belongs in AI Automation
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.