Key takeaways
- Anthropic published the result on September 25, 2026: Fable 5.1 in Claude Science computed the six-particle amplitude in planar N=4 super-Yang-Mills at nine loops, a target publicly specified on August 7.
- The team reports two routes—a form-factor calculation mapped by antipodal duality and a direct hexagon-function bootstrap—and says their compared coefficients agree. Lance Dixon independently validated the result, according to his signed addendum.
- The reported user cost is about $1,000–$2,000 for either route. The direct bootstrap also used 96 CPUs for a week, estimated at roughly $100. These are author estimates, not AccessAllGPT billing measurements.
- The result extends known methods on a deliberately simplified theory. The guest author says humans concurrently obtained most of the same result with GPT-6 assistance; this does not show a new physical principle or a shortcut around computation.
- Use the release as evidence that long-horizon research agents can execute fragile, checkable workflows. Do not generalize it to fields without machine-verifiable outputs, and do not approve scientific conclusions without independent domain validation.
Anthropic published the nine-loop result on September 25
Anthropic reported on September 25, 2026 that Fable 5.1, operating inside the paid Claude Science harness, computed the six-particle amplitude in planar N=4 super-Yang-Mills at nine loops. The initial instruction quoted in the post was simply: “The problem is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops.” Follow-ups largely told the agent to continue working and provide periodic updates.
This is a research result, not a new model or API launch. Anthropic does not announce a new endpoint, model identifier, token rate or general Claude Science price in the reviewed post. What is available now is the Claude Science product, the result page, machine-readable amplitude files, checksums and validation records. AccessAllGPT did not authenticate to Claude Science or reproduce the calculation.
The target was named publicly before Anthropic attempted it
On August 7, physicist and science writer Matt von Hippel challenged AI companies to compute either seven-loop N=8 supergravity or the six-particle N=4 super-Yang-Mills amplitude at nine loops using resources an academic could plausibly access. That pre-commitment matters: the successful target was not selected after seeing a favorable run.
The theory is nevertheless a toy model rather than a direct prediction for the Large Hadron Collider. Its high symmetry makes calculations easier than realistic particle physics. A nine-loop result is useful for testing amplitude methods and studying structure, but it should not be described as a nine-loop calculation of the Standard Model or as an experimental discovery.
Claude reached the answer by two established routes
The public account describes one route through the nine-loop form factor, mapped to the amplitude on a restricted surface with antipodal duality and then lifted. A second route directly bootstrapped the amplitude in the space of hexagon functions. The Cosmic9 page says the two representations agree on every coefficient compared.
That redundancy is stronger than a single fluent derivation, but it is not the same as two independent teams. Both routes were produced in the same Anthropic effort and rely on an established research program. The result page says assumptions and untested items are documented, while the programs used for the computation are not distributed.
The reported cost is meaningful, but not a universal research-agent price
Von Hippel estimates that either calculation route would cost an end user roughly $1,000–$2,000, mostly because Claude ran for a long time. He separately reports that the direct Python and SymPy bootstrap used 96 CPUs for one week, costing about $100. AccessAllGPT did not inspect invoices, token traces, retry counts or labor allocation, so these figures remain author-reported estimates.
The useful denominator is not tokens or wall time alone. A reproduction budget must include model usage, CPUs, storage, failed runs, expert setup, independent checking, artifact preparation and the value of specialist review. A validated result on this unusually checkable problem does not establish cost per accepted result in biology, materials science or less formal theoretical work.
A domain expert checked the output
Lance Dixon, a Stanford and SLAC physicist who helped develop the earlier loop calculations, says in a signed addendum that he validated Claude’s result, largely by converting it back through the related form factor. He describes the workflow as fragile and says Claude used the methods his collaborators developed. The repository adds machine-readable coefficients, multiple modular calculations, checksums and validation records.
The conflict disclosures remain important. Anthropic invited and compensated von Hippel for the guest post, Anthropic staff commented on drafts, and Dixon received Claude usage credits. Dixon’s technical validation is still substantive evidence, but the launch article is not independent journalism or anonymous peer review.
Humans were closer than the challenge assumed
The post says Song He and collaborators had independently obtained most of the result, using GPT-6 assistance for some constraints but not the overall framework. The guest author concludes that Claude used known methods with somewhat more compute and strong software engineering; it did not reveal a surprising way around the computational barrier he had hoped to probe.
That correction makes the result more useful, not less. It narrows the claim from “AI broke an assumed computational limit” to “a research agent executed a long, delicate, verifiable calculation with little ongoing scientific supervision.” The latter is a concrete capability claim with operational implications. It is not evidence of superintelligence or autonomous theory formation.
Public artifacts improve auditability, but reproduction is incomplete
Cosmic9 publishes representations of the symbol and function, sample coefficients, conventions, a SHA-256 manifest and records of checks. Large files are linked through a Zenodo archive. That lets specialists inspect outputs and compare alternative representations without relying only on screenshots or prose.
The computation code is not distributed, and AccessAllGPT did not download or verify the large archives. The reviewed material therefore supports artifact availability and reported cross-checks, not a clean-room reproduction. A formal paper from the human research teams was still future work in the September 25 account.
The transferable lesson is to design for falsification
This problem offered an unusually strong acceptance contract: prior loop results, a specialized function space, mathematical constraints, two routes to the answer and an expert able to validate a candidate. Long agent runs become safer when they produce deterministic intermediate artifacts and can fail against independent invariants instead of persuading a reviewer with polished prose.
For another scientific workflow, define the target before the run, hold out validation checks, preserve every program and environment, record model and harness versions, and require a specialist who did not steer the run to inspect the result. If the output cannot be falsified cheaply, autonomy should fall and review should rise.
What research teams should do next
Pilot one bounded problem whose answer has machine-checkable structure and whose failure cannot trigger a real-world side effect. Compare a research agent with the best human-plus-software baseline on accepted outputs, elapsed expert time, compute, model cost, reproducibility and severe errors. Pre-register what counts as success and publish negative runs rather than selecting only the striking completion.
Keep model-generated code, logs, checkpoints, dependency locks and result hashes. Require an independent rerun or a genuinely different validation path before a scientific claim leaves the team. A second prompt to the same model is not independent replication.
The decision: trial checkable research agents, not autonomous science claims
Trial Claude Science or another research harness when the task is bounded, artifacts are inspectable, intermediate constraints are strong and a qualified expert owns validation. Budget the full verification loop, not only token spend. Use this result as a demanding acceptance-test pattern: pre-specify the target, generate durable artifacts, verify by a second method and disclose human and vendor involvement.
Wait when the model or harness version cannot be frozen, the output has no independent checker, provenance is incomplete or the only reviewer shaped the run. Reject unsupervised publication, laboratory action or policy based on a persuasive final narrative alone. The nine-loop result shows that research agents can carry a brittle computation farther than many experts expected; it also shows why expert validation and public artifacts remain part of the product.
Copy-ready research-agent result review
Complete this record before treating an AI-assisted scientific output as an accepted result. One record covers one target, model-and-harness configuration and validation plan.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Exact question, prior state of the art, dated pre-registration, acceptance criteria, prohibited post-hoc target changes and accountable scientist.
Model and version, harness, prompts, tools, permissions, context, dependencies, code revision, environment and provider settings.
Model tokens and charges, CPUs/GPUs, wall time, storage, failed runs, specialist setup, review time and total cost per accepted result.
Programs, logs, checkpoints, outputs, manifests, cryptographic hashes, licenses, durable archive and retention owner.
Known invariants, baselines, dimensional or physical checks, held-out checks, numerical tolerances and automatic failure conditions.
Validator identity and conflicts, method distinctness, blind or held-out elements, result, unresolved discrepancies and sign-off.
Clean environment, independent implementation or team, rerun result, deterministic and stochastic differences, and archived evidence.
What was computed; what was not; toy-versus-real system; known-versus-new method; vendor claim, artifact fact and local observation labels.
Who may interpret, publish or act; required peer review; prohibited autonomous actions; disclosure text and correction owner.
Trial, use with review, wait or reject; approved scope; spend ceiling; stop conditions; expiry and triggers for re-validation.
Primary sources
Browse the publication-wide evidence index →
- Yes, Claude can do Nine LoopsAnthropic (guest post by Matt von Hippel, with addendum by Lance Dixon) · Reviewed: September 25, 2026 publication date; challenge; nine-loop computation; prompts; methods; compute and user-cost estimates; concurrent human result; interpretation; Lance Dixon validation addendum; downloadable materials; disclosure · Retrieved · Supports: Anthropic published the account on September 25. It says Fable 5.1 in Claude Science computed a six-particle planar N=4 super-Yang-Mills amplitude at nine loops by two routes, with limited human direction. The post estimates either route would cost an end user $1,000–$2,000 and says the direct bootstrap used about $100 of 96-CPU compute for one week. Independent physicist Lance Dixon describes how he validated the result. The post discloses that Anthropic compensated the guest author and gave Dixon Claude usage credits.
- Cosmic9: Cosmically Normalized Six-Point Amplitudes at Nine LoopsSiddharth Mishra-Sharma · Reviewed: Status; method summary; two representations; rational and modular certification; validation notes; conventions; manifests and checksums; downloadable amplitude, symbol-space and function artifacts; Zenodo archive reference · Retrieved · Supports: The result page publishes machine-readable files for the nine-loop amplitude, explains the form-factor and direct-bootstrap routes, says the two representations agree on every compared coefficient, identifies assumptions and untested items, and provides a SHA-256 manifest plus validation records. It also states that the computation programs are not distributed and hosts files over 100 MB separately on Zenodo.
- It Only Counts When AI Gets to My Field4 gravitons (Matt von Hippel) · Reviewed: August 7, 2026 challenge; definition of a meaningful result; named N=4 super-Yang-Mills nine-loop target; computational-difficulty rationale; expected academic-compute constraint; author biography and conflict context · Retrieved · Supports: Before Anthropic attempted the work, von Hippel publicly named the six-particle N=4 super-Yang-Mills amplitude at nine loops as one result that would count as meaningful AI progress in his former field. The post framed the target as computationally difficult but solvable in principle with known methods, and required resources available to an academic rather than extraordinary compute.
Limitations
AccessAllGPT did not use Claude Science, run Fable 5.1, inspect private prompts or traces, execute the bootstrap, download the large Zenodo artifacts, verify SHA-256 manifests, reconstruct a coefficient, compare the two representations, validate the physics, inspect billing or interview the participants. The central narrative is hosted by Anthropic and written at Anthropic’s invitation; the author was paid and the validating physicist received Claude usage credits. The public result page is maintained by an Anthropic researcher and does not distribute the computation programs. The concurrent human result and its degree of completion were not independently audited here. Reported $1,000–$2,000 user cost and $100 CPU cost are estimates, not measured by AccessAllGPT. This toy-theory calculation does not establish capability on real-world amplitudes, experimental physics, novel theory creation or scientific tasks without formal checks. Public artifacts, products, prices and documentation can change.
Disclosures
AccessAllGPT did not receive Claude access, credits, traces, result files, a briefing, review or compensation for this article. Anthropic, the authors, validators and artifact maintainers did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Anthropic, OpenAI, Stanford, SLAC or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Claude Fable 5.1: The Cache Math Changes the Agent Migration
- WeatherNext 3 Is Operational—but Managed Inference Still Specifies WeatherNext 2
- From Paper to Production Without Losing the Plot
- Design an Agent Benchmark That Predicts Production
- Choose a Model Without Chasing the Leaderboard
- Where Human Approval Belongs in AI Automation
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.