Key takeaways

  • NVIDIA published its cross-competition Nemotron account on October 7, 2026. The underlying specialist artifacts and papers were released in September, so this is a new synthesis—not a same-day foundation-model launch.
  • NVIDIA reports 535.4/600 for Nemotron-3-Ultra-CC plus GenCorrect on a single prospective IOI 2026 run and 30/42 for a three-checkpoint natural-language proof system at IMO 2026.
  • The IOI run was unofficial, unsupervised and outside the ranking. The IMO proofs were graded by official graders, but the result belongs to a multi-checkpoint, high-compute system rather than one checkpoint at ordinary sampling settings.
  • Public artifacts include coding, math-SFT and math-RL weights, training datasets, a 200-problem math benchmark, papers, prompts, inference code and submitted proofs. “Open weights” does not mean cheap or turnkey: the checkpoints are 550B-total, 55B-active models.
  • Teams should reproduce one bounded workload and account for the complete search-and-verification budget before treating the results as evidence for production coding, theorem proving or agent deployment.
01

NVIDIA connected two Nemotron competition systems on October 7

NVIDIA’s October 7 Hugging Face publication puts its 2026 mathematics and competitive-programming work under one claim: Nemotron 3 can be specialized with supervised fine-tuning, reinforcement learning where useful, and feedback-driven inference. It reports 535.4 out of 600 for an IOI system and 30 out of 42 for an IMO system, both above the cited gold thresholds of 361.12 and 29.

The availability date needs a qualifier. The linked papers and specialist repositories appeared in September; October 7 is the publication date of the unified account, not the first release date of a new base model. NVIDIA names no new managed API SKU or price in this announcement. What readers can access now is a set of public weights, datasets, code and papers under their respective terms.

02

The IOI result came from Ultra-CC plus GenCorrect

The coding checkpoint is nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4: a 550-billion-total, 55-billion-active specialist based on Nemotron 3 Ultra, with a stated context limit of 262,144 tokens. NVIDIA says it fine-tuned the model for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 competitive-programming problems, then quantized it to NVFP4.

The reported 535.4/600 is not a single-response checkpoint score. NVIDIA paired Ultra-CC with GenCorrect, which generates candidate programs, clusters diverse candidates, obtains evaluator feedback and refines later rounds within a submission budget. The paper describes one prospective run under the competition’s time, internet and submission constraints before public problem release. It was unofficial, unsupervised and not included in the IOI ranking; AccessAllGPT did not independently verify the setup or result.

03

The IMO result used three checkpoints and a large search

For mathematics, NVIDIA released Nemotron-3-Labs-Ultra-Math-SFT and Nemotron-3-Labs-Ultra-Math-RL alongside the general Nemotron 3 Ultra checkpoint. The SFT corpus contains 414,890 quality-filtered examples across 15,818 unique proof problems; the RL dataset contains 9,597 problems. NVIDIA says the final system generated, verified and refined natural-language proofs without a formal prover, external tools or internet access.

The released recipe makes the compute shape concrete. Its first round schedules eight prompts across three checkpoints with 16 samples each—384 proof attempts per problem. Every distinct proof can receive 16 verifier judgments; refinement can continue through round eight; and each finalist can receive 48 grading judgments. Official IMO graders awarded the submitted set 30/42, according to the authors. That supports a system result, not a claim that either specialist alone scores 30/42 or that a low-budget deployment will match it.

04

The open release is substantial but not turnkey

The IMO collection exposes two specialist checkpoints, SFT and RL datasets, Nemotron-IMO-Bench with 200 olympiad-level problems, and the paper. NeMo Skills supplies prompts, submitted proofs, configuration and an auditable pipeline with request identities, hashes, manifests, errors and resume behavior. The coding model card supplies weights, a vLLM deployment recipe and its reported training and benchmark details.

All three specialist repositories were public and ungated when reviewed. Their cards use the OpenMDW 1.1 license, so “open weights” should not be collapsed into “Apache licensed” or “unrestricted.” Hugging Face metadata showed roughly 352 GB stored for the NVFP4 coding repository and roughly 1.12 TB for each BF16 mathematics repository. The mathematics card recommends eight B200 GPUs for a single-node deployment; buyers still need infrastructure, serving, energy and inference-search cost estimates.

05

There is no single public price or hosted-access promise

NVIDIA’s release does not quote a managed price for either specialist system. The mathematics model records listed no hosted inference provider in the Hugging Face API at retrieval. The open recipe can target a hosted OpenAI-compatible endpoint or local vLLM servers, but an example endpoint in configuration documentation is not proof that every named specialist is currently offered there or covered by one commercial rate.

Procurement should therefore separate artifact access from service access. For self-hosting, calculate storage, GPU topology, model loading, long-context memory, candidate count, verifier count, retries and final judging. For a hosted service, require the exact model IDs, context and output limits, throughput, retention, region, rate limits and token pricing in current service documentation before estimating cost.

06

The results show specialization leverage, not broad superiority

Across both projects, the useful pattern is co-design. Domain data changed the checkpoint; specialist and general checkpoints supplied different candidates or judgments; and test-time search converted those capabilities into stronger final submissions. This is more informative than attributing the outcome to parameter count or fine-tuning alone.

Transfer remains unproven. Competitive-programming tasks have executable hidden tests; production software work includes changing repositories, requirements, tools, security boundaries and maintenance. Olympiad proofs are difficult but bounded; research mathematics and formal verification have different evidence requirements. Neither result establishes chat quality, agent reliability, factual accuracy, safety, latency or cost effectiveness on a buyer’s workload.

07

What builders should do next

Start by choosing the relevant artifact, not the medal headline. A coding team can compare the released Ultra-CC checkpoint with its existing model on frozen, uncontaminated tasks and then add GenCorrect one round at a time. A math team can first reproduce the published pipeline’s dry run, inspect the prompts and submitted proofs, and evaluate the SFT and RL checkpoints separately before attempting the full ensemble.

Adopt the artifacts for research when the license, hardware and bounded task fit; constrain any pilot to offline workloads with executable or expert review; wait if hosted access, cost or reproducibility is unresolved; and reject using the two medal labels as a general production-model ranking. Record both quality and total attempts, tokens, wall time, hardware, failures and human review so the decision reflects the system that produced the answer.

08

Copy-ready Nemotron specialist reproduction record

Use one record per artifact and workload. Keep checkpoint quality separate from search-system quality.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Exact model ID, revision hash, format, file inventory, license version, dataset revisions and code commit.

IOI-style coding, natural-language proof generation, verification or another bounded task; state what the medal result does not cover.

Runtime, container, GPU type and count, precision, context limit, KV-cache settings, concurrency, endpoint and tokenizer.

Frozen task IDs, contamination review, hidden tests or expert rubric, baselines, repetitions and confidence rule.

Checkpoints, prompts, candidates, rounds, clustering, verifiers, judges, retries, stopping conditions and submission budget.

Per-task outputs, test or grader results, invalid responses, failures, variance, human interventions and retained traces.

Downloads, storage, startup, GPU-hours, tokens, wall time, energy assumptions, hosted charges and engineering labor.

Representative production tasks, tool and repository boundaries, latency target, safety checks and gap from contest conditions.

Research, constrained pilot, wait or reject; approved scope, owner, expiry and evidence required to expand.

Primary sources

  1. One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMONVIDIA on Hugging Face · Reviewed: October 7 publication date; result table; reusable specialization recipe; coding and mathematics post-training; test-time compute; artifact links; competition-status caveats · Retrieved · Supports: NVIDIA’s October 7 synthesis connects the Nemotron 3 specialist systems, reports 535.4/600 on an unofficial prospective IOI run and 30/42 on IMO submissions, and links the released weights, data, papers and inference recipes.
  2. NVIDIA Nemotron Labs 3 Competitive Coding 550B-A55B NVFP4 model cardNVIDIA on Hugging Face · Reviewed: Immutable model revision; model summary; intended use; OpenMDW 1.1 terms; reported benchmarks; architecture; context; training data; SFT recipe; NVFP4 deployment; IOI methodology caveat · Retrieved · Supports: The public ungated coding checkpoint is a 550B-total, 55B-active specialist with a 262,144-token context limit, OpenMDW 1.1 terms, an NVFP4 deployment recipe and no claim that it is a general coding assistant.
  3. Nemotron Labs IMO 2026 collectionNVIDIA on Hugging Face · Reviewed: Collection identity and update date; SFT and RL checkpoint records; base-model record; SFT and RL datasets; Nemotron-IMO-Bench; public and gated status; provider availability metadata · Retrieved · Supports: The public collection exposes two 550B-total mathematics checkpoints, two training datasets, a 200-problem benchmark and the IMO paper; the specialist checkpoint records list no hosted inference provider.
  4. An Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsNVIDIA researchers on arXiv · Reviewed: Version 1 metadata and abstract; checkpoint roles; natural-language inference system; external-tool boundary; 30/42 result; released checkpoints, datasets, code, submitted solutions and benchmark · Retrieved · Supports: The authors report a three-checkpoint generate-verify-refine system that used no formal prover, external tools or internet, scored 30/42, and is accompanied by released specialist checkpoints, data, code and proofs.
  5. Post-Training Language Models for Gold-Medal Performance in Coding CompetitionsNVIDIA researchers on arXiv · Reviewed: Version 2 metadata and abstract; 22,000-problem corpus; Nano and Ultra specialization; GenCorrect; IOI 2025 development results; prospective IOI 2026 protocol, score and unofficial status · Retrieved · Supports: The authors report that Ultra-CC plus GenCorrect scored 535.4/600 in a single prospective IOI 2026 run under contest constraints, while the run remained outside the official ranking.
  6. Nemotron-IMO-TTS: the IMO 2026 ensemble proof pipelineNVIDIA NeMo Skills on GitHub · Reviewed: Immutable recipe revision; pipeline stages and derived request counts; requirements; dry run; configuration; verification and refinement rules; outputs, provenance, resume behavior and submitted proofs · Retrieved · Supports: The released recipe specifies 384 initial proof attempts per problem, 16 verification judgments per distinct proof, up to seven refinement rounds, 48 final-selection judgments per finalist and durable per-request audit outputs.

Limitations

This briefing reviews NVIDIA-authored release material and public registry metadata, not an independent reproduction. AccessAllGPT did not download or hash the weight shards, inspect every model or dataset file, accept or interpret the OpenMDW license, run vLLM, reproduce SFT or RL, execute GenCorrect, run the IMO ensemble, test contamination, confirm that the live IOI environment exactly matched contestant constraints, independently establish that problems were unseen, validate the unsupervised status, grade code or proofs, contact organizers or official graders, verify the 30/42 or 535.4/600 scores, benchmark latency or quality, or calculate a deployable cost. Repository sizes and provider fields are retrieval-time API observations and can change. Parameter, context, throughput, training and benchmark statements are vendor- or author-reported. Official-grader involvement does not make AccessAllGPT an independent verifier, and contest performance does not establish production fitness.

Disclosures

AccessAllGPT received no NVIDIA or Hugging Face access, weights, GPU time, hosted credits, dataset preview, briefing, source code, competition submission, review, payment or compensation for this article. NVIDIA and Hugging Face did not sponsor, review or endorse it. AccessAllGPT did not run, score or rank the systems. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with NVIDIA, Hugging Face, arXiv, IMO or IOI. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. Google Launches EmbeddingGemma 2 for Multimodal Retrieval
  2. OpenAI Releases 722 AI-Generated Math Manuscripts
  3. Claude Helps Report a Nine-Loop Physics Result
  4. Microsoft and Hugging Face Release ThinkingBox
  5. Design an Agent Benchmark That Predicts Production
  6. Build an LLM Evaluation Platform You Can Move
  7. AccessAllGPT Research methodology
  8. Evidence standards
  9. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.