Key takeaways

  • Decide for one workload and one configured system; “API” and “self-hosted” are not performance categories.
  • Reject any option that fails a mandatory data, security, licensing, reliability or audit gate before comparing economics.
  • Compare total cost per accepted production outcome under measured demand, including people, idle capacity, retries and review.
  • Use a bounded dual-path trial when utilization, quality or operating effort is still uncertain—and preserve an exit path.
01

The decision: buy the service, operate the stack or test both

Make the decision for a named workload, data class, traffic shape and consequence level. A managed API delegates much of model serving to a provider; self-hosting puts the selected weights, inference runtime and infrastructure inside an environment your team operates. Neither label determines outcome quality, privacy, cost or control by itself.

Record one of four outcomes: use a managed API behind named controls; self-host a pinned stack; run a time-bounded dual-path trial; or reject the LLM use case because neither path clears the gates. This four-way rule is AccessAllGPT guidance, not a conclusion from the cited sources.

02

Freeze two comparable configured systems

Name the exact managed endpoint, model version or snapshot, region, account controls, prompt, retrieval path, tools, output parser, retries and fallback. For self-hosting, also name the model artifact and license version, inference runtime and version, quantization, accelerator type and count, orchestration, autoscaling rule and deployment commit.

Hold the workload, acceptance criteria and consequence window constant, but do not pretend unlike systems have identical operations. Disclose different safety filters, context limits, batching, caching, model versions and fallback behavior. The comparison unit is an accepted production outcome, not a raw request.

03

Use non-compensable gates before weighted preferences

Translate mandatory requirements into pass, fail or unresolved evidence: permitted data classes; retention and training-use terms; residency and transfer; tenant isolation; identity and key handling; model and software license obligations; prohibited outputs; audit records; incident response; recovery objectives; and a supported security-patching path. A low price cannot offset a failed mandatory gate.

NIST describes AI 600-1 as a voluntary, cross-sector companion to AI RMF 1.0 for incorporating trustworthiness into design, development, use and evaluation. It does not select an architecture. The gates and decision order in this article are AccessAllGPT implementation guidance that each organization must adapt.

04

Verify the managed API boundary, account by account

Obtain dated evidence for the exact product, feature, account tier and region. Record whether prompts, outputs and metadata are retained; whether they may be used for training or service improvement; available retention controls; application-state behavior; subprocessors; encryption; private connectivity; identity controls; logs; rate and spend limits; version-change policy; and incident notice.

OpenAI currently states that data sent to its API is not used to train or improve its models unless the customer explicitly opts in. Its documentation also describes default abuse-monitoring logs retained for up to 30 days and feature- or eligibility-dependent controls. These are OpenAI vendor statements retrieved 2026-08-07—not independent verification, not a claim about consumer products, and not evidence for another provider or every OpenAI feature. Verify the configured service and contract rather than copying this example into an approval.

05

Price the self-hosted operating system, not only the GPU

Include accelerators, CPU, memory, storage, networking, images and registries, orchestration, observability, security scanning, backups, redundancy, capacity headroom and non-production environments. Add engineering and on-call work for model packaging, runtime upgrades, scaling, queueing, incident response, evaluation, security patches and license tracking. Separate one-time migration work from recurring operation.

The vLLM stable documentation exposes configuration across model loading, quantization, context length, parallelism, GPU-memory use, CPU offload, scheduling and observability. Kubernetes GPU scheduling requires vendor drivers and device plugins and documents accelerator selection for heterogeneous nodes. Those primary implementation sources establish that serving is a configured system with infrastructure prerequisites; they do not prove that vLLM, Kubernetes or self-hosting is cheaper, faster or appropriate for this workload.

06

Measure demand shape before claiming a cost crossover

Replay or shadow representative arrival patterns: steady load, bursts, long contexts, output-heavy tasks, quiet periods and failure retries. Measure accepted outcomes, input and output units, request concurrency, queue time, first-token and completion latency distributions, accelerator utilization, idle capacity, scaling delay, errors, retries and reviewer effort. Keep quality and safety gates unchanged across paths.

Calculate total cost per accepted outcome for each observed demand window. For a managed path include metered usage, ancillary tools, network charges, support, evaluation and human correction. For self-hosting include allocated infrastructure during idle and peak periods plus amortized platform and on-call labor. Do not publish a universal break-even token count: it changes with model, hardware, utilization, purchasing terms, region, latency target and staffing.

07

Test quality after every serving optimization

Quantization, shorter context, batching, caching, speculative methods and model substitution can alter quality, latency or both. Treat each material optimization as a new configured candidate. Run the same frozen cases, deterministic checks, blinded review where needed, and prohibited-action gates before accepting the apparent saving.

A managed API may change available versions or operational behavior; a self-hosted team may change weights, runtime, kernels or hardware. Pin what can be pinned, retain artifacts and prompts, and define what change forces re-evaluation. “Open-weight” does not mean the model, code and surrounding dependencies share one license or update policy.

08

Compare control with the burden required to exercise it

Self-hosting can give a team direct choices over network placement, runtime, scaling, logging and upgrade timing, subject to the model license and infrastructure. Those choices create value only when the organization can implement, test and operate them. A control that exists in configuration but lacks an owner, alert, runbook and recovery test is not production evidence.

A managed service can reduce serving work but creates a provider and contract dependency. Test rate limiting, regional failure, model retirement or substitution, quota exhaustion, malformed responses and loss of the primary endpoint. For either path, keep a safe degraded workflow that does not depend on the failing model.

09

Make portability a tested property

Inventory provider-specific request formats, tool schemas, safety behavior, fine-tunes, embeddings, caches, batch jobs, evaluation assets and observability integrations. Estimate data export, prompt and parser changes, model requalification, infrastructure provisioning, parallel operation and contract termination. A nominally compatible endpoint does not establish equivalent semantics or outcomes.

Run one exit drill before approval: route a representative subset through the fallback, validate the output contract and controls, and record the time, failures and manual work. Portability is local evidence from that drill, not a feature inferred from an SDK or API shape.

10

Run a bounded dual-path trial

When both paths clear mandatory gates but demand, outcome quality or operating effort remains uncertain, authorize a dual-path trial with a frozen case mix, maximum traffic, spend and infrastructure ceilings, eligible data, named owners, start and expiry dates, incident route and rollback. Do not send prohibited production data merely to make the comparison realistic.

At expiry, choose a path, design the next experiment around one unresolved question, or stop. Do not extend the trial only because infrastructure has been purchased or integration work has begun. Sunk work is not evidence that an option meets the declared decision threshold.

11

Apply the ship, trial or reject rule

Reject a path when any mandatory gate fails. Choose a bounded trial when gates pass but evidence does not meet the declared quality, uncertainty, latency, reliability, cost or operating-effort threshold. Ship only when the configured system passes every gate, meets the predeclared thresholds and has an accountable operating and exit plan.

If both paths qualify, apply weights approved before results were visible: accepted-outcome economics, operational fit, change control, portability and strategic constraints. Publish raw measures beside any composite and test whether small weight changes reverse the result. A fragile ranking supports a close-call record, not certainty.

12

Keep the decision current

Set re-evaluation triggers for a model, runtime, hardware, prompt, retrieval or tool change; provider deprecation; material price or contract change; demand shift; new data class; incident; license change; missed service objective; or security advisory. Preserve the previous evidence and explain why it no longer transfers.

The output is an architecture decision record, not a permanent declaration that APIs or self-hosting win. Date every vendor statement and local measurement. Keep verified facts, vendor claims, team assumptions and AccessAllGPT guidance visibly separate.

13

Copy-ready API vs self-hosting decision record

Complete this AccessAllGPT template for one workload and two named configurations. Entries stay in this browser tab. Replace prompts with dated evidence or mark them unresolved; unresolved mandatory gates cannot pass.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Workload; users; accountable owner; data classes; output contract; consequence window; demand window; decision date.

Provider; product and endpoint; model version; region; account tier and controls; prompt; retrieval; tools; retries; fallback; dated contract and documentation.

Model artifact and license version; runtime and version; quantization; accelerator and node configuration; orchestration; scaling; deployment commit; fallback.

Pass, fail or unresolved evidence for data use and retention, residency, security, licensing, safety, audit, incident response, recovery and patch ownership.

Case and holdout versions; demand trace; acceptance checks; reviewers; prohibited outcomes; repetition and uncertainty rules; stop conditions.

Accepted outcomes; critical failures; correction effort; latency distributions; queueing; errors; retries; recovery tests; evidence links.

Metered services; infrastructure and idle capacity; support; platform and on-call labor; evaluation; review; migration; parallel operation; cost per accepted outcome.

Change and version control; observability; security patching; scaling; capacity; runbooks; on-call owner; tested degraded mode.

Fallback configuration; provider-specific dependencies; export and termination terms; exit-drill date, outcome, elapsed time and manual work.

Managed API, self-host, bounded trial or reject; thresholds met; unresolved questions; rollout ceiling; rollback owner; expiry; re-evaluation triggers.

Primary sources

  1. NIST AI RMF: Generative Artificial Intelligence Profile (NIST AI 600-1)National Institute of Standards and Technology · Reviewed: Publication abstract, scope, citation and report metadata · Retrieved · Supports: NIST describes the profile as a voluntary, cross-sector companion to AI RMF 1.0 for incorporating trustworthiness considerations into design, development, use and evaluation.
  2. Data controls in the OpenAI platformOpenAI Developer Documentation · Reviewed: Data use, abuse-monitoring retention, application state, retention controls and data residency · Retrieved · Supports: OpenAI states that API data is not used to train its models unless the customer opts in, describes default abuse-monitoring retention of up to 30 days, and documents feature- and eligibility-dependent retention controls. These are vendor statements about one managed service, not findings about every API or account.
  3. Engine ArgumentsvLLM documentation · Reviewed: Model, load, parallel, cache, device, scheduler and observability configuration groups · Retrieved · Supports: The serving engine exposes choices including model and tokenizer resolution, data type, quantization, model length, tensor and pipeline parallelism, GPU memory utilization, CPU offload and observability. This establishes configuration surface, not a universal performance or cost result.
  4. Schedule GPUsKubernetes Documentation · Reviewed: Device plugin prerequisites, GPU resource requests and heterogeneous node selection · Retrieved · Supports: Kubernetes documents vendor drivers and device plugins as prerequisites for GPU scheduling, GPU resources specified through limits, and node labels or affinity for selecting accelerator types. It describes a deployment mechanism, not a complete inference operating model.

Limitations

This guide contains no original model, infrastructure, security or cost tests and supplies no universal break-even point. OpenAI documentation is vendor-authored and applies only to the stated API controls; vLLM and Kubernetes documentation describe implementation surfaces, not workload outcomes. Prices, products, licenses, model artifacts, documentation and service terms can change. Local trials cannot prove the absence of rare failures, and labor allocation is organization-specific. Qualified security, privacy, legal, licensing, finance and reliability review remains necessary.

Disclosures

AccessAllGPT did not run, price, score or rank a managed API, model, accelerator, inference runtime or cloud for this article. OpenAI documentation is included as labeled vendor evidence; OpenAI did not review or sponsor this work. vLLM and Kubernetes are cited as implementation documentation, not endorsements. No vendor supplied data, paid for placement or received an endorsement. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with OpenAI. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. Choose a Model Without Chasing the Leaderboard
  2. The AI Tooling Procurement Scorecard
  3. Design an Agent Benchmark That Predicts Production
  4. AccessAllGPT Research methodology
  5. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.