Key takeaways
- AWS announced the aws-ai-ml skill on October 5, 2026 for Kiro, Claude Code, Codex and other compatible coding agents. It is an agent interface to existing SageMaker inference optimization—not a new model or inference API.
- The skill generates SageMaker Python SDK v3 code for live-endpoint load tests, candidate instance recommendations and comparisons between benchmark jobs. AWS says deployment itself is out of scope for the featured optimization workflow.
- Local setup requires AWS CLI 2.35+ and uv. Generated operations run under the user’s AWS credentials, so the effective authority is determined by IAM—not by the skill’s prose guardrails.
- AWS lists no additional Agent Toolkit charge, but endpoints, benchmark and recommendation jobs, Studio spaces and storage can cost money. Pricing was not stated in the launch post.
- Pilot with a restricted role and disposable endpoint. Inspect generated code, require confirmation outside the model, cap duration and spend, preserve raw results, and delete endpoints, Studio spaces and S3 artifacts afterward.
AWS launched the aws-ai-ml skill on October 5
Amazon Web Services announced the aws-ai-ml agent skill on October 5, 2026 through the Agent Toolkit for AWS. AWS names Kiro, Claude Code and Codex, and says the skill can work with other compatible coding agents. The launch adds a conversational code-generation route for SageMaker inference optimization; it does not introduce a new foundation model, model ID or inference endpoint.
The underlying optimized generative AI inference recommendation service predates this release. AWS announced that capability on April 22 and described it as a way to benchmark candidate deployment configurations using NVIDIA AIPerf. The new event is the agent skill that turns natural-language intent into SageMaker Python SDK v3 code for that workflow.
The skill covers three decision tasks
For an existing SageMaker endpoint, the agent can generate a notebook that uses Workload.synthetic() and start_benchmark() to drive a load test. AWS says the report includes requests and output tokens per second, p50 and p99 latency, time to first token, inter-token latency and supported concurrency. Those are measurements of the selected endpoint and workload—not model-quality scores.
For a new deployment choice, the skill can generate code to test candidate instances and configurations for a model in S3, SageMaker JumpStart or the Hugging Face Hub. It can also compare two named benchmark jobs and calculate metric deltas. The launch post says deploying a model is outside this optimization workflow’s scope; generated configuration or recommendations should not be mistaken for a completed production deployment.
Setup requires AWS CLI 2.35+, uv and usable credentials
AWS’s local path starts with aws configure agent-toolkit and requires AWS CLI 2.35 or later plus uv. The launch post then shows npx skills add aws/agent-toolkit-for-aws/skills/aws-ai-ml. AWS also offers a preconfigured private SageMaker Studio JupyterLab space; its first boot is described as taking 5–10 minutes.
The public repository currently stores version 4 of the source at skills/core-skills/aws-ai-ml, while the launch post’s npx argument uses a shorter skills/aws-ai-ml selector. AccessAllGPT did not run the installer, so teams should verify what artifact and revision the command resolves before approving it. Pin the source commit or plugin version where reproducibility matters.
The agent acts with your IAM authority
AWS says credentials need permissions for SageMaker operations such as creating endpoints and running benchmark or recommendation jobs, and that generated code runs under those credentials. The Agent Toolkit product page separately describes IAM controls, CloudWatch visibility and a sandboxed MCP script environment. Those platform features do not make a broad local AWS credential least-privileged by default.
Use a dedicated role with only the required SageMaker, S3, logging and supporting actions; constrain account and Region; and deny unrelated mutation. Keep confirmation in a deterministic wrapper or reviewer workflow rather than relying only on an agent instruction to ask before load-testing. A benchmark intentionally sends traffic to a live endpoint and can affect capacity, latency and cost.
Toolkit access is free; the resources are not
AWS says the Agent Toolkit has no additional charge and customers pay for resources the agent provisions or interacts with. The October launch post does not provide endpoint, benchmark-job, recommendation-job, Studio, storage or data-transfer prices. “No additional charge” therefore does not mean a benchmark run is free.
Budget from the concrete plan: instance types and counts, endpoint uptime, load-test duration, recommendation candidates, Studio runtime, S3 artifacts, logs and retries. Apply spend alerts and resource tags before execution. The launch cleanup list explicitly calls for deleting created endpoints, stopping or deleting JupyterLab spaces and removing benchmark and recommendation objects from the default SageMaker bucket.
AWS’s example is not a model ranking
AWS compares Qwen3-8B on a four-GPU ml.g5.12xlarge with Qwen3-1.7B on a one-GPU ml.g6.4xlarge under a 512-input, 256-output-token workload at concurrency four. AWS reports roughly 44–47% better throughput and end-to-end latency for the larger setup, while the smaller model has better time to first token.
The post itself attributes much of the delta to approximately four times the compute. This is a product walkthrough, not a controlled model comparison: model size, hardware and configuration all change. Reuse the workflow, not the conclusion. Hold task, traffic, output constraints and acceptance criteria constant when comparing candidate systems, and report quality and complete cost beside speed.
What builders should do next
Adopt the skill for a bounded pilot when SageMaker is already the target and a reviewer will inspect every generated operation. Constrain it to recommendation and read-only comparison when endpoint mutation or live load is unnecessary. Wait if the resolved source revision, required IAM policy, regional service availability or expected cost is unknown. Reject autonomous execution with broad production credentials or an unbounded load test.
Use one disposable endpoint and one representative workload. Record the skill revision, agent and plugin version, generated notebook hash, IAM policy, Region, model artifact, container, instance, concurrency, token shape, duration, raw percentiles, throughput, failures and cost. Require a human or external policy gate before create, benchmark and delete actions, then verify cleanup from AWS—not from the agent’s completion message.
Copy-ready SageMaker agent-skill pilot record
Complete this record before authorizing generated code against an AWS account.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Adopt, constrain, wait or reject; workload, accountable owner, reviewer, start date and expiry.
Agent, version, plugin or skill revision, resolved source path, SKILL.md hash and installation evidence.
Account, Region, role, exact allowed and denied actions, resource tags, network boundary and CloudWatch destination.
Endpoint or model artifact, container, instance candidates, autoscaling state, production or disposable status and rollback.
Dataset or synthetic shape, input/output lengths, concurrency, duration, warm-up, repetitions, stop conditions and quality checks.
Notebook or script hash, SDK version, resources created, destructive calls, secrets handling and reviewer approval.
Endpoint, benchmark, recommendation, Studio, S3, logs and transfer assumptions; alert, hard stop and payer approval.
Raw metrics, errors, quality outcomes, complete cost, comparison method, limitations and approved configuration.
Deleted endpoints and jobs, stopped Studio space, removed S3 artifacts, residual logs, completion time and independent verification.
Primary sources
Browse the publication-wide evidence index →
- New agent skill: Amazon SageMaker optimized generative AI inference for your coding agentAmazon Web Services · Reviewed: October 5, 2026 publication date; availability and prerequisites; local and SageMaker Studio setup; endpoint benchmarking; instance recommendations; run comparison; example results; confirmation behavior; scope and cleanup · Retrieved · Supports: AWS announced an aws-ai-ml agent skill that generates SageMaker Python SDK v3 code for endpoint benchmarks, deployment recommendations and benchmark comparisons in Kiro, Claude Code, Codex and other compatible agents.
- Agent Toolkit for AWSAmazon Web Services · Reviewed: Setup command and prerequisites; charge statement; MCP Server, Agent Skills, Agent Plugins and Rule Files; supported agents; runtime discovery; IAM, CloudWatch and sandbox descriptions · Retrieved · Supports: AWS documents AWS CLI 2.35+ as a setup requirement, says the toolkit has no additional charge, and states that customers pay for AWS resources their agent provisions or uses.
- aws-ai-ml skill source (version 4)Amazon Web Services on GitHub · Reviewed: Version 4 frontmatter; scope; routing table; model-deployment route; progressive-disclosure, best-effort and usage-attribution rules; current repository path and September 15 inference-optimization commit history · Retrieved · Supports: The public source routes SageMaker real-time endpoint benchmarking, benchmark comparison and cost, latency, throughput or performance recommendations through its model-deployment references; the skill also covers broader AI/ML lifecycle tasks.
- Amazon SageMaker AI now supports optimized generative AI inference recommendationsAmazon Web Services · Reviewed: April 22, 2026 announcement; inference-recommendation scope; NVIDIA AIPerf basis; candidate configuration search; workload and metric descriptions; resource creation and pricing caveats · Retrieved · Supports: AWS previously launched the underlying optimized inference recommendation service, describing measured candidate configurations and NVIDIA AIPerf-based benchmarking. The October agent skill is a new conversational code-generation layer over that existing capability.
Limitations
AccessAllGPT reviewed public first-party AWS pages and repository metadata but did not install the Agent Toolkit or aws-ai-ml skill; resolve the launch post’s npx selector; inspect every source file or dependency; authenticate Kiro, Claude Code, Codex or another agent; access SageMaker Studio; verify regional availability; derive a minimum IAM policy; create an endpoint; execute generated SDK v3 code; run a benchmark or recommendation; reproduce AWS’s Qwen example; test confirmation behavior, sandboxing, CloudWatch visibility or failure handling; inspect charges; or verify cleanup. The public skill covers more AI/ML lifecycle tasks than the inference features highlighted in this launch. Repository paths, commands, features, permissions, prices and service behavior can change.
Disclosures
AccessAllGPT received no AWS account, credits, software access, briefing, demo, benchmark results, review or compensation for this article. AWS did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Amazon Web Services, Anthropic, OpenAI, Kiro, NVIDIA, Qwen or Hugging Face. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- AWS Adds GLM 5.3 to Bedrock at $1.68 Input and $5.28 Output
- AWS Reports Multi-Turn RL Gains for a Qwen3.6-27B Search Agent
- AWS Publishes a Secure Claude Desktop Route to AgentCore Web Search
- Design an Agent Benchmark That Predicts Production
- Managed LLM API vs Self-Hosting: Make the Production Decision
- Choose a Model Without Chasing the Leaderboard
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.