AivexaNewsSearch
AI news for builders and product teamsChecked every hour

New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent

Collected Oct 5, 2026

Amazon SageMaker AI optimized generative AI inference introduced a new skill called aws-ai-ml, available through the Agent Toolkit for AWS. According to the announcement, the skill gives coding agents including Kiro, Claude Code, and Codex expertise in inference optimization and benchmarking.

The skill plugs into any coding agent that supports the Model Context Protocol, turning it into what the post describes as a SageMaker AI inference optimization expert. It can benchmark endpoints, recommend deployment configurations, compare performance runs, and generate executable SageMaker Python SDK v3 code. The agent asks clarifying questions and produces code based on real benchmarks and measured performance data, per the post.

Installation is available two ways: locally through the Agent Toolkit for AWS, requiring AWS Command Line Interface 2.35+ and uv, or inside an Amazon SageMaker Studio JupyterLab space using a pre-configured image. The post states setup can take users from zero to a working conversation in 10 minutes. For Kiro and Claude Code, agents can discover and load skills at runtime through the AWS MCP Server without local installation. AWS credentials must have permissions to call SageMaker AI APIs; the generated code runs under the user's credentials, and no additional IAM configuration is needed for the skill itself.

Capabilities described include benchmarking an existing endpoint using Workload.synthetic() and start_benchmark() APIs, producing measured throughput, latency percentiles, time-to-first-token, inter-token latency, and concurrency values. The agent confirms an endpoint is safe to load-test before running a benchmark. It also evaluates models against candidate instances to present ranked deployment options, including models stored in Amazon S3, SageMaker JumpStart catalog models, and Hugging Face Hub models, with gated models requiring license acceptance and a Hugging Face token. Users can also compare two benchmark runs, with the agent computing metric deltas and offering to run a missing benchmark first.

The post shares benchmark results comparing Qwen3-8B on a 4-GPU ml.g5.12xlarge against Qwen3-1.7B on a single-L4 ml.g6.4xlarge at 512/256 tokens and concurrency 4. Qwen3-8B showed 44.1% higher output token throughput, 45.9% higher per-user throughput, 46.7% higher request throughput, 33.0% lower inter-token latency, and 32.0% lower request latency; time to first token was 146.5% higher. The post attributes the deltas largely to approximately 4x the compute, noting Qwen3-1.7B wins only on time-to-first-token. AWS advises deleting endpoints, JupyterLab spaces, and S3 objects created during benchmarking or recommendations to avoid ongoing charges.

Read at AWS Machine Learning Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Amazon SageMaker optimized generative AI inference introduces the aws-ai-ml skill through the Agent Toolkit for AWS, giving coding agents like Kiro, Claude Code, and Codex deep expertise in inference optimization and benchmarking. Describe what you want, and your agent generates executable SageMaker Python SDK v3 code to benchmark, recommend, and compare deployments.