AWS Adds a Coding-Agent Skill to Benchmark and Tune SageMaker AI Deployments
The aws-ai-ml skill generates reviewable Python code for performance tests and instance recommendations. Live-endpoint tests require confirmation, and deployment itself remains outside the agent’s scope.
For SageMaker AI users, AWS’s new aws-ai-ml skill offers a code-generating workflow for testing deployed models and evaluating configurations before deployment. Its outputs are Python code users can inspect and run; it can compare benchmark jobs but cannot deploy models itself. Launched October 5, 2026, the skill uses real endpoint traffic for live tests, so users must approve those tests and run the generated code with appropriately permissioned AWS credentials.
01
In AWS’s example, Qwen3-8B delivered 44.1% higher output-token throughput than Qwen3-1.7B, but used four GPUs instead of one; the smaller model returned its first token sooner.
02
Undeployed-model evaluations can use models from S3, SageMaker JumpStart, or Hugging Face, with candidate configurations ranked by performance.
03
The skill supports MCP-compatible agents including Kiro, Claude Code, and Codex; local setup requires AWS CLI 2.35 or later and uv.
AWS introduced aws-ai-ml on October 5, 2026, giving coding agents a new way to help engineers evaluate generative AI deployments on Amazon SageMaker AI. Available through the Agent Toolkit for AWS, the skill generates Python code to benchmark running models, recommend deployment configurations and compare performance tests. AWS says users can inspect, modify and run that code in their own environment.
The skill works with coding agents that support Model Context Protocol, or MCP, including Kiro, Claude Code and Codex. Engineers describe a goal in ordinary language rather than select a specific optimization workflow. AWS says the agent asks for missing details, such as an endpoint name or a model’s storage location, instead of guessing.
For an existing endpoint—the running service that handles model requests—the agent generates a Python notebook for a load test. Results cover throughput, or work handled per second; response delays, including time to the first generated token; and the number of simultaneous requests supported. AWS says these measurements come from real traffic on real infrastructure, not estimates.
For models not yet deployed, the skill generates code to evaluate candidate instances and configurations. It accepts models stored in Amazon S3, listed in SageMaker JumpStart or hosted on Hugging Face. The resulting options are ranked with performance metrics so users can choose against their cost and speed requirements. For gated Hugging Face models, the agent requests license acceptance and an access token.
The comparison workflow takes two benchmark job names and calculates changes in throughput and response delays. But the launch’s own example shows why the configuration still matters: Qwen3-8B delivered 44.1% higher output-token throughput than Qwen3-1.7B, while running on four GPUs rather than one. AWS attributes the throughput advantage largely to additional compute, not simply the model. The smaller model produced its first token sooner.
There is also a boundary between advice and execution. AWS says the agent cannot deploy a model itself; it can generate the deployment configuration instead. The generated code runs under the user’s AWS credentials, which need permission for the relevant SageMaker operations.
Local setup requires AWS CLI 2.35 or later and uv. After configuring the Agent Toolkit for AWS, users can add the skill with the command below. AWS also offers a preconfigured JupyterLab image in SageMaker Studio; that route requires a private space for skill syncing.
Testing also leaves resources to manage. AWS instructs users to delete endpoints created during benchmarking or recommendations, stop unused JupyterLab spaces and remove stored job objects to avoid ongoing charges.
Sources
aws.amazon.comNew agent skill: Amazon SageMaker optimized generative AI inference for your coding agent | Amazon Web Services
Reader comments
Newest comments first. Replies stay oldest first.