Building the Cloud for AI Agents | AWS CEO Matt Garman
a16z
AWS and agentic work 0:00
Agentic workflows tend to run better on AWS because the company has already built, or is actively building, the pieces they need. Those pieces include compute sandboxes, gateways, and permissions that separate what agents can do from what people can do. The focus is on making AWS a good fit for systems that act on their own.
Startups and scale 0:30
Startups have been central to AWS from the beginning. They bring the newest ideas, they push the platform hardest, and they often become major customers later. He says AWS still keeps capacity for them, while also buying 2 million NVIDIA GPUs over the next couple of years and continuing heavy spending because demand is so large.
Cloud growth path 3:00
AWS began with EC2, and he recalls helping study who the service would matter to before launch. The answer was startups, and that has held up. He says many workloads are still on premises, but cloud migration and AI keep adding demand, so AWS still feels early in its growth.
Adapting for agents 5:01
Startups now arrive larger, better funded, and more ambitious than before. They still care about architecture, security, performance, and identity setup, but they also want tools that work well for agents as well as people. AWS has added services like Agent Core, Bedrock, and an AWS context layer, and it has also tuned basics such as latency and throughput because agents can be blocked by delays that people would ignore.
Faster account setup 9:30
He says the core AWS building blocks are in good shape, but the first-use experience needed work. New customers can now create an account with a Gmail address, skip a credit card, and get running in under 30 seconds with defaults handled behind the scenes. If they later need deeper control, they can add it without migrating anywhere.
Database tradeoffs 12:30
The hardest shift is that agents often want short-lived resources, while AWS was built for durable systems. An agent may create a database, do a task, and delete it, so it does not need the same level of durability as a production Aurora database. That means AWS has to rethink where strong defaults are still useful and where they are more than an agent needs.
Agent-specific building blocks 14:00
The discussion shifts to what agents need from cloud systems. Some needs are familiar, like databases, but they have new trade-offs because an agent may want something short-lived or durable. That leads to new building blocks such as compute sandboxes, gateways, and permissions that are meant for agents, not people. The point is to give an agent short, narrow access for one task, rather than handing it broad service rights.
GPU supply and planning 16:30
The focus then turns to GPUs and the strain they create. Demand is so high that capacity is limited by data centers, power, chips, memory, and even construction labor. The response is to allocate carefully across frontier labs, large enterprises, and startups, while still saying yes to as many requests as possible. The same long view now applies to power and supply planning, which has to be done years ahead instead of quarter by quarter.
Supply limits 28:02
The discussion turns to bottlenecks in hardware supply. Power is only the first constraint. When that eases, memory, TSMC capacity, HBM, networking parts, or even small supply chain gaps like connectors can become the next limit. The point is to track every component, including where in the world it is needed, because capacity in one place does not solve a shortage in another.
Data center duties 30:30
The talk then shifts to how data centers are seen outside the company. There is a need to be clearer about the good they do for local areas, including renewable energy, low water use through free air cooling, community jobs, and lower taxes. The speaker says most operators are careful, but a few bad actors have hurt the industry’s reputation, so the good examples need to be named more openly.
Custom chips evolve 33:30
The last topic is the move into custom chips. It began with pushing network work and then storage work onto offload cards so servers felt closer to bare metal and security was stronger. That path led to Graviton, which is now widely used because it is cheaper and faster, and then to Trainium for AI work. Trainium 3 is sold out well into next year, Bedrock runs mostly on it, and it is now used for both training and inference because its design gives strong performance at lower cost.
Agent workflows rethink 41:31
An agent should not just copy what Bob does step by step. The better move is to step back and ask how the task could be solved differently, with more parallel work and many tries at once. That is where the value comes from, especially for customers who need a blank slate rather than a direct copy of today’s process.
Trust and safety 43:00
The hard problem is turning agent work into fully autonomous workflows you can trust. Enterprises want the right guardrails, permissions, and sandboxing so an agent does not touch the wrong data or delete production systems. They also need evals, labeled data, live measurement, and back testing so they can spot drift and prove the system works in production.
Data and model choices 45:31
Enterprise data is treated as the most valuable asset, so Bedrock is built so data stays in the customer’s VPC and does not go back to the model provider. That is part of why customers use it for open and closed models alike, and why more workloads are moving there. The same logic supports open weights models, where customers may fine-tune on their own data in SageMaker to get a better model at a lower cost.
AI security at speed 49:31
Customers also worry about attacks, model abuse, and whether agents can be launched safely. AWS is building controls for permissions, sandboxing, guardrails, and human-in-the-loop choices. At the same time, its Continuum service uses AI to find vulnerabilities, rank them with knowledge of the customer’s setup, and help security teams work at machine speed instead of human speed.
Agents inside AWS 52:00
AWS uses agents across its own business, from security and software development to HR and finance. Quick has been rolled out to every employee, and teams are using it to do work in hours that once took weeks. The biggest gains show up in software and product development, where frontier teams now manage teams of agents and ship new features much faster.
Small teams and pods 55:00
AWS is trying experiments with pods. A product group that once kept ten people on one capability can now be three or four people, build faster, and then move on to a new problem. The hard part is keeping what you built running while staying flexible inside a large company.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

