WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=pqlWNihgdjI

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/pqlWNihgdjI.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:12.400 --> 00:00:17.680
My name is Claire La Gory and I'm a senior principal engineer at AWS.

s3
00:00:17.680 --> 00:00:31.520
I mostly work on Kuro, our agent encoding assistant, but today I want to talk about some of the practices we've been seeing inside of Amazon and Amazon teams where we've been seeing really exciting results of

s4
00:00:31.520 --> 00:00:38.560
productivity increases that are step function improvements since what what we've been seeing with AI so far.

s5
00:00:38.560 --> 00:00:50.320
So, I've been working on agentic AI for over 3 years now and I've kind of seen the evolution that's happened in our industry when it comes to coding assistance with AI.

s6
00:00:50.320 --> 00:00:57.280
First, we had this inline code completion helping us to write the next line, maybe the next function.

s7
00:00:57.280 --> 00:01:01.040
We moved on to chat, asking questions about our code.

s8
00:01:01.040 --> 00:01:11.080
Everybody started doing vibe coding sometime last year, but now we're starting to see kind of an early adopter phase of what we've been calling frontier development.

s9
00:01:11.080 --> 00:01:22.520
And completely anecdotally, based on my own experience, I've really only felt maybe 10 to 20% more productive with all of these phases that have come before.

s10
00:01:22.520 --> 00:01:34.640
But now inside of Amazon, we've been running pilots with different teams across the company and we've been seeing a median of 4.5x productivity improvement and sometimes more than 10x.

s11
00:01:34.640 --> 00:01:40.840
So, something has really changed here now that we're seeing these step function improvements in productivity.

s12
00:01:40.840 --> 00:01:50.520
And I like to define what we've been calling frontier developers inside of Amazon by three behaviors that I've been seeing.

s13
00:01:50.520 --> 00:01:52.240
One is hands-off coding.

s14
00:01:52.240 --> 00:01:57.680
Frontier developers write maybe 1 to 2% of the code that they produce.

s15
00:01:57.680 --> 00:02:00.000
The rest is agents.

s16
00:02:00.000 --> 00:02:03.800
The second is that they interact with their agents infrequently.

s17
00:02:03.800 --> 00:02:10.600
They'll aim to get their coding assistant to run for up to hours at a time without their intervention.

s18
00:02:10.600 --> 00:02:13.520
And third is that they minimize idle time.

s19
00:02:13.520 --> 00:02:21.880
These frontier developers tend to run multiple agents in parallel churning through a backlog of tasks.

s20
00:02:22.240 --> 00:02:28.200
The first time that I saw a frontier developer team was the Bedrock Mantle team.

s21
00:02:28.200 --> 00:02:31.560
Bedrock is our model hosting service.

s22
00:02:31.560 --> 00:02:35.760
Hosts LLMs like Claude and GPT.

s23
00:02:35.760 --> 00:02:46.200
And sometime last year we knew or I say we but the Bedrock team knew that they were going to need to build a new inference data plane.

s24
00:02:46.200 --> 00:02:51.120
But they had estimated it at 30 people over 18 months.

s25
00:02:51.120 --> 00:02:59.640
This is a big big service and it was going to take time to build the new one, migrate customers over, migrate models over.

s26
00:02:59.640 --> 00:03:01.320
They decided to take a step back.

s27
00:03:01.320 --> 00:03:03.402
They took six people and they built

s28
00:03:03.402 --> 00:03:03.640
[snorts]

s29
00:03:03.640 --> 00:03:06.160
it in 76 days with Kiro.

s30
00:03:06.160 --> 00:03:08.280
So this was a huge achievement.

s31
00:03:08.280 --> 00:03:12.760
This was the first time we've we'd seen anything of the kind inside of Amazon.

s32
00:03:12.760 --> 00:03:20.239
So this was truly the pathfinder team that proved that it was possible to get up to 20X improvement.

s33
00:03:20.239 --> 00:03:27.840
Now they looked at commits and I'll talk about a couple of other ways that we are uh measuring productivity improvements.

s34
00:03:27.840 --> 00:03:33.440
But there was one problem with this story which was that yes, it was built with six people.

s35
00:03:33.440 --> 00:03:40.680
It was built with some of the top engineers literally in the company including two distinguished engineers.

s36
00:03:40.680 --> 00:03:43.880
So this was not just any team of six people.

s37
00:03:43.880 --> 00:03:51.000
These were experts in distributed systems, experts at LLMs and their architecture.

s38
00:03:51.000 --> 00:03:59.360
So this the story was amazing and it kind of spread like wildfire across Amazon, but it was also very unachievable for a lot of teams.

s39
00:03:59.360 --> 00:04:05.560
There were a lot of questions about can this actually be reproduced on another team?

s40
00:04:05.560 --> 00:04:13.360
So, another experiment that I want to talk about is an experimental sprint that was done in the Prime Video organization.

s41
00:04:13.360 --> 00:04:23.640
They took a 10-day sprint and they did an experiment where they put, again, six engineers in a room and they let them go wild with Kiro.

s42
00:04:23.640 --> 00:04:36.280
Uh they brought down the project delivery time estimate from what was going to be 90 weeks down to 24 based on all of the progress they had made in this 10-day sprint.

s43
00:04:36.280 --> 00:04:48.120
And they they looked at their commit history and they looked at what did they used to do prior to this 10-day sprint and how many commits did they produce just in this 10 days.

s44
00:04:48.120 --> 00:05:00.800
And so, this sprint really proved that we can achieve, again, at least something close to what the Bedrock Mantle team had uh had achieved with a different set of engineers.

s45
00:05:00.800 --> 00:05:13.040
But again, there was a challenge with this story, which was it was six engineers in a room, but they had no on-call duties, limited meetings, very few distractions, which we all know

s46
00:05:13.040 --> 00:05:16.840
are regular in the lives of an engineer.

s47
00:05:16.840 --> 00:05:32.000
And the senior engineer on the team had spent the previous 3 weeks creating very detailed, small, well-scoped tasks with detailed requirements for these six [clears throat] engineers to just go churn on for those 2 weeks.

s48
00:05:32.000 --> 00:05:35.080
So, this was again not necessarily real life.

s49
00:05:35.080 --> 00:05:47.040
This was a structured sprint, uh a a point in time that they were able to achieve this, but again, the question is is this achievable on real teams on day-to-day

s50
00:05:47.040 --> 00:05:48.440
work?

s51
00:05:48.440 --> 00:05:59.160
So, Amazon stores which encompasses amazon.com, all of our retail websites, as well as our physical stores, did a more structured pilot.

s52
00:05:59.160 --> 00:06:11.760
They watched 50 teams that were totally normal normal distribution of um early career folks, mid-career, senior engineers, and that worked on existing systems.

s53
00:06:11.760 --> 00:06:19.040
Nothing green field like the mantle team got to build from the ground up, but existing systems with existing code bases.

s54
00:06:19.040 --> 00:06:26.800
And they they watched them for the better part of last year, and they found something super interesting.

s55
00:06:26.800 --> 00:06:34.360
They found that there was a big difference in the productivity gains that they saw between half of the teams and the other half.

s56
00:06:34.360 --> 00:06:40.200
And in this case, they used a productivity metric of deployment velocity to production.

s57
00:06:40.200 --> 00:06:47.360
So, not just commits, how many commits are they producing, but how quickly are we getting changes out to customers?

s58
00:06:47.360 --> 00:06:50.440
How how quickly are we able to ship things?

s59
00:06:50.440 --> 00:06:55.640
And they saw that for half of the teams, they achieved less than 3x increase.

s60
00:06:55.640 --> 00:07:06.560
And what they found that was the difference between seeing less than 3x productivity increase, these teams that saw a median of 4.5x, and and in some cases more than 10,

s61
00:07:06.560 --> 00:07:08.919
was how they used the tools.

s62
00:07:08.919 --> 00:07:19.760
90% of these teams used Kiro, among other internal tools that we have, and what they found was it wasn't about the tools, it was about the way that they worked.

s63
00:07:19.760 --> 00:07:31.080
The teams that achieved step function improvements intentionally changed the way that they worked, and the other simply kind of sprinkled Kiro and some of the other tools that we have

s64
00:07:31.080 --> 00:07:33.760
on top of their existing way of working.

s65
00:07:33.760 --> 00:07:37.000
And for me at least, this was the big aha moment.

s66
00:07:37.000 --> 00:07:46.000
That why I hadn't been feeling potentially the massive gains that productive that in in productivity that AI has promised.

s67
00:07:46.000 --> 00:07:49.480
It's about changing the way that we work.

s68
00:07:49.480 --> 00:07:57.760
So, across this pilot, they went and interviewed uh the teams that were involved in the pilot as well as some of these other teams on the Bedrock mantel team,

s69
00:07:57.760 --> 00:08:02.400
on uh Prime Video, and they found five habits.

s70
00:08:02.400 --> 00:08:08.240
And and I use the word habits very specifically because again, it's not about that one sprint.

s71
00:08:08.240 --> 00:08:10.360
It's about doing this day-to-day.

s72
00:08:10.360 --> 00:08:17.560
And it And what they found when they interviewed with these teams was that it really was habits that they had to build day-to-day.

s73
00:08:17.560 --> 00:08:21.640
When we change our way of working, it's it's hard to build these habits.

s74
00:08:21.640 --> 00:08:24.040
It takes time to build these habits.

s75
00:08:24.040 --> 00:08:26.960
So, let's go through each of these one by one.

s76
00:08:26.960 --> 00:08:30.040
Habit number one is investing in agent context.

s77
00:08:30.040 --> 00:08:32.719
We have a lot of stuff in our head.

s78
00:08:32.719 --> 00:08:45.640
We tend to transfer all of that stuff in our head to other people through Slack conversations, through onboarding, mentors, things like that, through code reviews, through stand-ups and sprint planning,

s79
00:08:45.640 --> 00:08:47.880
and they had to write all of that down.

s80
00:08:47.880 --> 00:08:57.600
And the habit that they built was every time the agent makes a mistake or does something not the way that you would have done it, what am I missing in my skills files?

s81
00:08:57.600 --> 00:09:02.000
What am I missing in my steering files that the agent needed?

s82
00:09:02.000 --> 00:09:09.640
But then, as we know, across last year, we saw leaps and bounds in models' abilities and their behaviors.

s83
00:09:09.640 --> 00:09:18.680
Uh the Sonnet 3.7 in the middle of last year had a lot of quirks that we had to put a lot of do nots in our uh in our steering files,

s84
00:09:18.680 --> 00:09:28.360
and now we don't have to do that as much with Opus 4.5 as of last November, and then we've had 6 months more than 6 months of improvement since then

s85
00:09:28.360 --> 00:09:31.640
uh with all of the new versions of models that have come out since then.

s86
00:09:31.640 --> 00:09:39.240
And so, the question, the new habit, again, is do I still need this in my steering files or is this just bloating context?

s87
00:09:39.240 --> 00:09:42.520
The second one is slowing down to speed up.

s88
00:09:42.520 --> 00:09:52.680
In almost every team that was interviewed, they reported that their productivity actually went down as they intentionally adopted a new way of working.

s89
00:09:52.680 --> 00:09:54.600
That's counterintuitive, right?

s90
00:09:54.600 --> 00:10:02.520
You have to do intentional engineering work before you're going to see that hockey stick curve in productivity improvement.

s91
00:10:02.520 --> 00:10:10.760
Because we have to do real work in our code base first for agents to be successful there, especially in brownfield existing code bases.

s92
00:10:10.760 --> 00:10:13.040
So they had to build that agent context up.

s93
00:10:13.040 --> 00:10:19.360
They had to improve existing tools error messages so that the model knew what was going on when it failed.

s94
00:10:19.360 --> 00:10:26.560
They built new tools, new MCP servers for helping that model to actually get done what it needed to get done.

s95
00:10:26.560 --> 00:10:32.600
A lot of teams ended up restructuring their code base so that agents could actually navigate it more easily.

s96
00:10:32.600 --> 00:10:38.200
And I've even seen drastic changes like changing the programming language of the code base.

s97
00:10:38.200 --> 00:10:44.720
Um often I've seen teams struggle with Python, with JavaScript because they're untyped languages.

s98
00:10:44.720 --> 00:10:46.040
It's hard to test.

s99
00:10:46.040 --> 00:10:47.960
There's no compiler errors.

s100
00:10:47.960 --> 00:10:51.839
So the model kind of guesses and give it gives it back to you.

s101
00:10:51.839 --> 00:10:55.080
And so I've seen teams moving to TypeScript.

s102
00:10:55.080 --> 00:10:57.839
Um Rust has become very popular inside of Amazon.

s103
00:10:57.839 --> 00:11:00.880
The compiler gives great error messages.

s104
00:11:00.880 --> 00:11:09.480
Um you don't have to do that, but I've seen a lot of teams making those intentional changes for the productivity gains that they're able to see.

s105
00:11:09.480 --> 00:11:13.800
The third one is feeding agents, not babysitting agents.

s106
00:11:13.800 --> 00:11:21.160
And for me this was one of those aha moments of why we're seeing this step function improvement in productivity.

s107
00:11:21.160 --> 00:11:32.480
If you are vibe coding, if you are having a back-and-forth conversation with your agent all day long, of course you're not going to see four to five x productivity improvements

s108
00:11:32.480 --> 00:11:35.560
because you are in the loop the entire time.

s109
00:11:35.560 --> 00:11:44.600
You're probably sitting there for 30 seconds to a minute waiting for it to generate code and come back to you with with the code to review.

s110
00:11:44.600 --> 00:11:48.600
If you're sitting there waiting for it, then you can't go off and do other stuff.

s111
00:11:48.600 --> 00:11:51.640
It's really difficult to run agents in parallel.

s112
00:11:51.640 --> 00:11:57.320
It's very difficult to get to to clone yourself into multiple agents.

s113
00:11:57.320 --> 00:12:03.560
And so if your conversations look a bit like this on the left, then you're babysitting that agent.

s114
00:12:03.560 --> 00:12:09.800
As opposed to the right side where you're feeding it what it needs to do and how it can self-validate.

s115
00:12:09.800 --> 00:12:18.880
And that's really the key so that agents can self-correct and only come back to you when it meets a certain quality bar, when it when it actually runs and compiles

s116
00:12:18.880 --> 00:12:24.240
and passes tests, when it's testable, when it it actually has high coverage.

s117
00:12:24.240 --> 00:12:32.880
And of course the next level is put all of this content into your steering file so it does it every time without you having to prompt it.

s118
00:12:33.560 --> 00:12:37.880
The fourth habit is to make intent explicit.

s119
00:12:37.880 --> 00:12:41.040
At Amazon we practice a lot of behavior-driven development.

s120
00:12:41.040 --> 00:12:53.800
We've built that into the Q product and so it's very natural for Amazon engineers to adopt it in Q. Um what what I've typically seen with live coding as opposed to frontier engineering

s121
00:12:53.800 --> 00:13:06.640
is giving a very high-level prompt, letting the agent generate a ton of code, and then having a back-and-forth conversation saying, "Oh, that's not really what I meant.

s122
00:13:06.640 --> 00:13:10.440
That you haven't you haven't exactly gotten the the requirements right.

s123
00:13:10.440 --> 00:13:12.320
No, I didn't actually want to build it that way.

s124
00:13:12.320 --> 00:13:14.320
Here's a technical design."

s125
00:13:14.320 --> 00:13:23.640
And it is less I find less productive to iterate with the agent on code when the intent itself was incorrect.

s126
00:13:23.640 --> 00:13:34.440
So often will have will see Amazon engineers go through this process for for ambiguous complex features of writing the specification.

s127
00:13:34.440 --> 00:13:37.640
And in Kiro, of course, you don't have to write this whole specification.

s128
00:13:37.640 --> 00:13:53.440
You can have the model generate it, but it's a lot easier to to iterate with the model in kind of a back and forth conversation about a document than it is about code that's code changes that are spread across a code base.

s129
00:13:53.520 --> 00:13:57.840
The fifth one is shift testing left.

s130
00:13:57.840 --> 00:14:02.400
One of the keys here is to give the agent that fast feedback loop.

s131
00:14:02.400 --> 00:14:06.680
Because that's what lets it go off for hours at a time and self-correct.

s132
00:14:06.680 --> 00:14:10.080
The agent is going to make mistakes and that's fine.

s133
00:14:10.080 --> 00:14:16.520
But if you give it the right signals, it can self-correct and it can spend a while doing that.

s134
00:14:16.520 --> 00:14:22.800
So, I've seen teams adding linters, adding unit tests, integration tests, performance tests, security tests.

s135
00:14:22.800 --> 00:14:26.160
These are all things we all know we should have been doing all along.

s136
00:14:26.160 --> 00:14:29.680
This is good engineering hygiene and practices.

s137
00:14:29.680 --> 00:14:36.360
But now the ROI is, I think, finally high enough for actually us to actually invest in it.

s138
00:14:36.360 --> 00:14:41.120
Um one thing that I've been seeing a lot of teams do is mock out services.

s139
00:14:41.120 --> 00:14:48.000
Often with integration tests, we would test kind of end-to-end an entire system including live services.

s140
00:14:48.000 --> 00:14:58.400
But we've been investing a lot in in mock services that run entirely locally with deterministic responses because it lets the agent do everything locally.

s141
00:14:58.400 --> 00:15:12.800
Um doing everything on your laptop without having to spin up a bunch of other services and and connect to cloud services makes everything a lot faster because the the more that your agent can get fast feedback

s142
00:15:12.800 --> 00:15:19.600
means the more loops that it can can do and the more productive your own agent can be.

s143
00:15:19.640 --> 00:15:30.240
So, across all of these, these are some of the habits we've seen, but of course I would be remiss if I would tell you if you adopt all of these habits, you will

s144
00:15:30.240 --> 00:15:31.160
achieve nirvana.

s145
00:15:31.160 --> 00:15:35.880
You will be the most productive engineering organization the world has ever seen.

s146
00:15:35.880 --> 00:15:36.960
Things are still hard.

s147
00:15:36.960 --> 00:15:43.840
We are still very much in an early adopter phase and teams are still figuring it out.

s148
00:15:43.840 --> 00:15:49.800
So, one thing that we've been seeing across our teams just organizationally is the risk of burnout.

s149
00:15:49.800 --> 00:15:51.520
I did not coin this term.

s150
00:15:51.520 --> 00:15:55.920
I forget who did at what conference, but flow mat is real.

s151
00:15:55.920 --> 00:16:08.120
We've been seeing engineers staying up late late at night trying to get that perfect prompt that's going to make their agent run for hours overnight so that they wake up in the morning with a code change ready.

s152
00:16:08.120 --> 00:16:13.760
The cognitive load increases as you run these multiple agents in parallel.

s153
00:16:13.760 --> 00:16:17.240
You're constantly shifting between terminal tabs.

s154
00:16:17.240 --> 00:16:26.240
And then we do see that reviewing AI output is often harder for some than than actually writing it, especially early in career.

s155
00:16:26.240 --> 00:16:32.400
Senior engineers have have already spent a large portion of their career reviewing others code.

s156
00:16:32.400 --> 00:16:44.520
But early career engineers don't have that muscle yet and so reviewing it can can feel like a lot more cognitive load than they're used to and actually writing it.

s157
00:16:44.520 --> 00:16:47.839
The other one is organizational change.

s158
00:16:47.839 --> 00:16:51.240
So, it's already hard to change the way we work as engineers.

s159
00:16:51.240 --> 00:17:02.960
The way that we spend our entire day completely changes when we're frontier engineers, but also organizations have to change to enable frontier engineering teams.

s160
00:17:02.960 --> 00:17:09.120
One that I've seen very commonly is accepting slowing down to speed up.

s161
00:17:09.120 --> 00:17:11.040
And I've been guilty of this myself.

s162
00:17:11.040 --> 00:17:19.199
My my fellow leaders have been guilty of of this of saying, "Well, you have the AI tools now and the models are so amazing now.

s163
00:17:19.199 --> 00:17:22.120
Why are you not going faster?

s164
00:17:22.120 --> 00:17:35.120
Um and that's because you have to take those two months to invest in your code base, to figure out the best practices for your team, to make hard habit changes on your team.

s165
00:17:35.120 --> 00:17:48.880
Um and and if you're constantly expecting shipping features every month because now we have these amazing models and we're seeing um all of these these companies on X saying how they're shipping 20

s166
00:17:48.880 --> 00:17:54.360
PRs a day, um we have to slow down to speed up.

s167
00:17:54.360 --> 00:17:58.120
The second one is actually going too broad in the organization too fast.

s168
00:17:58.120 --> 00:18:11.920
I think that if we had um expected all teams in massive organizations to be frontier teams immediately, we would not have had the learnings that we had from the Pathfinder,

s169
00:18:11.920 --> 00:18:18.240
from the from the sprint experiment, from the pilot uh teams within Amazon.

s170
00:18:18.240 --> 00:18:21.000
And now the challenge for us is how do we scale it out?

s171
00:18:21.000 --> 00:18:31.040
And that's what 2026 is about for Amazon is how do we scale this out to more and more teams, to the next uh 2,000 teams instead of uh 50 teams.

s172
00:18:31.040 --> 00:18:38.000
Um and so I think that when you roll it out too quickly, you have a lot of teams who don't know what they're doing.

s173
00:18:38.000 --> 00:18:45.920
You haven't had time to find the best practices for your own organizations, the the context that your organization needs.

s174
00:18:45.920 --> 00:18:49.360
And the last one is that you're going to find new bottlenecks.

s175
00:18:49.360 --> 00:18:54.240
Previously, code writing code manually was the bottleneck.

s176
00:18:54.240 --> 00:19:01.600
Um I find that within Amazon, we've found um the speed of decision-making becomes a new bottleneck.

s177
00:19:01.600 --> 00:19:13.468
Um the more that you spend reviewing the decision to actually build a new product, the slower it is to build the product now because the code only takes 1 to two months to write.

s178
00:19:13.468 --> 00:19:13.480
[snorts]

s179
00:19:13.480 --> 00:19:20.840
Um all of the review processes associated with the launch of a product become the bottleneck.

s180
00:19:20.840 --> 00:19:34.680
When it used to take 9 to 12 months to build a new product, it didn't matter so much in the in the overall wash of things if it took two months to make the decision to build the product and then two months to approve the launch.

s181
00:19:34.680 --> 00:19:37.040
But now those are the bottlenecks.

s182
00:19:37.040 --> 00:19:38.760
Those are the long pole.

s183
00:19:38.760 --> 00:19:44.040
And so you find all of these all of these things that slow you down.

s184
00:19:44.040 --> 00:19:50.960
Often I find that frontier engineering teams spend more time making decisions than they do writing code.

s185
00:19:50.960 --> 00:19:57.680
And so the more that you can make fast decisions, especially ones that are easy to be reversed, the better.

s186
00:19:57.880 --> 00:20:07.080
So my one big takeaway for for everyone here is that frontier engineering is about intentionally changing the way that you work.

s187
00:20:07.080 --> 00:20:08.880
And that is difficult.

s188
00:20:08.880 --> 00:20:10.080
That takes time.

s189
00:20:10.080 --> 00:20:14.400
It is forming new habits and a new way of working.

s190
00:20:14.400 --> 00:20:20.160
And that goes across any engineering team as well as your organization.

s191
00:20:20.160 --> 00:20:31.360
Um so I encourage you to think about um how you're interacting with AI tools and how that can change to free yourself up from being in the loop.

s192
00:20:31.360 --> 00:20:31.960
Um thanks.

s193
00:20:31.960 --> 00:20:36.720
I'm going to I'll hang out uh a little bit if anyone has questions in the back.

s194
00:20:36.720 --> 00:20:39.960
Um but thanks for the time today.

s195
00:20:54.994 --> 00:20:56.994
[music]
