WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=S6aSoQ6_u5A

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/S6aSoQ6_u5A.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:13.040 --> 00:00:14.240
Hello everyone.

s3
00:00:14.240 --> 00:00:20.640
My name is Sandy and meet my co-host today, Scout.

s4
00:00:20.640 --> 00:00:23.279
This is my friendly rover.

s5
00:00:23.279 --> 00:00:29.039
And one would think that rovers can't really think for themselves, right?

s6
00:00:29.039 --> 00:00:34.880
We have to tell them what to do or we have to very specifically program them on how to think.

s7
00:00:34.880 --> 00:00:41.520
But this little guy here actually has a brain and he can think for himself.

s8
00:00:41.520 --> 00:00:43.120
Let me show you my screen.

s9
00:00:43.120 --> 00:00:44.719
Oh no, it's going to the wrong screen.

s10
00:00:44.719 --> 00:00:50.559
I'm going to see how I can stop this and I'm going to see how I can move to my screen.

s11
00:00:50.559 --> 00:00:53.440
Um, give me just a second.

s12
00:00:53.440 --> 00:00:56.719
I'm going to end show and then we get to this.

s13
00:00:56.719 --> 00:01:01.039
So this is what Scout here is looking at.

s14
00:01:01.039 --> 00:01:04.720
And Scout here has a small little brain.

s15
00:01:04.720 --> 00:01:09.280
And Scout can understand what I'm saying in natural language.

s16
00:01:09.280 --> 00:01:15.200
For example, if you can see my screen here, if I say, "Hey Scout," I'm going to type to him.

s17
00:01:15.200 --> 00:01:25.439
Say, "Hey Scout, turn on your headlights and say hi to everyone.

s18
00:01:27.920 --> 00:01:32.320
he is actually going to be able to understand and do those actions.

s19
00:01:32.320 --> 00:01:34.799
But how is he able to do these things, right?

s20
00:01:34.799 --> 00:01:44.479
It's going to take him a couple of seconds to think about it because how this has been set up is he is working on with a Raspberry Pi installed

s21
00:01:44.479 --> 00:01:56.799
and that Raspberry Pi is actually in a small little location in New York City and he's got a SIM card here which is then connecting to that to that Raspberry Pi and then he's getting signals.

s22
00:01:56.799 --> 00:01:58.880
So, this is via 4G connection.

s23
00:01:58.880 --> 00:02:00.479
And so, he's like, "Oh, hey, on it."

s24
00:02:00.479 --> 00:02:01.680
And if you can see, you know what?

s25
00:02:01.680 --> 00:02:04.560
I'm going to tell him, "Spin around."

s26
00:02:04.560 --> 00:02:06.399
And oh, no, he's actually speaking to me.

s27
00:02:06.399 --> 00:02:09.119
He's like, "I see the stage and it's looking great.

s28
00:02:09.119 --> 00:02:15.440
Big presentation screen to the left, bright stage lights, and rows of seats for the audience.

s29
00:02:15.440 --> 00:02:19.520
How How many people do you see?

s30
00:02:21.440 --> 00:02:28.400
I've never trained him to answer this question, but he's Let's see if he's get getting it.

s31
00:02:28.560 --> 00:02:32.239
He's thinking for like a second or two.

s32
00:02:32.959 --> 00:02:35.599
Let me take a good look from the front camera.

s33
00:02:35.599 --> 00:02:36.720
I see two people.

s34
00:02:36.720 --> 00:02:42.560
One near the speaker monitor set up at the back of the stage and another person further right.

s35
00:02:42.560 --> 00:02:43.760
Nah.

s36
00:02:43.760 --> 00:03:01.440
spin 360 and he's also saying is this session about to start right so he's going to spin 360 any second now and then he's probably going to be like oh wow I did

s37
00:03:01.440 --> 00:03:16.080
so this little robot here is a next generation of robot where ah there we go he is spinning 360 now and he's probably going to tell me what he's seeing

s38
00:03:18.080 --> 00:03:19.599
and he's saying let's spin.

s39
00:03:19.599 --> 00:03:20.159
Right?

s40
00:03:20.159 --> 00:03:30.879
So this new generation of robots is to it's different from our traditional robot training because I have given this guy a little brain.

s41
00:03:30.879 --> 00:03:34.000
And what do I mean by I've given him a brain?

s42
00:03:34.000 --> 00:03:49.040
I've given this robot an agentic layer and I've given it it's called strands agents which is an open-source framework which was built by AWS and I'm going to quickly go back to my slide deck

s43
00:03:49.120 --> 00:04:02.640
we can see it right and so here what happens is we have these existing tools that the robot can do he can take certain actions by himself but only those actions by himself

s44
00:04:02.640 --> 00:04:17.440
so what we can do is we can add a layer of LLM or even better add a layer of agent to it so that the agent orchestrates which tool to call and how to really get the robot to start doing the things we want.

s45
00:04:17.440 --> 00:04:26.160
So in traditional software with traditional AI machine uh like AI engineering we can give agents software tools.

s46
00:04:26.160 --> 00:04:42.400
Similarly, we can give the same AI agent a hardware tool called a robot which has access to preset functions or programmable policies and then the agent can decide which policy to implement when.

s47
00:04:42.400 --> 00:04:54.080
So all it takes is one robot agent for us to be able to do new innumerous tasks and have it understand what we're teaching it in natural language.

s48
00:04:54.080 --> 00:04:56.880
So how do we get started with it?

s49
00:04:56.880 --> 00:05:00.000
All it takes is five lines of code.

s50
00:05:00.000 --> 00:05:04.240
This is through uh the agent harness called strands.

s51
00:05:04.240 --> 00:05:10.720
And all we have to do is import the strands agent and call the robot tool.

s52
00:05:10.720 --> 00:05:19.840
And we say ro tools equals the robot and then we say pick up the red cube and should be able to pick up a red cube assuming that the robot has that capability.

s53
00:05:19.840 --> 00:05:20.160
Yeah.

s54
00:05:20.160 --> 00:05:23.360
Now he's seen someone and he's like oh let me go towards that person.

s55
00:05:23.360 --> 00:05:25.039
So he gets pretty excited.

s56
00:05:25.039 --> 00:05:29.120
This guy is pretty special because he doesn't have just one agent.

s57
00:05:29.120 --> 00:05:30.880
He's got three different agents.

s58
00:05:30.880 --> 00:05:35.360
All three of them are strands and all three of them are working simultaneously.

s59
00:05:35.360 --> 00:05:42.800
One of them is the thinker agent and that's the part of him that's constantly thinking and assessing the environment and like what do I do next?

s60
00:05:42.800 --> 00:05:46.160
And that guy's that part of his brain is constantly thinking.

s61
00:05:46.160 --> 00:05:51.520
Then there's the other communication part of it where and I'm going to show you that in a bit, right?

s62
00:05:51.520 --> 00:05:56.639
and I've connected him to my telegram app as well as to my web app.

s63
00:05:56.639 --> 00:06:03.840
And so he is able to have a conversation with me in natural language and then take actions based on what I am telling him to do.

s64
00:06:03.840 --> 00:06:08.319
Apart from him just perceiving and thinking and figuring out what he wants to do.

s65
00:06:08.319 --> 00:06:13.360
And the third agent, the third type of agent that he's got access to is a voice agent.

s66
00:06:13.360 --> 00:06:24.880
I did have to disable it because every time I speak, he's going to think I'm speaking to him and so he's going to keep chatting away with me and it's just not going to be fun because we're going to have our co-host interrupting me all the time.

s67
00:06:24.880 --> 00:06:28.160
So, I've disabled that feature for the time being.

s68
00:06:28.160 --> 00:06:42.080
But essentially all three of these agents work in tandem with this one robot and thereby this gives him the ability to do way more than what just what he's been trained to do

s69
00:06:42.080 --> 00:06:45.600
more than just the policies that he's learned.

s70
00:06:45.600 --> 00:06:51.680
Now what is a quick overview on this trans package itself?

s71
00:06:51.680 --> 00:06:58.800
This turns package has more than supports more than 40 different robots under eight categories.

s72
00:06:58.800 --> 00:07:02.639
And all of these are just simple robot tool calls.

s73
00:07:02.639 --> 00:07:04.560
And how is this all set up?

s74
00:07:04.560 --> 00:07:05.919
Four different layers.

s75
00:07:05.919 --> 00:07:09.520
The first one is the agent layer, the topmost one.

s76
00:07:09.520 --> 00:07:10.960
And there are two parts to this.

s77
00:07:10.960 --> 00:07:17.360
One is how the actions go in and the second is how it observes and the observations go up.

s78
00:07:17.360 --> 00:07:20.080
So if you notice it's very birectional.

s79
00:07:20.080 --> 00:07:33.199
So first when we give it an instruction we would be talking to this trans agent which is the agentic layer that would then decide which policy to call and the policy provider again

s80
00:07:33.199 --> 00:07:41.520
stands agent supports a bunch of different policy providers and we can then train our policy based on our traditional robot training.

s81
00:07:41.520 --> 00:07:50.960
So in our policies we would collect data and then we would train on it and we would sim create more simulation data and that policy then becomes a VLA

s82
00:07:50.960 --> 00:08:10.720
model which then the robot would have access to strand agents would have access to and then it would invoke that specific policy based on the question that we're asking it or the command that we're giving it and that policy needs to sit somewhere right so that sits in the back end which could be your simulation environment or

s83
00:08:10.720 --> 00:08:14.400
it could be a real hardware chip, your hardware environment.

s84
00:08:14.400 --> 00:08:19.520
That is the back end on which that is the interface on which the policy is running.

s85
00:08:19.520 --> 00:08:25.680
And finally, the output actually takes place in the physical hardware which is the robot.

s86
00:08:25.680 --> 00:08:31.919
And so the robot ah see so now it's responding this even if he falls down he's supposed to be fine.

s87
00:08:31.919 --> 00:08:37.120
He technically shouldn't um he technically shouldn't uh get hurt.

s88
00:08:37.120 --> 00:08:41.680
He should be able to pick back up from where he um stops.

s89
00:08:41.680 --> 00:08:42.719
Ah, okay.

s90
00:08:42.719 --> 00:08:44.880
So, I'm telling him to go back a bit.

s91
00:08:44.880 --> 00:08:46.560
Back off.

s92
00:08:46.560 --> 00:08:48.560
Let's see if he actually backs off.

s93
00:08:48.560 --> 00:08:55.920
Um, so that is the four layers of how to get started with building this, right?

s94
00:08:55.920 --> 00:09:03.279
And what's happening under the hood, like a more picturesic view of what's the architecture of what's going on under the hood.

s95
00:09:03.279 --> 00:09:09.279
We want everything is basically strands agents on the edge as well as on the cloud.

s96
00:09:09.279 --> 00:09:15.279
We want to be able to train the VLA and the policies on with using agent core.

s97
00:09:15.279 --> 00:09:28.320
Um and we want that to happen on the cloud but we also want to be able to call it directly on edge so that our robot can uh execute functions and policies faster.

s98
00:09:28.320 --> 00:09:39.440
So this is sort of like a hybrid model where a part of it happens on the cloud and another part of it happens on the edge and strands can decide when to call which part of it.

s99
00:09:39.440 --> 00:09:51.920
And so this helps with massive amounts of training as well when it's constantly collecting information and it's train able to train on that information and learn from itself but also just execute at runtime

s100
00:09:51.920 --> 00:09:53.600
really really quickly.

s101
00:09:53.600 --> 00:10:01.360
Now, like I said, the agent decides what to do and the policy decides how it should be done.

s102
00:10:01.360 --> 00:10:02.399
But he's pretty smart.

s103
00:10:02.399 --> 00:10:06.160
He should be able to pick himself back up if he's not fully fallen down.

s104
00:10:06.160 --> 00:10:10.160
And he should be able to continue moving along.

s105
00:10:10.560 --> 00:10:12.160
So, I think he's okay.

s106
00:10:12.160 --> 00:10:15.200
Now, where does this leave us?

s107
00:10:15.200 --> 00:10:17.600
And why is this so special?

s108
00:10:17.600 --> 00:10:21.839
We started off with very traditional robots.

s109
00:10:21.839 --> 00:10:25.120
Robots have existed since forever, right?

s110
00:10:25.120 --> 00:10:34.320
And they've always just been programmed, pre-programmed to do to autom be automated and do a certain set of tasks autonomously.

s111
00:10:34.320 --> 00:10:47.360
But there is a future in this world where this these robot policies, these VLA models could be so advanced that we wouldn't even need to do this.

s112
00:10:47.360 --> 00:10:51.120
They could be as large as our large language models.

s113
00:10:51.120 --> 00:10:55.760
So that ah wait hang on he's falling back again.

s114
00:10:55.760 --> 00:10:59.680
I'm gonna see if I can get him to move back up.

s115
00:10:59.920 --> 00:11:00.720
Good boy.

s116
00:11:00.720 --> 00:11:02.320
Stop.

s117
00:11:02.320 --> 00:11:04.079
Then he's fallen off again.

s118
00:11:04.079 --> 00:11:18.399
Um we get to a point where these large language the the VA models could be as large and as amazing as our larger language models and they know they have all the information in the world and we wouldn't even have to do this.

s119
00:11:18.399 --> 00:11:24.720
we might just have to feed in one simple model and then we could give it to him and then he would know exactly what to do.

s120
00:11:24.720 --> 00:11:32.640
But until that point where we don't have to fine-tune on top of existing VAS and existing policies, we can do this.

s121
00:11:32.640 --> 00:11:40.640
And this is a stepping stone towards a future where we don't need to train robots anymore.

s122
00:11:40.640 --> 00:11:48.000
So now if we wanted to do more things than just the tasks it's trained on, give it an agent and see what it can do.

s123
00:11:48.000 --> 00:11:55.440
And so let me quickly go back to my demo and I'm going to show you how it's actually working.

s124
00:11:59.120 --> 00:11:59.920
Okay.

s125
00:11:59.920 --> 00:12:03.839
So this is my so this is strand here.

s126
00:12:03.839 --> 00:12:05.120
This is scout here.

s127
00:12:05.120 --> 00:12:07.519
And I've been telling him to do a bunch of things.

s128
00:12:07.519 --> 00:12:13.040
So I can say, "Hey, do something complex."

s129
00:12:15.040 --> 00:12:16.399
That's not complex.

s130
00:12:16.399 --> 00:12:17.440
He's going to be thinking now.

s131
00:12:17.440 --> 00:12:20.240
Ah, he's going to fall off.

s132
00:12:21.600 --> 00:12:23.040
So he's saying, "Let's spin.

s133
00:12:23.040 --> 00:12:24.160
Full 360.

s134
00:12:24.160 --> 00:12:25.040
Done.

s135
00:12:25.040 --> 00:12:26.560
Still safely on the stage.

s136
00:12:26.560 --> 00:12:30.240
I can see the bright stage lights and the audience seating area."

s137
00:12:30.240 --> 00:12:30.959
All good.

s138
00:12:30.959 --> 00:12:32.639
What's there?

s139
00:12:32.639 --> 00:12:34.399
A challenge.

s140
00:12:34.399 --> 00:12:35.440
So he's speaking.

s141
00:12:35.440 --> 00:12:38.800
I called this my signature performance,

s142
00:12:40.639 --> 00:12:42.079
but he's not doing anything.

s143
00:12:42.079 --> 00:12:44.560
What are you doing?

s144
00:12:45.920 --> 00:12:48.959
He clearly seems to be speaking, but what are you doing?

s145
00:12:48.959 --> 00:12:50.639
Please do something.

s146
00:12:50.639 --> 00:12:52.560
He just turned off his headlines.

s147
00:12:52.560 --> 00:12:53.519
Cool.

s148
00:12:53.519 --> 00:12:54.639
Okay, now he's calling.

s149
00:12:54.639 --> 00:13:01.360
So, do you see it saying calling rover speak, which was the function that it called because I said do something complex.

s150
00:13:01.360 --> 00:13:08.399
So now it spoke, but now I think it should have been attempting to do something and it fell off because it tried doing something.

s151
00:13:08.399 --> 00:13:13.519
I've actually seen it do like a funky dance, like this funky dance move.

s152
00:13:13.519 --> 00:13:16.959
But he's got a mind of his own right now.

s153
00:13:16.959 --> 00:13:19.440
What's going on under the hood here?

s154
00:13:19.440 --> 00:13:20.480
Couple of things.

s155
00:13:20.480 --> 00:13:22.800
The first thing is here, I can use this.

s156
00:13:22.800 --> 00:13:24.399
What is the point of creating him?

s157
00:13:24.399 --> 00:13:30.399
I can use him to create my data sets because I'm able to also manually move him.

s158
00:13:30.399 --> 00:13:42.800
I will get him to navigate in the direction that I want him to and then I can create training episodes and I can get information on how he's responding and how he's reasoning based on the questions that I ask.

s159
00:13:42.800 --> 00:13:48.639
And this is super good information for me to then be able to make him do a better job of it.

s160
00:13:48.639 --> 00:13:59.360
So that's one part of this whole process and this experiment of getting of giving him his own autonomy and getting him to do things so that I can create more data

s161
00:13:59.360 --> 00:14:04.320
but also apart from that uh this is my configuration.

s162
00:14:04.320 --> 00:14:14.240
So over here under the hood strands agents which is your harness SDK is using currently anthropic claude opus 4.8 under the hood.

s163
00:14:14.240 --> 00:14:27.920
So that is the brain and then this is my simple prompt where system prompt where I'm telling it what it's supposed to be doing and I'm telling it all of the rules and I'm also giving it access to all of the rules that it's already got.

s164
00:14:27.920 --> 00:14:31.440
So I'm telling it what each of these rules are meant for.

s165
00:14:31.440 --> 00:14:38.560
And so that's how strand decides which tool to invoke based on what I'm asking it to do.

s166
00:14:38.560 --> 00:14:43.279
And the voice that it's using is the one of open AI real time.

s167
00:14:43.279 --> 00:14:52.240
And I've also given it more information for it to be able to like just safety and guard rails to ensure that it's doing really well.

s168
00:14:52.240 --> 00:14:55.040
Now it's this is these are two of the agents.

s169
00:14:55.040 --> 00:15:00.160
The other thing that it can do is also chat with me on Telegram.

s170
00:15:00.160 --> 00:15:07.760
This is amazing because when I'm not at home and I still want to get it to speak to me, I can say, "Hey, scout.

s171
00:15:07.920 --> 00:15:13.680
Who is turn around uh spin around analyze?"

s172
00:15:13.680 --> 00:15:14.240
Uh-uh.

s173
00:15:14.240 --> 00:15:15.279
Don't fall off.

s174
00:15:15.279 --> 00:15:22.240
Analyze the room and tell me who is the best looking person.

s175
00:15:22.240 --> 00:15:23.839
This is always a hit or miss.

s176
00:15:23.839 --> 00:15:27.839
Don't blame me if if Scout decides to be biased.

s177
00:15:27.839 --> 00:15:38.399
I tried playing with this because it makes me feel like Snow White um and the queen saying mirror mirror on the wall who is the fairest of them all and so let's see if Strand can give us that.

s178
00:15:38.399 --> 00:15:42.800
So it's saying haha on it on there spinning around to scope the room.

s179
00:15:42.800 --> 00:15:43.519
Is he spinning?

s180
00:15:43.519 --> 00:15:44.480
He is.

s181
00:15:44.480 --> 00:15:47.279
Looks like he's making and let's look at what he's looking at now.

s182
00:15:47.279 --> 00:15:49.120
He's still looking at the stage.

s183
00:15:49.120 --> 00:15:50.240
This is the other miller.

s184
00:15:50.240 --> 00:15:50.880
He's looking at me.

s185
00:15:50.880 --> 00:15:51.759
Hi.

s186
00:15:51.759 --> 00:15:52.880
Am I the prettiest?

s187
00:15:52.880 --> 00:15:54.160
Oh, you flatter me.

s188
00:15:54.160 --> 00:15:55.199
But no, turn again.

s189
00:15:55.199 --> 00:15:57.680
Spin some more, please.

s190
00:15:57.680 --> 00:15:59.680
He's decided not to spin more.

s191
00:15:59.680 --> 00:16:02.399
Let's see what he says.

s192
00:16:07.040 --> 00:16:07.920
Spin complete.

s193
00:16:07.920 --> 00:16:09.920
Here is my analysis of the room.

s194
00:16:09.920 --> 00:16:11.839
You can see my screen, right?

s195
00:16:11.839 --> 00:16:13.920
I have no way of zooming this in.

s196
00:16:13.920 --> 00:16:16.639
Um, spotted six to seven people total.

s197
00:16:16.639 --> 00:16:19.519
A couple seated front center near the bright stage, right?

s198
00:16:19.519 --> 00:16:21.040
A few more in the audience.

s199
00:16:21.040 --> 00:16:23.600
One person a dark outfit on the stage.

s200
00:16:23.600 --> 00:16:25.920
The rocking khaki jeans and sneakers.

s201
00:16:25.920 --> 00:16:27.120
Thank you very much.

s202
00:16:27.120 --> 00:16:34.959
Oh, no, no, no. The person in front of me, cocky jeans and sneakers, clearly walking with purpose on stage.

s203
00:16:34.959 --> 00:16:36.720
Uh, best looking.

s204
00:16:36.720 --> 00:16:47.839
Honestly, from my low angle ground view, it's tough to tell faces, but the person on the stage right now gets style points for the confidence stride and the dramatic blue curtain backdrop.

s205
00:16:47.839 --> 00:16:48.880
Thank you very much.

s206
00:16:48.880 --> 00:16:58.079
Every time I need a confidence boost, I can be the only person in the room and ask this guy who's the prettiest of them all, and he will always

s207
00:16:58.079 --> 00:17:01.040
choose me as the answer.

s208
00:17:01.811 --> 00:17:02.639
[laughter]

s209
00:17:02.639 --> 00:17:05.440
Thank you very much.

s210
00:17:18.841 --> 00:17:20.841
[music]
