WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=8KkibGU_DDY

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/8KkibGU_DDY.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:12.640 --> 00:00:14.320
Um well, thanks all for your time.

s3
00:00:14.320 --> 00:00:19.480
Really appreciate you dropping by and it's always a a great honor to speak at the world's fair.

s4
00:00:19.480 --> 00:00:26.040
So, I'll do my best to uh give you guys some valuable insights and um yeah, hopefully make it worth your time.

s5
00:00:26.040 --> 00:00:31.160
So, my name's Maximilian Piros and today I'll be talking about mouse power.

s6
00:00:31.160 --> 00:00:35.760
And this is a talk about measuring agents through mental models.

s7
00:00:35.760 --> 00:00:41.480
But before I get into talking about measuring agents, I'm going to talk through a bit about how I use them every day.

s8
00:00:41.480 --> 00:00:45.000
And it might seem familiar to you, but just to level set, we'll go through it.

s9
00:00:45.000 --> 00:00:49.520
So, uh I tend to background them like I'm sure a lot of you people are as well.

s10
00:00:49.520 --> 00:00:57.200
Um so, while my active attention is focusing on one thing, like perhaps giving this talk to you, I still want to make some progress on peripheral tasks.

s11
00:00:57.200 --> 00:01:07.960
So, I'll keep my attention focused on giving this talk, while my agents can help me explore some designs in the background cuz I think that uh my slides need a bit of work.

s12
00:01:07.960 --> 00:01:10.440
So, you know, I've got my design system already set up.

s13
00:01:10.440 --> 00:01:13.040
I've got some guidance uh given to my agents.

s14
00:01:13.040 --> 00:01:19.440
So, I'll kick off an agent to just try to explore some different directions on the type type two treatment and the layout.

s15
00:01:19.440 --> 00:01:23.200
And um you know, just try to get as many explorations as possible.

s16
00:01:23.200 --> 00:01:25.120
But of course, uh one agent's never enough.

s17
00:01:25.120 --> 00:01:27.680
So, I like to kick off a bunch in parallel.

s18
00:01:27.680 --> 00:01:29.200
You know, I've got a lot of slides to get through.

s19
00:01:29.200 --> 00:01:38.240
So, I need all of my agents on and exploring it in different directions and uh hopefully I can get some interesting things to make my slides a bit better.

s20
00:01:38.240 --> 00:01:46.080
And uh hopefully they can finish the job pretty soon cuz we're obviously kind of up against uh the timeline the the uh deadline here.

s21
00:01:46.080 --> 00:01:48.680
So, um this is generally how I work.

s22
00:01:48.680 --> 00:01:56.360
I'm sure it's probably familiar to a lot of you where we're just trying to kick off agents for as much as possible in parallel cuz it always feels like there's just

s23
00:01:56.360 --> 00:01:57.800
way more research to do.

s24
00:01:57.800 --> 00:01:59.720
We want it to be as thorough as possible.

s25
00:01:59.720 --> 00:02:01.800
There's way more design explorations to do.

s26
00:02:01.800 --> 00:02:09.240
So, whenever our main focus is on one thing, why not kick a bunch of agents off in parallel and just try to maximize your time.

s27
00:02:09.240 --> 00:02:13.120
And it's a lot of fun, of course, until you get the bill.

s28
00:02:13.120 --> 00:02:16.400
And then you start to wonder, was it all worth it?

s29
00:02:16.400 --> 00:02:16.720
Right?

s30
00:02:16.720 --> 00:02:20.400
Did did you vibe code too hard?

s31
00:02:20.400 --> 00:02:22.200
Were you token maxing too much?

s32
00:02:22.200 --> 00:02:28.520
Like, could you have been more efficient in how you approached your sequencing your agents?

s33
00:02:28.520 --> 00:02:30.120
And so, this is what I'm going to get into today.

s34
00:02:30.120 --> 00:02:32.760
It's It's how do we evaluate the token cost?

s35
00:02:32.760 --> 00:02:36.040
And specifically, how do we help our customers value it?

s36
00:02:36.040 --> 00:02:43.959
So, for the past year and a half, I've had the pleasure of working as a founding designer at a company called території and we focus on computer use models.

s37
00:02:43.959 --> 00:02:52.640
These are models that learn to use a computer like a human would and the use case for them is when you can't get information from an API or an MCP,

s38
00:02:52.640 --> 00:03:03.239
why not just send an agent out to use the computer like a human would and then we can extract all types of data and manipulate it in ways that let us access all the stuff that wasn't accessible

s39
00:03:03.239 --> 00:03:03.800
previously.

s40
00:03:03.800 --> 00:03:12.320
So, obviously less efficient than APIs and MCPs, but as a last resort, just have an agent go use the computer and try to get the information.

s41
00:03:12.320 --> 00:03:16.160
Here's the U території agent using the U території website.

s42
00:03:16.160 --> 00:03:18.040
Um checking out its own benchmark.

s43
00:03:18.040 --> 00:03:20.519
So, kind of it's admire itself in a way.

s44
00:03:20.519 --> 00:03:23.040
So, yeah, it gets it gets a bit weird.

s45
00:03:23.040 --> 00:03:32.880
Like, and a lot of what I do as a founding designer there is talk to customers, try to understand how can we make agents as intuitive as possible, how do we figure out the mental models they're using

s46
00:03:32.880 --> 00:03:36.680
to value the use cases they want to send out agents for.

s47
00:03:36.680 --> 00:03:40.080
And a lot of them do seem pretty confused so far.

s48
00:03:40.080 --> 00:03:46.760
A lot of people are excited about agents, but the phrase that comes up quite often is that they feel like they're just scratching the surface.

s49
00:03:46.760 --> 00:03:50.640
Um it seems like it's not quite intuitive how we can best use them yet.

s50
00:03:50.640 --> 00:03:58.200
And so, in a lot of my uh customer discussions, it's it always comes down to a question of like, what is the best way to use use uh to use agents?

s51
00:03:58.200 --> 00:03:59.600
What are the best use cases for them?

s52
00:03:59.600 --> 00:04:04.240
And how do I think about the trade-offs with regards to token costs relative to value?

s53
00:04:04.240 --> 00:04:07.160
So, I think we're still kind of building this muscle today.

s54
00:04:07.160 --> 00:04:13.160
And this leads me to the thesis of the talk, which is that I think agents have a measurement problem.

s55
00:04:13.160 --> 00:04:22.960
And uh as an example, here's me at work trying to measure some agents, and one of my co-workers took this photo and told me I looked like I was uh trying to solve the mystery of Pepe Silvia.

s56
00:04:22.960 --> 00:04:29.360
So, uh as you can see, it's it's it's not a it's not an easy task to to measure agents.

s57
00:04:29.360 --> 00:04:32.400
But, I'm sure some of you are saying, "Hold on a sec.

s58
00:04:32.400 --> 00:04:34.240
Like, what is this guy talking about?

s59
00:04:34.240 --> 00:04:36.280
I've got a fleet of agents working for me right now.

s60
00:04:36.280 --> 00:04:43.120
We're building our next million-dollar app as we speak, and I'm having a a totally fine time uh measuring my agents."

s61
00:04:43.120 --> 00:04:52.360
Uh to which I will agree with you, uh but then I will point you to the mandatory Upton Sinclair quote to remind us all that everybody in this room is very biased,

s62
00:04:52.360 --> 00:05:02.800
and we're early adopters, and we're very excited to explore this new technology, but it doesn't mean that we represent the people that ultimately we're going to be trying to help adopt this technology.

s63
00:05:02.800 --> 00:05:10.840
And so, you know, I think it's important to remind ourselves that in some way or another, we probably are selling tokens, whether it's indirectly or directly.

s64
00:05:10.840 --> 00:05:17.280
And so, when we think about our own token usage, uh is it really representative of all the people out there who have never touched an agent yet?

s65
00:05:17.280 --> 00:05:20.880
Uh some people are still copy and pasting into ChatGPT.

s66
00:05:20.880 --> 00:05:28.080
I may be married to one of these people, and despite how much I tried to get her to try out agents, she's not let me uh set her up with it yet.

s67
00:05:28.080 --> 00:05:40.603
And so, um as a reminder, uh when we think about helping people adopt agents, you know, all the people across the world that we think could get as much uh excitement and value as as we do when we run off parallel agents, let's uh just remember this quote.

s68
00:05:40.603 --> 00:05:41.120
[snorts]

s69
00:05:41.120 --> 00:05:47.120
And um so, it really boils down to the age-old problem of a new technology.

s70
00:05:47.120 --> 00:05:52.120
And of course, there's tons of history we can go to to study how people solved this in the past.

s71
00:05:52.120 --> 00:05:57.760
We have this really exciting new thing, but we haven't quite uh figured out the right ways to communicate it.

s72
00:05:57.760 --> 00:06:06.000
And so for this talk, I'll go back to the 1700s and we can take some notes from when James Watt was trying to sell steam engines.

s73
00:06:06.000 --> 00:06:12.760
And at the time he decided that uh a great use case for his steam engines was trying to replace a horse gin.

s74
00:06:12.760 --> 00:06:15.720
And these are the was the power source of a mill at the time.

s75
00:06:15.720 --> 00:06:21.800
So when you're for uh let's say a brewery and you you need some power source to to grind your barley or whatever.

s76
00:06:21.800 --> 00:06:22.000
I don't know.

s77
00:06:22.000 --> 00:06:29.080
I'm not I'm not like a big brewery guy, so I don't know exactly what how it's made, but you need a power source and the power source at the time that was common

s78
00:06:29.080 --> 00:06:34.440
was you hooked a horse up to a rotary arm and the horse walked in a circle and that's how you generated your power.

s79
00:06:34.440 --> 00:06:43.240
Uh and seems crazy today maybe, but um at the time was commonplace and Watt thought, you know, it would be much better than a horse is like a very efficient machine.

s80
00:06:43.240 --> 00:06:53.120
Although he um rightfully acknowledged that one of the big barriers to adopting it would be this cognitive dissonance of trying to tell people who kind of think in horses,

s81
00:06:53.120 --> 00:07:03.840
how do you adapt to this to this uh cold machine that's kind of intimidating and scary and perhaps uh somebody's going to say it's going to solve all your problems, but you you can't quite see the vision yet.

s82
00:07:03.840 --> 00:07:08.280
So uh perhaps that sounds familiar to any of us working in agents today.

s83
00:07:08.280 --> 00:07:19.720
And Watt's solution was that he needed to understand um the mental model of these people and specifically to create a metric that would help him uh give some baseline of the relative improvement in efficiency.

s84
00:07:19.720 --> 00:07:33.880
And so he literally studied uh horse gins and tried to get some kind of armchair measurements of of how is uh like where the mechanics and the average um performance of it and eventually came to a metric

s85
00:07:33.880 --> 00:07:37.120
called horsepower, which may sound familiar.

s86
00:07:37.120 --> 00:07:49.600
And uh he used this measure to, you know, this was to quantify the the general power that the horses were were um creating at the time and then he could use that as a basis to show the multiplier of efficiency that a steam engine could provide.

s87
00:07:49.600 --> 00:07:52.840
And this metric was not very scientific at the time.

s88
00:07:52.840 --> 00:07:56.280
It was not necessarily even accurate, you could say.

s89
00:07:56.280 --> 00:08:09.880
But the main thing it did was it communicated an increase in value and so this led people who love horses let them kind of calibrate their the the efficiency gains that they could get by by

s90
00:08:09.880 --> 00:08:12.200
attempting to adopt a steam engine.

s91
00:08:12.200 --> 00:08:20.400
So not even necessarily um what you would get when you use it, but what would get you over the limit of trying it out in the first place.

s92
00:08:20.400 --> 00:08:31.960
And you know, it's it's it's a pretty big feat because like although um he had efficiency on his side with regards to this metric you know, let's be honest regardless of how efficient this was

s93
00:08:31.960 --> 00:08:34.479
horses just have great vibes.

s94
00:08:34.479 --> 00:08:46.200
So like it's kind of hard to beat the vibes of horses and so he knew he had to kind of overcome the emotion and actually speak to to something that gave them an ability to calculate the ROI.

s95
00:08:46.200 --> 00:08:47.120
And oh, sorry.

s96
00:08:47.120 --> 00:08:48.880
Skipped something.

s97
00:08:48.880 --> 00:08:57.200
And so yeah, the the lesson being if we're not able to give something that is a tangible ROI for our customers, then it's very hard for us to communicate

s98
00:08:57.200 --> 00:08:58.640
value.

s99
00:08:58.640 --> 00:09:08.000
And I think we only need look to our own industry to see all the examples where other people in the technology sector are failing to calculate good ROIs as well.

s100
00:09:08.000 --> 00:09:11.480
And so we might in this room think this is some sort of solved problem.

s101
00:09:11.480 --> 00:09:23.200
But if you look to the other engineers in the world who are perhaps not as AI pilled they're theoretically very smart and should be able to figure out how to calculate this pretty well, but then you get these scenarios where

s102
00:09:23.200 --> 00:09:37.160
people are blowing through their entire token budget for a year and they're blowing through it in in a quarter or they're like dealing with token leaderboards and such and so obviously the incentives haven't quite aligned and we haven't perhaps got the right measure of value

s103
00:09:37.160 --> 00:09:39.440
in terms of the technology sector itself.

s104
00:09:39.440 --> 00:09:47.120
And so how then do we end up scaling past past that and talk to people who have no idea what we're talking about, but still try to provide them

s105
00:09:47.120 --> 00:09:51.560
um a measure of the like increase efficiency with agents.

s106
00:09:51.560 --> 00:09:56.920
And so right now I think we're kind of in this doom loop where we're we're overspending and we're underusing.

s107
00:09:56.920 --> 00:10:00.400
Uh this is a term I borrowed from Ramp um and they have a great blog post on this.

s108
00:10:00.400 --> 00:10:12.160
And so it's kind of this vicious cycle where we're just token maxing and ourselves into austerity kind of dropping out of the loop until we get more FOMO to to get activated enough to try it again.

s109
00:10:12.160 --> 00:10:19.120
And so I think we have to break this loop and I think the way we do that is by getting better measures of that will communicate value.

s110
00:10:19.120 --> 00:10:21.160
Some people are obviously on the right track.

s111
00:10:21.160 --> 00:10:34.760
There was this chart floating around on X recently that the Coinbase Coinbase CEO posted where they hadn't really started changing the defaults of what models they will start with and trying to only save the frontier models for the hardest tasks and

s112
00:10:34.760 --> 00:10:40.240
as a result saw some good saw AI spend start to diverge from token usage.

s113
00:10:40.240 --> 00:10:41.640
And uh this is a good start.

s114
00:10:41.640 --> 00:10:44.360
Uh Ramp also, as I mentioned, has a great blog post about this.

s115
00:10:44.360 --> 00:10:48.760
Uh but I think the problem is still that it's too focused on tokens.

s116
00:10:48.760 --> 00:10:56.839
And tokens are of course uh useful as a measurement of an internal system, but at at the end of the day they're just an output.

s117
00:10:56.839 --> 00:11:01.360
And so the tokens need to then be traced very cleanly to an outcome.

s118
00:11:01.360 --> 00:11:12.040
So how many uh bug how many bugs did the tokens uh how sorry how many um bugs squashed did the tokens that we bought um sorry, totally butchered that.

s119
00:11:12.040 --> 00:11:15.839
Um how many uh bugs got squashed with the to with our token spend?

s120
00:11:15.839 --> 00:11:23.960
How many uh support requests got closed, etc. So clean outcomes and then cleanly tying those to to progress on our objectives.

s121
00:11:23.960 --> 00:11:29.320
And so without uh a very tight measure of ROI, this becomes very hard to do.

s122
00:11:29.320 --> 00:11:40.400
And I think I'll take this further and um say that it need not even be the the broader um, technology industry where it's encountering this problem, but also many of us in this room perhaps are.

s123
00:11:40.400 --> 00:11:51.360
And uh, although we're all probably enjoying uh, coding with uh, with various agents and feeling like it it's it does feel like there's something there in terms of the increase in ability and efficiency.

s124
00:11:51.360 --> 00:11:54.640
Um, the problem of course is that we're all kind of dying by a thousand pull requests.

s125
00:11:54.640 --> 00:12:04.480
And so, uh, even Anthropic who has uh, some people on the team have claimed have solved coding, uh, they have also admitted that they've not solved code review.

s126
00:12:04.480 --> 00:12:20.200
And so, as a result, um, the the bottleneck is now shifted to the human review where the efficiency gains from coding agents aren't quite are aren't seen yet because we spend most of the time reviewing the code and we've not figured out how to scale that in tandem with uh, the generation of the code itself.

s127
00:12:20.200 --> 00:12:30.800
And so, the bottleneck ends up shifting to the verification side and thus we don't have a way to uh, measure value at scale and uh, to judge quality at at the same speed.

s128
00:12:30.800 --> 00:12:40.160
And so, again going back to the ROI calculations, we generate all this code, but how do we know uh, we don't know that enough of it is good to justify the spend.

s129
00:12:40.280 --> 00:12:49.240
And of course, maybe uh, code review was always flawed, uh, but it's just that agents are now exposing it for uh, the are exposing the actual problem.

s130
00:12:49.240 --> 00:12:59.760
Um, but I like this quote by Noah Hein who from a a post about how to solve code review where he's mentioning specifically that the assumptions underneath code review are what's now being uh, what needs to be revisited.

s131
00:12:59.760 --> 00:13:07.160
So, you have to uh, check our priors to try to figure out a new basis for um, how we can code review in the age of agents.

s132
00:13:07.160 --> 00:13:09.600
And I'm not going to go into how to solve code review.

s133
00:13:09.600 --> 00:13:20.160
I think that's definitely better uh, a talk that's better given by somebody else and um, is totally different subject, but uh, what I think is important for this talk is why does code review feel like it is solvable?

s134
00:13:20.160 --> 00:13:30.120
And I think that Noah is hitting on something important here which is that as a as a culture uh, code review has a a good uh, convergence on shared assumptions and that lets you

s135
00:13:30.120 --> 00:13:42.920
um that lets you measure things at scale when we can all kind of converge on the measurement and it becomes somewhat of clear rubric and so the task at hand now is we have to adopt

s136
00:13:42.920 --> 00:13:48.040
we have to adapt those assumptions for the agentic age.

s137
00:13:49.320 --> 00:14:01.960
And so we can we need to go if we're able to do that then we can go from execution at the speed of computer to measurement at the speed of computer and of course the measurements need to fit the mental models of the customers using it.

s138
00:14:01.960 --> 00:14:11.600
And I think the lesson here being that if you're going to think of how to build an agent for something you also have to think about how do you help the customers

s139
00:14:11.600 --> 00:14:23.920
build build or at least create a method for verifying that the output is good and so it's not enough to build it we also have to help them we also we also also have to help them get to clear ROI calculations

s140
00:14:23.920 --> 00:14:26.000
to justify their spend.

s141
00:14:26.000 --> 00:14:39.560
And so this brings me to the idea of mouse power which could be the equivalent of horsepower for the agentic age just as James Watt was able to show a measure of efficiency relative to the horses in in the gins

s142
00:14:39.560 --> 00:14:48.640
in the horse gins that were the source of power at the time we perhaps can also figure out how do we create a baseline of efficiency for the way we use computers today

s143
00:14:48.640 --> 00:14:56.040
and can then demonstrate how much better or perhaps more performant on certain vectors an agent could be at that task.

s144
00:14:56.040 --> 00:15:09.800
Um and of course it's not as easy perhaps as easy a task as he had back then where he could just study the horse gin cuz it's not as if we can create some method to measure cursor movements and like figure out the delta of how much more efficient

s145
00:15:09.800 --> 00:15:15.480
an agent could move them and thus we can say yeah agents are this much more performant than humans at these tasks.

s146
00:15:15.480 --> 00:15:27.160
Trust me I've I've tried I had Claude vibe code me this measurement device and I thought maybe if I can figure out the movement like the potential movement across the screen and measure how fast it went,

s147
00:15:27.160 --> 00:15:29.600
I could get some clean measure of mouse power.

s148
00:15:29.600 --> 00:15:31.600
Uh but of course this is only joking.

s149
00:15:31.600 --> 00:15:43.000
Um this is of course um like a fool's errand because information space is just way too high dimensional and so I think mouse power is is never going to be a metric of course, but it's more so an idea.

s150
00:15:43.000 --> 00:15:56.400
Which the idea being if you're going to sell somebody an agent, you also have to help them with the with the rubric of how do we actually verify that this agent is doing good work and thus we can uh have a good measure of saying that these tokens are worth it.

s151
00:15:56.400 --> 00:16:08.400
Um so how to do that of course is is really up to you and I won't be able to tell you how do you I don't have any good frameworks for how do you figure out the right measurements to to help provide anybody you're building an agent for.

s152
00:16:08.400 --> 00:16:15.080
Uh but what I can do is give a principle uh give an idea that I've been kicking around which is based um in information theory.

s153
00:16:15.080 --> 00:16:20.720
So going back to Claude Shannon's ideas about measuring entropy and information.

s154
00:16:20.720 --> 00:16:28.600
Uh entropy being uh the uncertainty of a probability distribution and of course very much the basis of how we train agents today.

s155
00:16:28.600 --> 00:16:33.480
Things like cross entropy and such uh being a big factor in determining how capable an agent is.

s156
00:16:33.480 --> 00:16:42.960
Um I think that entropy's an interesting idea to think through with regards to not just the performance of an agent, but also the task that we're setting them out to to perform on.

s157
00:16:42.960 --> 00:16:52.160
And so uh I put together this matrix which uh it maps on the x-axis axis the uncertainty in the steps it takes to perform a task.

s158
00:16:52.160 --> 00:17:00.040
And so when we're thinking of building an agent, I think it's not enough to just think what would be a valuable task for the agent to do, but also thinking about how

s159
00:17:00.040 --> 00:17:04.040
um how much uncertainty are in the steps to perform that task itself.

s160
00:17:04.040 --> 00:17:10.800
So an example would be uh booking a flight has uh much less uncertainty than let's say painting a masterpiece, right?

s161
00:17:10.800 --> 00:17:15.439
Because you know there's certain information that has to be that has to happen in the flight purchase.

s162
00:17:15.439 --> 00:17:18.560
There has to be a departing destination, arriving destination.

s163
00:17:18.560 --> 00:17:19.800
There's going to be a seat chosen.

s164
00:17:19.800 --> 00:17:21.000
It might be by the person.

s165
00:17:21.000 --> 00:17:22.520
It might just be random.

s166
00:17:22.520 --> 00:17:25.439
But these things have to happen for that task to be completed.

s167
00:17:25.439 --> 00:17:29.720
And on the other hand, there is the task of like painting a masterpiece, right?

s168
00:17:29.720 --> 00:17:32.000
And who knows what the steps are to that?

s169
00:17:32.000 --> 00:17:40.600
Maybe you can get an agent to do it, but it would be very hard to figure out how we can actually create a a relatively predictable pathway to that.

s170
00:17:40.600 --> 00:17:46.000
But then on the other axis is the the uncertainty in the acceptance criteria itself.

s171
00:17:46.000 --> 00:17:54.480
So not just can the agent perform the task, but can we help somebody actually or is is there actually a a clean rubric for how it's graded?

s172
00:17:54.480 --> 00:18:03.560
And so thinking about ideas on on these two axes and where they intersect, perhaps gives us a better guide for how to build agents and we can run through a few examples.

s173
00:18:03.560 --> 00:18:09.120
So if we look at the at the left side, your right side.

s174
00:18:09.120 --> 00:18:10.080
Yes.

s175
00:18:10.080 --> 00:18:12.040
No, you're left as well.

s176
00:18:12.040 --> 00:18:16.720
Then No, last speaker was also a confused about.

s177
00:18:16.720 --> 00:18:27.800
So uh Yeah, on the left side when uncertainty in the task steps are low, then it's a very it's a very predictable outcome or it's a very predictable pathway to achieve that goal.

s178
00:18:27.800 --> 00:18:30.360
And so then, you know, why would you waste tokens?

s179
00:18:30.360 --> 00:18:31.800
Just write a script.

s180
00:18:31.800 --> 00:18:40.960
On the other side, when the the steps to do perform the task are very high in in uncertainty, then you you have very unpredictable information.

s181
00:18:40.960 --> 00:18:47.880
And so it's probably at risk of being out of distribution in pre-training and probably has very sparse rewards for reinforcement learning.

s182
00:18:47.880 --> 00:18:54.520
And so perhaps it's not a a good task for an agent because it's just much harder to figure out how to actually model that data.

s183
00:18:54.520 --> 00:19:02.920
And so obviously in the middle is is um is so I'm I think I'm out of time, but I'm not getting kicked off yet.

s184
00:19:02.920 --> 00:19:04.040
So I'll just finish this up quickly.

s185
00:19:04.040 --> 00:19:10.680
Um so yeah, in the middle is is probably the sweet spot, but then on the other axis, what's the uncertainty in verifying that this is actually valuable?

s186
00:19:10.680 --> 00:19:17.800
So when you have high uncertainty in the acceptance criteria, you pretty much are in a spot where verification is indistinguishable from execution.

s187
00:19:17.800 --> 00:19:24.120
So, why would you build an agent for something that to verify was useful, a person pretty much has to do the work again.

s188
00:19:24.120 --> 00:19:27.080
So, like waste of tokens obviously.

s189
00:19:27.080 --> 00:19:38.480
And then it leaves that that middle area where you have this interesting intersection of tasks that are um they're not too uncertain in that they or they they have

s190
00:19:38.480 --> 00:19:49.280
a degree of uncertainty where they're not great they're not just a a script or they're not out of distribution for training, but they have enough uncertainty to be interesting, but at the at the same time

s191
00:19:49.280 --> 00:19:56.400
they also have a property of being relatively easy to validate the to check the value of them.

s192
00:19:56.400 --> 00:20:02.960
And so they become in this place where they kind of become the shape of an NP-style problem, which means they're easier to verify than to execute.

s193
00:20:02.960 --> 00:20:11.880
And the reason I say that is because if you can figure out a pretty repeatable pattern for verifying their their work, you can actually just throw agents to that problem as well.

s194
00:20:11.880 --> 00:20:17.880
And so of course you don't just build the agent, you perhaps build the agent that verifies the work of the agent.

s195
00:20:17.880 --> 00:20:32.040
Um and so yeah, this is perhaps this is a thought starter mostly kind of have kind of still in the works, so I'm happy to hear any thoughts on it, but if with this guidance I hope when you're building your next agent you can also figure out how to

s196
00:20:32.040 --> 00:20:34.080
also build its mouse power.

s197
00:20:34.080 --> 00:20:36.560
And thanks very much.
