Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori
AI Engineer · 20 min · 197 sentences · from YouTube's caption track
Each timecode opens YouTube at the start of that sentence. Line anchors (#s42) are the cue ids in the WebVTT, and every line carries its start and end seconds. All transcripts has every talk, and the whole corpus as one file.
- 00:01[music]
- 00:12Um well, thanks all for your time.
- 00:14Really appreciate you dropping by and it's always a a great honor to speak at the world's fair.
- 00:19So, I'll do my best to uh give you guys some valuable insights and um yeah, hopefully make it worth your time.
- 00:26So, my name's Maximilian Piros and today I'll be talking about mouse power.
- 00:31And this is a talk about measuring agents through mental models.
- 00:35But before I get into talking about measuring agents, I'm going to talk through a bit about how I use them every day.
- 00:41And it might seem familiar to you, but just to level set, we'll go through it.
- 00:45So, uh I tend to background them like I'm sure a lot of you people are as well.
- 00:49Um so, while my active attention is focusing on one thing, like perhaps giving this talk to you, I still want to make some progress on peripheral tasks.
- 00:57So, I'll keep my attention focused on giving this talk, while my agents can help me explore some designs in the background cuz I think that uh my slides need a bit of work.
- 01:07So, you know, I've got my design system already set up.
- 01:10I've got some guidance uh given to my agents.
- 01:13So, I'll kick off an agent to just try to explore some different directions on the type type two treatment and the layout.
- 01:19And um you know, just try to get as many explorations as possible.
- 01:23But of course, uh one agent's never enough.
- 01:25So, I like to kick off a bunch in parallel.
- 01:27You know, I've got a lot of slides to get through.
- 01:29So, I need all of my agents on and exploring it in different directions and uh hopefully I can get some interesting things to make my slides a bit better.
- 01:38And uh hopefully they can finish the job pretty soon cuz we're obviously kind of up against uh the timeline the the uh deadline here.
- 01:46So, um this is generally how I work.
- 01:48I'm sure it's probably familiar to a lot of you where we're just trying to kick off agents for as much as possible in parallel cuz it always feels like there's just
- 01:56way more research to do.
- 01:57We want it to be as thorough as possible.
- 01:59There's way more design explorations to do.
- 02:01So, whenever our main focus is on one thing, why not kick a bunch of agents off in parallel and just try to maximize your time.
- 02:09And it's a lot of fun, of course, until you get the bill.
- 02:13And then you start to wonder, was it all worth it?
- 02:16Right?
- 02:16Did did you vibe code too hard?
- 02:20Were you token maxing too much?
- 02:22Like, could you have been more efficient in how you approached your sequencing your agents?
- 02:28And so, this is what I'm going to get into today.
- 02:30It's It's how do we evaluate the token cost?
- 02:32And specifically, how do we help our customers value it?
- 02:36So, for the past year and a half, I've had the pleasure of working as a founding designer at a company called території and we focus on computer use models.
- 02:43These are models that learn to use a computer like a human would and the use case for them is when you can't get information from an API or an MCP,
- 02:52why not just send an agent out to use the computer like a human would and then we can extract all types of data and manipulate it in ways that let us access all the stuff that wasn't accessible
- 03:03previously.
- 03:03So, obviously less efficient than APIs and MCPs, but as a last resort, just have an agent go use the computer and try to get the information.
- 03:12Here's the U території agent using the U території website.
- 03:16Um checking out its own benchmark.
- 03:18So, kind of it's admire itself in a way.
- 03:20So, yeah, it gets it gets a bit weird.
- 03:23Like, and a lot of what I do as a founding designer there is talk to customers, try to understand how can we make agents as intuitive as possible, how do we figure out the mental models they're using
- 03:32to value the use cases they want to send out agents for.
- 03:36And a lot of them do seem pretty confused so far.
- 03:40A lot of people are excited about agents, but the phrase that comes up quite often is that they feel like they're just scratching the surface.
- 03:46Um it seems like it's not quite intuitive how we can best use them yet.
- 03:50And so, in a lot of my uh customer discussions, it's it always comes down to a question of like, what is the best way to use use uh to use agents?
- 03:58What are the best use cases for them?
- 03:59And how do I think about the trade-offs with regards to token costs relative to value?
- 04:04So, I think we're still kind of building this muscle today.
- 04:07And this leads me to the thesis of the talk, which is that I think agents have a measurement problem.
- 04:13And uh as an example, here's me at work trying to measure some agents, and one of my co-workers took this photo and told me I looked like I was uh trying to solve the mystery of Pepe Silvia.
- 04:22So, uh as you can see, it's it's it's not a it's not an easy task to to measure agents.
- 04:29But, I'm sure some of you are saying, "Hold on a sec.
- 04:32Like, what is this guy talking about?
- 04:34I've got a fleet of agents working for me right now.
- 04:36We're building our next million-dollar app as we speak, and I'm having a a totally fine time uh measuring my agents."
- 04:43Uh to which I will agree with you, uh but then I will point you to the mandatory Upton Sinclair quote to remind us all that everybody in this room is very biased,
- 04:52and we're early adopters, and we're very excited to explore this new technology, but it doesn't mean that we represent the people that ultimately we're going to be trying to help adopt this technology.
- 05:02And so, you know, I think it's important to remind ourselves that in some way or another, we probably are selling tokens, whether it's indirectly or directly.
- 05:10And so, when we think about our own token usage, uh is it really representative of all the people out there who have never touched an agent yet?
- 05:17Uh some people are still copy and pasting into ChatGPT.
- 05:20I may be married to one of these people, and despite how much I tried to get her to try out agents, she's not let me uh set her up with it yet.
- 05:28And so, um as a reminder, uh when we think about helping people adopt agents, you know, all the people across the world that we think could get as much uh excitement and value as as we do when we run off parallel agents, let's uh just remember this quote.
- 05:40[snorts]
- 05:41And um so, it really boils down to the age-old problem of a new technology.
- 05:47And of course, there's tons of history we can go to to study how people solved this in the past.
- 05:52We have this really exciting new thing, but we haven't quite uh figured out the right ways to communicate it.
- 05:57And so for this talk, I'll go back to the 1700s and we can take some notes from when James Watt was trying to sell steam engines.
- 06:06And at the time he decided that uh a great use case for his steam engines was trying to replace a horse gin.
- 06:12And these are the was the power source of a mill at the time.
- 06:15So when you're for uh let's say a brewery and you you need some power source to to grind your barley or whatever.
- 06:21I don't know.
- 06:22I'm not I'm not like a big brewery guy, so I don't know exactly what how it's made, but you need a power source and the power source at the time that was common
- 06:29was you hooked a horse up to a rotary arm and the horse walked in a circle and that's how you generated your power.
- 06:34Uh and seems crazy today maybe, but um at the time was commonplace and Watt thought, you know, it would be much better than a horse is like a very efficient machine.
- 06:43Although he um rightfully acknowledged that one of the big barriers to adopting it would be this cognitive dissonance of trying to tell people who kind of think in horses,
- 06:53how do you adapt to this to this uh cold machine that's kind of intimidating and scary and perhaps uh somebody's going to say it's going to solve all your problems, but you you can't quite see the vision yet.
- 07:03So uh perhaps that sounds familiar to any of us working in agents today.
- 07:08And Watt's solution was that he needed to understand um the mental model of these people and specifically to create a metric that would help him uh give some baseline of the relative improvement in efficiency.
- 07:19And so he literally studied uh horse gins and tried to get some kind of armchair measurements of of how is uh like where the mechanics and the average um performance of it and eventually came to a metric
- 07:33called horsepower, which may sound familiar.
- 07:37And uh he used this measure to, you know, this was to quantify the the general power that the horses were were um creating at the time and then he could use that as a basis to show the multiplier of efficiency that a steam engine could provide.
- 07:49And this metric was not very scientific at the time.
- 07:52It was not necessarily even accurate, you could say.
- 07:56But the main thing it did was it communicated an increase in value and so this led people who love horses let them kind of calibrate their the the efficiency gains that they could get by by
- 08:09attempting to adopt a steam engine.
- 08:12So not even necessarily um what you would get when you use it, but what would get you over the limit of trying it out in the first place.
- 08:20And you know, it's it's it's a pretty big feat because like although um he had efficiency on his side with regards to this metric you know, let's be honest regardless of how efficient this was
- 08:31horses just have great vibes.
- 08:34So like it's kind of hard to beat the vibes of horses and so he knew he had to kind of overcome the emotion and actually speak to to something that gave them an ability to calculate the ROI.
- 08:46And oh, sorry.
- 08:47Skipped something.
- 08:48And so yeah, the the lesson being if we're not able to give something that is a tangible ROI for our customers, then it's very hard for us to communicate
- 08:57value.
- 08:58And I think we only need look to our own industry to see all the examples where other people in the technology sector are failing to calculate good ROIs as well.
- 09:08And so we might in this room think this is some sort of solved problem.
- 09:11But if you look to the other engineers in the world who are perhaps not as AI pilled they're theoretically very smart and should be able to figure out how to calculate this pretty well, but then you get these scenarios where
- 09:23people are blowing through their entire token budget for a year and they're blowing through it in in a quarter or they're like dealing with token leaderboards and such and so obviously the incentives haven't quite aligned and we haven't perhaps got the right measure of value
- 09:37in terms of the technology sector itself.
- 09:39And so how then do we end up scaling past past that and talk to people who have no idea what we're talking about, but still try to provide them
- 09:47um a measure of the like increase efficiency with agents.
- 09:51And so right now I think we're kind of in this doom loop where we're we're overspending and we're underusing.
- 09:56Uh this is a term I borrowed from Ramp um and they have a great blog post on this.
- 10:00And so it's kind of this vicious cycle where we're just token maxing and ourselves into austerity kind of dropping out of the loop until we get more FOMO to to get activated enough to try it again.
- 10:12And so I think we have to break this loop and I think the way we do that is by getting better measures of that will communicate value.
- 10:19Some people are obviously on the right track.
- 10:21There was this chart floating around on X recently that the Coinbase Coinbase CEO posted where they hadn't really started changing the defaults of what models they will start with and trying to only save the frontier models for the hardest tasks and
- 10:34as a result saw some good saw AI spend start to diverge from token usage.
- 10:40And uh this is a good start.
- 10:41Uh Ramp also, as I mentioned, has a great blog post about this.
- 10:44Uh but I think the problem is still that it's too focused on tokens.
- 10:48And tokens are of course uh useful as a measurement of an internal system, but at at the end of the day they're just an output.
- 10:56And so the tokens need to then be traced very cleanly to an outcome.
- 11:01So how many uh bug how many bugs did the tokens uh how sorry how many um bugs squashed did the tokens that we bought um sorry, totally butchered that.
- 11:12Um how many uh bugs got squashed with the to with our token spend?
- 11:15How many uh support requests got closed, etc. So clean outcomes and then cleanly tying those to to progress on our objectives.
- 11:23And so without uh a very tight measure of ROI, this becomes very hard to do.
- 11:29And I think I'll take this further and um say that it need not even be the the broader um, technology industry where it's encountering this problem, but also many of us in this room perhaps are.
- 11:40And uh, although we're all probably enjoying uh, coding with uh, with various agents and feeling like it it's it does feel like there's something there in terms of the increase in ability and efficiency.
- 11:51Um, the problem of course is that we're all kind of dying by a thousand pull requests.
- 11:54And so, uh, even Anthropic who has uh, some people on the team have claimed have solved coding, uh, they have also admitted that they've not solved code review.
- 12:04And so, as a result, um, the the bottleneck is now shifted to the human review where the efficiency gains from coding agents aren't quite are aren't seen yet because we spend most of the time reviewing the code and we've not figured out how to scale that in tandem with uh, the generation of the code itself.
- 12:20And so, the bottleneck ends up shifting to the verification side and thus we don't have a way to uh, measure value at scale and uh, to judge quality at at the same speed.
- 12:30And so, again going back to the ROI calculations, we generate all this code, but how do we know uh, we don't know that enough of it is good to justify the spend.
- 12:40And of course, maybe uh, code review was always flawed, uh, but it's just that agents are now exposing it for uh, the are exposing the actual problem.
- 12:49Um, but I like this quote by Noah Hein who from a a post about how to solve code review where he's mentioning specifically that the assumptions underneath code review are what's now being uh, what needs to be revisited.
- 12:59So, you have to uh, check our priors to try to figure out a new basis for um, how we can code review in the age of agents.
- 13:07And I'm not going to go into how to solve code review.
- 13:09I think that's definitely better uh, a talk that's better given by somebody else and um, is totally different subject, but uh, what I think is important for this talk is why does code review feel like it is solvable?
- 13:20And I think that Noah is hitting on something important here which is that as a as a culture uh, code review has a a good uh, convergence on shared assumptions and that lets you
- 13:30um that lets you measure things at scale when we can all kind of converge on the measurement and it becomes somewhat of clear rubric and so the task at hand now is we have to adopt
- 13:42we have to adapt those assumptions for the agentic age.
- 13:49And so we can we need to go if we're able to do that then we can go from execution at the speed of computer to measurement at the speed of computer and of course the measurements need to fit the mental models of the customers using it.
- 14:01And I think the lesson here being that if you're going to think of how to build an agent for something you also have to think about how do you help the customers
- 14:11build build or at least create a method for verifying that the output is good and so it's not enough to build it we also have to help them we also we also also have to help them get to clear ROI calculations
- 14:23to justify their spend.
- 14:26And so this brings me to the idea of mouse power which could be the equivalent of horsepower for the agentic age just as James Watt was able to show a measure of efficiency relative to the horses in in the gins
- 14:39in the horse gins that were the source of power at the time we perhaps can also figure out how do we create a baseline of efficiency for the way we use computers today
- 14:48and can then demonstrate how much better or perhaps more performant on certain vectors an agent could be at that task.
- 14:56Um and of course it's not as easy perhaps as easy a task as he had back then where he could just study the horse gin cuz it's not as if we can create some method to measure cursor movements and like figure out the delta of how much more efficient
- 15:09an agent could move them and thus we can say yeah agents are this much more performant than humans at these tasks.
- 15:15Trust me I've I've tried I had Claude vibe code me this measurement device and I thought maybe if I can figure out the movement like the potential movement across the screen and measure how fast it went,
- 15:27I could get some clean measure of mouse power.
- 15:29Uh but of course this is only joking.
- 15:31Um this is of course um like a fool's errand because information space is just way too high dimensional and so I think mouse power is is never going to be a metric of course, but it's more so an idea.
- 15:43Which the idea being if you're going to sell somebody an agent, you also have to help them with the with the rubric of how do we actually verify that this agent is doing good work and thus we can uh have a good measure of saying that these tokens are worth it.
- 15:56Um so how to do that of course is is really up to you and I won't be able to tell you how do you I don't have any good frameworks for how do you figure out the right measurements to to help provide anybody you're building an agent for.
- 16:08Uh but what I can do is give a principle uh give an idea that I've been kicking around which is based um in information theory.
- 16:15So going back to Claude Shannon's ideas about measuring entropy and information.
- 16:20Uh entropy being uh the uncertainty of a probability distribution and of course very much the basis of how we train agents today.
- 16:28Things like cross entropy and such uh being a big factor in determining how capable an agent is.
- 16:33Um I think that entropy's an interesting idea to think through with regards to not just the performance of an agent, but also the task that we're setting them out to to perform on.
- 16:42And so uh I put together this matrix which uh it maps on the x-axis axis the uncertainty in the steps it takes to perform a task.
- 16:52And so when we're thinking of building an agent, I think it's not enough to just think what would be a valuable task for the agent to do, but also thinking about how
- 17:00um how much uncertainty are in the steps to perform that task itself.
- 17:04So an example would be uh booking a flight has uh much less uncertainty than let's say painting a masterpiece, right?
- 17:10Because you know there's certain information that has to be that has to happen in the flight purchase.
- 17:15There has to be a departing destination, arriving destination.
- 17:18There's going to be a seat chosen.
- 17:19It might be by the person.
- 17:21It might just be random.
- 17:22But these things have to happen for that task to be completed.
- 17:25And on the other hand, there is the task of like painting a masterpiece, right?
- 17:29And who knows what the steps are to that?
- 17:32Maybe you can get an agent to do it, but it would be very hard to figure out how we can actually create a a relatively predictable pathway to that.
- 17:40But then on the other axis is the the uncertainty in the acceptance criteria itself.
- 17:46So not just can the agent perform the task, but can we help somebody actually or is is there actually a a clean rubric for how it's graded?
- 17:54And so thinking about ideas on on these two axes and where they intersect, perhaps gives us a better guide for how to build agents and we can run through a few examples.
- 18:03So if we look at the at the left side, your right side.
- 18:09Yes.
- 18:10No, you're left as well.
- 18:12Then No, last speaker was also a confused about.
- 18:16So uh Yeah, on the left side when uncertainty in the task steps are low, then it's a very it's a very predictable outcome or it's a very predictable pathway to achieve that goal.
- 18:27And so then, you know, why would you waste tokens?
- 18:30Just write a script.
- 18:31On the other side, when the the steps to do perform the task are very high in in uncertainty, then you you have very unpredictable information.
- 18:40And so it's probably at risk of being out of distribution in pre-training and probably has very sparse rewards for reinforcement learning.
- 18:47And so perhaps it's not a a good task for an agent because it's just much harder to figure out how to actually model that data.
- 18:54And so obviously in the middle is is um is so I'm I think I'm out of time, but I'm not getting kicked off yet.
- 19:02So I'll just finish this up quickly.
- 19:04Um so yeah, in the middle is is probably the sweet spot, but then on the other axis, what's the uncertainty in verifying that this is actually valuable?
- 19:10So when you have high uncertainty in the acceptance criteria, you pretty much are in a spot where verification is indistinguishable from execution.
- 19:17So, why would you build an agent for something that to verify was useful, a person pretty much has to do the work again.
- 19:24So, like waste of tokens obviously.
- 19:27And then it leaves that that middle area where you have this interesting intersection of tasks that are um they're not too uncertain in that they or they they have
- 19:38a degree of uncertainty where they're not great they're not just a a script or they're not out of distribution for training, but they have enough uncertainty to be interesting, but at the at the same time
- 19:49they also have a property of being relatively easy to validate the to check the value of them.
- 19:56And so they become in this place where they kind of become the shape of an NP-style problem, which means they're easier to verify than to execute.
- 20:02And the reason I say that is because if you can figure out a pretty repeatable pattern for verifying their their work, you can actually just throw agents to that problem as well.
- 20:11And so of course you don't just build the agent, you perhaps build the agent that verifies the work of the agent.
- 20:17Um and so yeah, this is perhaps this is a thought starter mostly kind of have kind of still in the works, so I'm happy to hear any thoughts on it, but if with this guidance I hope when you're building your next agent you can also figure out how to
- 20:32also build its mouse power.
- 20:34And thanks very much.