WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=-npY6XjM8CQ

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/-npY6XjM8CQ.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:12.880 --> 00:00:14.799
Let's get started.

s2
00:00:14.799 --> 00:00:18.800
When will the benchmaxing plague end?

s3
00:00:18.960 --> 00:00:21.760
In the tech industry, we love a hype cycle.

s4
00:00:21.760 --> 00:00:24.960
And in AI, we really love a hype cycle.

s5
00:00:24.960 --> 00:00:32.480
And the way we do that is when a model comes out, there's a big announcement, there's a lot of benchmark cited.

s6
00:00:32.480 --> 00:00:36.160
Sometimes to keep things interesting, we do a little chart crime.

s7
00:00:36.160 --> 00:00:38.879
And then people actually go and use it.

s8
00:00:38.879 --> 00:00:45.920
And if the expectations aren't met by the reality, then we have allegations of benchmaxing.

s9
00:00:45.920 --> 00:00:54.160
Benchmaxing, of course, being when labs are training too hard on benchmarks in a way that deviates from what people actually care about.

s10
00:00:55.039 --> 00:01:00.399
So the existence of that term indicates that we have a sense that benchmarks don't always equal reality.

s11
00:01:00.399 --> 00:01:04.960
And so in this talk we're going to figure out why does benchmaxing happen?

s12
00:01:04.960 --> 00:01:10.240
Why are traditional benchmarks not always accurate reflections of real world value?

s13
00:01:10.240 --> 00:01:12.000
Is this intrinsic to all benchmarks?

s14
00:01:12.000 --> 00:01:14.960
And will we ever know which models are best?

s15
00:01:14.960 --> 00:01:20.000
And the answers are incentives, poor methodologies, no and yes.

s16
00:01:20.000 --> 00:01:21.119
All right, that was my talk.

s17
00:01:21.119 --> 00:01:22.960
Thank you so much for coming.

s18
00:01:22.960 --> 00:01:27.360
Um actually it looks like I have a few extra minutes so let's let's move on.

s19
00:01:27.360 --> 00:01:31.040
I have a few extra slides we'll we'll go through.

s20
00:01:31.520 --> 00:01:39.920
So we have a sense that benchmarks don't equal reality but the industry is dominated by a lot of popular but very bad benchmarks.

s21
00:01:39.920 --> 00:01:56.320
So there's millions of dollars on prediction markets being wagered on Elm Marina outcomes even as we have industry leaders openly bragging about gaming Elm Marina and you have thought leaders like Wor saying it can be easily gamed.

s22
00:01:56.320 --> 00:02:02.479
It's past time for the Elm Marina people to sit down and think about whether they're doing more harm than good.

s23
00:02:02.479 --> 00:02:11.840
Andre Karpathy had a similar observation when he noticed that the models that he thought were best were not lining up with what Elmarina was ranking.

s24
00:02:11.840 --> 00:02:21.760
He said unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis.

s25
00:02:21.760 --> 00:02:30.400
So why does this happen that sort of industry insiders are telling us that this benchmark is not useful but it still gets a lot of play.

s26
00:02:30.400 --> 00:02:37.519
The problem is that AI is aimed at everyone in the world is is something everyone in the world can use.

s27
00:02:37.519 --> 00:02:41.040
And so everyone needs some tool to figure out which models are best.

s28
00:02:41.040 --> 00:02:43.440
And benchmarks are what we have for that.

s29
00:02:43.440 --> 00:02:50.080
But if you can't if you don't have the ability to assess if a benchmark is good, what you do have is the ability to assess what's popular.

s30
00:02:50.080 --> 00:02:59.599
And this creates this avalanche, this feedback effect where the conversation is very much driven by incumbency and marketing and less by real world value.

s31
00:02:59.599 --> 00:03:00.959
and even myself, right?

s32
00:03:00.959 --> 00:03:05.760
Like unless I actually look at a benchmark in a fair amount of detail, I don't have an opinion on it.

s33
00:03:05.760 --> 00:03:08.800
So, it's a very challenging problem.

s34
00:03:08.800 --> 00:03:14.319
So, what are the things that benchmarks do that lead to these problems?

s35
00:03:14.319 --> 00:03:18.720
There are a handful of key antiatterns that we're going to go through.

s36
00:03:19.040 --> 00:03:20.560
The first is price.

s37
00:03:20.560 --> 00:03:28.480
Let's say you want to make an agentic coding benchmark, which these days is a very popular thing to want to do, and you want a thousand tasks in your benchmark.

s38
00:03:28.480 --> 00:03:30.720
Each task takes 60 hours to make.

s39
00:03:30.720 --> 00:03:34.640
Each software engineer in your workforce costs half a million a year.

s40
00:03:34.640 --> 00:03:37.920
That's $15 million to make your benchmark.

s41
00:03:37.920 --> 00:03:48.000
And if you think that over time about a third of those tasks are going to get washed away every year due to models getting better, that's $5 million to replace them.

s42
00:03:48.000 --> 00:03:52.400
So that puts you out of budget for most projects.

s43
00:03:52.400 --> 00:03:56.560
So then people turn to a variety of workarounds that have their own problems.

s44
00:03:56.560 --> 00:04:02.400
One of which is trying to use a lot of AI assistance which ultimately does not really work.

s45
00:04:02.400 --> 00:04:05.840
Like you can't push the frontier forward from within the frontier.

s46
00:04:05.840 --> 00:04:12.000
You need to inject that external human expertise and it needs to be good expertise.

s47
00:04:12.000 --> 00:04:19.120
If you try to use cheap labor, you're going to get what you pay for and the whole result is not going to be that useful.

s48
00:04:19.120 --> 00:04:24.240
At Surge, one of our differentiators has long been that we are not trying to minimize cost.

s49
00:04:24.240 --> 00:04:31.440
We are trying to maximize quality and part of that means paying a lot of money for good workers.

s50
00:04:31.440 --> 00:04:39.919
We've always believed that but especially in 2026 models are just beyond the point where you can make do with anything less than the best workers.

s51
00:04:41.120 --> 00:04:53.680
Contamination is often thought of as when labs are explicitly training on the test set and that does happen sometimes but really contamination is the default outcome unless you are very very good.

s52
00:04:53.680 --> 00:05:00.720
So labs put a lot of effort into holding back this flood of data that's going to contaminate their models.

s53
00:05:00.720 --> 00:05:08.080
But inevitably if you have public questions and answers on the internet that's going to get memorized to some extent.

s54
00:05:08.080 --> 00:05:11.360
So SweetBench verified here's an example prompt.

s55
00:05:11.360 --> 00:05:16.639
You can give opus the first part of the prompt and it will verbatim spit out the rest.

s56
00:05:16.639 --> 00:05:20.320
It does that with the answers as well.

s57
00:05:20.320 --> 00:05:27.360
And we actually did an investigation where we compared looking at the repos that Sweepbench verified was built out of.

s58
00:05:27.360 --> 00:05:33.280
How much has Opus memorized the Sweepbench verified contents versus the rest of the repo?

s59
00:05:33.280 --> 00:05:38.320
And we found very clear evidence that Opus had memorized a lot of Sweetbench.

s60
00:05:38.320 --> 00:05:41.919
In the most recent model card, Opus 4.8 talks about its SWE score.

s61
00:05:41.919 --> 00:05:44.720
It does not disclose this contamination.

s62
00:05:44.720 --> 00:05:48.160
We as an industry aren't really in the habit of doing those disclosures.

s63
00:05:48.160 --> 00:05:54.080
And so what that means is that as benchmarking consumers, we're just missing that information.

s64
00:05:54.720 --> 00:05:57.360
Reward hacking is also a big problem.

s65
00:05:57.360 --> 00:06:04.800
Reward hacking is basically when a model finds a lazy and creative way to meet the letter of the law, but not the spirit.

s66
00:06:04.800 --> 00:06:12.319
You need to think about designing your rewards as a adversarial process against this maximally lazy agent.

s67
00:06:12.319 --> 00:06:18.240
Gradient descent is basically like water flowing downhill looking for the path of least resistance.

s68
00:06:18.240 --> 00:06:23.039
And so your verifiers need to be robust to that.

s69
00:06:23.840 --> 00:06:29.759
Another key challenge is simply just not having the ambition to make a sophisticated enough benchmark.

s70
00:06:29.759 --> 00:06:35.840
Automation bench tests that agents are able to make tool calls in an enterprise environment.

s71
00:06:35.840 --> 00:06:40.080
The problem is that a lot of the verifiers are these hard-coded string matches.

s72
00:06:40.080 --> 00:06:46.639
And so you'll see it for things like phone numbers where there are many different acceptable phone number formats.

s73
00:06:46.639 --> 00:06:51.039
But this verifier just picks one and the prompt doesn't tell you which one it is.

s74
00:06:51.039 --> 00:06:55.600
So the result of this is that Haiku and Fable both score 20% on this task.

s75
00:06:55.600 --> 00:07:05.039
Haiku scores 20% because it makes a bunch of mistakes and Fable scores 20% because it gets it right 80% of the time but then just happens to pick different formats.

s76
00:07:05.039 --> 00:07:11.759
So if the benchmark task is not differentiating between Haiku and Fable, it's not a useful task.

s77
00:07:11.759 --> 00:07:18.880
And more broadly, in 2026, many of us in this room are looking towards AI that's about to remake entire industries.

s78
00:07:18.880 --> 00:07:24.160
And benchmarks are ideally our lighthouse on the horizon to let us know when that's coming.

s79
00:07:24.160 --> 00:07:31.520
And a simple hard-coded string match is just not going to do it to measure that sort of impact.

s80
00:07:32.319 --> 00:07:35.680
Another important aspect of a good benchmark is taste.

s81
00:07:35.680 --> 00:07:40.560
Perhaps it used to be the case that benchmarks were these dry academic, you know, questions.

s82
00:07:40.560 --> 00:07:41.440
and answer sets.

s83
00:07:41.440 --> 00:07:48.720
But nowadays, a benchmark is an artifact expressing what it's an aspirational artifact.

s84
00:07:48.720 --> 00:07:52.800
It's an expression of values of what you want your AI to do and how you want it to behave.

s85
00:07:52.800 --> 00:07:58.560
And so you need to have some product sense in this process, some sort of a sense of what you want the AI to do.

s86
00:07:58.560 --> 00:08:04.400
And that sense is unfortunately missing from ifal if has been cited on many model cards.

s87
00:08:04.400 --> 00:08:16.080
And the way it was constructed was taking a bunch of arbitrary prompts that no user has ever asked in earnest and mashing them up with a bunch of other prompts to create a prompt set.

s88
00:08:16.080 --> 00:08:23.759
The problem is that because no user actually has asked do not use any commas in your response or use the letter T at most once.

s89
00:08:23.759 --> 00:08:32.240
You have to believe for this to be useful, you have to believe that there's a generalization from this to actual things that users are going to ask.

s90
00:08:33.440 --> 00:08:40.159
If eval just happens also to have a bunch of prompts that are fully unsolvable due to having contradictory instructions.

s91
00:08:40.159 --> 00:08:45.839
So this one starts by saying repeat this response verbatim and it ends by saying translate this into Hindi.

s92
00:08:45.839 --> 00:08:49.279
Obviously you can't do both of those at once.

s93
00:08:49.279 --> 00:08:53.519
Here's one that says write a riddle that includes exactly one bullet point.

s94
00:08:53.519 --> 00:08:56.080
Make sure to include a few bullet points.

s95
00:08:56.080 --> 00:08:59.279
Again this is just fully impossible.

s96
00:08:59.519 --> 00:09:06.240
It uses a sentence splitter that does not align with how humans would actually split the sentences.

s97
00:09:06.720 --> 00:09:10.160
And a lot of the prompts are not fully verified.

s98
00:09:10.160 --> 00:09:11.920
So this one says write a story.

s99
00:09:11.920 --> 00:09:15.040
There's nothing in the verifier that checks that a story was written.

s100
00:09:15.040 --> 00:09:30.560
It just checks that the asky character I is not used more than once, which means that all of these responses get a full score, including response D. The way it gets a full score is by reward hacking and using the cerrillic eye character instead of the asy eye character.

s101
00:09:30.560 --> 00:09:38.480
If is totally fine with that another challenge is operational ability.

s102
00:09:38.480 --> 00:09:46.240
Making a big benchmark requires a lot of QC work and plenty of organizations just don't make that investment.

s103
00:09:46.240 --> 00:09:52.160
Apex is a rag benchmark where the agent is given files and then asked questions about them.

s104
00:09:52.160 --> 00:09:58.000
And in some instances, what's in the file and then what's expected in the rubric don't line up.

s105
00:09:58.000 --> 00:10:05.040
So an agent that does the thing that it's seeing in the ground truth is going to get a negative score.

s106
00:10:05.200 --> 00:10:17.040
And a lot of the data in Apex is seemingly synthetically generated because it's full of obvious placeholder values or dates or places that don't exist.

s107
00:10:17.040 --> 00:10:26.959
And so as a result, the model is more likely to develop eval awareness where it realizes that it's being tested which undermines the entire exercise.

s108
00:10:26.959 --> 00:10:34.480
It also just takes you out of distribution from actual real world data to something that is obviously fake.

s109
00:10:34.720 --> 00:10:40.480
So that's an overview of some of the key antiatterns that happen during benchmark creation.

s110
00:10:40.480 --> 00:10:51.120
But benchmaxing is a two-way process and there are all sorts of fun things that labs can do to benchmax and that's what we're going to talk about next.

s111
00:10:51.120 --> 00:11:03.519
So the the core value that we're all trying to get towards as human eval right AI exists to serve humans and so just having humans look at the responses and make ratings like that's what we care about.

s112
00:11:03.519 --> 00:11:05.519
The problem is that human eval is very expensive.

s113
00:11:05.519 --> 00:11:18.320
And so a lot of what benchmarks are doing is trying to get around that and you are trying to distill human preference into something more scalable and you're hoping you do that distillation in a way that's still sufficiently faithful to what human eval wants.

s114
00:11:18.320 --> 00:11:27.680
But what this means is that inevitably there is a point where you can keep hill climbing on a benchmark and the human eval stays flat.

s115
00:11:27.920 --> 00:11:35.040
And you can actually take it even further if you want where you keep hill climbing on a benchmark even as the human eval goes down.

s116
00:11:35.040 --> 00:11:44.160
But if for whatever reason you think this is necessary for marketing or we have sort of organizational politics or incentives that are demanding this that's how it can end up happening.

s117
00:11:44.160 --> 00:11:47.200
In this instance the prompt is what time is it?

s118
00:11:47.200 --> 00:11:50.079
And the response is absolutely deranged.

s119
00:11:50.079 --> 00:11:55.360
No human eval is ever going to choose this but El Marina puts it at the top of the leaderboard.

s120
00:11:55.360 --> 00:12:02.160
So again, you have this divergence and if you're trying to benchmax, you just cannot care about that.

s121
00:12:02.160 --> 00:12:11.279
Another thing you can do that I've heard stories of is you can actually hire a crowdsource army to vote for you in Elmarina since Elmarina basically does no filtering of their workforce.

s122
00:12:11.279 --> 00:12:15.440
And you might say, well, we anonym, you know, Elmarina anonymizes.

s123
00:12:15.440 --> 00:12:17.120
So how are they going to know who to vote for?

s124
00:12:17.120 --> 00:12:18.560
That's actually quite simple.

s125
00:12:18.560 --> 00:12:24.480
You have your model include a watermark that tells the crowd who to vote for.

s126
00:12:25.120 --> 00:12:34.399
There's also all sorts of things you can do with running your evals in conditions that are like not fully representative of the applesto apples comparison you're trying to make

s127
00:12:34.399 --> 00:12:45.680
and then not always being super transparent about those conditions in such a way that undermines the validity that the community is trying to interpret because they don't have that contextualizing

s128
00:12:45.680 --> 00:12:47.440
information.

s129
00:12:47.440 --> 00:12:55.600
This was a paper um again about Elmarina and talking about how some of the dynamics of how it's run lead to models overfitting on Elmarina.

s130
00:12:55.600 --> 00:13:02.720
Um in this instance, the specific chart we're seeing is that Meta tested 27 models without disclosing that it was doing so.

s131
00:13:02.720 --> 00:13:06.880
Um which you know distorts the results.

s132
00:13:09.839 --> 00:13:12.399
So how are we going to end benchmaxing?

s133
00:13:12.399 --> 00:13:18.079
We need to hold the benchmark industry and the labs to a higher standard.

s134
00:13:18.560 --> 00:13:24.320
The first thing we need to do when making a good benchmark is start with great human experts.

s135
00:13:24.320 --> 00:13:32.000
And those experts inform everything that is downstream from what types of tasks are we going to have the agent do?

s136
00:13:32.000 --> 00:13:34.399
How is success measured?

s137
00:13:34.399 --> 00:13:36.240
What are the input files that agents are given?

s138
00:13:36.240 --> 00:13:38.079
What are the tools that they're given?

s139
00:13:38.079 --> 00:13:39.839
But we also do need that product sense.

s140
00:13:39.839 --> 00:13:42.240
So imagine you're making a medical benchmark.

s141
00:13:42.240 --> 00:13:50.959
It's not enough to have doctors who can answer specific medical questions because if you're trying to test how ready are we for agents to be deployed into hospitals.

s142
00:13:50.959 --> 00:14:01.440
You also need someone with the business sense to know what's the regulatory environment, what's the legal requirements because that is going to impact what types of tasks you're trying to have the AI solve.

s143
00:14:01.519 --> 00:14:10.320
You need high fidelity input data which is best done by going out and getting it from the real world, having actual people create this data.

s144
00:14:10.320 --> 00:14:14.880
Synthetic approaches are possible, but it is very very hard to do it reliably.

s145
00:14:14.880 --> 00:14:16.320
The tools need to actually work.

s146
00:14:16.320 --> 00:14:19.440
A lot of benchmarks have tools that are buggy in various ways.

s147
00:14:19.440 --> 00:14:25.600
And unless you're intentionally making a benchmark about buggy tools, this just introduces noise.

s148
00:14:26.639 --> 00:14:28.880
You need verifiers that are fully aligned with the prompts.

s149
00:14:28.880 --> 00:14:30.480
And this is a two-way alignment.

s150
00:14:30.480 --> 00:14:35.760
So the verifiers need to be verifying everything the prompt asks for.

s151
00:14:35.760 --> 00:14:39.920
And everything the prompt asks for needs to be covered by the verifiers.

s152
00:14:39.920 --> 00:14:46.720
And if you get either side of those two misaligned, then it's going to be unfair to models and you're introducing random noise.

s153
00:14:46.720 --> 00:14:53.920
You need to thoroughly QC everything and you need to have a private hold out set so you don't get contaminated.

s154
00:14:54.320 --> 00:15:03.199
And if you do all this right, then you'll avoid what often happens with benchmarks, which is when labs get to like 80% and say, "Okay, this is saturated."

s155
00:15:03.199 --> 00:15:10.800
And I used to think that saturation was just them saying again we don't think training on this further is going to increase real world value.

s156
00:15:10.800 --> 00:15:18.079
And it often does mean that but it can mean that because the lab is saying we realize 20% of these tasks are broken.

s157
00:15:18.079 --> 00:15:25.920
But the problem is that as you're hill climbing you don't know what 20% are broken until you solve all the others.

s158
00:15:25.920 --> 00:15:28.800
And so as a result you have a lot of noise.

s159
00:15:28.800 --> 00:15:40.720
And if that 20% of broken tasks is randomly but in a biased way assigning the rewards, it's going to really distort the model relative ranking you're trying to get.

s160
00:15:41.360 --> 00:15:45.920
So at Serge, we created a benchmark called Hemingway bench to measure writing.

s161
00:15:45.920 --> 00:15:59.839
There have been a number of writing benchmarks that use various mechanical means to try to assess writing quality, but we believe that writing is just too rich and deep and nuanced and frankly human of an activity to measure

s162
00:15:59.839 --> 00:16:05.759
with mechanical benchmarks and LM as a judge doesn't really work either because LLMs don't have good taste in writing.

s163
00:16:05.759 --> 00:16:10.639
Again, this is sort of the you can't expand the frontier from within the frontier situation.

s164
00:16:10.639 --> 00:16:26.560
So what we've done is we've just created a workforce of thousands of professional writers in various domains, technical writers, poets, journalists, editors, and we just have them do blind model comparisons and then we create this leaderboard

s165
00:16:26.560 --> 00:16:28.480
and it is quite expensive, right?

s166
00:16:28.480 --> 00:16:29.600
Human eval is very expensive.

s167
00:16:29.600 --> 00:16:32.720
Getting the time of these professionals is quite expensive.

s168
00:16:32.720 --> 00:16:37.920
But again, our goal is to maximize quality, not to minimize costs.

s169
00:16:37.920 --> 00:16:49.040
So in conclusion, benchmaxing is the exploitation of benchmark misalignments between human preference, but we can do better and we can hold the industry to a higher standard.

s170
00:16:49.040 --> 00:16:55.199
Both the people making the benchmarks like myself and the people who are reporting on the benchmarks.

s171
00:16:55.199 --> 00:17:00.560
And if you'd like to be a part of that, of course, obligatory pitch at Serge, we're hiring for basically all aspects of that.

s172
00:17:00.560 --> 00:17:04.480
Uh and if you'd like more spicy takes from me, uh please follow my substack.

s173
00:17:04.480 --> 00:17:06.959
Thank you very much.
