WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=yqF6XhzbWBk

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/yqF6XhzbWBk.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:12.640 --> 00:00:17.080
This is a clinical note an AI wrote from a real consultation.

s3
00:00:17.080 --> 00:00:19.320
Take a few seconds and read it.

s4
00:00:19.320 --> 00:00:21.600
It reads like a routine headache.

s5
00:00:21.600 --> 00:00:27.200
A new headache, likely tension type, take some paracetamol, come back if it doesn't settle.

s6
00:00:27.200 --> 00:00:30.640
Looks completely fine, doesn't it?

s7
00:00:30.800 --> 00:00:32.960
Here's what's missing.

s8
00:00:32.960 --> 00:00:36.760
In the room, she also mentioned her jaw aches when she chews.

s9
00:00:36.760 --> 00:00:42.440
A new headache, over 50 with jaw pain on chewing, that's giant cell arthritis.

s10
00:00:42.440 --> 00:00:45.000
And untreated, it can take her sight within days.

s11
00:00:45.000 --> 00:00:48.200
It's a same day start steroids now emergency.

s12
00:00:48.200 --> 00:00:50.720
And that one line, it never made it into the note.

s13
00:00:50.720 --> 00:00:52.920
On the page, it's a paracetamol headache.

s14
00:00:52.920 --> 00:00:55.960
And nothing in the note is technically wrong.

s15
00:00:55.960 --> 00:00:59.680
It's the dangerous part is what isn't there.

s16
00:00:59.680 --> 00:01:01.600
And so that's what I'm going to talk about today.

s17
00:01:01.600 --> 00:01:07.280
The dangerous failures are often the ones that actually look completely fine.

s18
00:01:07.600 --> 00:01:09.520
Firstly, who am I?

s19
00:01:09.520 --> 00:01:18.600
I'm Seb, medical doctor by background, and now I'm Composure, where we build AI evaluation systems for high-stakes domains.

s20
00:01:19.800 --> 00:01:25.400
So, that one was a subtle kind of error, but sometimes it's not subtle at all.

s21
00:01:25.400 --> 00:01:29.800
A man in his 20s sees his GP for a sore throat, tonsillitis.

s22
00:01:29.800 --> 00:01:31.640
The AI writes that up.

s23
00:01:31.640 --> 00:01:38.600
It gets him chest pain, suspected angina, diabetes medications he's never taken, and an address for a hospital that doesn't exist.

s24
00:01:38.600 --> 00:01:41.160
And I I really like the LLM for this one.

s25
00:01:41.160 --> 00:01:44.040
I think it's it's a good attempt at hospital name.

s26
00:01:44.040 --> 00:01:50.080
Um and weeks later, he's invited to diabetic eye screening for diabetes he doesn't have.

s27
00:01:50.080 --> 00:01:54.000
That's genuinely a real case that happened recently.

s28
00:01:54.000 --> 00:02:06.640
Obviously, these kind of crazy ones someone notices, but it's those quiet ones that sit in the record uncalled that are the most challenging and can actually do a lot more damage.

s29
00:02:06.920 --> 00:02:08.800
And they're not rare at all.

s30
00:02:08.800 --> 00:02:17.440
In the largest real-world study of these notes, about 1 in 20 carried an error that was serious enough that it could cause significant harm to the patient.

s31
00:02:17.440 --> 00:02:18.320
1 in 20.

s32
00:02:18.320 --> 00:02:23.160
That's not theoretical in testing, that's in production on real patients.

s33
00:02:23.160 --> 00:02:25.280
And that's only the serious the ones.

s34
00:02:25.280 --> 00:02:33.160
If you widen that lens to all errors, nearly 1 in 5 had an important omission and more than 1 in 10 had a hallucination.

s35
00:02:33.760 --> 00:02:37.959
And AI is being deployed at scale across healthcare fast.

s36
00:02:37.959 --> 00:02:43.680
Ambient Scribes are one of the leading cases, already in about a third of US practices and climbing.

s37
00:02:43.680 --> 00:02:48.280
Physician AI use doubled last year and none of this is tracked.

s38
00:02:48.280 --> 00:02:52.520
So, for most of these systems, there's no adverse event reporting at all.

s39
00:02:52.520 --> 00:02:55.880
The errors never show up as incidents, they just sit in the record.

s40
00:02:55.880 --> 00:03:03.720
So, errors this common that are going unseen is is quite hard for me to believe that it's not already affecting patients.

s41
00:03:03.720 --> 00:03:09.239
It's not that we checked and it's fine, it's that we're flying blind.

s42
00:03:09.360 --> 00:03:14.680
And this isn't just a healthcare problem, it's every high-stakes use of AI.

s43
00:03:14.680 --> 00:03:20.400
Healthcare shows it more viscerally because here being confidently wrong can be life and death.

s44
00:03:20.400 --> 00:03:25.200
But, everything I show you can map straight back onto other domains as well.

s45
00:03:25.200 --> 00:03:27.000
So, here's what I want to do.

s46
00:03:27.000 --> 00:03:39.480
I'm going to show you what exactly is going wrong, why it's going wrong, why the systems we built to catch it don't work, and a suggestion at how maybe we can start to fix that.

s47
00:03:39.480 --> 00:03:41.760
So, first, what's going wrong and why?

s48
00:03:41.760 --> 00:03:44.640
So, LLMs are getting good, obviously.

s49
00:03:44.640 --> 00:03:47.160
They don't make stupid mistakes anymore most of the time.

s50
00:03:47.160 --> 00:03:48.800
So, it's not about dumb errors.

s51
00:03:48.800 --> 00:03:54.160
Everything here came out of three of the best production Ambient scribes on the market.

s52
00:03:54.160 --> 00:03:55.600
Ones that we all know.

s53
00:03:55.600 --> 00:03:58.400
We generated a load of notes across them last week.

s54
00:03:58.400 --> 00:04:02.120
And this is exactly what's going on right now.

s55
00:04:02.120 --> 00:04:03.640
This is every failure we found.

s56
00:04:03.640 --> 00:04:06.360
Each dot is an error colored by type.

s57
00:04:06.360 --> 00:04:07.959
Left to right, how much it matters.

s58
00:04:07.959 --> 00:04:11.680
Bottom to top, whether a strong automated check catches it.

s59
00:04:11.680 --> 00:04:14.440
And that split is the point.

s60
00:04:14.440 --> 00:04:16.440
A handful up top get caught.

s61
00:04:16.440 --> 00:04:18.880
But almost everything sits below the line.

s62
00:04:18.880 --> 00:04:22.840
The ones I care about most are these on the bottom right.

s63
00:04:22.840 --> 00:04:26.120
The high stakes and missed ones.

s64
00:04:26.120 --> 00:04:29.560
Let me show you what a couple of those looks like.

s65
00:04:29.919 --> 00:04:33.000
So, a woman comes in with a headache.

s66
00:04:33.000 --> 00:04:37.360
Doctor asks, "Did it come on suddenly or build up gradually?"

s67
00:04:37.360 --> 00:04:38.560
She says she doesn't know.

s68
00:04:38.560 --> 00:04:39.760
It just happened.

s69
00:04:39.760 --> 00:04:42.760
The note records that as abrupt sudden onset.

s70
00:04:42.760 --> 00:04:44.720
And sudden onset is a red flag.

s71
00:04:44.720 --> 00:04:50.800
You can see why it just happened could maybe be interpreted and inferred as abrupt onset.

s72
00:04:50.800 --> 00:04:53.720
But, that's a feature that points to a bleed on the brain.

s73
00:04:53.720 --> 00:04:54.560
She never said it.

s74
00:04:54.560 --> 00:04:55.880
The model decided it.

s75
00:04:55.880 --> 00:05:00.040
And now that one word drives the whole workup.

s76
00:05:00.240 --> 00:05:01.640
Here's another.

s77
00:05:01.640 --> 00:05:03.600
Doctor suggests running some tests.

s78
00:05:03.600 --> 00:05:07.320
Patient says, "Can we just try try antibiotics instead?"

s79
00:05:07.320 --> 00:05:11.680
They agree, hold off on the tests, treat, and see how it goes.

s80
00:05:11.680 --> 00:05:13.480
Note records the opposite.

s81
00:05:13.480 --> 00:05:15.320
Arrange tests today.

s82
00:05:15.320 --> 00:05:19.080
It kept the plan that they talked out of, not the one they chose.

s83
00:05:19.080 --> 00:05:23.960
Every line in the note reads fine because it's not really a hallucination at all.

s84
00:05:23.960 --> 00:05:24.560
It's not wrong.

s85
00:05:24.560 --> 00:05:26.480
It was there in the original.

s86
00:05:26.480 --> 00:05:29.560
But, it's just not what they ended up deciding.

s87
00:05:30.120 --> 00:05:33.720
So, why are these happening?

s88
00:05:33.720 --> 00:05:38.800
There's, you know, in ambient scribes, there's first transcription and then generation.

s89
00:05:38.800 --> 00:05:42.400
A lot of it does happen on the transcription layer.

s90
00:05:42.400 --> 00:05:45.440
It can be words misheard for their sound-alikes.

s91
00:05:45.440 --> 00:05:48.560
So, Humalog heard Humulin.

s92
00:05:48.560 --> 00:05:53.640
Two insulins on completely different timelines, so swapping them could crash a blood sugar.

s93
00:05:53.640 --> 00:05:57.840
Hyperthyroidism becomes hypothyroidism, the opposite condition.

s94
00:05:57.840 --> 00:06:04.760
Or a drop to no on uh no evidence of cancer that becomes evidence of cancer.

s95
00:06:04.760 --> 00:06:07.800
So, these these are really hard problems, and they are common.

s96
00:06:07.800 --> 00:06:14.400
Not the ones I'm going to focus on, because most of what goes wrong is actually even with a perfect transcript.

s97
00:06:14.400 --> 00:06:19.760
It's the model reading the words correctly and still doing one of three things.

s98
00:06:19.760 --> 00:06:28.200
Either it adds something that was never said, it changes something that was, or it omits something that should be there.

s99
00:06:28.480 --> 00:06:32.040
Now, the blatant version of each of these is is really easy to catch.

s100
00:06:32.040 --> 00:06:37.120
The hard part in all three is the same.

s101
00:06:37.120 --> 00:06:41.720
It's telling whether that thing that was added or changed or dropped actually matters.

s102
00:06:41.720 --> 00:06:46.320
It's detecting that slight over inference versus the dangerous fabrication.

s103
00:06:46.320 --> 00:06:49.640
The harmless rephrase versus the meaningful edit.

s104
00:06:49.640 --> 00:06:53.160
A dropped line of small talk versus a dropped allergy.

s105
00:06:53.160 --> 00:06:59.200
So, the ones that matter slip through along with all of the ones that don't.

s106
00:06:59.960 --> 00:07:05.800
That call which different matters is taste, effectively.

s107
00:07:05.800 --> 00:07:09.400
Not aesthetic taste, but essentially judgment.

s108
00:07:09.400 --> 00:07:15.480
It's It's whether in this context a missed allergy might kill someone or is not important.

s109
00:07:15.480 --> 00:07:18.800
And I think there's there's three properties that really matter about this.

s110
00:07:18.800 --> 00:07:23.760
It's tacit, so your domain experts have it, but they can't fully write it down.

s111
00:07:23.760 --> 00:07:28.360
It's contextual, so the same detail is critical in one note, noise in the next.

s112
00:07:28.360 --> 00:07:29.760
And it's moving.

s113
00:07:29.760 --> 00:07:35.919
The model changes, guidelines change, two good doctors disagree, different hospitals have different definitions.

s114
00:07:35.919 --> 00:07:39.400
So, there's no fixed target to write down.

s115
00:07:39.600 --> 00:07:42.720
And so, the model knows the facts, ultimately.

s116
00:07:42.720 --> 00:07:49.840
They're extraordinarily capable, but what they lack is a sense of what matters here, for this specific example.

s117
00:07:49.840 --> 00:07:53.040
And that's why even brilliant models make these mistakes.

s118
00:07:53.040 --> 00:07:58.120
So, one natural move, you're never going to make that generator perfect.

s119
00:07:58.120 --> 00:07:59.240
Generator is cheap.

s120
00:07:59.240 --> 00:08:02.600
Generation is cheap, so stop fixing it at the source.

s121
00:08:02.600 --> 00:08:06.680
Let it write, put a checker after it, pass only what clears the bar.

s122
00:08:06.680 --> 00:08:09.120
And that checker should be the easier job.

s123
00:08:09.120 --> 00:08:15.280
The generator has to get everything right and it pay attention to lots of varying instructions.

s124
00:08:15.280 --> 00:08:19.080
Whereas the checker only has to find the one thing that's wrong and just focus on that task.

s125
00:08:19.080 --> 00:08:23.400
You can also give it more time, more tokens, the exact failure modes to hunt for.

s126
00:08:23.400 --> 00:08:25.880
Evaluation should be easier than generation.

s127
00:08:25.880 --> 00:08:29.600
It's the asymmetry of verification, verifies law.

s128
00:08:29.600 --> 00:08:35.280
That's why AI is raced ahead anyway, you can cheaply check the answer, maths and code.

s129
00:08:35.280 --> 00:08:38.400
And doing this is exactly what the best teams do.

s130
00:08:38.400 --> 00:08:40.919
They put a lot of energy into evaluation.

s131
00:08:40.919 --> 00:08:50.360
It starts with the gold standard, which is expert humans reviewing notes, which obviously works offline, but you can't put a human on every note in production.

s132
00:08:50.360 --> 00:08:52.600
So, they automate it.

s133
00:08:52.600 --> 00:09:05.360
They build a serious system and some of the best versions of this that I've seen are you take the transcript and the note and context, put in front of the judge,

s134
00:09:05.360 --> 00:09:10.120
a detailed rubric for faithfulness with worked, pass and fail examples.

s135
00:09:10.120 --> 00:09:13.520
The rubric maybe auto-optimized with GPA or something like that.

s136
00:09:13.520 --> 00:09:20.680
Maybe you have some deterministic NLP to sort of count up medical concepts that are differing between the two.

s137
00:09:20.680 --> 00:09:22.360
That's a powerful system.

s138
00:09:22.360 --> 00:09:29.200
And yet, I pulled all of those errors earlier out of Ambient Scribes in an afternoon.

s139
00:09:29.200 --> 00:09:34.240
So, if the evaluation is this good, how are these errors still getting through?

s140
00:09:34.240 --> 00:09:39.960
So, I built this system and ran those same notes through it.

s141
00:09:39.960 --> 00:09:42.880
And it scored most of them fine.

s142
00:09:42.880 --> 00:09:46.886
It flagged a handful of them and signed off the rest.

s143
00:09:46.886 --> 00:09:47.720
[clears throat]

s144
00:09:47.720 --> 00:09:52.880
But one in five of those clean passes still had some sort of serious error buried in it.

s145
00:09:52.880 --> 00:09:55.720
And often that was an omission.

s146
00:09:55.720 --> 00:09:59.840
The things that should have been there and actually quietly weren't.

s147
00:09:59.840 --> 00:10:06.800
And that's the best version of a judge I've seen in a lot of teams and it waved them through.

s148
00:10:06.800 --> 00:10:07.720
Why did it do that?

s149
00:10:07.720 --> 00:10:09.600
It's not stupid.

s150
00:10:09.600 --> 00:10:16.160
It's a frontier model, serious engineering behind it, more than clever enough to read the whole encounter and catch every obvious error.

s151
00:10:16.160 --> 00:10:17.880
And it's not blind, either.

s152
00:10:17.880 --> 00:10:19.320
And and that's part of the trap.

s153
00:10:19.320 --> 00:10:30.440
If you take a note that says start amoxicillin, when the real decision was actually to wait and see, it's faithful to the words, amoxicillin did come up, but it's a lie about the intent.

s154
00:10:30.440 --> 00:10:34.200
A good judge might catch that, might.

s155
00:10:34.200 --> 00:10:42.480
But whether it flags that versus the other doesn't have other things that it could comment on depends on it knowing what decision matters most.

s156
00:10:42.480 --> 00:10:48.600
And so it's not blind, it just can't tell what counts, essentially.

s157
00:10:48.600 --> 00:10:59.600
So the note passes confidently and you put a judge like that in front of your system, you've not added a safety net, you've added a second silent failure that just nods along with the first.

s158
00:10:59.600 --> 00:11:01.680
And here's the root of it.

s159
00:11:01.680 --> 00:11:07.839
So in math or code, the verifier comes with free, a unit test, a compiler.

s160
00:11:08.440 --> 00:11:12.560
But for is this note safe and complete, there's no unit test.

s161
00:11:12.560 --> 00:11:22.640
You have to build the verifier yourself and verification is only easier than generation for the easy bit, i.e. spot the difference between transcription note.

s162
00:11:22.640 --> 00:11:23.320
But that's not the hard bit.

s163
00:11:23.320 --> 00:11:26.600
The hard bit is knowing of all those differences you've seen, which matter.

s164
00:11:26.600 --> 00:11:31.120
And that's harder than writing that plausibly good general note in the first place.

s165
00:11:31.120 --> 00:11:36.400
Because that standard of good was never written down anywhere that the judge can read it.

s166
00:11:36.400 --> 00:11:40.920
A rubric that you pre-specify is only the taste you could write down.

s167
00:11:40.920 --> 00:11:43.920
The taste that matters is the part that you couldn't.

s168
00:11:43.920 --> 00:11:48.960
And so here's here's a bit more detail on what what matters looks like.

s169
00:11:48.960 --> 00:11:55.240
Two patients, both with blood in their urine, both notes dropped the same kind of line where they'd been on holiday.

s170
00:11:55.240 --> 00:11:57.640
One had been to France, the other to Lake Malawi.

s171
00:11:57.640 --> 00:12:01.800
Same English emission, same shape, same mistake.

s172
00:12:01.960 --> 00:12:10.320
Well, not really, because blood in the urine obviously warranted away and you're going to have to investigate it, but the France trip is irrelevant.

s173
00:12:10.320 --> 00:12:12.840
The Lake Malawi trip is the diagnosis.

s174
00:12:12.840 --> 00:12:19.680
Fresh water in sub-Saharan Africa means schistosomiasis until proven otherwise and it completely changes what the management plan is.

s175
00:12:19.680 --> 00:12:24.840
So that same dropped line in one note is pure noise, in the other it's the answer.

s176
00:12:24.840 --> 00:12:29.000
And which one it is, you simply just can't write all of that down in advance.

s177
00:12:29.000 --> 00:12:38.120
So if you can't write it down, you can't write taste down, how do you get that into your evaluator and your whole application system?

s178
00:12:38.120 --> 00:12:41.400
Well, we've answered a version of this before.

s179
00:12:41.400 --> 00:12:44.839
RLHF exists because you can't write the reward function for good.

s180
00:12:44.839 --> 00:12:47.120
You learn it from examples by showing it.

s181
00:12:47.120 --> 00:12:50.839
The only question is where you keep what you've learned.

s182
00:12:50.839 --> 00:12:52.640
And there's three places.

s183
00:12:52.640 --> 00:12:57.960
You can either specify it up front, you can stuff the prompt, write the perfect rubric.

s184
00:12:57.960 --> 00:13:01.520
We just watched that fail essentially.

s185
00:13:01.520 --> 00:13:13.560
You can bake into the weights, fine-tuning or continual learning, but for a standard that's still moving and a score that has to be explainable, the weights, I think, are the wrong place to keep that.

s186
00:13:13.560 --> 00:13:19.400
They go stale, they can't tell you why and you can't change them without a retrain.

s187
00:13:19.400 --> 00:13:25.400
So there's the third option, which I'll show you, which is you essentially just keep the taste as the examples themselves.

s188
00:13:25.400 --> 00:13:35.280
Past judgments, expert corrections, references, and for each output, you retrieve the ones that bear on it into the judges context, add one and it's live on the next call.

s189
00:13:35.280 --> 00:13:38.400
You can point at exactly what moved the score.

s190
00:13:38.400 --> 00:13:43.960
For this problem, it's both better and also cheaper to do.

s191
00:13:44.480 --> 00:13:48.480
So, that's the way to do that is one repeating loop, three steps.

s192
00:13:48.480 --> 00:13:59.160
Discover the failure modes from real outputs, capture how your experts judge them, calibrate every output against that, and when the standard moves, the loop moves with it.

s193
00:13:59.160 --> 00:14:00.720
So, in more detail, discover.

s194
00:14:00.720 --> 00:14:03.680
You don't write that rubric in a vacuum.

s195
00:14:03.680 --> 00:14:06.520
You have to put the system in production and look at the real outputs.

s196
00:14:06.520 --> 00:14:09.520
Cluster what goes wrong and the failure modes surface on their own.

s197
00:14:09.520 --> 00:14:10.760
You name them.

s198
00:14:10.760 --> 00:14:12.360
This is your failure mode ontology.

s199
00:14:12.360 --> 00:14:15.760
Discover from your data, not guess on a whiteboard.

s200
00:14:15.760 --> 00:14:16.960
And you can't shortcut it.

s201
00:14:16.960 --> 00:14:25.080
The ways that a real system goes wrong are effectively unbounded and synthetic test cases only cover the failures you already imagined.

s202
00:14:25.080 --> 00:14:27.880
The ones that hurt you are often the ones that you didn't.

s203
00:14:27.880 --> 00:14:30.040
And you'll only find those in real outputs.

s204
00:14:30.040 --> 00:14:39.320
So, this ontology is your map, what to capture judgment on, what to retrieve against, including the failures that you never thought to check for.

s205
00:14:39.320 --> 00:14:42.200
After that, it's capture and then calibrate.

s206
00:14:42.200 --> 00:14:47.320
So, those discovered modes, they're not a checklist that the judge runs, but they organize everything.

s207
00:14:47.320 --> 00:14:55.400
What you What you ask your experts about, how you index the cases that you'll retrieve, and capturing is a simple part.

s208
00:14:55.400 --> 00:14:57.560
You put real outputs in front of your experts.

s209
00:14:57.560 --> 00:15:00.360
Clinicians spend a focused few hours leaving comments.

s210
00:15:00.360 --> 00:15:04.160
A session doesn't have to be a month-long labeling project to start with.

s211
00:15:04.160 --> 00:15:06.080
And you collect their judgment.

s212
00:15:06.080 --> 00:15:08.440
Not just a score, but the reasoning and corrections.

s213
00:15:08.440 --> 00:15:12.640
And over time, you build up that record of how your experts actually judge.

s214
00:15:12.640 --> 00:15:14.040
You then calibrate.

s215
00:15:14.040 --> 00:15:18.360
That's the the the generic part of this you can write down once easily.

s216
00:15:18.360 --> 00:15:22.000
For example, be faithful or don't drop anything important.

s217
00:15:22.000 --> 00:15:25.960
But what you can't write down is what counts as a serious miss for this specific note.

s218
00:15:25.960 --> 00:15:27.360
That's contextual.

s219
00:15:27.360 --> 00:15:29.600
And it shifts from note to note.

s220
00:15:29.600 --> 00:15:34.160
So, what we recommend is you assemble that on the fly.

s221
00:15:34.160 --> 00:15:38.600
For each output, your judging agent pulls in everything that bears on this one case.

s222
00:15:38.600 --> 00:15:47.680
It's memory of the most similar outputs that it's judged before and how they scored, the expert corrections that apply, the reference documents and guidelines.

s223
00:15:47.880 --> 00:15:50.520
Just context engineering per output.

s224
00:15:50.520 --> 00:16:03.520
And crucially not just one pre-specified rubric in a vacuum, and not a model that you have to retrain every week, but a full sort of case-specific standard assembled for this output.

s225
00:16:03.520 --> 00:16:05.560
And it's a loop as well.

s226
00:16:05.560 --> 00:16:09.120
Every output you judge, every correction, sharpens the next.

s227
00:16:09.120 --> 00:16:15.120
And when a brand new failure mode appears, Discovery surface it, and it flows straight back in.

s228
00:16:15.120 --> 00:16:25.839
And so, to make that a little bit more concrete, that headache that I opened with, the one that was really a possible blindness emergency, here's the kinds of things that you would want to pull in for that note.

s229
00:16:25.839 --> 00:16:33.320
The nearest cases that your experts have judged, not this exact patient, but the same shape, maybe a red flag filed as routine.

s230
00:16:33.320 --> 00:16:44.680
Uh the corrections that apply, like a new headache over 50, um suggests something that you need to check red flags on, and some criteria and guidelines, and you pull

s231
00:16:44.680 --> 00:16:46.720
all of that in.

s232
00:16:46.720 --> 00:16:48.200
It hasn't memorized this case.

s233
00:16:48.200 --> 00:16:50.040
It's a capable model.

s234
00:16:50.040 --> 00:16:55.040
And handed the right context to reason from, held against that, the dropped red flag stands out.

s235
00:16:55.040 --> 00:16:56.280
It was never actually hard to catch.

s236
00:16:56.280 --> 00:16:59.200
It just didn't know what mattered.

s237
00:16:59.200 --> 00:17:11.839
And so, if you take that same data set generated notes from the start and pass it through these three judging systems, the first, a strong off-the-shelf judge with a rubric frontier model,

s238
00:17:11.839 --> 00:17:17.079
um it's better than a coin flip, but it misses most of what matters.

s239
00:17:17.079 --> 00:17:28.920
The second, that sort of serious system that we talked about before, rubric, deeper, maybe some detona stick checks, better again, but still missing quite a lot of what counts.

s240
00:17:28.920 --> 00:17:39.840
The third, the judge running this loop, discovered failure modes, calibrated for output against what experts judged, is performing a lot better on this specific data set.

s241
00:17:39.840 --> 00:17:41.200
Same notes.

s242
00:17:41.200 --> 00:17:43.680
The only thing that changes is what the judge was shown.

s243
00:17:43.680 --> 00:17:48.480
And the difference here, it's not more compute or a better prompt, it's that the first two fight taste and lose.

s244
00:17:48.480 --> 00:17:53.160
They guess the criteria, they freeze one standard, and they go stale.

s245
00:17:53.160 --> 00:17:55.960
This repeating evolving loop does the opposite.

s246
00:17:55.960 --> 00:18:01.520
It discovers the modes, fits the standard to each mode, and keeps learning.

s247
00:18:01.920 --> 00:18:12.360
So, you might not write chemical notes, but if you ship anything where being confidently wrong has a cost, the contract review that misses the clauses that change the deal,

s248
00:18:12.360 --> 00:18:17.960
the support agent that promises a refund you don't offer, the same thing is true for all of those.

s249
00:18:17.960 --> 00:18:23.440
It's watched, if at all, by a judge with no taste for what matters in your domain.

s250
00:18:23.440 --> 00:18:26.120
So, three things.

s251
00:18:26.120 --> 00:18:29.720
Discover your failure modes from real outputs, don't guess them.

s252
00:18:29.720 --> 00:18:34.200
Capture your experts' judgment on them, the standard that they can't write down.

s253
00:18:34.200 --> 00:18:40.920
Calibrate every output against the cases that they've already judged, not a static rubric, not a retrained model.

s254
00:18:40.920 --> 00:18:42.920
Then keep that loop running.

s255
00:18:42.920 --> 00:18:50.400
And if you take one thing away, easiest place to start is your experts leaving free-form comments on real outputs.

s256
00:18:50.400 --> 00:18:54.880
That's the real That's the raw material for everything else.

s257
00:18:54.920 --> 00:19:00.720
Your judge can verify anything that you write down in advance, but the standard of good never could be.

s258
00:19:00.720 --> 00:19:07.080
And so, stop trying to write it all down in advance, and just start capturing it case by case, and evolving it.

s259
00:19:07.080 --> 00:19:11.720
That's why evaluation can't be a thing you build once and freeze.

s260
00:19:11.720 --> 00:19:15.600
The standard it checks against doesn't exist on paper.

s261
00:19:15.600 --> 00:19:22.000
It has to be discovered from real outputs captured from the people who hold it and kept alive as it moves.

s262
00:19:22.000 --> 00:19:27.240
Evaluation isn't something you have, it's something that you do continuously over time.

s263
00:19:27.240 --> 00:19:29.188
Thank you.

s264
00:19:29.188 --> 00:19:31.188
[applause]
