WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=dQ-_i1tZiws

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/dQ-_i1tZiws.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:13.519 --> 00:00:14.400
Hello everyone.

s3
00:00:14.400 --> 00:00:18.560
Um, this is a practitioner report uh from real production work.

s4
00:00:18.560 --> 00:00:21.520
So, let's get into it.

s5
00:00:21.520 --> 00:00:25.519
Um, I'll skip the generic uh yet another loop agent intro.

s6
00:00:25.519 --> 00:00:30.080
This is about the hard part most agent demos skip.

s7
00:00:30.080 --> 00:00:37.280
and about turning messy operational knowledge into something an agent can execute safely.

s8
00:00:37.280 --> 00:00:47.360
This comes from real work uh in my company I'm working for supporting global shipping operations and grounded in production.

s9
00:00:48.399 --> 00:00:56.640
On paper it's uh one workflow usually but uh in reality every shipment is an orchestration of many parallel state machines.

s10
00:00:56.640 --> 00:01:05.280
While they agree the happy paths work the moment one drifts you get exception work.

s11
00:01:06.240 --> 00:01:10.880
The easy majority is already automated in many companies.

s12
00:01:10.880 --> 00:01:18.479
What's left is the long tail and more exceptions than system built uh to handle them.

s13
00:01:18.479 --> 00:01:22.560
That tail is uh the expensive part.

s14
00:01:23.200 --> 00:01:25.280
And then there's my favorite category.

s15
00:01:25.280 --> 00:01:30.320
And it comes with a special uh plate here.

s16
00:01:30.320 --> 00:01:35.119
See for EI builder dreams and their laptops.

s17
00:01:35.119 --> 00:01:40.799
This what you can find outside of AI bubble in San Francisco.

s18
00:01:42.240 --> 00:01:47.280
The signal process uh depends on many systems being coherent at once.

s19
00:01:47.280 --> 00:02:00.079
If any step uh can't complete the happy path breaks and then it takes expert uh archist expert orchestration across uh multiple incomplete systems.

s20
00:02:00.399 --> 00:02:06.399
All these uh variations um path pathways should be captured in SOPs.

s21
00:02:06.399 --> 00:02:10.640
SOPs is a standard operating procedure common and regulated industries.

s22
00:02:10.640 --> 00:02:16.160
So an expert and the model read them uh the same way.

s23
00:02:16.959 --> 00:02:19.040
That gap is the hard part.

s24
00:02:19.040 --> 00:02:30.800
Stable intent detection tool calls you can guarantee are safe integrating with legacy back ends and results evaluated with experts.

s25
00:02:32.319 --> 00:02:34.959
Uh I call this uh tribal dungeons.

s26
00:02:34.959 --> 00:02:43.200
Uh the knowledge exists but not in a form uh agent can execute and you can safely run a process.

s27
00:02:43.200 --> 00:02:44.560
You can't safely run a process.

s28
00:02:44.560 --> 00:02:59.680
The organization cannot represent standard legacy SOPs [clears throat] bunch of bunch of screenshots organized in sequence and but screenshots not uh a process.

s29
00:02:59.680 --> 00:03:03.360
A legacy SOPs explain what a person sees and clicks.

s30
00:03:03.360 --> 00:03:18.560
And an agent SOP needs a more complex uh setup, preconditions, uh decisions, identifiers, back end calls, validation, recovery, and evidence of uh successful execution.

s31
00:03:21.040 --> 00:03:25.280
Experts own the what, agents own the how.

s32
00:03:25.280 --> 00:03:28.159
And exception becomes a guardrail.

s33
00:03:28.159 --> 00:03:36.959
Most of the effort is the translation and negotiation between them to align on common sense.

s34
00:03:38.319 --> 00:03:49.840
Three parts here um in this architecture it's SOP memory uh organized as SOP corpus execution runtime and theme feedback capture.

s35
00:03:49.840 --> 00:03:52.000
The agent loop is not the system.

s36
00:03:52.000 --> 00:03:58.640
The refining loop around the agent is the system and it's the most complex part.

s37
00:03:58.640 --> 00:04:00.799
Oh, sorry SAP is okay.

s38
00:04:00.799 --> 00:04:03.280
It's this slide for UK.

s39
00:04:03.280 --> 00:04:05.040
This is correct one.

s40
00:04:05.040 --> 00:04:19.680
So and it's good illustration why the the same thing is means different and uh describing differently in different countries and it's creating a lot of variations between each country

s41
00:04:19.680 --> 00:04:36.240
and that corpus is a asset the company company's process memory uh modified and aligned with every country um conditions and far bigger than than than runtime you could see the proportion 20 to1

s42
00:04:36.240 --> 00:04:55.759
So and this is concurrently operating system and this is the scale we run in production today over 200 instances and spikes and latencies deviates from few minutes to up to 10 minutes.

s43
00:04:56.960 --> 00:05:09.440
Um and mainly yeah the mainly main reason for it that u we depending on many legacy system which is uh so cannot be faster than agent loop itself.

s44
00:05:11.440 --> 00:05:14.240
Expert time is the bottleneck.

s45
00:05:14.240 --> 00:05:19.199
So the theme bench uh does the triage for us.

s46
00:05:19.199 --> 00:05:23.759
It clusters the failures and hands back something you can act on.

s47
00:05:23.759 --> 00:05:27.440
Not just look at look at it.

s48
00:05:28.880 --> 00:05:38.479
The trace is the shared evidence that lets an expert and an engineer review the same case and agree on what happened.

s49
00:05:39.280 --> 00:05:43.280
A correction only counts when it becomes an executable change.

s50
00:05:43.280 --> 00:05:49.919
And that's the line between an opinion and a production fix.

s51
00:05:52.880 --> 00:05:55.840
And and this is where quality comes from.

s52
00:05:55.840 --> 00:06:11.919
not from vibes uh not from a bigger model from replaying real examples with u disabled rights to uh protect the production systems and checking whether behavior improved.

s53
00:06:13.440 --> 00:06:22.160
You can see here on the uh cognitive proportion u or this effort ratio uh between each activity in our project.

s54
00:06:22.160 --> 00:06:27.600
So usually uh pipe coding ends here.

s55
00:06:27.600 --> 00:06:39.360
Here there ends um specdriven development because it cannot uh grow improve accuracy more than this stage on this scale.

s56
00:06:39.360 --> 00:06:44.080
And this is uh where the real work starts.

s57
00:06:44.080 --> 00:06:44.880
Nothing exotic.

s58
00:06:44.880 --> 00:06:50.160
It's engineering common engineering sense applied at scale.

s59
00:06:50.160 --> 00:07:05.599
So if uh you don't know all this uh terminology which developed over lastuh 30 years in software development argument to check because this is what every AI agent uh AI coding agent should know uh to help you

s60
00:07:05.599 --> 00:07:23.360
develop reliable production systems and accuracy it's uh wasn't designed uh in one diagram up front it was earned one small correction at the time at the scale you see here.

s61
00:07:23.360 --> 00:07:41.440
So we have over 100,000 corrections over last 9 months in the system when we developing it [clears throat] and this um heat maps uh turned thousands of traces into priorities.

s62
00:07:41.440 --> 00:07:51.360
is how we keep experts and engineers uh looking at the same problems and prioritize where the the most beneficial work for them.

s63
00:07:51.360 --> 00:08:05.680
Every cell is a group of tracked scenarios we have and uh usually to turn one block in red it's around one two months of force for the whole team

s64
00:08:06.960 --> 00:08:20.160
whole team of engineers and also AI agents um the agent failed is uh where the investigation starts not where it ends each failure maps to a specific uh fix

s65
00:08:22.319 --> 00:08:27.360
discovery needs agent freedom and production needs a cage.

s66
00:08:27.360 --> 00:08:30.560
Uh the harness isn't there to give the agent more room.

s67
00:08:30.560 --> 00:08:34.640
It's there to make the dumb mistakes impossible.

s68
00:08:36.640 --> 00:08:40.240
So on this scale please be careful is not a guard guard.

s69
00:08:40.240 --> 00:08:43.760
Uh if we have wrong workflow then classifier eval.

s70
00:08:43.760 --> 00:08:46.640
If it's wrong right then right gate.

s71
00:08:46.640 --> 00:08:49.760
If it's wrong assumption then it's a mere view.

s72
00:08:49.760 --> 00:08:56.720
A preventive measure eliminates the unsafe path on critical paths.

s73
00:08:56.720 --> 00:08:59.200
U review and approval stay in the loop.

s74
00:08:59.200 --> 00:09:06.560
The engine engineering focus is uh to build safe hands offs and a trail you can trust.

s75
00:09:08.880 --> 00:09:13.680
The real outcome uh wasn't the agent in the system.

s76
00:09:13.680 --> 00:09:17.360
It was the [clears throat] methodology we built around it.

s77
00:09:17.360 --> 00:09:22.080
If you want the blueprint, then it's uh these five moves.

s78
00:09:22.080 --> 00:09:23.519
Make work representable.

s79
00:09:23.519 --> 00:09:26.000
Make exe execution bounded.

s80
00:09:26.000 --> 00:09:31.920
Make behavior observable for every agent and make correction cheap.

s81
00:09:31.920 --> 00:09:34.560
And last thing is make improvement compound.

s82
00:09:34.560 --> 00:09:39.680
So gradually systematically improve the quality of the system.

s83
00:09:41.920 --> 00:09:47.040
AI native um operation is more than agents in workflow.

s84
00:09:47.040 --> 00:09:57.760
It's a system that learns from what works and fold folds it back into code as new composite tools adapting to the applications and the people around it.

s85
00:09:57.760 --> 00:10:03.040
The best AI models um oriented intelligence for us.

s86
00:10:03.040 --> 00:10:23.600
The adaptive architecture we built is the asset, the final asset and we aggregating all um repeatable sequences of steps successful scenarios and uh merging them into bigger tools which uh

s87
00:10:23.600 --> 00:10:28.720
combine the disproven scenarios into the reusable snippets by other agents.

s88
00:10:28.720 --> 00:10:37.600
So and then um it's possible to roll out them not only for one country but for hundreds country in one go.

s89
00:10:38.640 --> 00:10:46.079
So this is um um all for the talk and little time for questions and I'll be around afterwards.

s90
00:10:46.079 --> 00:11:01.200
And the final reminder you know if you you know if you are AI builder if you emotionally attached to tools not MCPS we're not using MCPS because uh for us it's uh always

s91
00:11:01.200 --> 00:11:02.640
not the best choice.

s92
00:11:02.640 --> 00:11:16.640
So because all all systems usually really bloated and we have to distill responses and uh tune the tools through function calling uh to our agents then we can control

s93
00:11:16.640 --> 00:11:25.600
quality of um our software and ensure that uh it's correctly processing assigned tasks.

s94
00:11:27.200 --> 00:11:27.600
Thank you.

s95
00:11:27.600 --> 00:11:29.839
Any questions?

s96
00:11:33.519 --> 00:11:37.760
Okay, then um thanks for your attent u attention.

s97
00:11:37.760 --> 00:11:42.880
Then I will be around so you can ask me questions if you want.

s98
00:11:44.133 --> 00:11:46.133
[applause]

s99
00:12:00.508 --> 00:12:02.508
[music]
