WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=_PdK6x7PQNM

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/_PdK6x7PQNM.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:13.040 --> 00:00:15.519
Good morning everybody.

s2
00:00:15.759 --> 00:00:16.880
My name is Ari Marcos.

s3
00:00:16.880 --> 00:00:19.359
I'm the CEO and co-founder of Datlogy AI.

s4
00:00:19.359 --> 00:00:23.119
Uh and really excited to kick off the data quality track uh today.

s5
00:00:23.119 --> 00:00:26.720
Uh data quality is is what we live and breathe uh at Datlogy.

s6
00:00:26.720 --> 00:00:27.599
It's all we think about.

s7
00:00:27.599 --> 00:00:31.119
In fact, the company's name literally means the science or study of data.

s8
00:00:31.119 --> 00:00:38.879
Um, so very excited to see the increasing excitement and interest in this area and a amazing lineup of talks uh today.

s9
00:00:38.879 --> 00:00:47.920
So today I'm going to tell you about why data quality is a compute multiplier that we're all overlooking and where we can make massive gains by just working on better data.

s10
00:00:47.920 --> 00:00:55.199
Um, we've seen that compute availability over the last six months has become extremely scarce and is only getting worse.

s11
00:00:55.199 --> 00:01:03.680
We saw H100 prices reverse their several year-long drop which is normal for hardware and all of a sudden come up where now they're about 40% up from their lows

s12
00:01:03.680 --> 00:01:05.439
um at the end of last year.

s13
00:01:05.439 --> 00:01:15.600
Um as test time compute has become a critical part of models and as we have put more and more thinking tokens in we're seeing that the number the token usage is absolutely

s14
00:01:15.600 --> 00:01:16.640
skyrocketing.

s15
00:01:16.640 --> 00:01:23.680
Um reasoning models use eight times as many tokens um as non-reasoning models and that's projected to 5x again in the next year or so.

s16
00:01:23.680 --> 00:01:28.880
Um, so the number of tokens we're pushing through goes higher and higher and that constrains compute even further.

s17
00:01:28.880 --> 00:01:36.159
And this has led actually to a world where it's not implausible today that we might see access to some of the frontier APIs actually get limited or go away.

s18
00:01:36.159 --> 00:01:41.119
Um, as one example, Google just capped uh Meta's Gemini usage because of inference constraints.

s19
00:01:41.119 --> 00:01:47.680
OpenAI has effectively started selling token futures or you can guarantee token capacity some amount of time into the future.

s20
00:01:47.680 --> 00:01:59.920
This is only necessary because they are legitimately wondering there might be a world where access to frontier API tokens is limited not as a business decision but because there's just simply not enough inference and first party products will be prioritized.

s21
00:01:59.920 --> 00:02:04.079
So in a world where compute is increasingly scarce and you need to make models better.

s22
00:02:04.079 --> 00:02:05.119
Well, what do you do?

s23
00:02:05.119 --> 00:02:06.399
Well, we work on data.

s24
00:02:06.399 --> 00:02:08.479
Um I you're going to hear me say this over and over again.

s25
00:02:08.479 --> 00:02:13.280
Data quality is a compute multiplier because what it does is it it makes the learning curve steeper.

s26
00:02:13.280 --> 00:02:18.959
So this is just a very simple schematic of performance on the y-axis as a function of data on the x-axis.

s27
00:02:18.959 --> 00:02:23.599
Note that this x-axis, you could swap it out for data, for compute, for time, for dollars.

s28
00:02:23.599 --> 00:02:26.239
They're all the same x-axis fundamentally.

s29
00:02:26.239 --> 00:02:32.400
Um, and if you can make data quality better, you can turn this gray curve um into this blue curve.

s30
00:02:32.400 --> 00:02:39.760
And that means that now you can get dramatically better performance for the same compute budget as if you had trained with far more compute.

s31
00:02:39.760 --> 00:02:44.480
Um and similarly you can get the same performance for a much smaller compute budget.

s32
00:02:44.480 --> 00:02:53.920
Um which is exactly showing you how you can get performance as if you had spent 10 times this 10 times as much on compute and really shows this compute multiplier

s33
00:02:53.920 --> 00:02:55.120
point.

s34
00:02:55.120 --> 00:02:56.640
Um so how do you actually do this?

s35
00:02:56.640 --> 00:03:02.080
Well fundamentally the idea is we want to make it so that we get the maximum signal per token and per batch.

s36
00:03:02.080 --> 00:03:09.200
Um for those that are a little bit more technically what we want to want to do here is maximize the marginal information gain per data point when we show it to the model.

s37
00:03:09.200 --> 00:03:11.519
what data is going to teach the model the most.

s38
00:03:11.519 --> 00:03:16.800
Um, and that's all about finding data that's relevant to the use cases that you want.

s39
00:03:16.800 --> 00:03:23.200
One thing that you know is very important is that there's no one golden data set to rule them all that's good for everything no matter what you want to do.

s40
00:03:23.200 --> 00:03:28.720
A data set's only going to be optimal with respect to a particular set of output tasks that you want the model to do.

s41
00:03:28.720 --> 00:03:33.519
So if you want a great legal model, you're going to want legal data more than healthcare data and vice versa.

s42
00:03:33.519 --> 00:03:34.959
It needs to be diverse.

s43
00:03:34.959 --> 00:03:40.400
A lot of the issues we see with model robustness and brittleleness comes from try training on data that's not diverse enough.

s44
00:03:40.400 --> 00:03:49.760
So the model can answer a question correctly if it's presented just so but if it's presented a little bit differently now everything breaks um needs to be information and you have to mix the data correctly.

s45
00:03:49.760 --> 00:03:52.400
This is a hugely difficult part of you now have many different sources.

s46
00:03:52.400 --> 00:03:56.959
How do you combine them to actually drive the largest improvement in performance?

s47
00:03:56.959 --> 00:04:00.080
So um this is a high level of what we do at here.

s48
00:04:00.080 --> 00:04:02.400
Um you can think of us as the oil refinery for data.

s49
00:04:02.400 --> 00:04:04.319
We don't source new tokens like many data providers.

s50
00:04:04.319 --> 00:04:10.959
Rather we take existing tokens coming from public data sets, proprietary data sets and licensed data sets and make them way better.

s51
00:04:10.959 --> 00:04:11.760
And how do we do that?

s52
00:04:11.760 --> 00:04:13.360
Well, we do that through these four C's.

s53
00:04:13.360 --> 00:04:15.840
Um, clean, curate, create, and compose.

s54
00:04:15.840 --> 00:04:17.759
Um, so cleaning is fairly straightforward.

s55
00:04:17.759 --> 00:04:21.199
This is doing things like heruristic filters, all a gopher and things like that.

s56
00:04:21.199 --> 00:04:25.840
Removing documents that have, you know, only 10 characters in them or all winging.

s57
00:04:25.840 --> 00:04:27.840
That's kind of basic table stakes.

s58
00:04:27.840 --> 00:04:30.560
Um, benchmark decontamination is incredibly important.

s59
00:04:30.560 --> 00:04:35.919
As I'm sure you all know, benchmaxing has become a real problem and makes it very difficult to interpret model results.

s60
00:04:35.919 --> 00:04:43.520
So, we rigorously decontaminate all of our training data with respect to all downstream benchmarks with a pretty low engram to ensure that that's the case.

s61
00:04:43.520 --> 00:04:49.040
That gets you to a point where now you can feed the model into the data, but it's still sorry, feed the data into the model, but it's still not very good.

s62
00:04:49.040 --> 00:04:50.080
So then how do you make it better?

s63
00:04:50.080 --> 00:04:57.680
Well, it's a combination of many things ranging from quality classifiers and tonomy across different topics and balancing that redundancy reduction.

s64
00:04:57.680 --> 00:05:14.240
So removing data points that are not the same that are semantically similar but convey very similar information even if they're not the same pixels themselves say upsampling and downsampling data points based off the quality and the relevance and then task distribution matching identifying what data do you actually need in order to solve this given task.

s65
00:05:14.240 --> 00:05:19.280
That now gives you a data set that is very high quality but is typically still too small.

s66
00:05:19.280 --> 00:05:20.880
And that's where synthetic data comes in.

s67
00:05:20.880 --> 00:05:33.520
Now we can go and rephrase that data as effectively a very fancy form of data augmentation to produce dramatically more data in many different formats and this helps a lot both with kind of uh data size and with diversity

s68
00:05:33.520 --> 00:05:46.800
um because we can really inject a lot of diversity in through this and then finally you know have these data sets how do you combine them and how do you com and how do you sequence them across different training stages it's now become table stakes that any large model is generally trained for at least three phases of

s69
00:05:46.800 --> 00:05:54.080
data um how do you do that um and can you actually even do continuous uh curricula and things like that which is a lot of what we work on at Dtology.

s70
00:05:54.080 --> 00:05:56.960
Um and that ultimately gets you a much better data set out.

s71
00:05:56.960 --> 00:05:59.600
All right, so that's a high level of kind of what we need to do.

s72
00:05:59.600 --> 00:06:01.199
What can you actually get out of this?

s73
00:06:01.199 --> 00:06:03.919
Can this actually really make a massive difference?

s74
00:06:03.919 --> 00:06:07.600
Um so we've about half of our team at Datlogy are just researchers.

s75
00:06:07.600 --> 00:06:12.160
Um and you know we uh do all of our own research on how we do data curation effectively.

s76
00:06:12.160 --> 00:06:22.000
um because this is such a critical part of the uh model building pipeline uh there's very little published here um because there's a very strong disincentive not to share how you do this

s77
00:06:22.000 --> 00:06:37.360
um the kind of foundational paper for daty was one I wrote when I was at meta called beyond scaling laws um which was fortunate to get a best paper at nurips a couple years ago which showed that if you choose your data correctly you can actually bend the scaling laws itself you can change the exponent

s78
00:06:37.360 --> 00:06:44.560
um and that's because you're now not wasting your time looking at redundant or unnecessary data That was very much the proof of principle for all of Dtology.

s79
00:06:44.560 --> 00:06:47.120
And we've since expanded this into many public research releases.

s80
00:06:47.120 --> 00:06:50.720
We've shared of various ways to improve models just through data curation.

s81
00:06:50.720 --> 00:06:55.840
I'm going to go through a couple of those results now um and show you what we've been able to to to achieve.

s82
00:06:55.840 --> 00:06:58.720
Um so first let's talk about vision language models.

s83
00:06:58.720 --> 00:07:03.599
How can we improve um VLMs just through data curation alone?

s84
00:07:03.599 --> 00:07:06.639
Um so in this case what we did is we took the mammoth data set.

s85
00:07:06.639 --> 00:07:20.240
This is a fairly small data set about 25 billion um tokens uh that we use for the purposes of training the fusion adapter layer between your um text uh your text your text model and your vision model.

s86
00:07:20.240 --> 00:07:27.199
Um and what you can see this is a scaling plot where we have error on the y-axis as a function of log flops um on the x-axis.

s87
00:07:27.199 --> 00:07:32.479
Um so there's about a scale of a thousand from the left uh most part of this plot to the rightmost part.

s88
00:07:32.479 --> 00:07:39.520
Um and what you can see is that if you look at the parto frontier defined by many of the best uh public VLMs like the Quen 3 series and 3.5

s89
00:07:39.520 --> 00:07:48.800
intern VL etc you can see that a model trained on daties data are able to go well well beyond uh that frontier and I'll note this is actually without any post- training

s90
00:07:48.800 --> 00:07:49.360
as well.

s91
00:07:49.360 --> 00:07:52.720
Uh so you can get very strong performance across many different benchmarks.

s92
00:07:52.720 --> 00:07:59.440
Um so just looking at kind of taking the input data set that gray diamond there that's the input data set we use to do our curation.

s93
00:07:59.440 --> 00:08:08.000
Um you can see that just through curation you're able to get around a 14 absolute percentage point improvement um holding everything else constant just through better data alone.

s94
00:08:08.000 --> 00:08:16.879
Um and not only that you can also see that we can roughly match the performance of quen 3.54b come with about a percentage of it um while using 145x

s95
00:08:16.879 --> 00:08:20.400
less training compute in a world with no with less compute.

s96
00:08:20.400 --> 00:08:21.120
How do you do more?

s97
00:08:21.120 --> 00:08:25.199
You make data better and now it's as if you had 100 times um the compute.

s98
00:08:25.199 --> 00:08:30.160
Um, interestingly I mentioned that reasoning models are be are using tokens at a very high rate as well.

s99
00:08:30.160 --> 00:08:37.760
Well, another thing that we found is that data curation can also lead uh to more concise answers um depending on how you represent the data.

s100
00:08:37.760 --> 00:08:45.760
So what's plotted here is the mean number of tokens per response um across all the same set of models for the large most part um that we just showed.

s101
00:08:45.760 --> 00:08:51.200
Um and you can see that that models trained on data those three blue lines right at the top are all extremely concise.

s102
00:08:51.200 --> 00:08:58.959
Um and if we do kind of the same sort of plot um but now on the x-axis instead of log training flops this is now um log flops per response.

s103
00:08:58.959 --> 00:09:00.640
Um so this is inference efficiency.

s104
00:09:00.640 --> 00:09:09.360
Um you can still see that we go well on the ex by beating that router frontier um and you know roughly get uh similar performance to quen 35 with 35 fewer

s105
00:09:09.360 --> 00:09:12.399
uh times fewer flops uh per correct answer.

s106
00:09:12.399 --> 00:09:14.800
So data curation can make a huge impact in VLMs.

s107
00:09:14.800 --> 00:09:16.000
What about text models?

s108
00:09:16.000 --> 00:09:24.560
Um, one of the most uh challenging things about many models is that they work very well on English data, but they don't work well um for non-western use cases in general.

s109
00:09:24.560 --> 00:09:29.680
The internet is an extremely biased view of the world that does not represent the world uniformly at all.

s110
00:09:29.680 --> 00:09:34.640
And this has major implications for fairness and for the usability of these models across the world.

s111
00:09:34.640 --> 00:09:38.880
I don't want to live in a future where only developed countries can access this very effectively.

s112
00:09:38.880 --> 00:09:40.320
Um, so how do you do this?

s113
00:09:40.320 --> 00:09:42.640
Well, curation again can be a massive lever here.

s114
00:09:42.640 --> 00:09:50.399
Um so what I'm plotting here now is a similar plot error on the y- axis as a function of log flops about 100x going from left to right here.

s115
00:09:50.399 --> 00:09:53.360
Um this is highlighting multilingual MMLU performance.

s116
00:09:53.360 --> 00:09:56.560
You can see we have a parto frontier here defined by many models.

s117
00:09:56.560 --> 00:09:59.360
The quen models um the liquid some of the liquid models.

s118
00:09:59.360 --> 00:10:03.279
Um the green square is tiny coher's best multilingual model.

s119
00:10:03.279 --> 00:10:07.519
You can see again that we're well off the predto frontier with a couple things I really want to highlight.

s120
00:10:07.519 --> 00:10:11.440
First off we only use 8% of the data here as multilingual tokens.

s121
00:10:11.440 --> 00:10:15.920
Um, so most languages actually only had um at max 6 billion tokens here.

s122
00:10:15.920 --> 00:10:20.880
So these are not massive amounts of data uh in the non-English languages that are going in here.

s123
00:10:20.880 --> 00:10:24.079
Um you can again see we get the same sort of compute multiplier effect.

s124
00:10:24.079 --> 00:10:29.519
We're a little better than Quen 3 um while while having roughly 8x less compute budget um here.

s125
00:10:29.519 --> 00:10:31.360
So you can make a huge improvement.

s126
00:10:31.360 --> 00:10:39.600
One last thing I want to show here is that if you look at the two blue points on the upper left here um those are both dense llama style models trained for a trillion tokens

s127
00:10:39.600 --> 00:10:40.959
um on curated data.

s128
00:10:40.959 --> 00:10:49.600
The point on the lower right here I'll come back to but is a model trained by one of our customers rci trendy large that was trained on 17 trillion tokens and is a hypersparse.

s129
00:10:49.600 --> 00:10:58.880
Um and what you can see is that if you take the line defined by the two smaller models, um it goes mostly right through that blue star uh which is training large despite it being trained

s130
00:10:58.880 --> 00:11:01.200
um with 50x more training compute.

s131
00:11:01.200 --> 00:11:13.920
So if you use your data correctly and you simulate token scarcity appropriately, you can also get um very predictable scaling to much larger models and you can derisk a run with 50 or 100 times less compute effectively

s132
00:11:13.920 --> 00:11:19.519
um before you actually go and scale up the hero run um and find that maybe it doesn't end up where you want it to be.

s133
00:11:19.519 --> 00:11:26.480
Um an interesting scientific result I want to share here is that we also see very strong cross-lingual benefits um from curation.

s134
00:11:26.480 --> 00:11:36.800
Um so what's plotted here is the non-English accuracy um where in the the the left bar is showing not curating anything at all and then the right bar just curating the English.

s135
00:11:36.800 --> 00:11:39.920
Um we also curate all the non-English data and that leads to much better performance.

s136
00:11:39.920 --> 00:11:46.000
But I just want to show this because I think it's quite interesting that curating English data benefits non-English performance.

s137
00:11:46.000 --> 00:11:58.079
Um and that's because we see this cross-lingual transfer where the model understands how English relates to say Spanish and and so therefore making it better at English would also make it better at Spanish to some extent and interestingly we see that the the um the

s138
00:11:58.079 --> 00:12:11.360
size of magnitude of that transfer um is strongly correlated uh with the similarity between English and that language um and we also see it go the other way although the effects a bit smaller um where curating the non-English data also helps to benefit

s139
00:12:11.360 --> 00:12:17.440
English data um performance um okay And let me talk a little bit about synthetic data.

s140
00:12:17.440 --> 00:12:20.560
Um we take an approach to synthetic data that we call rephrasing.

s141
00:12:20.560 --> 00:12:26.399
This is something that our team pioneered several years ago and has now become um basically table stakes for building a very strong model.

s142
00:12:26.399 --> 00:12:30.639
Um in anyway I think we'll hear a lot about the synthetic data in various forms uh throughout the day.

s143
00:12:30.639 --> 00:12:40.399
Um but fundamentally with beyond web our goals is how can we define a synthetic data platform that works extremely well um and can be applied to anyone's proprietary doc data and documents.

s144
00:12:40.399 --> 00:12:44.320
Fundamentally we want to help folks build models that wouldn't be able to do so otherwise.

s145
00:12:44.320 --> 00:12:47.680
And that's what where data quality can make an absolute difference.

s146
00:12:47.680 --> 00:12:54.480
Um so to give an example, we might take a document like this about a corporate takeover and we might convert that into one of hundreds of templates.

s147
00:12:54.480 --> 00:12:57.519
One of which might be a series of true false questions.

s148
00:12:57.519 --> 00:13:00.160
Um by doing this, there's a couple things that are really great.

s149
00:13:00.160 --> 00:13:05.200
Number one, because all the information is coming from the document on the left, you don't have any issue with model collapse.

s150
00:13:05.200 --> 00:13:13.440
Um and you can actually train models um that are much better than the rephrasing model because the rephrasing model doesn't actually have to teach and understand all the concepts.

s151
00:13:13.440 --> 00:13:20.240
All it needs to do is transform the left document into a true false questions accurately which is a much easier task.

s152
00:13:20.240 --> 00:13:23.519
And then you do this into many many different formats um throughout data.

s153
00:13:23.519 --> 00:13:29.200
This effectively increases diversity and it makes it so you learn a lot more from the highest quality data points.

s154
00:13:29.200 --> 00:13:32.320
One thing that's really critical here, what do you rephrase?

s155
00:13:32.320 --> 00:13:34.320
All documents are not created equal for rephrasing.

s156
00:13:34.320 --> 00:13:37.839
If you just pick random sets of documents to rephrase, you will not get a great result.

s157
00:13:37.839 --> 00:13:41.680
Um but if you find the high quality documents and rephrase them, it can make a big difference.

s158
00:13:41.680 --> 00:13:48.399
Um and we've seen if you compare this to lots of other public synthetic corpora, we can get much better performance much faster.

s159
00:13:48.399 --> 00:13:56.000
Um and critically this can be applied to any any proprietary data um in one of our customers own environments.

s160
00:13:56.000 --> 00:13:56.320
All right.

s161
00:13:56.320 --> 00:14:03.519
Uh let me spend the last few minutes just quickly talking about a couple um of what we've seen uh the things we've seen with our customers where this can actually drive really gains.

s162
00:14:03.519 --> 00:14:10.480
Um so first one of our customers Thompson Reuters um has really focused on post training quite a bit.

s163
00:14:10.480 --> 00:14:19.360
um and they have a very sophisticated post- trainining infrastructure with the goal of uh building better legal models on their proprietary high-quality legal data that they have.

s164
00:14:19.360 --> 00:14:28.320
So we partnered with them to mid-train a model first on a combination of their data um and uh public data um to then make much better legal reasoning models.

s165
00:14:28.320 --> 00:14:28.959
So what do we see?

s166
00:14:28.959 --> 00:14:30.320
Well, first off, look at the left here.

s167
00:14:30.320 --> 00:14:37.600
Um in this case, we took a an open source model um and then uh just did continued pre-trainer and mid-training on a 100 billion tokens.

s168
00:14:37.600 --> 00:14:47.519
What you'll see here is that we see that legal capabilities go up about five percentage points um as measured by uh legal bench after you do this 100 billion mid-training which was less than 1% of the pre-training budget.

s169
00:14:47.519 --> 00:14:49.600
Um but you don't get catastrophic forgetting.

s170
00:14:49.600 --> 00:14:51.920
You also see the general capabilities go up as well.

s171
00:14:51.920 --> 00:14:59.120
I'm sure that many of you have seen or experienced um when you try to adapt a model to a particular domain, you lose general performance.

s172
00:14:59.120 --> 00:15:00.800
Not if you use the data correctly.

s173
00:15:00.800 --> 00:15:06.959
Um, the key here is actually the majority of the data we showed the model was actually data that was representative of the pre-training distribution.

s174
00:15:06.959 --> 00:15:14.560
There was only some of the domain specific data in and that's necessary to prevent the model from losing the capabilities that it had before.

s175
00:15:14.560 --> 00:15:17.440
Um, and you can solve this entirely through better data.

s176
00:15:17.440 --> 00:15:19.839
Um, but this actually isn't the most exciting part here.

s177
00:15:19.839 --> 00:15:23.600
Um, as I mentioned, the TR team had done a lot of work on post- training.

s178
00:15:23.600 --> 00:15:25.920
Um, and had a very sophisticated post-training harness.

s179
00:15:25.920 --> 00:15:30.480
Well, they then applied that to the mid-train model versus just the default instruction tune model.

s180
00:15:30.480 --> 00:15:33.360
And what they found was that the gain deriving from post-training.

s181
00:15:33.360 --> 00:15:42.880
So the y- axis here is a delta um as a result of post-training um almost tripled uh when you applied it to the mid-train model versus to the um just default instruction tune model.

s182
00:15:42.880 --> 00:15:48.079
And that's because its policy um when it starts is now much more accurate um and it can make much better inference.

s183
00:15:48.079 --> 00:15:58.880
Um, so you can even even if you don't change the post- training data at all, showing your model better domain specific data can actually make post- training two to three times more effective out of the box.

s184
00:15:58.880 --> 00:16:07.920
Um, which I think really goes to show not only how important data can be in these uh factors, but it also actually goes to show how we really should be thinking about all these stages synergistically

s185
00:16:07.920 --> 00:16:17.440
rather than as uh three completely independent stages of pre-training and then I hand it off to somebody else who mid-trains who then I hand off to somebody else um who post trains.

s186
00:16:17.440 --> 00:16:31.839
Um okay uh in the last minute or so I just want to quickly talk about um one other one which is rci who I mentioned that large model trendy large um you'll actually hear from verun who was the pre-training lead for this model later today um so look forward to that talk um

s187
00:16:31.839 --> 00:16:47.920
but in this case they trained a model this is a fully open um US-made model um uh on 17 trillion tokens that we curated from public uh data sets no proprietary data involved here um and no closed model usage so no asking Claude uh to do this for you Um,

s188
00:16:47.920 --> 00:16:51.680
and with that, RC was able to train a model that is competitive with the Open Frontier.

s189
00:16:51.680 --> 00:16:55.040
Um, matches uh, GLM5 and and Kimmy on many tasks.

s190
00:16:55.040 --> 00:16:57.279
Um, and even outperforms Claude on a couple tasks.

s191
00:16:57.279 --> 00:17:05.280
Um, but I think what's most exciting about this is that the RC team um, had not trained a model prior to uh, the middle of last year when they started working with us.

s192
00:17:05.280 --> 00:17:17.600
Um and critically in total across salaries, across compute, across R&amp;D, um across everything for this and several other models, they were able to get to a model that's competitive with the open frontier,

s193
00:17:17.600 --> 00:17:20.640
um for less than $20 million total.

s194
00:17:20.640 --> 00:17:23.839
That includes all the repetitions, that includes compute, that includes everything.

s195
00:17:23.839 --> 00:17:28.400
So if you hear this story over and over again, oh, if I want to customize a model, it's going to cost hundreds of millions of dollars.

s196
00:17:28.400 --> 00:17:29.360
That's just not true.

s197
00:17:29.360 --> 00:17:35.120
you can train an immensely powerful model especially in a narrow domain for high six figures million dollars.

s198
00:17:35.120 --> 00:17:37.919
It's very doable to get a model that's extremely performant.

s199
00:17:37.919 --> 00:17:39.200
Um this is for a general purpose.

s200
00:17:39.200 --> 00:17:41.120
So this is kind of the upper bound of that.

s201
00:17:41.120 --> 00:17:43.600
Um and data quality is how you can do that.

s202
00:17:43.600 --> 00:17:44.080
All right.

s203
00:17:44.080 --> 00:17:51.840
So um last slide here just kind of say um summarizing um focus on kind of what's going to give you the most signal per token.

s204
00:17:51.840 --> 00:17:53.840
That's the thing that matters a lot more than more tokens.

s205
00:17:53.840 --> 00:18:01.679
that it's almost always better to repeat highquality data than it is to show lowquality data at a certain point um up to a threshold.

s206
00:18:01.679 --> 00:18:03.840
Um but focus on kind of how can you get that?

s207
00:18:03.840 --> 00:18:07.840
This is data quality remains the single most underleveraged compute multiplier.

s208
00:18:07.840 --> 00:18:12.880
If you're sitting in a world where you want to build a model or customize a model and you're limited on compute, how do you get past that?

s209
00:18:12.880 --> 00:18:16.320
Invest in data and that's something that can do a tremendous amount of effort.

s210
00:18:16.320 --> 00:18:19.360
Um and then finally, this is a frontier research engineering problem.

s211
00:18:19.360 --> 00:18:22.640
You need to be able to score and understand data across many different axes.

s212
00:18:22.640 --> 00:18:23.760
because that's a frontier research problem.

s213
00:18:23.760 --> 00:18:31.600
And then have that scale up to pabytes of data, massive scale, um, and can be very critical and it can lead to tremendous leverage.

s214
00:18:31.600 --> 00:18:33.280
Um, and with that, I'll say thank you.

s215
00:18:33.280 --> 00:18:36.400
Um, I will note that we're hiring for a bunch of different roles on the left here.

s216
00:18:36.400 --> 00:18:44.799
Um, if you're interested in building or customizing your own model and um, would like to to get much better data out of the box or apply it to your own data, um, would love to chat with you.

s217
00:18:44.799 --> 00:18:48.720
Um, and uh, thank you very much.
