WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=5Cxe5dv2Xlw

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/5Cxe5dv2Xlw.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:12.640 --> 00:00:20.120
Joining us on stage is the co-founder and chief science officer at Hugging Face, Thomas Wolf.

s3
00:00:20.305 --> 00:00:22.305
[music]

s4
00:00:26.200 --> 00:00:28.200
[music]

s5
00:00:32.000 --> 00:00:33.560
Hello everyone.

s6
00:00:33.560 --> 00:00:36.040
Hello Olive, nice to have you on stage.

s7
00:00:36.040 --> 00:00:37.000
Hi, nice to meet you.

s8
00:00:37.000 --> 00:00:38.640
Thanks for having me, yeah.

s9
00:00:38.640 --> 00:00:48.440
So I think you're on for a treat today because you just saw a GLM uh which is current number two on the artificial intelligence leaderboard.

s10
00:00:48.440 --> 00:00:51.640
I take Fable out because nobody can use it.

s11
00:00:51.640 --> 00:00:52.680
And now we have number four.

s12
00:00:52.680 --> 00:00:57.760
So you basically you will have all the top models, at least the top open source model in a row.

s13
00:00:57.760 --> 00:01:04.720
And we're very lucky to have Olive who has a pretty amazing path in life.

s14
00:01:04.720 --> 00:01:06.960
Uh so she came to the US, Pennsylvania.

s15
00:01:06.960 --> 00:01:17.520
She was studying, doing PhD at uh NYU uh in the lab of Jan LeCun working on J Pa, but we decided we won't talk about J Pa today, right?

s16
00:01:17.520 --> 00:01:19.080
Something for another day.

s17
00:01:19.080 --> 00:01:28.040
Um and then instead of joining Hugging Face, which was in New York also at that time she decided to go join MiniMax.

s18
00:01:28.040 --> 00:01:38.480
So for those who who maybe don't know all the all the neo labs around the world and you're you're forgiven because I think there's like 64 neo labs right now.

s19
00:01:38.480 --> 00:01:43.240
MiniMax is one of the top of what we call the AI dragons in China.

s20
00:01:43.240 --> 00:01:53.240
So these are the new there's there's Deep Seek which is very well known now, Moonshot who does Kimi Z and GLM that you just saw and now we have a MiniMax.

s21
00:01:53.240 --> 00:01:59.000
They're all extremely good, extremely a team uh fighting for the first spot.

s22
00:01:59.000 --> 00:02:09.080
Uh so the the the latest uh release of MiniMax was M3 uh just earlier earlier in June, which was the the top model at the time, top open-source model.

s23
00:02:09.080 --> 00:02:10.280
Uh very impressive.

s24
00:02:10.280 --> 00:02:15.120
There's a lot of very interesting things about this model, so we'll quickly dive in them.

s25
00:02:15.120 --> 00:02:22.040
And then talk a little bit about uh what's what's what's specific about MiniMax, what's what's great there.

s26
00:02:22.040 --> 00:02:25.040
So uh maybe Olive to to start a little bit.

s27
00:02:25.040 --> 00:02:33.200
Can you Can you give us, you know, a a little bit of your your view of of M3, what you like about this model, how was the release?

s28
00:02:33.200 --> 00:02:34.080
Mhm.

s29
00:02:34.080 --> 00:02:45.480
Yeah, M3 we released M3 earlier this month, and it is a smaller model with 400 around 400 billion total parameters and 20 billion activated.

s30
00:02:45.480 --> 00:02:52.880
Um but it is very capable in terms of both coding performances, and also it understands vision.

s31
00:02:52.880 --> 00:02:58.920
So um that's uh what open-source models don't usually have.

s32
00:02:58.920 --> 00:03:09.760
It's that they can the model can only deal with coding, but it can also understand videos, um images, and it has a super uh long context of 1 million.

s33
00:03:09.760 --> 00:03:14.280
Um with our new architecture called MSA, MiniMax Sparse Attention.

s34
00:03:14.280 --> 00:03:25.400
So we we really put these three things together uh because we know that they are they will be very important in future AI applications.

s35
00:03:25.400 --> 00:03:32.320
Coding capabilities, agentic capabilities, longer context, and multimodal understanding.

s36
00:03:32.320 --> 00:03:35.560
Um yeah, I think that would be very interesting about the model.

s37
00:03:35.560 --> 00:03:46.120
Yeah, so so there's a lot to unpack unpack in this model, and it's um it's it's still, I think, the only top five model open-source model that is actually multimodal, so we need to talk about that.

s38
00:03:46.120 --> 00:03:59.160
But maybe first about the long context, because that was also the first one that really had this real 1 million token long context is actually functional and you guys had also the the Minimax pass attention

s39
00:03:59.160 --> 00:04:05.240
which is this one technique to to make that efficient that you also published and and share extensively.

s40
00:04:05.240 --> 00:04:07.160
So can you can you talk a little bit about this?

s41
00:04:07.160 --> 00:04:12.080
Maybe how the project went from from the attention how to make this long context.

s42
00:04:12.080 --> 00:04:25.080
Yeah, I would say the story about long context went back to even Minimax M1 and Minimax 01 where the model was actually was able to perform tasks of 10 million

s43
00:04:25.080 --> 00:04:26.080
token context.

s44
00:04:26.080 --> 00:04:26.880
10 million?

s45
00:04:26.880 --> 00:04:28.919
10 million, yes.

s46
00:04:28.919 --> 00:04:31.680
But then it was not an agentic model, right?

s47
00:04:31.680 --> 00:04:37.680
It was just 10 for example, you dump in a book, it would be able to give reviews on it and stuff like that.

s48
00:04:37.680 --> 00:04:54.560
So what we realized was that, you know, longer context actually unlocks a lot of capabilities especially when interacting with users and now when, you know, the agent is interacting with the whole environment and getting all the tool responses,

s49
00:04:54.560 --> 00:05:04.480
getting multi rounds the like shorter context wouldn't be enough to perform the complex task.

s50
00:05:04.480 --> 00:05:10.640
So for this version we said, "Oh, we have to have our longer context backs."

s51
00:05:10.640 --> 00:05:15.919
So what we pursued was with our Minimax sparse attention.

s52
00:05:15.919 --> 00:05:24.120
Um which you know, was the architecture that was scalable and had a simple design.

s53
00:05:24.120 --> 00:05:27.520
So I would say from a higher level, right?

s54
00:05:27.520 --> 00:05:43.280
It has an index branch that, you know, selects on a higher level what is what matters more in the context and then we have a sparse attention branch that calculates performs the calculation on the selected blocks

s55
00:05:43.280 --> 00:05:45.480
to actually performs the task.

s56
00:05:45.480 --> 00:05:57.600
Um and so yeah, like that we really designed um an elegant architecture so that we can scale the length and then scale the model size in the future with that.

s57
00:05:57.600 --> 00:05:58.400
That's beautiful.

s58
00:05:58.400 --> 00:06:03.440
I like how for those who've been in the field for quite some time, we we had a lot of work on attention, right?

s59
00:06:03.440 --> 00:06:06.280
This N square and there was a lot of linear attention.

s60
00:06:06.280 --> 00:06:06.760
Yeah.

s61
00:06:06.760 --> 00:06:11.200
And then some that somehow all of these disappeared at some point when flash attention came around.

s62
00:06:11.200 --> 00:06:13.720
We discovered we just needed more efficient camera.

s63
00:06:13.720 --> 00:06:18.360
Now, I like how we come back to thinking, you know, first principle, what is attention?

s64
00:06:18.360 --> 00:06:20.840
How can we make that more efficient?

s65
00:06:20.840 --> 00:06:23.280
So, 1 million token is crazy, right?

s66
00:06:23.280 --> 00:06:27.000
GPT-2 was 1,024 and and everyone was like, "Oh, that's really big.

s67
00:06:27.000 --> 00:06:29.160
We we we never need more."

s68
00:06:29.160 --> 00:06:30.480
Where do you see this coming?

s69
00:06:30.480 --> 00:06:31.600
Like, going in the future?

s70
00:06:31.600 --> 00:06:35.400
Like, Jeff Dean was pitching me the other day a trillion token attention.

s71
00:06:35.400 --> 00:06:38.360
You think we should go to a trillion token attention?

s72
00:06:38.360 --> 00:06:42.000
So, that's definitely something we can explore towards, right?

s73
00:06:42.000 --> 00:06:44.760
Ultra lengths of the context.

s74
00:06:44.760 --> 00:06:47.800
Definitely, it's something that's very exciting to explore with.

s75
00:06:47.800 --> 00:06:54.520
And something that architecture design along with hardware um would require a lot of research onto that.

s76
00:06:54.520 --> 00:06:54.919
Yeah.

s77
00:06:54.919 --> 00:06:57.280
You think there's still a lot of low-hanging fruits?

s78
00:06:57.280 --> 00:07:09.080
So, typically today we saw Open AI really reducing I mean, we don't know how as of firm, but like reducing their their inference bill by half, by probably having some more efficient processing around tensions of any type.

s79
00:07:09.080 --> 00:07:14.440
You think there's still a lot of low-hanging fruit that can be getting um how we can process that.

s80
00:07:14.440 --> 00:07:23.160
So, so one one thing we're still very interesting about M3 is how cheap it is in particular because of this part attention and part because of its small one, but it's also very efficient.

s81
00:07:23.160 --> 00:07:23.600
Right.

s82
00:07:23.600 --> 00:07:25.960
You think we can go even way further?

s83
00:07:25.960 --> 00:07:30.120
And maybe how did you guys invented uh Min Max Fast Attention?

s84
00:07:30.120 --> 00:07:32.120
Was it an agent coming up with the idea?

s85
00:07:32.120 --> 00:07:34.440
Was it a human still coming up with the idea?

s86
00:07:34.440 --> 00:07:35.880
Tell us a little bit about

s87
00:07:35.880 --> 00:07:36.760
Yeah.

s88
00:07:36.760 --> 00:07:49.800
Um, so we do think there's still a lot of work that can get into architecture and inference optimization so that the model can be more efficient, especially if there are tasks that are very task sensitive

s89
00:07:49.800 --> 00:07:52.880
but require very strong capabilities, right?

s90
00:07:52.880 --> 00:07:57.000
And for that those kind of tasks we really want model to be efficient.

s91
00:07:57.000 --> 00:08:03.200
Um and who came up with this part of actually I think an intern from our team worked on that.

s92
00:08:03.200 --> 00:08:05.440
That yeah, an intern.

s93
00:08:05.440 --> 00:08:15.720
Uh that doesn't usually happen in a lot of labs because I think in some labs interns don't have access to the data, the work, and stuff.

s94
00:08:15.720 --> 00:08:20.600
Uh but yeah, we are open to anyone who would like to contribute to our models.

s95
00:08:20.600 --> 00:08:24.360
So, um the architecture was actually designed by an intern.

s96
00:08:24.360 --> 00:08:25.040
It's very good.

s97
00:08:25.040 --> 00:08:27.120
Still some work for interns here.

s98
00:08:27.120 --> 00:08:28.680
Good good news.

s99
00:08:28.680 --> 00:08:32.599
Um that's also a good segue to also how Min Max is working internally.

s100
00:08:32.599 --> 00:08:38.280
So, so we were discussing before coming on stage I was saying everyone can propose a project.

s101
00:08:38.280 --> 00:08:42.159
Can you tell us a little bit about how you are organized, how you do your research?

s102
00:08:42.159 --> 00:08:42.560
Mhm.

s103
00:08:42.560 --> 00:08:43.320
Mhm.

s104
00:08:43.320 --> 00:08:54.000
I think that is very different from uh even in school or even in earlier you know, the earlier tech companies is pretty pretty different.

s105
00:08:54.000 --> 00:09:08.360
It's that um what we what we make sure is that we have good foundation and good um infrastructure so that anyone can play with the model and can think of what they can improve with the model.

s106
00:09:08.360 --> 00:09:11.880
And then after model releases when they are free, right?

s107
00:09:11.880 --> 00:09:13.640
They can play with the model.

s108
00:09:13.640 --> 00:09:16.240
They can think of their own evaluations.

s109
00:09:16.240 --> 00:09:21.800
They can find their own weaknesses and propose a thing that they want to improve on the model.

s110
00:09:21.800 --> 00:09:30.280
And then other people who are interested in that would, you know, propose to join the project and they will work on for a couple of weeks or even a couple of months.

s111
00:09:30.280 --> 00:09:34.240
And when they work out, the final thing is shipped to our model.

s112
00:09:34.240 --> 00:09:39.480
It it is, you you we use that in our final training and it's shipped out to the audience.

s113
00:09:39.480 --> 00:09:39.800
Interesting.

s114
00:09:39.800 --> 00:09:42.440
So, you can have people working for a really long time on project.

s115
00:09:42.440 --> 00:09:46.040
When you say a couple of months, it can be like really deep exploration of

s116
00:09:46.040 --> 00:09:46.840
Yes.

s117
00:09:46.840 --> 00:09:48.040
Yes.

s118
00:09:48.040 --> 00:09:58.400
I would say, for example, architecture might require longer time of investigation, research, experiments, even redoing the evaluations for pre-training.

s119
00:09:58.400 --> 00:10:00.600
Yes, so it might require longer time.

s120
00:10:00.600 --> 00:10:01.400
Very nice, yeah.

s121
00:10:01.400 --> 00:10:03.320
And I know you're also very big on evaluation.

s122
00:10:03.320 --> 00:10:03.800
I agree.

s123
00:10:03.800 --> 00:10:04.800
We could talk about that.

s124
00:10:04.800 --> 00:10:12.280
I think what One thing probably related to that is this unique specificity that M3 and your team has um around multimodality.

s125
00:10:12.280 --> 00:10:16.720
So, not just text, but the similar can also understand image and video.

s126
00:10:16.720 --> 00:10:28.320
And as I understand, but but please explain explain better, when you read the model card on Hugging Face, it say the model was trained from the first step as a multimodal, not just have a like user one as after solved, right?

s127
00:10:28.320 --> 00:10:37.400
Can you tell us a little bit more about that and why you think it's important and and and why starting from the first step on multimodal training and not just just training this.

s128
00:10:37.400 --> 00:10:40.839
Um so, we call it native multimodality.

s129
00:10:40.839 --> 00:10:54.080
Um and so, it is somehow typical for model labs to train the multimodal, let's say, vision understanding capabilities after the text pre-training is done.

s130
00:10:54.080 --> 00:10:57.560
Um they put adapters and then train that part.

s131
00:10:57.560 --> 00:11:03.240
But what we found out was that that would actually harm the text performance.

s132
00:11:03.240 --> 00:11:12.280
And the vision vision understanding performance wouldn't converge that well because the model is kind of converges towards the model the text understanding.

s133
00:11:12.280 --> 00:11:15.440
Um and it's just not the most optimal.

s134
00:11:15.440 --> 00:11:18.560
And also not the most scalable, if you think about it.

s135
00:11:18.560 --> 00:11:20.839
We want to scale the data, right?

s136
00:11:20.839 --> 00:11:27.960
And also, we can also some labs um train this capability from halfway through the pre-training.

s137
00:11:27.960 --> 00:11:30.000
For example, continued pre-training.

s138
00:11:30.000 --> 00:11:36.560
But what we found that this would be very, you know, uh recipe sensitive.

s139
00:11:36.560 --> 00:11:44.600
It is different for the recipe would be different for different architectures, different, you know, data mixtures, different learning rates.

s140
00:11:44.600 --> 00:11:55.560
It's hard to control, hard to, you know, scale to you can't really scale your experiment results and conclusions to a larger model.

s141
00:11:55.560 --> 00:12:01.960
And so you know, what we thought was why not just training from the very first step?

s142
00:12:01.960 --> 00:12:04.200
That comes to the most natural.

s143
00:12:04.200 --> 00:12:08.400
We know that a lot of labs run into problems doing that.

s144
00:12:08.400 --> 00:12:18.880
The model would collapse after a couple of steps of training, you know, both text and vision understanding, but we managed to solve that problem.

s145
00:12:18.880 --> 00:12:26.240
We did a lot of work on VIT and we did a lot of work on the data that we actually training.

s146
00:12:26.240 --> 00:12:31.200
For example, we do interleave the data, what we call interleave the data.

s147
00:12:31.200 --> 00:12:45.480
It's actually natural data, but we keep the images and videos in instead of masking it out and we do some pretty good cleaning and masking on the data and we do very good reward modeling

s148
00:12:45.480 --> 00:12:50.720
so that we train it from the first step and scales up a lot.

s149
00:12:50.720 --> 00:12:53.080
Yeah, it does does not collapse.

s150
00:12:53.080 --> 00:12:54.560
That's really impressive.

s151
00:12:54.560 --> 00:12:55.240
Impressive.

s152
00:12:55.240 --> 00:12:58.280
Should we Should we expect much larger model in the future?

s153
00:12:58.280 --> 00:13:00.160
So this one is still fairly small, right?

s154
00:13:00.160 --> 00:13:05.640
It's It's 428 billion parameters, 23 active billion.

s155
00:13:05.640 --> 00:13:09.200
Well, do you think you will go past the trillion?

s156
00:13:09.200 --> 00:13:10.440
Definitely.

s157
00:13:10.440 --> 00:13:13.120
Yeah, definitely in the future.

s158
00:13:13.120 --> 00:13:22.120
There are many tasks that wouldn't be able to the more model wouldn't be able to perform very good at with smaller parameters.

s159
00:13:22.120 --> 00:13:25.640
We are definitely going more ambitious than this.

s160
00:13:25.640 --> 00:13:26.160
It's great.

s161
00:13:26.160 --> 00:13:27.520
Looking forward.

s162
00:13:27.520 --> 00:13:36.520
Um another interesting thing I I always been find fascinating about Min Max is how how you also have this whole range of of apps and product, right?

s163
00:13:36.520 --> 00:13:43.760
So, I remember already So, so Min Max started to open source things on the on the hugging face platform in in January last year.

s164
00:13:43.760 --> 00:13:52.240
So, that 18 month ago and and we were chatting a little bit about the team to understand what you were doing and I remember you So, you were already having a huge usage

s165
00:13:52.240 --> 00:13:55.000
on some of these of some of these apps.

s166
00:13:55.000 --> 00:13:58.000
Um can you tell us a little bit how how this started, right?

s167
00:13:58.000 --> 00:14:05.400
So, was it basically you had a lot of apps and then you thought we have we have all this data, why not training a model and then they build up research team.

s168
00:14:05.400 --> 00:14:07.800
How is how is the story there?

s169
00:14:07.800 --> 00:14:12.400
Um our our story is modeled from the first day.

s170
00:14:12.400 --> 00:14:26.080
So, um I believe that multi-modality model a model that can understand all visions and outputs all modalities was the first thing that our um CEO planned on the first day even before the company even started.

s171
00:14:26.080 --> 00:14:27.600
So, that was the dream of AGI.

s172
00:14:27.600 --> 00:14:31.240
I think that was very very early even before ChatGPT came out.

s173
00:14:31.240 --> 00:14:31.720
Wow.

s174
00:14:31.720 --> 00:14:33.120
Um yeah.

s175
00:14:33.120 --> 00:14:40.960
And then apps were something that comes along because you have some model capabilities you want people to experience it well.

s176
00:14:40.960 --> 00:14:43.560
Not many people can use it with API, right?

s177
00:14:43.560 --> 00:14:46.680
We can't expect everyone to experience with API.

s178
00:14:46.680 --> 00:14:58.360
So, we need good um user interaction, you know, interfaces, good apps, good scenarios that people can can you know, experience experience model with.

s179
00:14:58.360 --> 00:15:07.160
I think actually those apps covered more than 300 million people around 200 countries globally.

s180
00:15:07.160 --> 00:15:10.200
And I think over a million companies as well.

s181
00:15:10.200 --> 00:15:17.960
Yeah, this was a mind-blowing when I heard about the the size and we we don't often realize the size of of of this type of usage already.

s182
00:15:17.960 --> 00:15:34.360
And and that I kind of brings me to the question around um open source business model, and and all of that, which is the always existing question, which is right now it's nice to open source model, but you you also need to have some revenue stream, right?

s183
00:15:34.360 --> 00:15:41.400
So, I guess M3 is something you you decided, for instance, to be for free, and I'm I think it's it's it's great for the world.

s184
00:15:41.400 --> 00:15:42.880
Um how do you see this?

s185
00:15:42.880 --> 00:15:45.960
Do you also have some specific models you use for the app?

s186
00:15:45.960 --> 00:15:52.680
Do you think about Do you think in the future you'll keep It's probably hard to say for sure, but do you think you'll keep open sourcing models?

s187
00:15:52.680 --> 00:15:55.680
How is the culture around open sourcing right now?

s188
00:15:55.680 --> 00:16:01.440
Personally, and also for the model research team, we always hope to open source the models.

s189
00:16:01.440 --> 00:16:03.120
Um that is our plan.

s190
00:16:03.120 --> 00:16:09.680
Because we really see how the open source community together can help the model build better.

s191
00:16:09.680 --> 00:16:20.520
For example, we receive a lot of um feedbacks on the model performance from the great community, and we receive PRs on uh wh- whatever we open source.

s192
00:16:20.520 --> 00:16:24.480
And those are very very valuable and come comes to our later versions.

s193
00:16:24.480 --> 00:16:27.760
So, definitely open sourcing is great.

s194
00:16:27.760 --> 00:16:28.120
That's great.

s195
00:16:28.120 --> 00:16:33.680
And actually, do you have some ask for the audience, people who are using uh M3 or MiniMax?

s196
00:16:33.680 --> 00:16:38.440
Is there something you would love them to send back to you as feedback?

s197
00:16:38.440 --> 00:16:44.000
Do you Do you, for instance, do you read when people try to modify the models or play around, you know, tweaks?

s198
00:16:44.000 --> 00:16:50.880
Or what is the best thing you you think you can take from the community uh for future models, for instance?

s199
00:16:50.880 --> 00:16:51.960
Mhm.

s200
00:16:51.960 --> 00:16:57.560
I would say whatever um issues that people are running into, especially with multimodality, right?

s201
00:16:57.560 --> 00:17:00.120
This is the first time that we're combining it together.

s202
00:17:00.120 --> 00:17:03.560
We are definitely going more ambitious on that in the future.

s203
00:17:03.560 --> 00:17:06.360
It might have some flaws right now, but we are improving on that.

s204
00:17:06.360 --> 00:17:13.880
So, whatever that's uh feedback that model is not good doing that great, we will definitely improve that in future versions.

s205
00:17:13.880 --> 00:17:17.680
And also, whatever features that uh people want.

s206
00:17:17.680 --> 00:17:21.439
Say, you know, for example, thinking effort.

s207
00:17:21.439 --> 00:17:21.680
Right?

s208
00:17:21.680 --> 00:17:23.560
Some people ask for that.

s209
00:17:23.560 --> 00:17:29.240
Um like everyone can ask, and we will try to accomplish that in the future models.

s210
00:17:29.240 --> 00:17:29.480
Yeah.

s211
00:17:29.480 --> 00:17:34.160
Do you see a lot of usage right now already in multimodality in terms of coding agents?

s212
00:17:34.160 --> 00:17:38.120
I feel like it's it's a little bit un- underexplored.

s213
00:17:38.640 --> 00:17:47.800
It is It is, but um it can actually unlocks a lot of capabilities and a lot of uh agent applications.

s214
00:17:47.800 --> 00:17:55.840
Say that for example, you want a model to read a PP- PowerPoint, uh or to read some report that is not very structured.

s215
00:17:55.840 --> 00:18:05.680
Um and you want it to understand a very long video, say that you dump in a long-playing video, and then you want the model to act uh using some tools

s216
00:18:05.680 --> 00:18:07.520
uh after understanding it.

s217
00:18:07.520 --> 00:18:12.640
And it unlocks a wide variety of um agent use cases.

s218
00:18:12.640 --> 00:18:18.880
So, like the agent could finally watch my YouTube tutorial and understand how to use my coding tools, how I described it?

s219
00:18:18.880 --> 00:18:20.520
Is it something like that?

s220
00:18:20.520 --> 00:18:21.200
Uh

s221
00:18:21.200 --> 00:18:25.200
Could the agent finally watch YouTube tutorials and understand things from them?

s222
00:18:25.200 --> 00:18:26.520
Yeah, yeah, yeah.

s223
00:18:26.520 --> 00:18:27.400
I think so.

s224
00:18:27.400 --> 00:18:30.680
Do you use a lot of uh agent coding tools internally?

s225
00:18:30.680 --> 00:18:35.240
Is it like I mean, coding for sure, but like is it also already in terms of research?

s226
00:18:35.240 --> 00:18:37.600
Is it automated part or not?

s227
00:18:37.600 --> 00:18:38.720
How How does this uh

s228
00:18:38.720 --> 00:18:39.280
Yes.

s229
00:18:39.280 --> 00:18:39.640
work?

s230
00:18:39.640 --> 00:18:42.840
Um we have our own research harnesses.

s231
00:18:42.840 --> 00:18:46.600
Um we build our own research harnesses that automate our workflows.

s232
00:18:46.600 --> 00:18:49.800
I would say a lot of our workflows are automated.

s233
00:18:49.800 --> 00:18:57.760
You can see how um the latest frontier models all pursues capability less kernel optimization, right?

s234
00:18:57.760 --> 00:19:01.400
Like let the model post string other models.

s235
00:19:01.400 --> 00:19:03.720
Um let the model build data.

s236
00:19:03.720 --> 00:19:05.840
Auto data, stuff like that.

s237
00:19:05.840 --> 00:19:11.000
Um you can see how more and more models are capable of doing those, including M3.

s238
00:19:11.000 --> 00:19:16.680
Actually, we were very good at those cases, longer horizons and coronal organizations.

s239
00:19:16.680 --> 00:19:27.320
Um and so we can use that model capability, harness it together, and help with our um daily routine, and make our iterations even faster.

s240
00:19:27.320 --> 00:19:29.640
Is M3 building M4 already?

s241
00:19:29.640 --> 00:19:32.120
Um building M3.1.

s242
00:19:32.120 --> 00:19:33.773
M3.1, okay.

s243
00:19:33.773 --> 00:19:34.000
[laughter]

s244
00:19:34.000 --> 00:19:34.880
Let's hit the gym.

s245
00:19:34.880 --> 00:19:36.000
Already.

s246
00:19:36.000 --> 00:19:45.120
Um I I would love to finish on what what you find exciting in the coming month, what what do you think it it can be Asia in terms of feature

s247
00:19:45.120 --> 00:19:54.200
or or things you want to see happening in in AI or more generally in terms of whatever whatever really is top of your mind I would say and it's going to

s248
00:19:54.200 --> 00:19:55.480
happen.

s249
00:19:55.480 --> 00:19:58.280
A lot of things are very exciting.

s250
00:19:58.280 --> 00:20:12.000
Um but what I recently find the most exciting would be a multi-agents that I think a lot of AI applications are using, model routing, multi-agents um that allows even more capabilities,

s251
00:20:12.000 --> 00:20:23.000
even more complex tasks, and also it tells us what the models are capable and not capable of, and you can, you know, do a lot of things with that.

s252
00:20:23.000 --> 00:20:24.760
It's pretty exciting.

s253
00:20:24.760 --> 00:20:26.000
Thanks a lot, Alif.

s254
00:20:26.000 --> 00:20:27.400
Pleasure to have you.

s255
00:20:27.400 --> 00:20:28.840
Thanks for having me.

s256
00:20:28.840 --> 00:20:30.104
Thanks, everyone.

s257
00:20:30.104 --> 00:20:32.104
[applause]
