WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=KLDdXOw6jIc

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/KLDdXOw6jIc.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:12.639 --> 00:00:16.560
and welcome back for those on the stream and those those in person.

s3
00:00:16.560 --> 00:00:31.119
um we take tend to basically take these longer sessions between uh all the sort of mainstage keynotes to reflect on things that um you know are particularly important but like don't have like a significant like sort of launch moments.

s4
00:00:31.119 --> 00:00:39.920
Today we're very lucky to have people working on Omni and VO Nano Banana like the you know the world's best generative models here with us.

s5
00:00:39.920 --> 00:00:45.195
Uh, Demetrio, I I I first saw you when you were posting about your office.

s6
00:00:45.195 --> 00:00:46.719
[laughter]

s7
00:00:46.719 --> 00:00:53.600
Um, I think you're you're probably number one uh Google Google's number one office influencer at least in in San Francisco.

s8
00:00:53.600 --> 00:00:55.280
I think you like you like to bike as well.

s9
00:00:55.280 --> 00:00:56.879
You like to take photos of

s10
00:00:56.879 --> 00:00:57.840
bike here.

s11
00:00:57.840 --> 00:00:58.559
Yeah.

s12
00:00:58.559 --> 00:01:01.760
Um, but you know, but also you work on video models.

s13
00:01:01.760 --> 00:01:02.239
That's right.

s14
00:01:02.239 --> 00:01:05.680
Um, Shane, I I met you I think at like a dinner.

s15
00:01:05.680 --> 00:01:06.080
Yeah.

s16
00:01:06.080 --> 00:01:13.200
Um and uh and uh and I I remember you were trying to get me invested in like one of the companies.

s17
00:01:13.200 --> 00:01:14.799
I forget forget which one.

s18
00:01:14.799 --> 00:01:15.980
Forget about that.

s19
00:01:15.980 --> 00:01:17.759
[laughter]

s20
00:01:17.759 --> 00:01:21.759
But now but now you're um now you're working on Omni Thinking.

s21
00:01:21.759 --> 00:01:24.320
Um and and just you know a bunch of other

s22
00:01:24.320 --> 00:01:25.520
Gemini RL.

s23
00:01:25.520 --> 00:01:25.840
Yeah.

s24
00:01:25.840 --> 00:01:26.560
Yeah.

s25
00:01:26.560 --> 00:01:34.960
Uh and Nicole also uh the rest of the gen media models uh nano banana and uh all and everything you just launched actually even this week.

s26
00:01:34.960 --> 00:01:35.520
Uh,

s27
00:01:35.520 --> 00:01:35.920
yeah.

s28
00:01:35.920 --> 00:01:37.360
We launched some APIs.

s29
00:01:37.360 --> 00:01:37.600
Yeah.

s30
00:01:37.600 --> 00:01:37.840
Yeah.

s31
00:01:37.840 --> 00:01:38.479
Yeah.

s32
00:01:38.479 --> 00:01:41.600
And I haven't tried to convince you to invest in anything, but maybe I should.

s33
00:01:41.600 --> 00:01:44.079
I mean, so I try not to be an investor.

s34
00:01:44.079 --> 00:01:45.360
People just convince me anyway.

s35
00:01:45.360 --> 00:01:51.600
I'm like just, okay, well, I'm not that rich, but know like you can't not try to invest in some of these things.

s36
00:01:51.600 --> 00:01:56.720
And, you know, for those of us who are not working at a Frontier Lab, this is the best this closest we'll ever get.

s37
00:01:56.720 --> 00:02:02.320
Um, so yeah, actually, let's kind of recap since you're closest to it and we just did it, like what was launched this week?

s38
00:02:02.320 --> 00:02:03.680
What should people go try out?

s39
00:02:03.680 --> 00:02:04.000
Yeah.

s40
00:02:04.000 --> 00:02:07.040
Um so yesterday we had two launch moments.

s41
00:02:07.040 --> 00:02:16.239
Uh one of them we launched NanoBanana 2 light uh which is our fastest, cheapest um image model in the nano banana model family.

s42
00:02:16.239 --> 00:02:18.959
Um and it's better than the original NanoBanana.

s43
00:02:18.959 --> 00:02:28.319
Um so really for most people um that model replaces what you you know used and love the original Nano Banana for across like generation and editing and it gets really close

s44
00:02:28.319 --> 00:02:32.080
to the frontier quality of of the kind of mainland bigger models.

s45
00:02:32.080 --> 00:02:33.519
So that that's really exciting.

s46
00:02:33.519 --> 00:02:43.440
I think if you look at some of the demos or like things that people have been trying like getting kind of that like 3 second latency just unlocks a whole bunch of things that you can do with like ideation and iteration

s47
00:02:43.440 --> 00:02:54.720
and it's just really fun and the model's getting to a point where like the quality is really good um where um it you know you can use it for iteration but you can also use some of those outputs as just kind of like ready um production output.

s48
00:02:54.720 --> 00:02:55.840
So that's really exciting.

s49
00:02:55.840 --> 00:03:03.040
Um and then second launch we finally um launched the Gemini Omni Flash APIs um that we pre-announced at IO.

s50
00:03:03.040 --> 00:03:04.800
So thank you for waiting.

s51
00:03:04.800 --> 00:03:16.480
Um and that you know is the first time that we're making the APIs available for developers and it's basically really exciting kind of video generation and editing and we're pricing it the same as Y31

s52
00:03:16.480 --> 00:03:16.879
fast.

s53
00:03:16.879 --> 00:03:21.040
So we're getting you kind of like really really good quality for a really awesome price hopefully.

s54
00:03:21.040 --> 00:03:22.159
Um

s55
00:03:22.159 --> 00:03:24.319
yeah, I mean that that's incredible.

s56
00:03:24.319 --> 00:03:33.280
I'm actually really So when you guys launched Omni for the first time, you also did a podcast uh with Logan who couldn't be here today uh and you added like a sloth

s57
00:03:33.280 --> 00:03:35.760
uh and and Ramen and all these all these things.

s58
00:03:35.760 --> 00:03:37.680
I actually really want to do that to our videos.

s59
00:03:37.680 --> 00:03:40.560
I just didn't have an API for it because obviously I have to automate the whole thing.

s60
00:03:40.560 --> 00:03:41.760
So thank you for the API.

s61
00:03:41.760 --> 00:03:43.440
Uh that is my favorite use case.

s62
00:03:43.440 --> 00:03:44.560
Everybody should do that.

s63
00:03:44.560 --> 00:03:48.159
Um I got a cat which is probably like the most boring of the animals.

s64
00:03:48.159 --> 00:03:50.080
Um if you don't know what we're talking about, you should look it up.

s65
00:03:50.080 --> 00:03:50.560
It's very funny.

s66
00:03:50.560 --> 00:03:51.440
Feurer.

s67
00:03:51.440 --> 00:03:54.560
um Furer who's um you know on on the team did that.

s68
00:03:54.560 --> 00:04:00.239
Furer is the number one guy you should follow for you should follow ideas on okay what can this thing do?

s69
00:04:00.239 --> 00:04:00.720
Yes.

s70
00:04:00.720 --> 00:04:01.040
Right.

s71
00:04:01.040 --> 00:04:01.519
Yes.

s72
00:04:01.519 --> 00:04:03.360
He he's he's amazing at that.

s73
00:04:03.360 --> 00:04:06.720
I've tried to get him for the last two years to come to AIE.

s74
00:04:06.720 --> 00:04:07.519
He hasn't made it yet.

s75
00:04:07.519 --> 00:04:08.879
He's actually come in person.

s76
00:04:08.879 --> 00:04:11.040
He just didn't want to speak because he's anonymous.

s77
00:04:11.040 --> 00:04:11.920
I know.

s78
00:04:11.920 --> 00:04:13.840
I I want to say his real name but I can't say his real name.

s79
00:04:13.840 --> 00:04:14.142
No

s80
00:04:14.142 --> 00:04:14.319
[laughter]

s81
00:04:14.319 --> 00:04:15.680
no we won't we won't do that to him.

s82
00:04:15.680 --> 00:04:16.799
But you should really follow him.

s83
00:04:16.799 --> 00:04:17.759
He's amazing.

s84
00:04:17.759 --> 00:04:18.799
He did all that work.

s85
00:04:18.799 --> 00:04:24.400
I actually met him uh in the office uh when we did the podcast I think and I didn't realize it was him.

s86
00:04:24.400 --> 00:04:27.440
So his badge doesn't say Popers.

s87
00:04:27.440 --> 00:04:27.759
Yeah,

s88
00:04:27.759 --> 00:04:28.400
I know.

s89
00:04:28.400 --> 00:04:33.840
So he used to be part of uh Replicate and Replicate had this joke where like everyone was Deep Fates.

s90
00:04:33.840 --> 00:04:36.320
Deep Fates is this like kind of mysterious character and replicate.

s91
00:04:36.320 --> 00:04:38.720
Replicate is very cool company and both was part of it.

s92
00:04:38.720 --> 00:04:49.919
Um, so, okay, one thing I want to get on there before I go into like sort of the the the the sort of omniper is we added cats, we added sloths,

s93
00:04:49.919 --> 00:04:51.759
very cool, very cute, very fun.

s94
00:04:51.759 --> 00:04:59.360
Uh, what are the, you know, inspire people as to like what are the more sort of workhorse use cases that maybe are not just demos, you know?

s95
00:04:59.360 --> 00:04:59.600
Yeah.

s96
00:04:59.600 --> 00:05:07.840
So, so obviously the hero capability of the model or maybe there's two like one is the ability to kind of take in anything as input and then get video on the other side.

s97
00:05:07.840 --> 00:05:18.080
Obviously in the future and and we've kind of talked about this as a pre-announce like we want to get the other output modalities out as well but basically what that means is you know you can take a set of images that you have as maybe a storyboard.

s98
00:05:18.080 --> 00:05:24.800
You can take like an audio track as a reference of you know like a voice that you want a character to speak and then you can get a video on the other side.

s99
00:05:24.800 --> 00:05:32.400
So like that just unlocks a whole bunch of things that you can do in like you know short film production or you know shorts we've launched on YouTube as well

s100
00:05:32.400 --> 00:05:36.400
um to help creators kind of like create um content more easily.

s101
00:05:36.400 --> 00:05:38.880
Um and then the other one is obviously video editing.

s102
00:05:38.880 --> 00:05:46.960
Like that's another thing that we're really excited about that we're just making easier because now you can use natural language to take a video, you know, add something, remove something.

s103
00:05:46.960 --> 00:05:49.120
Sloth is obviously like fun example.

s104
00:05:49.120 --> 00:05:59.520
Um, but there there's obviously kind of there's consumer use cases that we kind of had in mind where, you know, you could take your beach vacation video that was too noisy and you want to clean up that noise.

s105
00:05:59.520 --> 00:06:04.960
Maybe in the past you wouldn't have because you didn't have the tools or you didn't know what the tools were that you needed to go to.

s106
00:06:04.960 --> 00:06:07.600
So, that's one use case that you can, you know, go to.

s107
00:06:07.600 --> 00:06:16.479
We've seen a lot of folks use it for kind of marketing ad campaign creation and I'm excited to see more of those use cases as we launch the APIs.

s108
00:06:16.479 --> 00:06:23.759
um because obviously like we don't we don't see all of it in the first party products but I'm really excited for people to start to explore that um in the API.

s109
00:06:23.759 --> 00:06:27.520
So those are just some of the kind of like high level um things that have come up.

s110
00:06:27.520 --> 00:06:30.880
U people also use it to create like education materials.

s111
00:06:30.880 --> 00:06:31.199
Yes.

s112
00:06:31.199 --> 00:06:33.840
Um and like like that's really exciting.

s113
00:06:33.840 --> 00:06:43.280
I think we're all we've all kind of talked about being excited about the future of education where like everything can be kind of customized to you and personalized to your knowledge level

s114
00:06:43.280 --> 00:06:47.919
and the style that you prefer and and so this is kind of just like a step in that direction.

s115
00:06:47.919 --> 00:06:48.160
Yeah.

s116
00:06:48.160 --> 00:06:54.400
I I I sort of actually used just none of yesterday, but my my parents are visiting and there was there was a very fun sort of use case.

s117
00:06:54.400 --> 00:07:04.479
They I bought some gadget off from Amazon that they wanted and the instructions to use it was were only in English and there was plenty of diagrams or whatever and I took a picture of it and said, you know, translate this into Romanian.

s118
00:07:04.479 --> 00:07:04.880
Yes.

s119
00:07:04.880 --> 00:07:06.479
And keep everything else the same, right?

s120
00:07:06.479 --> 00:07:07.599
So it was amazing, right?

s121
00:07:07.599 --> 00:07:12.960
Like it was just like, yeah, it looks identical and it has, you know, it's perfectly translated.

s122
00:07:12.960 --> 00:07:14.160
I mean, more or less, right?

s123
00:07:14.160 --> 00:07:22.080
But it's it's you know using Gemini under the hood obviously to kind of do the translation and so you can you can see this use case for video as well right like

s124
00:07:22.080 --> 00:07:26.880
the the power of text rendering in in in Omni is is quite next level.

s125
00:07:26.880 --> 00:07:38.400
So and you could you could you could think about plenty of use cases of like both text rendering translation internalization all sorts of things that would be actually genuinely useful to a lot of different people and sort of broader access to

s126
00:07:38.400 --> 00:07:42.080
either you could like redub a video or whatever it is that you wanted to do.

s127
00:07:42.080 --> 00:07:45.520
like there's plenty of different things that you could you could think about doing.

s128
00:07:45.520 --> 00:07:45.919
Yeah.

s129
00:07:45.919 --> 00:07:54.560
Um one of the most enlightening conversations I have on my podcast is with uh just people researchers at the frontier of these things.

s130
00:07:54.560 --> 00:08:05.840
Um I had one with um Ethan from the XAI video team, the Grock video team who was basically saying like you know the next trend is actually not just like single model, it's more like video agents.

s131
00:08:05.840 --> 00:08:11.680
Um, and I don't know if that terminology resonates uh obviously for for very relevant for RL.

s132
00:08:11.680 --> 00:08:17.120
Uh, but it was it was basically kind of like giving up on like trying to do everything in in effectively one pass.

s133
00:08:17.120 --> 00:08:24.355
Um, do you feel that same way or is it still an open research question which way the trends are going?

s134
00:08:24.355 --> 00:08:24.720
[snorts]

s135
00:08:24.720 --> 00:08:25.199
Yeah.

s136
00:08:25.199 --> 00:08:38.320
So um what kind of excite me most is really when the symbolic kind of foundational models and this kind of like video foundational model can actually kind of really work together and u in a way the if you look at the beginning of the generative

s137
00:08:38.320 --> 00:08:48.880
sort of like image generation video generation a lot of it kind of started when the language model got good enough to provide a very detailed captioning like from stable fusion days or kind of dowi 2 days.

s138
00:08:48.880 --> 00:09:03.760
So um so basically like language is extremely u helpful representation uh one is that it's kind of universal but the other kind of more um technical thing like kind of my hypothesis is like um one very difficult thing about machine learning

s139
00:09:03.760 --> 00:09:06.880
is um this sort of like spirious coordination.

s140
00:09:06.880 --> 00:09:13.040
So you don't know you know if the if this kind of feature right that's kind of predictive is actually causal factor or not.

s141
00:09:13.040 --> 00:09:13.920
There are two ways.

s142
00:09:13.920 --> 00:09:19.680
One is we can have really diverse data training data like from every intervention of the causal graph.

s143
00:09:19.680 --> 00:09:28.320
The other is you condition the causal information and conditioning the language is kind of like conditioning like a coal information of the of the kind of world.

s144
00:09:28.320 --> 00:09:29.360
So um

s145
00:09:29.360 --> 00:09:32.240
which is a prompt or a concept what

s146
00:09:32.240 --> 00:09:42.959
yeah exactly so if you look at like you know how we going to describe this video how this kind of image is actually very close to you know how would describe this kind of causality you know behind this like how this is kind of generated.

s147
00:09:42.959 --> 00:09:50.640
So one is like that can really allow for very rich generalization and then uh very kind of just like a good model.

s148
00:09:50.640 --> 00:09:58.399
Um the other is so eight months ago uh we put the evaluation paper called video models zero shot learners and reasoners.

s149
00:09:58.399 --> 00:09:58.959
Yes.

s150
00:09:58.959 --> 00:10:10.399
So that was a kind of you know it's it's a confirmed paper and then later on actually the N banana team followed up with a vision banana paper that basically used n banana to do but essentially the idea is

s151
00:10:10.399 --> 00:10:16.240
uh video model is extremely good sort of a foundation model for space and time kind of information.

s152
00:10:16.240 --> 00:10:29.200
So um classic computer vision tasks a lot of could be kind of zero shorted and when you like say feed in some like a visual quiz uh it can you know there's definitely like a lot to improve it can kind of solve

s153
00:10:29.200 --> 00:10:45.200
and it can um like robotics kind of like seeing it has really good kind of physical intuitions like word model uh and I think the the key is really the kind of mix of the visual kind of reasoning and then the text kind of reasoning kind of all tied together

s154
00:10:45.200 --> 00:10:59.120
Um obviously you know like whether doing it you know as kind of unified model versus like just kind of agent coation I think that's more like uh it's going to be more kind of incremental you know how it's going to I imagine everything's going to go into like a single model eventually

s155
00:10:59.120 --> 00:11:12.000
but right now there's like a lot you can do if you uh basically take like really good video understanding image understanding Gemini agentically with anomy and that's actually gonna yeah our team is like exploring a lot

s156
00:11:12.000 --> 00:11:24.240
yeah okay that there's a there's a lot in there um I I think uh one question I I am increasingly starting to wonder is does it all trend towards one product for you guys right like now you have multiple models out

s157
00:11:24.240 --> 00:11:33.440
the naming of omni does imply that eventually everything will go away and it just goes into omnis um is that the plan

s158
00:11:33.920 --> 00:11:34.307
is it

s159
00:11:34.307 --> 00:11:34.399
[laughter]

s160
00:11:34.399 --> 00:11:49.440
I don't know I I think I think uh maybe I mean I think eventually I I think there's sort of different trade-offs engineering research product trade-offs in like it's like for the same reason like

s161
00:11:49.440 --> 00:11:52.800
the the sorry how is it called nano banana light I don't know what the product name

s162
00:11:52.800 --> 00:11:53.920
nanob banana tite

s163
00:11:53.920 --> 00:12:09.040
nano banana too light yeah right it's it's it's it serves a particular niche right and it probably doesn't necessarily fit immediately in the same model literally checkpoint as uh something that can do 4K

s164
00:12:09.040 --> 00:12:15.120
you know uh 30 secondond videos right like they're probably not like trainable in the same quite way, right?

s165
00:12:15.120 --> 00:12:16.480
Like, so I I don't know.

s166
00:12:16.480 --> 00:12:18.720
It depends on how how far into the future you look like.

s167
00:12:18.720 --> 00:12:21.360
Sure, in five years from now, will they all be the same model?

s168
00:12:21.360 --> 00:12:22.320
Probably.

s169
00:12:22.320 --> 00:12:35.200
Uh but like, you know, six months from now, we'll we'll probably still have, you know, multiple different models doing different things because kind of from pragmatically the trade-offs are such that we we should have multiple different kinds of models.

s170
00:12:35.200 --> 00:12:36.160
Yeah, I

s171
00:12:36.160 --> 00:12:36.720
I think that's right.

s172
00:12:36.720 --> 00:12:46.320
And and just on that note, I mean, we did call it Gemini Omni because we wanted to hint at the future where Gemini just becomes fully multimodal in and out, right?

s173
00:12:46.320 --> 00:12:48.959
And so so it's definitely a move in that direction.

s174
00:12:48.959 --> 00:12:54.639
I think we'll probably see a move in the direction where Omni also generates images and edits images and all those kinds of things.

s175
00:12:54.639 --> 00:13:02.320
But Doo is right that I think on the way there, there's a bunch of really really useful applications of some of these more specialized models.

s176
00:13:02.320 --> 00:13:10.240
And so we we will probably continue to work on those as well because like that serves a certain need at this point in time that may not exist you know a year from now.

s177
00:13:10.240 --> 00:13:16.240
There's also like a research question about like just how much transfer there is between different kinds of modalities, right?

s178
00:13:16.240 --> 00:13:31.920
I think you may believe that there's some transfer between coding and video generation and I think most people don't necessarily believe that but they you know you could try to think that there is some some there something there or it could be a waste right to put them together to try to learn these both tasks at the

s179
00:13:31.920 --> 00:13:40.000
same time right so I think it's it's it's interesting sort of question to which extent like image and video obviously kind of there's some transfer like kind of not that different

s180
00:13:40.000 --> 00:13:47.839
there's value in in learning to output video and audio at the same time because joint audio visual is you know that's how that's how it is.

s181
00:13:47.839 --> 00:14:01.440
Um and then there's you know other kind of intersections of modalities that are not super obvious right like 3D representation coding I don't know maybe uh things like that right so like I think it's worth sort of exploring the different corners there and we are actively doing that

s182
00:14:01.440 --> 00:14:05.440
um with a focus towards like what people actually want to do with these models

s183
00:14:05.440 --> 00:14:18.000
yeah um what one thing I feel I feel like uh I'm surprised by but also I feel like it's insufficiently answered is what is the correct intermediate representation Um,

s184
00:14:18.000 --> 00:14:20.079
so captioning, right?

s185
00:14:20.079 --> 00:14:21.839
XI does captioning.

s186
00:14:21.839 --> 00:14:23.279
Omni does captioning.

s187
00:14:23.279 --> 00:14:28.160
Um, and I I I understand how captioning works for images.

s188
00:14:28.160 --> 00:14:33.600
Um, and I understand that you can extend it into to video and and sort of guide it across time.

s189
00:14:33.600 --> 00:14:35.519
It just feels very inefficient.

s190
00:14:35.519 --> 00:14:38.399
It there's got to be I feel like there should be something better.

s191
00:14:38.399 --> 00:14:49.839
Uh maybe it's code and maybe we generate you know and obviously I think a lot of um ffmpeg and mapplot um what's the three blue one brown one manm

s192
00:14:49.839 --> 00:15:01.279
um a lot of like video is generated through code and maybe that's like the optimal representation uh any hypothesis as to like is is it better or is just English all you need

s193
00:15:01.279 --> 00:15:10.240
well as so I'm in the Gemini and they know we do like a lot of RL agent and of course kind of coding so yeah We we're definitely exploring the coding representations.

s194
00:15:10.240 --> 00:15:10.720
Yeah.

s195
00:15:10.720 --> 00:15:13.120
As kind of better kind of way to represent.

s196
00:15:13.120 --> 00:15:13.440
Yeah.

s197
00:15:13.440 --> 00:15:17.665
But you know like do you what's your probability estimate on like

s198
00:15:17.665 --> 00:15:18.000
[laughter]

s199
00:15:18.000 --> 00:15:23.440
if we just output binaries like we just you know like just it's just ones and zeros.

s200
00:15:23.440 --> 00:15:33.680
Um I I guess maybe a kind of similar discussion was like um basically is the language the right representation like right.

s201
00:15:33.680 --> 00:15:42.320
So uh one kind of question for example uh professor you know like some ask is like you know why why does the channel of thought need to be in the natural language?

s202
00:15:42.320 --> 00:15:42.639
Yes.

s203
00:15:42.639 --> 00:15:48.800
Can it just be the kind of any kind of like continuous tokens just any amount of you know additional computations.

s204
00:15:48.800 --> 00:15:56.079
Um so one is like obviously the test like adaptive compute is going to give like you know better results.

s205
00:15:56.079 --> 00:16:04.079
So it's that but what really kind of made CH thought so you know like four years ago I wrote you know the larger model zero sort reasoner and then self-improvement.

s206
00:16:04.079 --> 00:16:16.880
So I kind of know from the very early day but the reason like it works really well is um right now the recipe that works is the pre-training that scales a lot and then that basically like learns a lot of intelligence.

s207
00:16:16.880 --> 00:16:30.720
there are a lot of you know scaling RL but those are still like extremely kind of comput incent intensive to extract the information and um you really want to rely the intelligence on that so basically by tying the

s208
00:16:30.720 --> 00:16:39.839
sort of like a reasoning in the natural language you basically directly use the intelligence of the pre-training to it while if you remove that kind of constraints then you're not

s209
00:16:39.839 --> 00:16:52.800
um and these days uh I feel the a lot of advancements in the texts but also doing this kind of multimodal space is very driven by this uh kind of text as a kind of great

s210
00:16:52.800 --> 00:16:54.480
uh sort of representation.

s211
00:16:54.480 --> 00:16:56.480
Yeah, it's a good backbone.

s212
00:16:56.480 --> 00:16:57.199
Yeah,

s213
00:16:57.199 --> 00:16:59.199
I I think to me it's even simpler than that.

s214
00:16:59.199 --> 00:17:01.279
It's text is is how we communicate.

s215
00:17:01.279 --> 00:17:11.520
So I think fundamentally if you're building kind of products that humans will be interfacing with um like like that we will be using text somehow if it's a text interface, right?

s216
00:17:11.520 --> 00:17:12.480
Not not for everything.

s217
00:17:12.480 --> 00:17:15.280
So I think it's it's natural to default to that.

s218
00:17:15.280 --> 00:17:15.919
Yeah.

s219
00:17:15.919 --> 00:17:24.880
Obviously there's like a conf discussion you know some arrow like arrow maximalists is like oh we don't care about you know kind of channel those kind of like stuff it's just just additional compute

s220
00:17:24.880 --> 00:17:26.880
sure but I personally yeah

s221
00:17:26.880 --> 00:17:32.720
ro maximalists I wonder I wonder who who qualifies in that description David silver

s222
00:17:32.720 --> 00:17:48.960
ah okay yeah I mean they they've just left to to start their thing um interesting okay so uh I I mean I think I'm very interested in just like better representations because I think that's one of our themes that we're curating today uh at the world fair is world models.

s223
00:17:48.960 --> 00:17:52.640
You mentioned the word world models but it's not something that's like super well defined.

s224
00:17:52.640 --> 00:17:57.760
I think everyone's like sort of converging on some version of it that is like the ideal.

s225
00:17:57.760 --> 00:17:58.400
Sure.

s226
00:17:58.400 --> 00:17:59.600
Everything is a world model now.

s227
00:17:59.600 --> 00:18:00.640
It's sort of a

s228
00:18:00.640 --> 00:18:02.880
it's not it's not that useful, right?

s229
00:18:02.880 --> 00:18:06.480
So I just gave a keynote at the IER world model workshop.

s230
00:18:06.480 --> 00:18:06.720
Yeah.

s231
00:18:06.720 --> 00:18:12.000
And then uh yeah essentially uh I definitely encourage to check out the definition by Jandra Matic.

s232
00:18:12.000 --> 00:18:15.200
He's like the you know OG computer vision professor UC Berkeley.

s233
00:18:15.200 --> 00:18:18.559
Uh he has pretty you know bit of word to say about world model

s234
00:18:18.559 --> 00:18:28.480
but also kind of Schmidt Herburver's kind of how he defined the world model from 2019 like 1990 sort of uh uh you know like Wayne was just basically just that kind of model base.

s235
00:18:28.480 --> 00:18:39.039
Uh for me the word model is basically just the model in the model based RL and I feel that has sufficient to describe but obviously you know there are like a lot of uh FE had a kind of nice blog post about what

s236
00:18:39.039 --> 00:18:41.600
about yeah this kind of broken down

s237
00:18:41.600 --> 00:18:43.440
um but yeah

s238
00:18:43.440 --> 00:18:55.840
yeah I mean so you know I I'll end this part of the conversation but like I I do think that language to me relying on language as like the sort of like the narrow pipe through which everything goes through

s239
00:18:55.840 --> 00:18:58.320
um still is like a lossy compression.

s240
00:18:58.320 --> 00:18:59.919
No, no, no. But we're not seeing that, right?

s241
00:18:59.919 --> 00:19:02.960
We're basically saying the video model and the language together.

s242
00:19:02.960 --> 00:19:06.640
So, so I think the language alone is uh not sufficient.

s243
00:19:06.640 --> 00:19:09.679
That's why we feel like the video is a very complement model.

s244
00:19:09.679 --> 00:19:09.919
Right?

s245
00:19:09.919 --> 00:19:18.480
Now the um you know kind of v omni many people feel as uh you know generating kind of pretty videos but I think our vision it's it's much more than that.

s246
00:19:18.480 --> 00:19:25.120
It's a missing foundational model that's absolutely required if you want to make the AGI that match to humans not just a jacked one.

s247
00:19:25.120 --> 00:19:26.080
Yeah.

s248
00:19:26.080 --> 00:19:27.200
Um okay.

s249
00:19:27.200 --> 00:19:38.640
So one one other thing you know you you mentioned on the vision side um and I'm kind of curious how sort of uh parallel you know in terms of your research careers

s250
00:19:38.640 --> 00:19:50.480
um this development is like I think basically a lot of vision people have crossed over into more model people um a lot of vision people also become generative video and image people

s251
00:19:50.480 --> 00:19:57.353
and is it just as simple as you know reversing uh image to text and then now it's text to image like

s252
00:19:57.353 --> 00:19:58.559
[laughter]

s253
00:19:58.559 --> 00:20:02.080
is is that if I mean that effectively was the diffusion process.

s254
00:20:02.080 --> 00:20:16.160
Um I I just you know I I just see the career paths of the people that I talked to and and see and I I I see this overall trend of research directions and I just wanted you to guys to sort of reflect on on that.

s255
00:20:16.160 --> 00:20:24.160
I mean I certainly went that way right I I started long time ago uh doing computer vision sort of object detection recognition things like that.

s256
00:20:24.160 --> 00:20:39.679
Uh I think just that's just simpler problem right just generation is just harder like it's a it's a different kind of mapping right you map from the the inverse mapping is not as simple as just inverting the the kind of network you use right it's it's a it's it's more ambiguous right to go from cat to image

s257
00:20:39.679 --> 00:20:46.640
of a cat and in some ways it's also a loop because your vision work creates the synthetic labels that then continues

s258
00:20:46.640 --> 00:20:47.789
I mean sure

s259
00:20:47.789 --> 00:20:48.400
[laughter]

s260
00:20:48.400 --> 00:20:55.360
I don't know I don't know I try to validate my my sort of theories about how fields develop how how careers has progressed through this

s261
00:20:55.360 --> 00:21:03.919
I mean for like the the the better the understanding side gets like we have seen that the generation side also gets better right so like like

s262
00:21:03.919 --> 00:21:05.840
it's completely bootstrapping yeah it's

s263
00:21:05.840 --> 00:21:18.960
and so like like like there's definitely they're there to that thesis and I think yeah I think a lot of people have kind of like I I definitely worked with a lot of um image understanding people who became image generation people you know and then some of them have moved on to video because it's kind of like

s264
00:21:18.960 --> 00:21:24.960
the next thing where you have so many more dimensions to work with so yeah I'm curious about you spec as your

s265
00:21:24.960 --> 00:21:34.240
so I definitely like recommend start with understanding recognition because that's basically discriminator and then that's going to lead to better generation and that's what the bridge is basically reinforcement learning

s266
00:21:34.240 --> 00:21:42.640
so my um my kind of journey is I initially kind of worked on the algorithmic research in the gent model against some like you know amnest kind of generation

s267
00:21:42.640 --> 00:21:53.360
and then I worked on like RL and robotics um and then like six years ago I was like leading like a moonshot on the dexterity it was pretty early but I see now everyone's kind of doing

s268
00:21:53.360 --> 00:22:02.799
uh four years ago I basically kind of figured out that this like symbolic AGI is going to accelerate much faster than the kind of physical AGI kind of counterpart.

s269
00:22:02.799 --> 00:22:06.640
So uh I decided to kind of like language models and then those things.

s270
00:22:06.640 --> 00:22:13.039
Um and then recently kind of work with Doomi and then like omni team I quite enjoy kind of collaboration there.

s271
00:22:13.039 --> 00:22:27.039
the what I quite enjoy uh what I recommend definitely to the researcher is to uh definitely kind of explore or at least like get exposure to what the top people in each of the community are like looking at how they kind of think about problems.

s272
00:22:27.039 --> 00:22:43.039
So when I look at the video model to me it kind of reminds me like pretty early on sort of like language model where like very early language model was a kind of creative sort of demo right you kind of like try to write like a story like mobile and then like you know GBD2 and then those

s273
00:22:43.039 --> 00:22:53.679
kind of days like LTM kind of days right and then you know uh instruction tuning you actually kind of make it usable as a chatbot but then at the chatbot stage it still had so much hallucinations

s274
00:22:53.679 --> 00:23:05.919
and instruction for wasn't good enough so it couldn't use for reasoning and when I got good enough um in pre-training and post- trainining for reasoning then you know this kind of test time scaling the RL really took off

s275
00:23:05.919 --> 00:23:15.679
to like many of the kind of best performing models and right now I think the video model is as we mentioned it's it is a complimentary foundational model and I can imagine it's going to follow a similar path

s276
00:23:15.679 --> 00:23:27.600
it's going to be very uh it's going to improve a lot instruction following a lot of uh this it's going to improve a lot in reducing coordinations to extend that it become a very reliable world model so we can kind of like intermixed

s277
00:23:27.600 --> 00:23:33.440
video like space-time simulation was a text simulation to solve like arbitrary AGI problems.

s278
00:23:33.440 --> 00:23:42.799
Also like I think the difference still is between sort of text models and like image video models is that like we haven't quite unified understanding and generation in in multimedia

s279
00:23:42.799 --> 00:23:53.919
I'd say yet like I mean I think I think without going to the details of course there's like it depends on on at which level you're thinking about this but generally like there's not that many as far as I know models

s280
00:23:53.919 --> 00:24:04.000
sot kind of you know frontier models that are genuinely kind of good at both understanding and generation of of let's videos, right?

s281
00:24:04.000 --> 00:24:06.159
Like it's a it's a it's an interesting challenge.

s282
00:24:06.159 --> 00:24:07.840
I'm not saying that we should do this.

s283
00:24:07.840 --> 00:24:14.559
Uh but but I think uh it kind of stands to reason that like you know understanding and generation are two sides of the same coin.

s284
00:24:14.559 --> 00:24:17.279
So they they kind of should be in the same model in some ways.

s285
00:24:17.279 --> 00:24:19.360
Uh but we don't necessarily always do that.

s286
00:24:19.360 --> 00:24:20.480
So yeah.

s287
00:24:20.480 --> 00:24:22.960
Uh you mentioned audio as well, right?

s288
00:24:22.960 --> 00:24:23.360
Yeah.

s289
00:24:23.360 --> 00:24:29.760
Uh is that as hard as video or qualitatively different?

s290
00:24:29.760 --> 00:24:31.679
If if so, in what way?

s291
00:24:31.679 --> 00:24:44.480
Uh, one of the interesting directions three years ago was people using um, I guess diffusion to do audio uh, as in like the the sort of refusion approach.

s292
00:24:44.480 --> 00:24:45.919
I don't know if you you guys saw that.

s293
00:24:45.919 --> 00:24:59.279
Um, and I just think it's like very interesting if a modality that we perceive which is audio is different than video actually two machines is exactly the same like there's they see no difference.

s294
00:24:59.279 --> 00:25:03.919
I mean I think on a technical level there are some differences but I think they're like relatively minor.

s295
00:25:03.919 --> 00:25:12.720
I think from my perspective audio came into into my life when we shipped V3 which was I believe the first model that did like a joint

s296
00:25:12.720 --> 00:25:14.240
with the slicing of the

s297
00:25:14.240 --> 00:25:14.400
Yeah.

s298
00:25:14.400 --> 00:25:14.960
Yeah.

s299
00:25:14.960 --> 00:25:16.480
the gold bars or whatever.

s300
00:25:16.480 --> 00:25:20.400
Um it it was the first model that did this sort of joint audiovisisual generation.

s301
00:25:20.400 --> 00:25:20.960
Yes.

s302
00:25:20.960 --> 00:25:28.880
uh like in the in the I mean there are there were other models that did kind of you know kind of kind of agentic hacking under the hood but this one was truly sort of

s303
00:25:28.880 --> 00:25:44.720
you know generating everything at once and we the reason we did that is because we felt and I think was the right choice we felt that like uh it only makes sense to generate them at the same time because there sort of kind of like from a machine learning perspective there's one latent kind of you know causal

s304
00:25:44.720 --> 00:25:56.720
kind of you know generative process right like there's something that generates you speaking it's not the pixels and then the the audio or somehow somehow generated by some other process like the lips have to move in sync with with the with the audio, right?

s305
00:25:56.720 --> 00:26:07.679
So, I think that that solved a lot of the issues that previous models had or the way that people did video generation before where it was like, okay, we generate the pixels and then we're going to hack something on top of it that like moves the lips

s306
00:26:07.679 --> 00:26:08.960
with the audio that we generate.

s307
00:26:08.960 --> 00:26:10.708
And that's was very bad.

s308
00:26:10.708 --> 00:26:11.440
[laughter]

s309
00:26:11.440 --> 00:26:18.960
And so, I think I think that was that's to me that's the the I mean after V3 like you know people were like what do you mean like there's no audio in your model?

s310
00:26:18.960 --> 00:26:22.480
like that makes no sense like once it's there like you you have to have it.

s311
00:26:22.480 --> 00:26:28.720
So I think that was that was the right choice and doing it to one single generative model I think was was the right choice.

s312
00:26:28.720 --> 00:26:38.640
One thing I kind of want to also can ask you guys an opinion as well once one difference I find the audio and then against the image and video is like the audio information is less verbalized.

s313
00:26:38.640 --> 00:26:47.520
I mean of course the TTS and stuff is trivial right but the when you get her outside like how to describe music how do you describe this like this person's

s314
00:26:47.520 --> 00:26:56.960
tone kind of pitch I feel the sort of the verbalization is insufficient and the interesting thing is that you kind of see that in two other things like taste

s315
00:26:56.960 --> 00:27:00.960
taste sense and also uh say um touch

s316
00:27:00.960 --> 00:27:14.640
like smell and then the another interesting thing is the skin color so skin color the the language is pretty limited to describe the skin color and the reason is that we're extremely uh sensitive to the small difference perturvations

s317
00:27:14.640 --> 00:27:23.840
or not skin color because that basically shows us is this person going to kill me or is can I befriend this person kind of those kind of information and then I feel the smell tastes

s318
00:27:23.840 --> 00:27:35.679
um skin color and like sound kind of stuff is very very tied into primitive a like survival kind of stuff and so our sort of sensory system is so sensitive

s319
00:27:35.679 --> 00:27:49.039
that it's intractable to Um, so for example, I asked like one the wine sort of taster and then like professional and then he basically said he kind of use like a language from like a dating, you know, describing like a, you know, partner

s320
00:27:49.039 --> 00:27:54.880
as a way to describe the taste because there's no sufficient vocab to describe.

s321
00:27:54.880 --> 00:27:56.559
Um, so I'm kind of curious.

s322
00:27:56.559 --> 00:27:56.960
Yeah.

s323
00:27:56.960 --> 00:27:58.480
Do you guys feel that?

s324
00:27:58.480 --> 00:28:04.720
I think well to some extent I think the same is true for visual information, right?

s325
00:28:04.720 --> 00:28:08.720
when you think about like a certain style or a certain aesthetic, right?

s326
00:28:08.720 --> 00:28:15.919
Like like there are some people who just have a much more kind of developed like whether it's palette or kind of visual taste and aesthetic, right?

s327
00:28:15.919 --> 00:28:25.679
Like I I think language just tends to be a bit of a limiting factor when you are trying to describe any of these things that like we experience with sensory information.

s328
00:28:25.679 --> 00:28:37.120
And to your point earlier, I think that is the kind of the reason why we are investing in world models and why we are pushing on kind of the like perception and like generation side of things because

s329
00:28:37.120 --> 00:28:41.919
it it is such a large part of how we as humans navigate the world.

s330
00:28:41.919 --> 00:28:46.399
It's a large part of how like embodied AI navigates the world.

s331
00:28:46.399 --> 00:28:56.159
Um, and and I do I do think language like does have a lot of it's it's gotten us very far and it can probably get us really far, but it it feels limiting in a lot of these kind of areas.

s332
00:28:56.159 --> 00:29:00.480
And yeah, I don't I don't really know how to describe, you know, like sense and taste.

s333
00:29:00.480 --> 00:29:03.279
Um, but yeah, I'm curious to me.

s334
00:29:03.279 --> 00:29:07.440
Um, I I yeah, I don't know that I have thought that deeply about this yet.

s335
00:29:07.440 --> 00:29:12.960
So, uh, yeah, I mean yeah, I don't have a good answer about audio.

s336
00:29:12.960 --> 00:29:22.080
I mean like I don't know the limit because I'm thinking about like well what is what is Omni bad at in terms of audio but they're all like solvable problems I find

s337
00:29:22.080 --> 00:29:34.640
uh so like with more data or better data or whatever it is so I don't know like that we have pushed the frontier so much that like we are have hit some sort of limits that are rooted in evolutionary

s338
00:29:34.640 --> 00:29:37.200
uh kind of you know limits imposed by humans.

s339
00:29:37.200 --> 00:29:38.399
I don't know.

s340
00:29:38.399 --> 00:29:41.039
He's feeling the limits of captioning which is the the thing I was

s341
00:29:41.039 --> 00:29:41.620
Yeah, exactly.

s342
00:29:41.620 --> 00:29:42.000
[laughter]

s343
00:29:42.000 --> 00:29:47.360
There there's a lot of information in the world and it connects to basically why we do work modeling you mentioned.

s344
00:29:47.360 --> 00:29:52.399
You just need srefs sref476 and then that's your that's what your journey does, right?

s345
00:29:52.399 --> 00:29:54.720
I guess maybe I can't describe this vibe but

s346
00:29:54.720 --> 00:29:59.279
well well I think that that's kind of the point of providing some of these references, right?

s347
00:29:59.279 --> 00:30:08.159
Because because like even just describing how someone talks and like their tone and and like procity and all of these things like I think I think some of these terms even like

s348
00:30:08.159 --> 00:30:10.080
I didn't used to know what they mean, right?

s349
00:30:10.080 --> 00:30:10.880
Well, now

s350
00:30:10.880 --> 00:30:11.600
yes.

s351
00:30:11.600 --> 00:30:22.960
Dispuencuencies ex like like there there's kind of an entire vocabulary that even if you're not kind of steeped in a domain, which is true for actually like most human domains that like you don't even know what it means.

s352
00:30:22.960 --> 00:30:29.840
Um and sometimes it's also a question of like if we haven't focused on those things, you know, with the large language models that they may also have gaps in those areas, right?

s353
00:30:29.840 --> 00:30:38.559
And then we feel them on the other side with generation because we're like fundamentally relying on on the language models understanding of the world to then be able to like represent it.

s354
00:30:38.559 --> 00:30:44.320
Um, so I yeah, it all kind of goes back to your question about like the the language as an intermediary.

s355
00:30:44.320 --> 00:30:54.240
Um, but yeah, I think to De's point like some of these might just be like focus areas and things that we haven't necessarily pushed on as much as we can and like as we will we will discover what the actual

s356
00:30:54.240 --> 00:30:55.919
ceiling is.

s357
00:30:55.919 --> 00:30:59.600
Yeah, as a podcaster I think a lot about sound.

s358
00:30:59.600 --> 00:31:05.520
Um, and and I I'll just offer a couple things for discussion in case in case it triggers anything with you guys.

s359
00:31:05.520 --> 00:31:19.679
Um I have three domains of rough audio which is like a music voice SFX you know is that rough okay covers everything and then also even within voice let's just let's just focus on voice forget the other two um room sound like the the echoiness

s360
00:31:19.679 --> 00:31:33.840
of like big room small room in person in a car over a phone all these like are labelable but we experience them very differently and I I often think like one of the tells of a AI video is that it is studio quality

s361
00:31:33.840 --> 00:31:48.960
because it was recorded in a studio video because that's your training data and like and and to me that's one thing actually like the most interesting thing is just uh when I tell this is how I convince people who are kind of skeptical about the need for world models because you need it even for audio

s362
00:31:48.960 --> 00:31:59.840
about well I'm further away from you so I should sound a little bit softer or more diffused and like the the video models need to pick that up because if they're going to do immersive video and audio

s363
00:31:59.840 --> 00:32:01.440
you need that

s364
00:32:01.440 --> 00:32:14.799
I I I love that example of basically like studio quality or not in a way like we don't have enough language to really describe like like this kind of echoing or like some kind of noise kind of happening we just like don't have precise enough

s365
00:32:14.799 --> 00:32:25.200
and uh if you um you know basically the reason that I think it's quite important to have like relatively information rich like kind of captioning is that we kind of rely on the natural language as a representation

s366
00:32:25.200 --> 00:32:40.399
but if you basically don't have enough uh representation that basically means the condition on the language the generation is very multimodal and if you anything can learn from the BAE kind of like you know very old you know GMBA kind of research the idea is we really want to capture most of the stoasticity

s367
00:32:40.399 --> 00:32:46.720
in the later representation and then the the X given the Z should be kind of like deterministic so yeah

s368
00:32:46.720 --> 00:32:51.760
yeah yeah um well I hope I hope there's more uh progress there and I'm sure you guys are doing

s369
00:32:51.760 --> 00:33:04.672
I even actually like facial expressions right and maybe this gets to your point about like things that we're very sensitive to right I think you can tell a lot of AI content also just by from like people's facial expressions stressful.

s370
00:33:04.672 --> 00:33:04.720
[laughter]

s371
00:33:04.720 --> 00:33:06.080
Yes.

s372
00:33:06.080 --> 00:33:06.399
Yes.

s373
00:33:06.399 --> 00:33:11.840
And we try not to contribute to it, but you know, um and or or like skin textures, right?

s374
00:33:11.840 --> 00:33:15.519
Like like the things that kind of make things look real in real life.

s375
00:33:15.519 --> 00:33:22.640
Like I you know, I can tell from the way you're nodding or from the way like your micro expressions are kind of changing of like how you're reacting to what I'm saying.

s376
00:33:22.640 --> 00:33:25.279
Like we haven't quite crossed that chasm.

s377
00:33:25.279 --> 00:33:28.399
I think like we're we're so much better than we were a year ago.

s378
00:33:28.399 --> 00:33:28.960
Yeah.

s379
00:33:28.960 --> 00:33:34.480
Um, but there's so much more headroom kind of in a lot of those things that like we as humans are super sensitive to.

s380
00:33:34.480 --> 00:33:46.240
And like I think image arguably probably is there because there's there's a lot of kind of images that I will see that like really do look indistinguishable from reality and I can't tell if they're generated or not.

s381
00:33:46.240 --> 00:33:47.360
Better than reality

s382
00:33:47.360 --> 00:33:49.919
um or well that's a different

s383
00:33:49.919 --> 00:33:51.833
No, I I think that one of the parad

s384
00:33:51.833 --> 00:33:51.840
[laughter]

s385
00:33:51.840 --> 00:33:54.320
better than what I would take on my vacation as a photo.

s386
00:33:54.320 --> 00:33:54.880
Yes.

s387
00:33:54.880 --> 00:34:01.679
One of the one of the fun experiments that we did a while ago in the team is is like can we generate videos that are better than than real videos, right?

s388
00:34:01.679 --> 00:34:06.640
So you just take the same caption from like oh yeah some video and then

s389
00:34:06.640 --> 00:34:07.279
recycle it.

s390
00:34:07.279 --> 00:34:07.600
Yeah.

s391
00:34:07.600 --> 00:34:13.760
Just just try to like describe a real video and then generate the equivalent version with omni and then do a human eval.

s392
00:34:13.760 --> 00:34:14.720
How does how does it do?

s393
00:34:14.720 --> 00:34:18.960
And then humans largely prefer AI generated

s394
00:34:18.960 --> 00:34:20.079
margin.

s395
00:34:20.079 --> 00:34:22.800
But because it's because it's the RL process,

s396
00:34:22.800 --> 00:34:23.839
that's the RL process working.

s397
00:34:23.839 --> 00:34:25.119
It's however you want to rationalize it.

s398
00:34:25.119 --> 00:34:26.320
It's not necessarily the old process.

s399
00:34:26.320 --> 00:34:29.760
It's just like I think it's just I'm not saying this is a good result.

s400
00:34:29.760 --> 00:34:41.440
I'm just saying is we have optimized in a way that like kind of potentially sort of, you know, triggers something in the human brain that like, oh, it looks it looks all a lot of the videos just look

s401
00:34:41.440 --> 00:34:41.919
better.

s402
00:34:41.919 --> 00:34:42.720
Like I'm not Yeah.

s403
00:34:42.720 --> 00:34:42.879
Yeah.

s404
00:34:42.879 --> 00:34:43.359
Yeah.

s405
00:34:43.359 --> 00:34:53.760
on on inspection on on deeper inspection they they would not actually be more useful or whatever but like if you just say side by side random YouTube video versus

s406
00:34:53.760 --> 00:35:06.320
generated version of it will you will just have a it will just look better because it's more it's a sharper more HDR uh you know the skin tone is is is better it's not again it's not more realistic

s407
00:35:06.320 --> 00:35:09.839
uh it doesn't solve your problem necessarily but it it looks better

s408
00:35:09.839 --> 00:35:13.440
I I since also depend on the sensitivity of the people.

s409
00:35:13.440 --> 00:35:23.040
Uh I was born raised in Japan and I think one thing I kind of know is like they're extremely extremely like sensitive about like you know that's why you know like architecture like food and stuff like they have.

s410
00:35:23.040 --> 00:35:32.560
Um so I talked to like a mangar like like artist there and he's like he's kind of disgusted by like the generation AI and one kind of thing he mentioned is like the eye gaze

s411
00:35:32.560 --> 00:35:35.200
eye gaze that slight difference

s412
00:35:35.200 --> 00:35:38.640
makes me makes him kind of feel creepy about like unnatural

s413
00:35:38.640 --> 00:35:40.480
like if you're looking a little bit off.

s414
00:35:40.480 --> 00:35:40.800
Yeah.

s415
00:35:40.800 --> 00:35:42.240
It's just uh Yeah.

s416
00:35:42.240 --> 00:35:44.160
just like uh it looks too fake.

s417
00:35:44.160 --> 00:35:44.480
Yeah.

s418
00:35:44.480 --> 00:35:44.960
To the point.

s419
00:35:44.960 --> 00:35:47.440
So So I think it does depend on the sensitivity and

s420
00:35:47.440 --> 00:35:47.599
Yeah.

s421
00:35:47.599 --> 00:35:48.560
Yeah.

s422
00:35:48.560 --> 00:36:01.040
All I'm saying is like you know human preferences are like a not particularly like uh reliable barometer of like what you should be optimizing for like if you just ask people do you like this or not you not necessarily get what you wanted.

s423
00:36:01.040 --> 00:36:01.359
Yeah.

s424
00:36:01.359 --> 00:36:14.240
Let let me just kind of add one thing but like four years ago there was a like debate that if the prompt engineering is going to disappear and uh my my like you know some very powerful people say you know it's going to disappear but I basically said like it shouldn't

s425
00:36:14.240 --> 00:36:25.280
because the prompt engineering like sort of you know specifying that is like the the only way you can sort of control the output sort of you know when you have like sort of control by the AI

s426
00:36:25.280 --> 00:36:29.520
and what allows you to prompt engineer is really that sensitivity.

s427
00:36:29.520 --> 00:36:42.560
So sure maybe like right now the AI can do a lot of autoprompting and that and it can generate something that's sufficient but uh if it's like that never be satisfied like never be satisfied with the AI's generated content

s428
00:36:42.560 --> 00:36:47.440
always fine-tune your sensitivity and always kind of keep prompting the differences.

s429
00:36:47.440 --> 00:37:02.000
I I think to the there's also a big difference between like the average human untrained eye which I I would put myself in that bucket you know like I have I have some aesthetic sensibilities and I've done this long enough that you know like I have I have a preference

s430
00:37:02.000 --> 00:37:10.480
um but you know like your example of a manga artist like that's somebody who has honed a craft like over possibly many decades.

s431
00:37:10.480 --> 00:37:13.760
Um, and anybody who does that, whether it's like design, architecture, right?

s432
00:37:13.760 --> 00:37:21.280
Like you you you just have a very different level of like expertise and you see things that like the average human will not see.

s433
00:37:21.280 --> 00:37:22.000
But Doom is right.

s434
00:37:22.000 --> 00:37:33.040
Like when we look at if you were to just, you know, um, poll 10 people on the street, they would probably prefer the like overly smooth like very saturated

s435
00:37:33.040 --> 00:37:33.839
kind of

s436
00:37:33.839 --> 00:37:35.520
It's called the Instagram filter.

s437
00:37:35.520 --> 00:37:35.920
It is.

s438
00:37:35.920 --> 00:37:37.156
It is the Yeah.

s439
00:37:37.156 --> 00:37:37.599
[laughter]

s440
00:37:37.599 --> 00:37:44.880
Um, and you know, and and so there's also a little bit of a question of like what does your default aesthetic look like if you don't specify?

s441
00:37:44.880 --> 00:37:59.680
But then to Shane's point, one of the things we always try to get these models better at is instruction follow so that like when you want to get them to a different outcome like you should be able to whether that's through language or whether that's through references because language is sometimes too limiting.

s442
00:37:59.680 --> 00:38:03.839
Um, and so like these models continue to get better at it but they so much at work.

s443
00:38:03.839 --> 00:38:09.251
Do do you feel pressure as a as a product director to set the default for the world like I mean

s444
00:38:09.251 --> 00:38:10.720
[laughter]

s445
00:38:10.720 --> 00:38:11.280
kind of

s446
00:38:11.280 --> 00:38:14.079
maybe I should I don't know I haven't thought about this

s447
00:38:14.079 --> 00:38:18.640
you know you know it's like someone has to have a default the default has to exist

s448
00:38:18.640 --> 00:38:29.680
actually I will say like we have thought about this um and I I think one of the so for example actually like if you look at nanobanana generations we had like an explosion of nanobanana infographics

s449
00:38:29.680 --> 00:38:31.280
when nanobanana pro came out

s450
00:38:31.280 --> 00:38:32.240
I tried it yeah

s451
00:38:32.240 --> 00:38:38.560
um yeah I think Nurb's papers were like all you know so so many had like infographics generated.

s452
00:38:38.560 --> 00:38:41.839
Can you run your uh watermarking on it and see how many

s453
00:38:41.839 --> 00:38:48.800
uh we pro we probably could we have we haven't done that but I saw so like my Twitter was maybe this is just also like the bias of my algorithm

s454
00:38:48.800 --> 00:39:01.119
but they were everywhere um and it was actually very painful because um I think our default aesthetic was a little bit too it was too cluttered like I think that the the model is like a bit of an overeager

s455
00:39:01.119 --> 00:39:09.520
student that just like learned you know it was like oh I know all these like I know all this information about this concept let me like shove into the same image.

s456
00:39:09.520 --> 00:39:12.132
Japanese infographics 5x that

s457
00:39:12.132 --> 00:39:13.440
[laughter]

s458
00:39:13.440 --> 00:39:17.280
or maybe it was you know um but it just and and

s459
00:39:17.280 --> 00:39:20.720
wait so same prompt same content if it's in Japanese it's

s460
00:39:20.720 --> 00:39:21.920
density density

s461
00:39:21.920 --> 00:39:22.880
oh wow

s462
00:39:22.880 --> 00:39:24.880
because that's the style in Japan

s463
00:39:24.880 --> 00:39:27.852
yeah some like very you know bureaucrat and

s464
00:39:27.852 --> 00:39:28.000
[laughter]

s465
00:39:28.000 --> 00:39:29.839
there's a famous word for it yeah

s466
00:39:29.839 --> 00:39:41.040
no but we do do go through this process with Omni we did it together right like where like we had like a bunch of like we like at the very end okay like this is we did some tuning and like okay what kind of style do we prefer right like you know

s467
00:39:41.040 --> 00:39:43.280
is it more muted more saturated

s468
00:39:43.280 --> 00:39:44.560
we had a lot of saturation

s469
00:39:44.560 --> 00:39:59.599
yeah there was there were I think Nicole just has PTSD so has forgotten about it but she was very much involved in this of like okay which which kind of color palette do we basically prefer right and it's you know it's it's it's not something that like you have to make a a trade-off there like

s470
00:39:59.599 --> 00:40:00.000
uh

s471
00:40:00.000 --> 00:40:13.680
and and and it's because it ends up being us right like actually it is true like it it ends up being the modeling teams and you could ask the question legitimately of like are we the best people to do that or should we actually work with someone who like has a really creative point of view and is

s472
00:40:13.680 --> 00:40:19.280
more of like you know an art director and like has like and we kind of go back and forth on this um

s473
00:40:19.280 --> 00:40:21.040
we have the trusted testers I'm on

s474
00:40:21.040 --> 00:40:25.040
we do we have trusted testers who give us a lot of feedback and we take that serious

s475
00:40:25.040 --> 00:40:29.359
very well organized by the way to have these like weekly calls and stuff like it's it's amazing

s476
00:40:29.359 --> 00:40:39.200
um Logan's team does a lot of that so kudo kuda kudos kudos to Logan um who couldn't be here today um and we have a lot of people actually internally at Google like Fulfur

s477
00:40:39.200 --> 00:40:51.680
who give us like a ton of No, no, no. Truly like who give us a ton of feedback on like when we when we release new checkpoints and like sometimes it will be stuff that we like don't see right like we would be like oh yeah this optimization

s478
00:40:51.680 --> 00:40:57.359
seems okay and then they would come back what have you done like you completely ruined my grass you know because now the detail is all blurry.

s479
00:40:57.359 --> 00:41:04.240
I think he just noticed not not a super secret at this point but like that our model tends to put rings wedding rings on on on hand.

s480
00:41:04.240 --> 00:41:04.880
That's yeah

s481
00:41:04.880 --> 00:41:05.440
very strange.

s482
00:41:05.440 --> 00:41:09.839
I had never noticed that but he's like he I just saw it and there's a faux fur channel basically.

s483
00:41:09.839 --> 00:41:13.200
Uh he posted I was like why is there wedding ring in every hand?

s484
00:41:13.200 --> 00:41:14.160
I'm like that's strange.

s485
00:41:14.160 --> 00:41:15.839
That sounds very common reward hacking.

s486
00:41:15.839 --> 00:41:16.079
Yeah.

s487
00:41:16.079 --> 00:41:16.319
Yeah.

s488
00:41:16.319 --> 00:41:16.560
Yeah.

s489
00:41:16.560 --> 00:41:22.319
So but you know something that we would not have we would not have noticed necessarily while while developing this right is an oral artifact or

s490
00:41:22.319 --> 00:41:32.640
I I don't know you do have like a lot of preference based and then you know you may can prefer that sperious correlation reward hacking it can happen like in many weird ways.

s491
00:41:32.640 --> 00:41:33.040
Yeah

s492
00:41:33.040 --> 00:41:33.440
it does.

s493
00:41:33.440 --> 00:41:34.240
It is

s494
00:41:34.240 --> 00:41:41.839
uh this was related to another topic that again I I try to use these mainstage things as introductions or ties in.

s495
00:41:41.839 --> 00:41:48.079
Uh we have the eval track we have character AI and YouTube talking about how they evaluate videos.

s496
00:41:48.079 --> 00:41:51.040
Um how do you evaluate videos

s497
00:41:51.040 --> 00:41:52.609
apart from furer

s498
00:41:52.609 --> 00:41:53.359
[laughter]

s499
00:41:53.359 --> 00:41:57.440
not everyone has a fauxur but also you know I think there needs to be something more quantitative

s500
00:41:57.440 --> 00:42:02.960
well I mean it's you improve Gemini to improve the evaluation for video.

s501
00:42:02.960 --> 00:42:08.240
Um that's that's no no that's that's definitely one way uh it's actually very hard.

s502
00:42:08.240 --> 00:42:08.800
It's very hard.

s503
00:42:08.800 --> 00:42:25.920
It's very hard um to get like you know audators to evaluate things in a video like including especially things like aesthetics right like that it's like there are some things that are a little bit more objective like especially when we talk like let's say we talk about images and we look at like infographics text rendering that's actually

s504
00:42:25.920 --> 00:42:36.480
fine right because like you can kind of OCR things out and then you can look at like okay this letter is like messed up and then the whole thing is actually useless because if like literally if a letter is off

s505
00:42:36.480 --> 00:42:38.640
in render text you just can't use that asset.

s506
00:42:38.640 --> 00:42:39.119
Right.

s507
00:42:39.119 --> 00:42:42.400
So th those things are like a little bit more auto ratable.

s508
00:42:42.400 --> 00:42:49.040
Um from what we found we do rely a lot on humans looking at things and so we do do a lot of human evals.

s509
00:42:49.040 --> 00:42:50.400
We do a lot of human evals.

s510
00:42:50.400 --> 00:43:03.440
Do a lot of human ev and every time Jane is like um and every time we have a new model we like want to do more things and we want to like gem in more capabilities and then we have like more emails that we have to run.

s511
00:43:03.440 --> 00:43:14.000
Um, and then at some point you do get two models that are like kind of close to each other and then like we literally make decisions based on like looking at outputs side by side.

s512
00:43:14.000 --> 00:43:24.240
Sometimes like in a room like I've been in rooms where there's like 10 of us and we're just like looking at video side by side and we're like do you prefer this or do you prefer that?

s513
00:43:24.240 --> 00:43:25.680
like oh wow it's

s514
00:43:25.680 --> 00:43:37.119
I mean but it is it is genuinely very complicated the more capabilities you add like you know even just the one capability but it's like almost AGI complete capabilities like video editing right like think about video editing as a and like

s515
00:43:37.119 --> 00:43:38.640
editing with audio and

s516
00:43:38.640 --> 00:43:40.400
my editor will be very happy to hear this

s517
00:43:40.400 --> 00:43:43.040
edit the hardest problem in g media

s518
00:43:43.040 --> 00:43:55.280
I mean I don't know if it's the hardest but it's definitely there right like uh in terms of like complexity of of evaluation like free form video editing is you can do anything like

s519
00:43:55.280 --> 00:43:59.839
yes uh and like I I spent a lot of money on that and it's very hard to tell me

s520
00:43:59.839 --> 00:44:04.560
like adding those we don't have like add a sloth eval right like uh that we

s521
00:44:04.560 --> 00:44:05.200
well now we should

s522
00:44:05.200 --> 00:44:10.240
now we should yeah yeah yeah but like things like that like it's it's it's not that easy to track

s523
00:44:10.240 --> 00:44:20.160
I think I'm just surprised at the sample size that you have right like to to test the entire surface of your models you still rely on a magnitude of hundreds

s524
00:44:20.160 --> 00:44:26.960
no no no so we do like yeah well we do we do a ton of human evals on like on like you know thousands of things.

s525
00:44:26.960 --> 00:44:34.960
Um I I think there's also like an element of you know we can talk about things like live experiments right like which which is also where you get signal on like

s526
00:44:34.960 --> 00:44:45.520
like some of these more minute differences at like much larger scale then there's auto raers which is definitely kind of a more it's a very well defined space I think for LLMs

s527
00:44:45.520 --> 00:44:57.200
much more nent for media models and then like sometimes you still do rely on human judgment and we do rely on things like feedback from people who just like have a very owned

s528
00:44:57.200 --> 00:45:02.160
like aesthetic and and people who just like use these models in their workflows dayto-day, right?

s529
00:45:02.160 --> 00:45:09.520
Because we could also like you could have a model that does really well on some slice of human evals, but then it like really breaks a workflow for somebody.

s530
00:45:09.520 --> 00:45:15.920
And so this is why we do like early access programs and we try to get feedback and then we like try to incorporate it before we release something more broadly.

s531
00:45:15.920 --> 00:45:19.359
I feel like Shane had a hot take based on his

s532
00:45:19.359 --> 00:45:21.599
expression always when we were talking about this

s533
00:45:21.599 --> 00:45:33.200
every kind of human sort of you know work should be gradually kind of amortized and then the interesting thing is the video understanding especially like against like AI gener like detecting air stuff is extremely

s534
00:45:33.200 --> 00:45:35.359
interesting uh visual task

s535
00:45:35.359 --> 00:45:47.920
and then like some of it kind of aesthetics or this kind of visual quality but for some of the kind of cases like semantically doesn't make sense for example you're taking like some like a famous scene from a movie and try to sort of

s536
00:45:47.920 --> 00:45:58.720
um construct that and then if you kind of generate it uh it can generate something there but at some point some of the semantic information doesn't make sense like it's actually inconsistent.

s537
00:45:58.720 --> 00:46:01.760
So can the AI actually detect that?

s538
00:46:01.760 --> 00:46:10.800
So when I evaluate the AI videos like oh I feel I'm so smart you know like like AI is still kind of behind but we should make like a lot of effort.

s539
00:46:10.800 --> 00:46:18.000
I think the video understanding is extremely uh important intelligence task uh beyond just the pure aesthetics or the preference.

s540
00:46:18.000 --> 00:46:22.720
Um and yeah we we should always try to amatize the human

s541
00:46:22.720 --> 00:46:23.760
human label.

s542
00:46:23.760 --> 00:46:24.480
Yeah.

s543
00:46:24.480 --> 00:46:25.359
Yeah.

s544
00:46:25.359 --> 00:46:28.800
Um, what data do you need?

s545
00:46:28.800 --> 00:46:33.520
A lot of people I talked to wanted to get in front of you actually.

s546
00:46:33.520 --> 00:46:36.319
Uh, they I mean they want to be nice about it.

s547
00:46:36.319 --> 00:46:37.520
They have a lot of video data.

s548
00:46:37.520 --> 00:46:38.480
They have gaming data.

s549
00:46:38.480 --> 00:46:39.839
They have real world video data.

s550
00:46:39.839 --> 00:46:41.119
They have images.

s551
00:46:41.119 --> 00:46:42.319
They have labelers.

s552
00:46:42.319 --> 00:46:44.640
What do you want?

s553
00:46:44.880 --> 00:46:46.079
Are you like offering?

s554
00:46:46.079 --> 00:46:49.359
I'm just like this is your request for like Okay.

s555
00:46:49.359 --> 00:46:49.760
Okay.

s556
00:46:49.760 --> 00:46:52.319
We get I'm sure you get a lot of pitches, right?

s557
00:46:52.319 --> 00:46:54.240
You get a lot of people want to talk to you.

s558
00:46:54.240 --> 00:47:08.800
what's like I think actually it's the signal is this problem this sorting out signal from noise is the main problem so creating a nice API of like okay if you actually do a b and c we are interested in that

s559
00:47:09.680 --> 00:47:19.280
um loaded question there so uh I don't know that there's like an easy like you know did you do I think we we do already have a lot of data I think it's it's

s560
00:47:19.280 --> 00:47:21.280
hard to talk about this

s561
00:47:21.280 --> 00:47:23.839
you know you want to talk about the public I don't want to get you in trouble Yeah,

s562
00:47:23.839 --> 00:47:24.640
but like I think

s563
00:47:24.640 --> 00:47:35.920
No, no. What I just want to say is like hard to talk about this in a sort of you know without trying to without I have to think about the what I am revealing about our project and what where we're going.

s564
00:47:35.920 --> 00:47:40.640
Um generally high quality data I think maybe maybe let's just put it this way right it's not not the secret

s565
00:47:40.640 --> 00:47:42.319
embodied I'm sorry

s566
00:47:42.319 --> 00:47:43.760
embodied data

s567
00:47:43.760 --> 00:47:44.720
I mean

s568
00:47:44.720 --> 00:47:52.800
yeah sure I mean we have sort of announced I think publicly right that we we have some sort of robotics collaboration right like so I think it's like a

s569
00:47:52.800 --> 00:48:08.720
like or or but you because we have a robotics team at GDM so you know they're always interested in things like that um I mean for Omni specifically I think we're just quite interested just high quality data right like you know it it's not some sort of not necessarily like oh random

s570
00:48:08.720 --> 00:48:17.280
YouTube video but like you know some a some more professional shop things like that right the things that those are those are things that we're always on the lookout for like

s571
00:48:17.280 --> 00:48:18.800
uh and yeah

s572
00:48:18.800 --> 00:48:28.160
and I think for you know maybe this is easier to some extent to answer for like some of the agentic work as well like like like actual kind of

s573
00:48:28.160 --> 00:48:40.640
like what are the tests that people are trying to do right these things are actually kind of difficult to manufacture if you're doing it yourself or if you're like doing it with a vendor, like what is the actual like if you're creating a marketing campaign,

s574
00:48:40.640 --> 00:48:42.160
like what does that look like, right?

s575
00:48:42.160 --> 00:48:55.760
Like do do you start from here's like a picture of my new product and then I want to turn that into a video ad and I want to turn that into a bunch of assets that like fit fit all these different ad formats that I need to push onto various platforms to promote

s576
00:48:55.760 --> 00:49:06.000
and then like so you kind of go from this to that and like what is that kind of trajectory of tasks that you're that you're like you know experiencing along the way like that is really useful

s577
00:49:06.000 --> 00:49:23.839
and that is actually kind of difficult to get right u because like we don't always have the right firstparty surface where people are actually doing some of these things or like you might work with someone who's a vendor but they don't also don't have that product surface right like like a lot of this kind of information lives

s578
00:49:23.839 --> 00:49:31.280
in the places where people are doing these tasks and so that's kind of difficult to get like if anyone's figured that out you should reach out to us

s579
00:49:31.280 --> 00:49:33.079
every channel of thought yeah every

s580
00:49:33.079 --> 00:49:33.599
[laughter]

s581
00:49:33.599 --> 00:49:33.839
thought

s582
00:49:33.839 --> 00:49:38.559
every thought yeah and maybe the data the Chinese lab is using

s583
00:49:38.559 --> 00:49:43.359
yes yeah uh you know as a media person myself, right?

s584
00:49:43.359 --> 00:49:52.319
Like there's so many podcasters and people in in marketing departments and all these like they would happy to be your data like you know just like put a BCI on my head

s585
00:49:52.319 --> 00:49:53.384
and podcast

s586
00:49:53.384 --> 00:49:53.680
[laughter]

s587
00:49:53.680 --> 00:50:09.440
watch my things uh because like you know there's just endless amount of work to do like there's so much work and this is all like this needs to somewhat be commodity like obviously you can be an art like an artisan like you can be Hollywood for like the really high quality stuff but actually a lot of work

s588
00:50:09.440 --> 00:50:13.689
is commodity and like should be modelable and we want you to do

s589
00:50:13.689 --> 00:50:13.839
[laughter]

s590
00:50:13.839 --> 00:50:18.480
And but we we want the high quality to Demi's point right like we do want we want the high quality.

s591
00:50:18.480 --> 00:50:19.280
We want commodity.

s592
00:50:19.280 --> 00:50:19.440
Yeah.

s593
00:50:19.440 --> 00:50:19.760
Yes.

s594
00:50:19.760 --> 00:50:20.000
Yes.

s595
00:50:20.000 --> 00:50:21.599
You want on both sides.

s596
00:50:21.599 --> 00:50:24.319
Um I I just

s597
00:50:24.319 --> 00:50:26.789
Thank you for the solicitation.

s598
00:50:26.789 --> 00:50:27.440
[laughter]

s599
00:50:27.440 --> 00:50:31.280
Uh I you know we we we also I also added a data quality track.

s600
00:50:31.280 --> 00:50:38.079
I I think that uh people want to understand like what uh at AI like how to raise the bar, right?

s601
00:50:38.079 --> 00:50:53.359
like like the and a lot of it is just educating the market and educating researchers and engineers and founders on like this is where we're going a lot of this is stop doing that do this do this instead and I'm like people will listen

s602
00:50:53.359 --> 00:50:56.000
yeah I don't know uh to that extent you know

s603
00:50:56.000 --> 00:51:09.839
but I think to that to that point like there's a lot of again just like craft that goes into this right and there's a lot of process like you even to the marketing campaign example you don't create that in like five minutes right you like go you go through a process and you iterate and you like pick

s604
00:51:09.839 --> 00:51:21.119
something over something else because you liked it for whatever reason like maybe the eye gaze was correct right like we just we don't know these things right because none of us are marketing directors and like the models don't know these things

s605
00:51:21.119 --> 00:51:33.359
I even kind of say this for the natural like a language as well like I I always kind of say 99% of information is inside people you can only extract it through active dialogue and befriending them so most of the stuff on the internet

s606
00:51:33.359 --> 00:51:41.680
is like sort of the outcome the output of that yes but you know what are what are all the trajectories you know how did this person have this inspiration to write this paper

s607
00:51:41.680 --> 00:51:53.599
what is the starting point what is the inspiration what are the dialogue that sparked it those kind of stuff is kind of inside people so even you know those kind of like even the language space is kind of that I think the creative is kind of similar as well there's a lot of dark knowledge

s608
00:51:53.599 --> 00:52:08.160
yeah it's like when you write a novel right like a novel speaks to you because like usually there's some sort of like a personal connection that you feel to like the story or the trajectory or the characters right like if you read most of the stuff that's written by LLM's today like it's,

s609
00:52:08.160 --> 00:52:16.160
you know, it's it's it starts it falls into these like default par patterns and like the language starts to feel really similar and all the descriptions sound really similar.

s610
00:52:16.160 --> 00:52:22.160
You can kind of like quickly read it as like, oh, this is not that interesting because like I can't connect to it, right?

s611
00:52:22.160 --> 00:52:26.240
Um, and again, that's that's kind of like a human expertise.

s612
00:52:26.240 --> 00:52:33.280
One nice thing recently is the Google Cloud and the Google Deep Mind are kind of starting to invest a lot more in the FTEEs for the product engineers.

s613
00:52:33.280 --> 00:52:38.240
And I also kind of saw some uh recruiting for the creative you know gem media kind of space as well.

s614
00:52:38.240 --> 00:52:50.480
So I think those are kind of really the effort because we we kind of feel that you know what we can kind of do with a lot of public data there's limits but really you know partnering with that we can provide kind of better models and products and yeah we kind of feedback

s615
00:52:50.480 --> 00:52:54.160
uh we have an FD track here for the first time every lab is announcing it.

s616
00:52:54.160 --> 00:52:55.359
It's it's crazy.

s617
00:52:55.359 --> 00:53:05.440
Um, one thing I'm actually very keen on doing and I push I push for this at Cognition as well is to turn the FDES not just into sales and solutions

s618
00:53:05.440 --> 00:53:08.319
but also to EVAL's uh eval workers.

s619
00:53:08.319 --> 00:53:11.680
FD is not the sales FD is way way bigger than that.

s620
00:53:11.680 --> 00:53:14.800
How do you frame FDs then?

s621
00:53:14.960 --> 00:53:15.126
Because

s622
00:53:15.126 --> 00:53:15.200
[laughter]

s623
00:53:15.200 --> 00:53:20.480
I do think about it as sales like you're you know the more the more you customize the solution for

s624
00:53:20.480 --> 00:53:28.480
so I define post training as anything between the pre-training and the final user experience anything anything is a post training

s625
00:53:28.480 --> 00:53:41.520
and to me when I first sort of you know learned a lot about I mean FD kind of I guess originally you know came from like path here and then that so I guess the kind of history is different but yeah I think the key is really that um you know the key is like not only to

s626
00:53:41.520 --> 00:53:50.800
kind of work uh with them and ensure that they kind of know how to but also to sort of code like derive kind of insights that can basically kind of help both parties.

s627
00:53:50.800 --> 00:53:53.760
They can put the like a lot of harness how they use the model.

s628
00:53:53.760 --> 00:53:55.520
We can improve like very upstream.

s629
00:53:55.520 --> 00:54:02.960
So how to get the customer feedback to the modeling I feel is the kind of more the the role I I kind of want for the fds.

s630
00:54:02.960 --> 00:54:03.440
Yeah.

s631
00:54:03.440 --> 00:54:03.839
Yeah.

s632
00:54:03.839 --> 00:54:04.079
Yeah.

s633
00:54:04.079 --> 00:54:11.359
and and even for sorry just on that like if you want to talk to us or at least me um I I'm not going to offer up your time

s634
00:54:11.359 --> 00:54:25.040
um but I it's really helpful for us to actually talk to people who are using our models and like understand where they're struggling uh because again that just like it's it's the real world task that you're actually trying to use them for right like I will talk to people

s635
00:54:25.040 --> 00:54:41.040
who do kind of interior inter interior design with some of our image models um you know and they will say hey like I really want to take this pattern pattern, but then I want to scale it across like 10 different ruck sizes and sometimes I have like a very custom ruck size and then the model fails at

s636
00:54:41.040 --> 00:54:51.839
like replicating the pattern the same way or you know I want to do a try on for these earrings and then the earrings have a certain size and then like my head has a certain size like it has to make sense

s637
00:54:51.839 --> 00:54:59.839
if you're actually trying to try things on and like the models kind of fail at a bunch of these things that like actually happen in the real world, right?

s638
00:54:59.839 --> 00:55:06.559
Um and so that that's like useful for us because for some of these things like we don't think about because we don't you know we don't use the models for those tasks

s639
00:55:06.559 --> 00:55:13.040
or like um you know I think to your point about ad campaigns or whatever like people have like notions of brand languages or whatever like which is

s640
00:55:13.040 --> 00:55:13.440
yes

s641
00:55:13.440 --> 00:55:26.000
like a a bunch of images or PDFs saying things you know it's a pretty kind of you know ambiguous question as well what is the IKEA brand language you know is it is it blue and yellow I mean that's that's not a very like

s642
00:55:26.000 --> 00:55:27.280
but like what shade of blue you know.

s643
00:55:27.280 --> 00:55:27.440
Yeah.

s644
00:55:27.440 --> 00:55:27.599
Yeah.

s645
00:55:27.599 --> 00:55:27.760
Yeah.

s646
00:55:27.760 --> 00:55:32.720
So there there's like, you know, and the brands are pretty spec, you know, pretty, you know, like they they do care about the shade of blue.

s647
00:55:32.720 --> 00:55:35.200
It's not shouldn't just be a random blue and a random yellow.

s648
00:55:35.200 --> 00:55:36.319
That's not going to be IKEA, right?

s649
00:55:36.319 --> 00:55:37.760
I'm just thinking about an example.

s650
00:55:37.760 --> 00:55:47.760
But like this is the kind of stuff that, you know, it's not necessarily part of our like, you know, developing frontier models kind of, you know, necessarily mandate, but it's something that we do want to we do want to fundamentally like build

s651
00:55:47.760 --> 00:55:53.359
products that people will use to solve concrete tasks, not just not just research artifacts, right?

s652
00:55:53.359 --> 00:55:57.040
So I think it's useful to understand what people do care about.

s653
00:55:57.040 --> 00:56:12.240
Uh well, I'm sure a lot of people are very grateful for your work and there's a lot more to do that you've made so much progress over the last like even just couple years of like Nano Banana and Theo and Omni and uh I don't know what else you got cooking but we're very excited like you this

s654
00:56:12.240 --> 00:56:21.200
is one of those things where like I was very disappointed you know when Sora shut down and and I think like there needs to be more general exploration of

s655
00:56:21.200 --> 00:56:25.078
uh you know generative models and not just you know coding.

s656
00:56:25.078 --> 00:56:25.200
[laughter]

s657
00:56:25.200 --> 00:56:26.400
I think I think that is

s658
00:56:26.400 --> 00:56:27.680
we obviously like this.

s659
00:56:27.680 --> 00:56:28.559
We love coding.

s660
00:56:28.559 --> 00:56:32.240
Love coding and and uh yes uh but thank you so much for your time.

s661
00:56:32.240 --> 00:56:35.200
Uh it's been a real pleasure and I can't wait to see what this looks like next.

s662
00:56:35.200 --> 00:56:36.000
Thank you for having us.

s663
00:56:36.000 --> 00:56:36.559
Great question.

s664
00:56:36.559 --> 00:56:37.524
Thank you everyone.

s665
00:56:37.524 --> 00:56:39.524
[applause]
