WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=2bvtay8wGYI

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/2bvtay8wGYI.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:12.480 --> 00:00:15.320
Uh so this talk is called scaling to long horizons.

s2
00:00:15.320 --> 00:00:16.000
My name is Ross.

s3
00:00:16.000 --> 00:00:17.680
Uh I'm the CEO of GR.

s4
00:00:17.680 --> 00:00:20.320
We're a London-based reinforcement learning company.

s5
00:00:20.320 --> 00:00:27.240
Uh before GR, I was the reasoning lead at Meta AI working on Llamas, uh Galactica, lots of other models back in the day.

s6
00:00:27.240 --> 00:00:30.560
I'm joined by Chengxi, uh co-founder and president of GR.

s7
00:00:30.560 --> 00:00:41.200
Uh and yeah, hit today we're going to talk about algorithms, environments, compute, all the things you need to do to get agents scaling uh to kind of long with tasks.

s8
00:00:41.240 --> 00:00:42.840
So we're going to have two parts of this talk today.

s9
00:00:42.840 --> 00:00:52.080
I'm going to first of all start with a personal perspective about, you know, the early days, the golden age of language modeling in between like maybe 2020 and 2023.

s10
00:00:52.080 --> 00:00:56.720
Uh I'll talk about, like I said, all these models and some of our early reinforcement learning efforts for LLMs.

s11
00:00:56.720 --> 00:01:00.560
And then Chengxi is going to talk about, you know, what's ahead, you know, what the next frontiers.

s12
00:01:00.560 --> 00:01:06.160
And yeah, that's going to be a really talk with a lot of alpha, so I'd encourage you to stick around for that.

s13
00:01:06.160 --> 00:01:07.640
So the journey so far.

s14
00:01:07.640 --> 00:01:10.200
So my journey started here.

s15
00:01:10.200 --> 00:01:12.880
Um so this was the Papers With Code team.

s16
00:01:12.880 --> 00:01:15.360
I'm sure many of you used Papers With Code back in the day.

s17
00:01:15.360 --> 00:01:18.160
So we were a London-based startup 2019.

s18
00:01:18.160 --> 00:01:20.360
Uh we were acquired by Meta later that year.

s19
00:01:20.360 --> 00:01:23.680
And then we had a crazy transition within Meta to do research.

s20
00:01:23.680 --> 00:01:26.280
Um so we did, like I said, Galactica.

s21
00:01:26.280 --> 00:01:34.560
Then after ChatGPT came out, we started the post-training for Llama 2, Llama 3. So all the great work you uh saw there was folks in this room.

s22
00:01:34.560 --> 00:01:37.080
And lots of other interesting stuff that never got published as well.

s23
00:01:37.080 --> 00:01:39.200
Uh reasoning Llama and lots of other things.

s24
00:01:39.200 --> 00:01:44.400
Um so yeah, this this small team I'd like to think like the open weight kind of revolution started in this room.

s25
00:01:44.400 --> 00:01:54.760
And you know, it really like hit home this idea to me that kind of small focused teams, like even in like the age of scaling, can do amazing things if people are aligned.

s26
00:01:54.960 --> 00:01:59.040
Now for me, uh things got particularly crazy in 2022.

s27
00:01:59.040 --> 00:02:00.560
Uh so let me tell you a story.

s28
00:02:00.560 --> 00:02:10.360
Um the media perception is that ChatGPT came out of nowhere, you know, shocked the world, and that's how kind of the modern AI wave started.

s29
00:02:10.360 --> 00:02:18.959
But, you know, I have a different personal perspective on this because 2 weeks before ChatGPT came along, there was another language model called Galactica.

s30
00:02:18.959 --> 00:02:20.840
So let's talk about Galactica.

s31
00:02:20.840 --> 00:02:25.480
Galactica and, you know, ChatGPT, you know, they were both, you know, in some respects quite similar.

s32
00:02:25.480 --> 00:02:27.920
They're both based on pretty good base models.

s33
00:02:27.920 --> 00:02:33.240
Galactica itself was a base model, and then ChatGPT was based on GPT-3.5.

s34
00:02:33.240 --> 00:02:35.000
But there was a clear difference in outcomes.

s35
00:02:35.000 --> 00:02:39.000
So Galactica at the time shipped with this slight base model demo.

s36
00:02:39.000 --> 00:02:42.560
And as you guys know now, like base models, they come with a lot of quirks.

s37
00:02:42.560 --> 00:02:43.920
You know, they hallucinate.

s38
00:02:43.920 --> 00:02:47.360
You prompt them to do, you know, silly things, they would do silly things.

s39
00:02:47.360 --> 00:02:54.080
Whereas ChatGPT wasn't just a base model, but had this like crucial reinforcement learning from human feedback pipeline.

s40
00:02:54.080 --> 00:02:58.320
And this was the key thing that made LLMs like really products for the first time.

s41
00:02:58.320 --> 00:03:05.680
So I like to think in a weird kind of way, this is like the first like kind of natural experiment showing you that kind of RL like provides value, right?

s42
00:03:05.680 --> 00:03:11.800
Uh and to my misfortune, it was like a very personal like kind of a natural experiment and you know, Galactica blew up.

s43
00:03:11.800 --> 00:03:13.600
Uh but that's like a good like kind of lesson there.

s44
00:03:13.600 --> 00:03:15.760
A good base model is not enough.

s45
00:03:15.760 --> 00:03:18.640
So I took that lesson quite early on.

s46
00:03:18.640 --> 00:03:21.560
So like I said, RLHF made LLMs products.

s47
00:03:21.560 --> 00:03:27.320
They were the thing that kind of made LLMs cross the Rubicon into something that wasn't just a toy, but used by now billions of people.

s48
00:03:27.320 --> 00:03:29.920
But you didn't have to wait until ChatGPT to see this.

s49
00:03:29.920 --> 00:03:34.920
Like even at the time, like InstructGPT in 2022 had these pretty stunning results.

s50
00:03:34.920 --> 00:03:40.560
Like a 1 billion parameter model with RLHF was outperforming 175 billion models.

s51
00:03:40.560 --> 00:03:44.880
So two orders of magnitude fewer parameters, but getting better results.

s52
00:03:44.880 --> 00:03:45.800
So that was astonishing.

s53
00:03:45.800 --> 00:03:52.760
So if you were paying attention closely, you know, maybe we should have been as well, but we were focused on a, you know, bloody base model, which is hard work in 2022.

s54
00:03:52.760 --> 00:03:57.560
But that shows you how you know important even basic RL is.

s55
00:03:57.560 --> 00:03:59.480
And the Galactica demo itself, I mean, it set off a storm.

s56
00:03:59.480 --> 00:04:01.960
So, ancient history now, but we put out a demo.

s57
00:04:01.960 --> 00:04:05.240
We let people play around with it um because it's kind of cool.

s58
00:04:05.240 --> 00:04:07.040
And at the time people got scared.

s59
00:04:07.040 --> 00:04:16.200
So, it was like, you know, like I said, you prompt it on like a research paper on Dyson spheres or a report on the benefits of eating crushed glass and people are like, "Oh my god."

s60
00:04:16.200 --> 00:04:18.799
Um so, that was the state of things in 2022.

s61
00:04:18.799 --> 00:04:22.200
And yeah, I'll be honest, the meta association didn't you know help us either.

s62
00:04:22.200 --> 00:04:28.160
Um and yeah, the tragic story in a way was a lot of the, you know, novel work was maybe overshadowed.

s63
00:04:28.320 --> 00:04:32.720
But you know, the paradox of this whole thing is that at the time Galactica was actually a bloody good model.

s64
00:04:32.720 --> 00:04:38.360
Um like it outperformed Palm, Chinchilla, GPT-3.5 with a lot less compute in scientific domains.

s65
00:04:38.360 --> 00:04:39.560
It was state-of-the-art.

s66
00:04:39.560 --> 00:04:42.320
So again, that also shows you how powerful RL is.

s67
00:04:42.320 --> 00:04:45.280
You can have a sota base model, but that is not enough.

s68
00:04:45.280 --> 00:04:49.040
Um so, here you see on like a math who's kind of beating Chinchilla.

s69
00:04:49.040 --> 00:04:54.880
Uh kind of latex equations, you know, science, you know, it was getting around 68% compared to GPT-3.5, 49%.

s70
00:04:54.880 --> 00:04:56.680
So, crushing there.

s71
00:04:56.680 --> 00:04:57.840
And chain of thought as well.

s72
00:04:57.840 --> 00:05:01.640
So, you know, a Palm at the time which was a Google Brain model, 540 billion.

s73
00:05:01.640 --> 00:05:05.040
You know, that's 30 billion in Galactica was getting like 36 versus 19%.

s74
00:05:05.040 --> 00:05:07.600
So, double the performance, order of magnitude less results.

s75
00:05:07.600 --> 00:05:12.200
So, that again reinforces base good base models not enough.

s76
00:05:12.280 --> 00:05:14.640
But it introduced some key ideas which I think are very important.

s77
00:05:14.640 --> 00:05:17.960
I mean, Galactica was the first LM to really crack data efficiency.

s78
00:05:17.960 --> 00:05:23.360
105 billion uh token corpus compared to you know, trillion uh tokens in Chinchilla.

s79
00:05:23.360 --> 00:05:27.440
And it was a really contrarian at the time cuz you know, at the time everyone was like, "Okay, we just need more tokens."

s80
00:05:27.440 --> 00:05:28.840
And Galactica said, "No.

s81
00:05:28.840 --> 00:05:31.960
High quality, you know, curated data sets really matter."

s82
00:05:31.960 --> 00:05:34.560
And that was a real driver of those results you just saw.

s83
00:05:34.560 --> 00:05:38.360
And it was also like the first major LM to really crack multi-epoch training.

s84
00:05:38.360 --> 00:05:42.360
It sounds ridiculous now, but at the time the consensus was you don't do more than an epoch.

s85
00:05:42.360 --> 00:05:47.840
Um but this kind of rule of thumb, you may have heard of it like four uh epochs repeated data, that was formalized later.

s86
00:05:47.840 --> 00:05:51.080
but Galactica was the first like real empirical result for that.

s87
00:05:51.080 --> 00:05:54.080
Now, perhaps more importantly, there's this idea of thinking tokens.

s88
00:05:54.080 --> 00:05:57.240
And some of you might remember this, but it was like really quite buried within the paper.

s89
00:05:57.240 --> 00:05:59.760
So, around this time there were like different ideas for reasoning.

s90
00:05:59.760 --> 00:06:01.360
There was chain of thought, which was one idea.

s91
00:06:01.360 --> 00:06:03.640
It's where you prompt for like kind of like the steps.

s92
00:06:03.640 --> 00:06:08.760
There were scratch pads where you just like put like very like numerical kind of intermediate steps.

s93
00:06:08.760 --> 00:06:13.680
But, Galactica was really this first idea which said, "No, this is an internal working memory process.

s94
00:06:13.680 --> 00:06:15.000
This is an internal thinking.

s95
00:06:15.000 --> 00:06:19.120
You should be inside these tags, and you should spend the inference computer before you get to an answer, right?"

s96
00:06:19.120 --> 00:06:23.080
So, these are all like quite like pressing ideas.

s97
00:06:23.080 --> 00:06:30.800
But, you know, Galactica came and went, blew up, and then, you know, we were kind of tasked as a team to kind of spin up the post-training effort for Llama.

s98
00:06:30.800 --> 00:06:34.000
But, I had like a personal obsession, which was like reasoning.

s99
00:06:34.000 --> 00:06:42.320
And I had like a really simple idea at the time, which is, "What if we applied kind of reinforcement learning pressure to this like kind of thinking tags as work?"

s100
00:06:42.320 --> 00:06:45.919
Like, what if we just optimize the thing in between the thinking?

s101
00:06:45.919 --> 00:06:51.160
And if that sounds familiar, then this is like kind of what Deep Seek 01 ended up doing 2 years later.

s102
00:06:51.160 --> 00:06:52.400
But, there's a key difference.

s103
00:06:52.400 --> 00:06:58.480
So, at the time we only had Llama 2 base models, terrible mathematics corpus, terrible results on math.

s104
00:06:58.480 --> 00:07:03.440
And, you know, the context window, you know, we're all like, you know, very context-rich now, 1 million tokens.

s105
00:07:03.440 --> 00:07:05.560
You know, the back in the day it was 4,000.

s106
00:07:05.560 --> 00:07:07.800
It wasn't too fun.

s107
00:07:07.800 --> 00:07:11.600
But, we still had a recipe at the time, and this is unpublished, but it was a really good for the meta.

s108
00:07:11.600 --> 00:07:14.120
So, RSP was this.

s109
00:07:14.120 --> 00:07:18.520
Number one, continue pre-training of Llama 2 towards mathematics and science data.

s110
00:07:18.520 --> 00:07:19.440
So, that's the first thing.

s111
00:07:19.440 --> 00:07:22.680
Llama 2, math corpus, so let's fix that.

s112
00:07:22.680 --> 00:07:25.800
Number two, PPO with verifiable rewards.

s113
00:07:25.800 --> 00:07:27.480
But, notice this isn't GRPO, right?

s114
00:07:27.480 --> 00:07:31.800
So, we had a time and a strong outcome reward model to initialize the value model.

s115
00:07:31.800 --> 00:07:34.919
That was a key thing, lots of data on value models at the time.

s116
00:07:34.919 --> 00:07:39.640
And internally at the time, this kind of recipe led to state-of-the-art results on math and reasoning.

s117
00:07:39.640 --> 00:07:44.040
So, we were kind of like, "Wow, this is like really shows the power of having the right objective."

s118
00:07:44.040 --> 00:07:51.000
But, the really fascinating thing is like we had great results, but we didn't have like inference time scaling.

s119
00:07:51.000 --> 00:07:55.240
We didn't have this reflective behavior that became like the hallmark of R1 and O1.

s120
00:07:55.240 --> 00:07:58.480
You know, back weights, you know, back tracking, all this kind of stuff.

s121
00:07:58.480 --> 00:08:00.160
So, it begs like the question like, "Why?

s122
00:08:00.160 --> 00:08:02.920
Well, why didn't we have that moment?"

s123
00:08:02.920 --> 00:08:05.080
And we got an answer around like 2 years later.

s124
00:08:05.080 --> 00:08:11.480
So, there's a couple of like things going on here, but essentially better base models were the thing that really got RL cooking.

s125
00:08:11.480 --> 00:08:13.800
And when DeepMind came out, I was kind of shocked at the time.

s126
00:08:13.800 --> 00:08:16.040
I was like, "Holy we just like tried the same thing.

s127
00:08:16.040 --> 00:08:16.680
We didn't have this.

s128
00:08:16.680 --> 00:08:17.720
What's going on here?"

s129
00:08:17.720 --> 00:08:23.640
And in a weird kind of way, the real lesson was it was just like the bitter lesson, like the most purest form of bitter lesson possible.

s130
00:08:23.640 --> 00:08:31.240
Like, better base models, more RL computes, bigger context windows, and that's all you need for this kind of emergent behavior.

s131
00:08:31.240 --> 00:08:41.560
It also like says something like quite important about the sociology of like research because the fact that OpenAI had this model, you know, GPT-4 level model before anyone else, it allowed them to see further, right?

s132
00:08:41.560 --> 00:08:42.840
So, that's a really interesting point.

s133
00:08:42.840 --> 00:08:49.680
Like, the the age of scaling means that if you have certain prerequisites in place, you become smarter, you see further, you see more ideas.

s134
00:08:49.680 --> 00:08:52.200
So, really interesting point.

s135
00:08:52.200 --> 00:08:53.480
So, this was me in 2024.

s136
00:08:53.480 --> 00:08:57.920
I was a very sad panda, uh defeated by ChatGPT and O1.

s137
00:08:57.920 --> 00:08:59.520
But, I wasn't deterred.

s138
00:08:59.520 --> 00:09:05.040
Um so, I wanted to seek the next wave, and I still was like convinced that kind of reasoning hadn't been solved.

s139
00:09:05.040 --> 00:09:08.480
So, we started GR uh to take on truly like big tasks.

s140
00:09:08.480 --> 00:09:10.400
And with that in mind, I'm going to hand over to Chengxi.

s141
00:09:10.400 --> 00:09:13.520
He's going to talk about what we're kind of thinking about now.

s142
00:09:13.520 --> 00:09:14.120
Ross.

s143
00:09:14.120 --> 00:09:14.640
Ross.

s144
00:09:14.640 --> 00:09:16.640
Ross.

s145
00:09:20.440 --> 00:09:21.200
Hi, everyone.

s146
00:09:21.200 --> 00:09:30.360
I'm Chengxi Taylor, co-founder and president of General Intelligence Inc. I'm going to share what it takes to scale to long horizon.

s147
00:09:31.960 --> 00:09:33.880
First, I want to make it clear.

s148
00:09:33.880 --> 00:09:38.040
Long horizon task is not just an engineering problem.

s149
00:09:38.040 --> 00:09:40.120
It is a mindset.

s150
00:09:40.120 --> 00:09:51.200
If we want to solve humanity's biggest problems, such as cure cancer, solve millennium's prize problem, or go into Mars, we have to be patient.

s151
00:09:51.200 --> 00:09:52.680
It take time.

s152
00:09:52.680 --> 00:10:00.400
And if we want AI to move us towards that level impact, we have to think about long horizon.

s153
00:10:02.160 --> 00:10:04.480
But here's the first problem.

s154
00:10:04.480 --> 00:10:07.040
We have a scarce context window.

s155
00:10:07.040 --> 00:10:18.680
If you take Fermat's Last Theorem as example, what it take for the mathematician was over 10 years time of reading paper, writing down thoughts in the scratch pad, or taking a walk

s156
00:10:18.680 --> 00:10:21.360
to generate the creative ideas.

s157
00:10:21.360 --> 00:10:26.960
If we convert to token, that's probably tens of a billions of even hundreds billions.

s158
00:10:26.960 --> 00:10:32.760
But where we are now, just a 1 million token context window.

s159
00:10:32.760 --> 00:10:35.600
So one solution is use compaction.

s160
00:10:35.600 --> 00:10:45.080
So what it essentially does is generate the token until the end of the context window, summarize, and then on top of that generate more tokens.

s161
00:10:45.080 --> 00:10:50.680
And the beauty of applying RL in the situation is kind of like kill two birds with one stone.

s162
00:10:50.680 --> 00:10:56.000
You apply RL to the compaction and also the task.

s163
00:10:56.000 --> 00:10:57.640
But here's the problem.

s164
00:10:57.640 --> 00:11:00.080
With long horizon, there are three issues.

s165
00:11:00.080 --> 00:11:04.120
The first is the gradient variance scales with the length.

s166
00:11:04.120 --> 00:11:06.720
And the second is a sparse reward.

s167
00:11:06.720 --> 00:11:09.720
And you have this credit assignment problem.

s168
00:11:09.720 --> 00:11:16.960
And finally, there's also variable length of the trajectory that adds to the problem of optimization.

s169
00:11:17.200 --> 00:11:22.520
So to solve this issue, we can apply critics, which is the value model.

s170
00:11:22.520 --> 00:11:34.520
And value model can reduce the variance and also have a couple advantages, such as on the trajectory level that fits compaction very well, and also encourage the batch diversity.

s171
00:11:34.520 --> 00:11:37.120
And also, I'll talk later on bootstrapping.

s172
00:11:37.120 --> 00:11:40.840
Basically, get signal before the end of the episode.

s173
00:11:40.840 --> 00:11:52.400
But, the downside for this is that um it's more complicated than GRPO, and basically, you have to train another value model alongside with the policy model.

s174
00:11:52.640 --> 00:12:02.920
And there's some tools to help with the context limitations, such as a file system tools, which essentially like a scratchpad for AI to write to the uh reasoning thought.

s175
00:12:02.920 --> 00:12:08.880
And self-search tools, which allows agent to search over the previous trajectory.

s176
00:12:08.880 --> 00:12:11.120
And then you have archive tools.

s177
00:12:11.120 --> 00:12:15.480
In the case like auto research, you can build upon your previous result.

s178
00:12:15.480 --> 00:12:16.800
But, we have to be careful.

s179
00:12:16.800 --> 00:12:23.880
In other scenario, you don't want an AI to cheat by just to grab the previous answer without thinking.

s180
00:12:24.320 --> 00:12:31.040
So, how good are the current model on this long horizon task?

s181
00:12:31.040 --> 00:12:39.240
In the general reasoning, we construct the benchmark called a Kelly bench, where we're actually featured in the front page of the Financial Times.

s182
00:12:39.240 --> 00:12:43.680
Kind of a caught us off guard how much the mainstream have interest in this.

s183
00:12:43.680 --> 00:12:54.360
So, basically, what we did is that we allowed the agents to build machine learning models to um trade in the football matches over 1-year horizon.

s184
00:12:54.360 --> 00:12:58.680
In this case, it's a Premier League, if you're interested in football.

s185
00:12:58.680 --> 00:13:06.400
And we're so fascinated by this because there's a real money to be made, and if it was a successful, couldn't make a billions.

s186
00:13:06.400 --> 00:13:08.960
There's a whole industry on sports betting.

s187
00:13:08.960 --> 00:13:15.480
And unlike things like a cargo competition, this has a real-world implication.

s188
00:13:15.760 --> 00:13:18.400
But, here's the result.

s189
00:13:18.400 --> 00:13:23.920
As you can see, we gave all the frontier models a 100K to start.

s190
00:13:23.920 --> 00:13:28.120
All of them lost her Sad.

s191
00:13:28.120 --> 00:13:30.960
And that captured the public's imagination.

s192
00:13:30.960 --> 00:13:34.480
Oh, AI is not as great as they thought.

s193
00:13:34.480 --> 00:13:37.800
And why are models so bad at long horizon?

s194
00:13:37.800 --> 00:13:45.000
First, I believe now the AI industry is a little bit too biased towards coding and procedure task.

s195
00:13:45.000 --> 00:13:51.160
What I mean is that the current task is a most formulated like do this and fix that.

s196
00:13:51.160 --> 00:13:54.320
Normally that limits the solution like one or two.

s197
00:13:54.320 --> 00:13:57.760
There isn't just too much space for creativity.

s198
00:13:57.760 --> 00:14:01.840
And second of all, not enough focus on open-ended task.

s199
00:14:01.840 --> 00:14:10.160
We live in the real world with a lot of a complexity, uncertainty, and that's not fully captured by the current benchmark.

s200
00:14:10.160 --> 00:14:13.920
And also, there isn't enough simulation of the real world.

s201
00:14:13.920 --> 00:14:16.160
We live in the world that there are other players.

s202
00:14:16.160 --> 00:14:23.600
Like in today's conference room, there are other real people who have a different thought, a different games than you have in your mind.

s203
00:14:23.600 --> 00:14:27.440
That's the complexity that's not fully captured.

s204
00:14:28.360 --> 00:14:33.880
And another thing I want to talk about is the long horizon impact on compute.

s205
00:14:33.880 --> 00:14:37.600
We know that GPUs are scarce and precious resources.

s206
00:14:37.600 --> 00:14:46.160
And in this case of a long horizon reasoning, you have to be careful about how to optimize your use between training and inference.

s207
00:14:46.160 --> 00:14:50.120
And pipeline RL is a quite popular technique nowadays.

s208
00:14:50.120 --> 00:14:56.640
So basically, it's a trade-off between off-policy and the GPU utilization.

s209
00:14:56.640 --> 00:15:02.280
So traditionally, you let inference run towards the end and then you start to train the model.

s210
00:15:02.280 --> 00:15:07.120
But in the case of long horizon, you have to wait until the inference finish.

s211
00:15:07.120 --> 00:15:17.200
What the pipeline RL does is that you let the sequence to be generated and you start to train the model while there's a still more sequences being generated.

s212
00:15:17.200 --> 00:15:19.520
And you see this created off-policy.

s213
00:15:19.520 --> 00:15:24.760
But from the our experience, normally off-policy up to eight steps is okay.

s214
00:15:24.760 --> 00:15:31.320
So, essentially, we made a trade-off between the off-policy and the GPU utilization.

s215
00:15:31.440 --> 00:15:33.200
But, here comes the issue.

s216
00:15:33.200 --> 00:15:39.480
As we the long horizon indicates, sometimes the inference would take weeks or even more.

s217
00:15:39.480 --> 00:15:45.960
In that case, inevitably, it will goes beyond the constraint of the eight steps of our policy.

s218
00:15:45.960 --> 00:15:50.800
So, your GPU have just to sit there idle and wait for it to finish.

s219
00:15:50.800 --> 00:15:57.400
And if you don't want to wait, as I mentioned before, applying the value model allows you to bootstrap.

s220
00:15:57.400 --> 00:16:02.360
What it means is that before the end of the episode, you generate expectation.

s221
00:16:02.360 --> 00:16:04.560
It's like a dopamine in human brain.

s222
00:16:04.560 --> 00:16:06.920
And that allows you to train the model.

s223
00:16:06.920 --> 00:16:08.760
But, here's another trade-off.

s224
00:16:08.760 --> 00:16:13.440
While you utilize the GPU fully, you introduce the value model bias.

s225
00:16:13.440 --> 00:16:17.800
So, there's always a bit of trade-off in those solutions.

s226
00:16:17.920 --> 00:16:25.400
And I want to also mention that in the long horizon, infrastructure is important, especially for the environment.

s227
00:16:25.400 --> 00:16:29.440
And Open Review was a product is a platform by General Reasoning.

s228
00:16:29.440 --> 00:16:31.280
If you're interested, you can check it out.

s229
00:16:31.280 --> 00:16:33.080
openreview.ai.

s230
00:16:33.080 --> 00:16:40.880
So, it's a place where host over 350 environments and with a single API endpoint.

s231
00:16:40.880 --> 00:16:47.839
And we use this for our internal RL and also some frontier labs and new labs are using this.

s232
00:16:48.760 --> 00:16:58.000
So, to summarize both Ross and my speech, it has been a long journey as long horizon indicate.

s233
00:16:58.000 --> 00:17:06.040
We as a team have seen the paradigms in AI reasoning on pre-training and agents in the past few years.

s234
00:17:06.040 --> 00:17:10.800
But, looking ahead, what makes us really excited is the long horizon.

s235
00:17:10.800 --> 00:17:19.280
And it requires us to think, have a new thinking on the algorithm, environments, and compute.

s236
00:17:19.280 --> 00:17:31.840
There are a lot of challenges and trade-offs, but we find it's really exciting to take on this journey because, as I mentioned in the very beginning, long horizon is not just engineering problem.

s237
00:17:31.840 --> 00:17:39.560
It is a mindset, and if we really are ambitious to solve humanity's biggest problems, this is the journey for everyone.

s238
00:17:39.560 --> 00:17:42.120
And that's also the mission for general reasoning.

s239
00:17:42.120 --> 00:17:45.000
So, if you are interested, follow us.

s240
00:17:45.000 --> 00:17:48.560
General Reasoning, we're London-based AI research company.

s241
00:17:48.560 --> 00:17:50.760
Thank you.

s242
00:18:04.037 --> 00:18:06.037
[music]
