Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning

AI Engineer · 18 min · 242 sentences · from YouTube's caption track

Each timecode opens YouTube at the start of that sentence. Line anchors (#s42) are the cue ids in the WebVTT, and every line carries its start and end seconds. All transcripts has every talk, and the whole corpus as one file.

  1. 00:12Uh so this talk is called scaling to long horizons.
  2. 00:15My name is Ross.
  3. 00:16Uh I'm the CEO of GR.
  4. 00:17We're a London-based reinforcement learning company.
  5. 00:20Uh before GR, I was the reasoning lead at Meta AI working on Llamas, uh Galactica, lots of other models back in the day.
  6. 00:27I'm joined by Chengxi, uh co-founder and president of GR.
  7. 00:30Uh and yeah, hit today we're going to talk about algorithms, environments, compute, all the things you need to do to get agents scaling uh to kind of long with tasks.
  8. 00:41So we're going to have two parts of this talk today.
  9. 00:42I'm going to first of all start with a personal perspective about, you know, the early days, the golden age of language modeling in between like maybe 2020 and 2023.
  10. 00:52Uh I'll talk about, like I said, all these models and some of our early reinforcement learning efforts for LLMs.
  11. 00:56And then Chengxi is going to talk about, you know, what's ahead, you know, what the next frontiers.
  12. 01:00And yeah, that's going to be a really talk with a lot of alpha, so I'd encourage you to stick around for that.
  13. 01:06So the journey so far.
  14. 01:07So my journey started here.
  15. 01:10Um so this was the Papers With Code team.
  16. 01:12I'm sure many of you used Papers With Code back in the day.
  17. 01:15So we were a London-based startup 2019.
  18. 01:18Uh we were acquired by Meta later that year.
  19. 01:20And then we had a crazy transition within Meta to do research.
  20. 01:23Um so we did, like I said, Galactica.
  21. 01:26Then after ChatGPT came out, we started the post-training for Llama 2, Llama 3. So all the great work you uh saw there was folks in this room.
  22. 01:34And lots of other interesting stuff that never got published as well.
  23. 01:37Uh reasoning Llama and lots of other things.
  24. 01:39Um so yeah, this this small team I'd like to think like the open weight kind of revolution started in this room.
  25. 01:44And you know, it really like hit home this idea to me that kind of small focused teams, like even in like the age of scaling, can do amazing things if people are aligned.
  26. 01:54Now for me, uh things got particularly crazy in 2022.
  27. 01:59Uh so let me tell you a story.
  28. 02:00Um the media perception is that ChatGPT came out of nowhere, you know, shocked the world, and that's how kind of the modern AI wave started.
  29. 02:10But, you know, I have a different personal perspective on this because 2 weeks before ChatGPT came along, there was another language model called Galactica.
  30. 02:18So let's talk about Galactica.
  31. 02:20Galactica and, you know, ChatGPT, you know, they were both, you know, in some respects quite similar.
  32. 02:25They're both based on pretty good base models.
  33. 02:27Galactica itself was a base model, and then ChatGPT was based on GPT-3.5.
  34. 02:33But there was a clear difference in outcomes.
  35. 02:35So Galactica at the time shipped with this slight base model demo.
  36. 02:39And as you guys know now, like base models, they come with a lot of quirks.
  37. 02:42You know, they hallucinate.
  38. 02:43You prompt them to do, you know, silly things, they would do silly things.
  39. 02:47Whereas ChatGPT wasn't just a base model, but had this like crucial reinforcement learning from human feedback pipeline.
  40. 02:54And this was the key thing that made LLMs like really products for the first time.
  41. 02:58So I like to think in a weird kind of way, this is like the first like kind of natural experiment showing you that kind of RL like provides value, right?
  42. 03:05Uh and to my misfortune, it was like a very personal like kind of a natural experiment and you know, Galactica blew up.
  43. 03:11Uh but that's like a good like kind of lesson there.
  44. 03:13A good base model is not enough.
  45. 03:15So I took that lesson quite early on.
  46. 03:18So like I said, RLHF made LLMs products.
  47. 03:21They were the thing that kind of made LLMs cross the Rubicon into something that wasn't just a toy, but used by now billions of people.
  48. 03:27But you didn't have to wait until ChatGPT to see this.
  49. 03:29Like even at the time, like InstructGPT in 2022 had these pretty stunning results.
  50. 03:34Like a 1 billion parameter model with RLHF was outperforming 175 billion models.
  51. 03:40So two orders of magnitude fewer parameters, but getting better results.
  52. 03:44So that was astonishing.
  53. 03:45So if you were paying attention closely, you know, maybe we should have been as well, but we were focused on a, you know, bloody base model, which is hard work in 2022.
  54. 03:52But that shows you how you know important even basic RL is.
  55. 03:57And the Galactica demo itself, I mean, it set off a storm.
  56. 03:59So, ancient history now, but we put out a demo.
  57. 04:01We let people play around with it um because it's kind of cool.
  58. 04:05And at the time people got scared.
  59. 04:07So, it was like, you know, like I said, you prompt it on like a research paper on Dyson spheres or a report on the benefits of eating crushed glass and people are like, "Oh my god."
  60. 04:16Um so, that was the state of things in 2022.
  61. 04:18And yeah, I'll be honest, the meta association didn't you know help us either.
  62. 04:22Um and yeah, the tragic story in a way was a lot of the, you know, novel work was maybe overshadowed.
  63. 04:28But you know, the paradox of this whole thing is that at the time Galactica was actually a bloody good model.
  64. 04:32Um like it outperformed Palm, Chinchilla, GPT-3.5 with a lot less compute in scientific domains.
  65. 04:38It was state-of-the-art.
  66. 04:39So again, that also shows you how powerful RL is.
  67. 04:42You can have a sota base model, but that is not enough.
  68. 04:45Um so, here you see on like a math who's kind of beating Chinchilla.
  69. 04:49Uh kind of latex equations, you know, science, you know, it was getting around 68% compared to GPT-3.5, 49%.
  70. 04:54So, crushing there.
  71. 04:56And chain of thought as well.
  72. 04:57So, you know, a Palm at the time which was a Google Brain model, 540 billion.
  73. 05:01You know, that's 30 billion in Galactica was getting like 36 versus 19%.
  74. 05:05So, double the performance, order of magnitude less results.
  75. 05:07So, that again reinforces base good base models not enough.
  76. 05:12But it introduced some key ideas which I think are very important.
  77. 05:14I mean, Galactica was the first LM to really crack data efficiency.
  78. 05:17105 billion uh token corpus compared to you know, trillion uh tokens in Chinchilla.
  79. 05:23And it was a really contrarian at the time cuz you know, at the time everyone was like, "Okay, we just need more tokens."
  80. 05:27And Galactica said, "No.
  81. 05:28High quality, you know, curated data sets really matter."
  82. 05:31And that was a real driver of those results you just saw.
  83. 05:34And it was also like the first major LM to really crack multi-epoch training.
  84. 05:38It sounds ridiculous now, but at the time the consensus was you don't do more than an epoch.
  85. 05:42Um but this kind of rule of thumb, you may have heard of it like four uh epochs repeated data, that was formalized later.
  86. 05:47but Galactica was the first like real empirical result for that.
  87. 05:51Now, perhaps more importantly, there's this idea of thinking tokens.
  88. 05:54And some of you might remember this, but it was like really quite buried within the paper.
  89. 05:57So, around this time there were like different ideas for reasoning.
  90. 05:59There was chain of thought, which was one idea.
  91. 06:01It's where you prompt for like kind of like the steps.
  92. 06:03There were scratch pads where you just like put like very like numerical kind of intermediate steps.
  93. 06:08But, Galactica was really this first idea which said, "No, this is an internal working memory process.
  94. 06:13This is an internal thinking.
  95. 06:15You should be inside these tags, and you should spend the inference computer before you get to an answer, right?"
  96. 06:19So, these are all like quite like pressing ideas.
  97. 06:23But, you know, Galactica came and went, blew up, and then, you know, we were kind of tasked as a team to kind of spin up the post-training effort for Llama.
  98. 06:30But, I had like a personal obsession, which was like reasoning.
  99. 06:34And I had like a really simple idea at the time, which is, "What if we applied kind of reinforcement learning pressure to this like kind of thinking tags as work?"
  100. 06:42Like, what if we just optimize the thing in between the thinking?
  101. 06:45And if that sounds familiar, then this is like kind of what Deep Seek 01 ended up doing 2 years later.
  102. 06:51But, there's a key difference.
  103. 06:52So, at the time we only had Llama 2 base models, terrible mathematics corpus, terrible results on math.
  104. 06:58And, you know, the context window, you know, we're all like, you know, very context-rich now, 1 million tokens.
  105. 07:03You know, the back in the day it was 4,000.
  106. 07:05It wasn't too fun.
  107. 07:07But, we still had a recipe at the time, and this is unpublished, but it was a really good for the meta.
  108. 07:11So, RSP was this.
  109. 07:14Number one, continue pre-training of Llama 2 towards mathematics and science data.
  110. 07:18So, that's the first thing.
  111. 07:19Llama 2, math corpus, so let's fix that.
  112. 07:22Number two, PPO with verifiable rewards.
  113. 07:25But, notice this isn't GRPO, right?
  114. 07:27So, we had a time and a strong outcome reward model to initialize the value model.
  115. 07:31That was a key thing, lots of data on value models at the time.
  116. 07:34And internally at the time, this kind of recipe led to state-of-the-art results on math and reasoning.
  117. 07:39So, we were kind of like, "Wow, this is like really shows the power of having the right objective."
  118. 07:44But, the really fascinating thing is like we had great results, but we didn't have like inference time scaling.
  119. 07:51We didn't have this reflective behavior that became like the hallmark of R1 and O1.
  120. 07:55You know, back weights, you know, back tracking, all this kind of stuff.
  121. 07:58So, it begs like the question like, "Why?
  122. 08:00Well, why didn't we have that moment?"
  123. 08:02And we got an answer around like 2 years later.
  124. 08:05So, there's a couple of like things going on here, but essentially better base models were the thing that really got RL cooking.
  125. 08:11And when DeepMind came out, I was kind of shocked at the time.
  126. 08:13I was like, "Holy we just like tried the same thing.
  127. 08:16We didn't have this.
  128. 08:16What's going on here?"
  129. 08:17And in a weird kind of way, the real lesson was it was just like the bitter lesson, like the most purest form of bitter lesson possible.
  130. 08:23Like, better base models, more RL computes, bigger context windows, and that's all you need for this kind of emergent behavior.
  131. 08:31It also like says something like quite important about the sociology of like research because the fact that OpenAI had this model, you know, GPT-4 level model before anyone else, it allowed them to see further, right?
  132. 08:41So, that's a really interesting point.
  133. 08:42Like, the the age of scaling means that if you have certain prerequisites in place, you become smarter, you see further, you see more ideas.
  134. 08:49So, really interesting point.
  135. 08:52So, this was me in 2024.
  136. 08:53I was a very sad panda, uh defeated by ChatGPT and O1.
  137. 08:57But, I wasn't deterred.
  138. 08:59Um so, I wanted to seek the next wave, and I still was like convinced that kind of reasoning hadn't been solved.
  139. 09:05So, we started GR uh to take on truly like big tasks.
  140. 09:08And with that in mind, I'm going to hand over to Chengxi.
  141. 09:10He's going to talk about what we're kind of thinking about now.
  142. 09:13Ross.
  143. 09:14Ross.
  144. 09:14Ross.
  145. 09:20Hi, everyone.
  146. 09:21I'm Chengxi Taylor, co-founder and president of General Intelligence Inc. I'm going to share what it takes to scale to long horizon.
  147. 09:31First, I want to make it clear.
  148. 09:33Long horizon task is not just an engineering problem.
  149. 09:38It is a mindset.
  150. 09:40If we want to solve humanity's biggest problems, such as cure cancer, solve millennium's prize problem, or go into Mars, we have to be patient.
  151. 09:51It take time.
  152. 09:52And if we want AI to move us towards that level impact, we have to think about long horizon.
  153. 10:02But here's the first problem.
  154. 10:04We have a scarce context window.
  155. 10:07If you take Fermat's Last Theorem as example, what it take for the mathematician was over 10 years time of reading paper, writing down thoughts in the scratch pad, or taking a walk
  156. 10:18to generate the creative ideas.
  157. 10:21If we convert to token, that's probably tens of a billions of even hundreds billions.
  158. 10:26But where we are now, just a 1 million token context window.
  159. 10:32So one solution is use compaction.
  160. 10:35So what it essentially does is generate the token until the end of the context window, summarize, and then on top of that generate more tokens.
  161. 10:45And the beauty of applying RL in the situation is kind of like kill two birds with one stone.
  162. 10:50You apply RL to the compaction and also the task.
  163. 10:56But here's the problem.
  164. 10:57With long horizon, there are three issues.
  165. 11:00The first is the gradient variance scales with the length.
  166. 11:04And the second is a sparse reward.
  167. 11:06And you have this credit assignment problem.
  168. 11:09And finally, there's also variable length of the trajectory that adds to the problem of optimization.
  169. 11:17So to solve this issue, we can apply critics, which is the value model.
  170. 11:22And value model can reduce the variance and also have a couple advantages, such as on the trajectory level that fits compaction very well, and also encourage the batch diversity.
  171. 11:34And also, I'll talk later on bootstrapping.
  172. 11:37Basically, get signal before the end of the episode.
  173. 11:40But, the downside for this is that um it's more complicated than GRPO, and basically, you have to train another value model alongside with the policy model.
  174. 11:52And there's some tools to help with the context limitations, such as a file system tools, which essentially like a scratchpad for AI to write to the uh reasoning thought.
  175. 12:02And self-search tools, which allows agent to search over the previous trajectory.
  176. 12:08And then you have archive tools.
  177. 12:11In the case like auto research, you can build upon your previous result.
  178. 12:15But, we have to be careful.
  179. 12:16In other scenario, you don't want an AI to cheat by just to grab the previous answer without thinking.
  180. 12:24So, how good are the current model on this long horizon task?
  181. 12:31In the general reasoning, we construct the benchmark called a Kelly bench, where we're actually featured in the front page of the Financial Times.
  182. 12:39Kind of a caught us off guard how much the mainstream have interest in this.
  183. 12:43So, basically, what we did is that we allowed the agents to build machine learning models to um trade in the football matches over 1-year horizon.
  184. 12:54In this case, it's a Premier League, if you're interested in football.
  185. 12:58And we're so fascinated by this because there's a real money to be made, and if it was a successful, couldn't make a billions.
  186. 13:06There's a whole industry on sports betting.
  187. 13:08And unlike things like a cargo competition, this has a real-world implication.
  188. 13:15But, here's the result.
  189. 13:18As you can see, we gave all the frontier models a 100K to start.
  190. 13:23All of them lost her Sad.
  191. 13:28And that captured the public's imagination.
  192. 13:30Oh, AI is not as great as they thought.
  193. 13:34And why are models so bad at long horizon?
  194. 13:37First, I believe now the AI industry is a little bit too biased towards coding and procedure task.
  195. 13:45What I mean is that the current task is a most formulated like do this and fix that.
  196. 13:51Normally that limits the solution like one or two.
  197. 13:54There isn't just too much space for creativity.
  198. 13:57And second of all, not enough focus on open-ended task.
  199. 14:01We live in the real world with a lot of a complexity, uncertainty, and that's not fully captured by the current benchmark.
  200. 14:10And also, there isn't enough simulation of the real world.
  201. 14:13We live in the world that there are other players.
  202. 14:16Like in today's conference room, there are other real people who have a different thought, a different games than you have in your mind.
  203. 14:23That's the complexity that's not fully captured.
  204. 14:28And another thing I want to talk about is the long horizon impact on compute.
  205. 14:33We know that GPUs are scarce and precious resources.
  206. 14:37And in this case of a long horizon reasoning, you have to be careful about how to optimize your use between training and inference.
  207. 14:46And pipeline RL is a quite popular technique nowadays.
  208. 14:50So basically, it's a trade-off between off-policy and the GPU utilization.
  209. 14:56So traditionally, you let inference run towards the end and then you start to train the model.
  210. 15:02But in the case of long horizon, you have to wait until the inference finish.
  211. 15:07What the pipeline RL does is that you let the sequence to be generated and you start to train the model while there's a still more sequences being generated.
  212. 15:17And you see this created off-policy.
  213. 15:19But from the our experience, normally off-policy up to eight steps is okay.
  214. 15:24So, essentially, we made a trade-off between the off-policy and the GPU utilization.
  215. 15:31But, here comes the issue.
  216. 15:33As we the long horizon indicates, sometimes the inference would take weeks or even more.
  217. 15:39In that case, inevitably, it will goes beyond the constraint of the eight steps of our policy.
  218. 15:45So, your GPU have just to sit there idle and wait for it to finish.
  219. 15:50And if you don't want to wait, as I mentioned before, applying the value model allows you to bootstrap.
  220. 15:57What it means is that before the end of the episode, you generate expectation.
  221. 16:02It's like a dopamine in human brain.
  222. 16:04And that allows you to train the model.
  223. 16:06But, here's another trade-off.
  224. 16:08While you utilize the GPU fully, you introduce the value model bias.
  225. 16:13So, there's always a bit of trade-off in those solutions.
  226. 16:17And I want to also mention that in the long horizon, infrastructure is important, especially for the environment.
  227. 16:25And Open Review was a product is a platform by General Reasoning.
  228. 16:29If you're interested, you can check it out.
  229. 16:31openreview.ai.
  230. 16:33So, it's a place where host over 350 environments and with a single API endpoint.
  231. 16:40And we use this for our internal RL and also some frontier labs and new labs are using this.
  232. 16:48So, to summarize both Ross and my speech, it has been a long journey as long horizon indicate.
  233. 16:58We as a team have seen the paradigms in AI reasoning on pre-training and agents in the past few years.
  234. 17:06But, looking ahead, what makes us really excited is the long horizon.
  235. 17:10And it requires us to think, have a new thinking on the algorithm, environments, and compute.
  236. 17:19There are a lot of challenges and trade-offs, but we find it's really exciting to take on this journey because, as I mentioned in the very beginning, long horizon is not just engineering problem.
  237. 17:31It is a mindset, and if we really are ambitious to solve humanity's biggest problems, this is the journey for everyone.
  238. 17:39And that's also the mission for general reasoning.
  239. 17:42So, if you are interested, follow us.
  240. 17:45General Reasoning, we're London-based AI research company.
  241. 17:48Thank you.
  242. 18:04[music]