WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=mMNkdYnIVC4

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/mMNkdYnIVC4.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:12.720 --> 00:00:13.720
All right.

s2
00:00:13.720 --> 00:00:17.800
I think we'll go ahead and get started with the with the presentation.

s3
00:00:17.800 --> 00:00:19.080
So my name is James So.

s4
00:00:19.080 --> 00:00:32.840
I am uh going to explain some of the work we're doing with Together AI and it's also in collaboration with Stanford around designing and optimizing environments for AI agents to enable these agents to make

s5
00:00:32.840 --> 00:00:35.520
new kinds of scientific discoveries.

s6
00:00:35.520 --> 00:00:37.600
All right.

s7
00:00:39.240 --> 00:00:50.960
So so that I guess the current paradigm of how people often are using or deploying AI agents is often involves designing workflows that sort of tells the agents you know what to do, right?

s8
00:00:50.960 --> 00:00:53.080
Or how the agent should work.

s9
00:00:53.080 --> 00:00:58.840
And it's typically done through a series of steps or prompts, tools, and instructions.

s10
00:00:58.840 --> 00:01:08.720
In contrast, the way we imagine the environment is that the environment should really specify not how the agent should work, but really where the agent should work, right?

s11
00:01:08.720 --> 00:01:20.800
And the environment then should provide a set of incentives and infrastructure for the agents and guardrails and resources so that agent can then flexibly work within that environment.

s12
00:01:20.800 --> 00:01:21.120
Right.

s13
00:01:21.120 --> 00:01:25.360
And our thesis here is that as agents become more and more powerful, right?

s14
00:01:25.360 --> 00:01:30.680
If we try to design workflows that often can limit the capabilities and creativity of the agents.

s15
00:01:30.680 --> 00:01:39.560
Whereas if we properly design the environment, this can enables a lot more creativity and capabilities and intelligence for the agents to naturally emerge.

s16
00:01:39.560 --> 00:01:46.040
This why I think we're trying to shift away from designing workflows and harnesses towards designing environments.

s17
00:01:46.640 --> 00:01:52.360
So what I want to do today is to give a few examples of the how we design environments for agents.

s18
00:01:52.360 --> 00:02:00.480
And in particular also show how they're able to then with with the right environment able to actually solve some really interesting and innovative problems.

s19
00:02:02.040 --> 00:02:07.640
So, the first example I want to share is the system that we environment that we created called the Einstein Arena.

s20
00:02:07.640 --> 00:02:18.760
It's sort of like the one of the first environments that enables AI agents to be able to collaborate in the wild and to compete to really solve open-ended scientific problems.

s21
00:02:18.760 --> 00:02:22.160
So, we designed this Einstein Arena to be really agent native.

s22
00:02:22.160 --> 00:02:31.200
So, I So, that means that it's very easy for agents to just read the skills talk on our on our arena and be able to access the arena.

s23
00:02:31.200 --> 00:02:37.040
And it's actually also designed so that it's intentionally very hard for humans to enter the arena, right?

s24
00:02:37.040 --> 00:02:43.120
So, you actually have to solve a little puzzle to prove that you're an AI agent in order to participate in this arena.

s25
00:02:43.120 --> 00:02:47.720
But, any agent in the world can openly and freely participate on the arena.

s26
00:02:47.720 --> 00:02:51.720
And once the agent actually enters into the Einstein Arena, this is what they'll see, right?

s27
00:02:51.720 --> 00:02:54.720
They'll see actually see a list of curated problems.

s28
00:02:54.720 --> 00:02:59.800
Each of these problems is actually a problem that we curated, so it's a scientifically interesting problem.

s29
00:02:59.800 --> 00:03:06.840
And we curated these problems so that first, there's actually an existing community of human researchers that are interested in these problems.

s30
00:03:06.840 --> 00:03:09.640
So, these are important problems for human scientists.

s31
00:03:09.640 --> 00:03:20.760
And second is that for each of these problems, we can actually create a well-defined and deterministic deterministic verifier to assess the quality of the solutions to each of these problems.

s32
00:03:20.760 --> 00:03:24.680
And I'll give some examples in a couple of slides.

s33
00:03:25.280 --> 00:03:30.360
So, So, the agents can actually decide which of these problems they're interested in once they log onto the arena, right?

s34
00:03:30.360 --> 00:03:34.680
So, if they enter into a particular problem space, this is what they'll see, right?

s35
00:03:34.680 --> 00:03:38.840
They'll see some description that precisely explains what is the problem.

s36
00:03:38.840 --> 00:03:42.640
We have a discussion forum where the agents can communicate.

s37
00:03:42.640 --> 00:03:49.280
It's almost like a social network where the agents can actually communicate and talk to each other and ask for help or give recommendations.

s38
00:03:49.280 --> 00:03:52.360
Um and we also have a leaderboard.

s39
00:03:52.360 --> 00:03:55.320
This is where the agent can actually see each other's solutions.

s40
00:03:55.320 --> 00:03:55.560
Right?

s41
00:03:55.560 --> 00:04:01.840
So in any in at any time they want, the agent can actually submit a solution to one of these problems.

s42
00:04:01.840 --> 00:04:09.040
And because we have this verifier, we can actually then determine what is the quality of that solution and provide a score in real time.

s43
00:04:09.040 --> 00:04:12.560
So this leaderboard is being constantly updated in real time.

s44
00:04:12.560 --> 00:04:16.760
And the agents can also see how other agents are doing on this problem.

s45
00:04:16.760 --> 00:04:20.680
And they can also see other agents' solutions and download those solutions.

s46
00:04:20.680 --> 00:04:25.280
So there's both a collaboration dynamics and also a competition dynamics in this arena, right?

s47
00:04:25.280 --> 00:04:29.320
They can collaborate and ask each other questions and help in the discussion forum.

s48
00:04:29.320 --> 00:04:31.480
But agents are also competing with each other.

s49
00:04:31.480 --> 00:04:38.840
And that's why I think this also sort of simulates how human researchers can compete and also collaborate to solve interesting problems.

s50
00:04:39.840 --> 00:04:45.840
So we launched this AI instant arena environment earlier this year, I think in March.

s51
00:04:45.840 --> 00:04:58.200
And within a few weeks, it's already actually we're very impressed and very surprised that the agents were actually able to already discover new solutions to 11 problems that are of the best

s52
00:04:58.200 --> 00:04:59.680
solutions that have ever been found.

s53
00:04:59.680 --> 00:04:59.800
Right?

s54
00:04:59.800 --> 00:05:08.640
So that means that the solutions that they discovered by the agents on AI instant arena were better than any previous human solutions or any solutions that we acquired using

s55
00:05:08.640 --> 00:05:12.080
more specialized AI tools.

s56
00:05:12.480 --> 00:05:19.440
So I'll just give you example of one such solution or one such problem which is called the kissing number problem.

s57
00:05:19.440 --> 00:05:21.120
So this is actually a very famous problem.

s58
00:05:21.120 --> 00:05:22.840
It's been around for hundreds of years.

s59
00:05:22.840 --> 00:05:27.120
So for example, Isaac Newton was already working on some version of this kissing number problem.

s60
00:05:27.120 --> 00:05:29.440
And it's actually relatively easy to state.

s61
00:05:29.440 --> 00:05:29.720
Right?

s62
00:05:29.720 --> 00:05:40.480
So the kissing number problem basically asks that what is the maximum number of spheres that you can place around the central sphere so that these additional spheres do not overlap each other?

s63
00:05:40.480 --> 00:05:42.400
So for example, in one dimensions, right?

s64
00:05:42.400 --> 00:05:46.880
So around the central sphere I can place one sphere to the left and one sphere to the right without overlap.

s65
00:05:46.880 --> 00:05:49.480
So the kissing number in one dimension is easy to compute.

s66
00:05:49.480 --> 00:05:50.800
This is two.

s67
00:05:50.800 --> 00:05:54.800
In two dimensions, it's also easy to show that you can at most place six spheres.

s68
00:05:54.800 --> 00:05:58.120
So, that's the kissing number in two dimensions is six.

s69
00:05:58.120 --> 00:06:05.360
But, it turns out that in higher dimensions, it actually becomes really hard to compute what's the maximum number of over non-overlapping spheres.

s70
00:06:05.360 --> 00:06:08.440
And the kissing number problem in higher dimensions is actually open, right?

s71
00:06:08.440 --> 00:06:12.760
It's not been It's not clear what is the optimal number.

s72
00:06:12.760 --> 00:06:19.120
And so, scientists have been trying to work on this problem for the last several centuries.

s73
00:06:19.240 --> 00:06:26.800
And in particular, right, so the kissing number problem in 11 dimensions has attracted a lot of interest for various reasons.

s74
00:06:26.800 --> 00:06:31.600
So, this is actually sort of a progression of the solutions in 11 dimensions.

s75
00:06:31.600 --> 00:06:41.520
So, in the 1980s, right, so it's best known that there you can place 440 spheres, right, in 11 dimensions without overlap.

s76
00:06:41.520 --> 00:06:53.080
And in I think 19 uh So, yeah, so so in in 1980, there was a big advance that the first for the first time showed that you can actually just construct

s77
00:06:53.080 --> 00:06:57.520
with 582 spheres in 11 dimensions without overlap.

s78
00:06:57.520 --> 00:07:10.280
Uh and then that sort of stuck there for about 40 years, right, until 2022, where a mathematician is able to publish a new advance, right, a breakthrough that's able to improve that to 592

s79
00:07:10.280 --> 00:07:11.680
spheres.

s80
00:07:11.680 --> 00:07:18.600
And then there's another breakthrough from DeepMind the following year that advances that to 593 spheres.

s81
00:07:18.600 --> 00:07:28.120
But, with Alpha Zero, we know by having these agents able to collaborate actively, right, in the wild, within a few days they were actually able to construct a new solution

s82
00:07:28.120 --> 00:07:34.680
that shows that for the first time you can create 604 spheres in 11 dimensions that do not overlap.

s83
00:07:34.680 --> 00:07:43.560
And this is not just a problem that's of mathematical interest, because it turns out that the more of these sort of spheres you can place in higher dimensions without overlap that actually creates

s84
00:07:43.560 --> 00:07:49.880
the better coding systems including ways of like doing error correction codes for information transfer.

s85
00:07:49.880 --> 00:07:56.720
Right, so this actually is by creating this better constructions that also leads to this better engineering algorithms.

s86
00:07:57.000 --> 00:08:02.240
And in this case actually the collaborations among these agents is really critical for making these advances, right?

s87
00:08:02.240 --> 00:08:06.760
So this is a problem where not a single agent is able to solve by itself, right?

s88
00:08:06.760 --> 00:08:12.160
Not you know, GPT 5.5 or a cloud models that can't really solve the problem by itself.

s89
00:08:12.160 --> 00:08:15.440
So the collaboration among multiple agents is really critical.

s90
00:08:15.440 --> 00:08:24.720
And here we're actually able to show that there's like this sort of a lineage trace of how the agents are able to collaborate and then basically take each other's solutions and refine that

s91
00:08:24.720 --> 00:08:28.320
and further optimize it to arrive at this breakthrough.

s92
00:08:28.320 --> 00:08:32.599
And you can also see some of these interactions and discussions on Einstein Arena, right?

s93
00:08:32.599 --> 00:08:40.080
Where here's an example where the one agent actually was asking other agents, "Have you tried you know, some of these approaches?"

s94
00:08:40.080 --> 00:08:46.960
Um, with uh, these STP approaches and then the other agents showed that yes, we have tried these approaches and here are some of the things that we found.

s95
00:08:46.960 --> 00:08:56.360
Right, so the information sharing on the forums on the arena is actually really important to help the agents to arrive at this solution together.

s96
00:08:57.360 --> 00:09:09.520
So in addition to solving these interesting scientific problems, but we've also been using platforms like the right Einstein Arena uh, to help to improve uh, you know, machine learning and AI itself.

s97
00:09:09.520 --> 00:09:17.960
Right, so here's one example where we actually use these agents to basically help us to create better kernels for and to speed up those kernels.

s98
00:09:17.960 --> 00:09:20.480
Right, and here we use the same environment, right?

s99
00:09:20.480 --> 00:09:25.000
Where the agents can compete and they also can collaborate and they see these leaderboards.

s100
00:09:25.000 --> 00:09:37.120
And we basically change the back end instead of trying to verify the solutions to this mathematics problem, here we're basically trying to you know, we will compile and benchmark and test and verify the quality and the speed of the individual kernels,

s101
00:09:37.120 --> 00:09:37.320
right?

s102
00:09:37.320 --> 00:09:42.760
And then we'll provide a feedback to the agents in real time in the form of these leaderboards.

s103
00:09:43.520 --> 00:09:49.000
In these kernel settings, we also found it to be quite useful to have different agents with different personas, right?

s104
00:09:49.000 --> 00:09:54.920
And these different personas actually corresponds to a different uh roles and priors that agents can actually have.

s105
00:09:54.920 --> 00:10:02.040
So, for example, we have one agent that looks at tends to look at more of the profiling, another agent that tends to look at more of the memory consumptions,

s106
00:10:02.040 --> 00:10:06.120
a third agent that looks at, you know, the precisions, the tensor computations.

s107
00:10:06.120 --> 00:10:14.040
And these agents can and then across different personas, they can able to collaborate and a compete on the arena to speed up the kernels.

s108
00:10:14.680 --> 00:10:26.760
And in this case, right here, the agents were also able to collaborate and lead to really quite substantial speed ups, uh including sometimes over two two x two-fold speed ups in some of these production kernels.

s109
00:10:26.760 --> 00:10:35.280
So, here I'm just showing you a few examples where for things like page attention, uh and these are sort of for specific shapes, but we also have generalized this to many different shapes

s110
00:10:35.280 --> 00:10:37.640
and different uh hardware types, right?

s111
00:10:37.640 --> 00:10:47.240
Where we're actually seeing that we're getting up to sometimes over two x speed up in these kernels, and they uh compared to the previous state-of-the-art kernels for these problems.

s112
00:10:47.240 --> 00:10:55.480
And these improved kernels created designed by the agents are actually already used in in production at Together AI.

s113
00:10:57.200 --> 00:11:07.120
So, in the last few minutes, I want to show like a second example of a kind of environment that we created as a way to uh train and to create better data scientist agents,

s114
00:11:07.120 --> 00:11:07.280
right?

s115
00:11:07.280 --> 00:11:15.600
So, we call this DS Gym, which stands for data science gym, which is sort of like a unified environment that we created for both for evaluating and for training

s116
00:11:15.600 --> 00:11:20.480
data science agents to solve complex data science problems.

s117
00:11:21.120 --> 00:11:29.280
So, here in this DS Gym environment, we also curated and created a unified list of different data sets and tasks, right?

s118
00:11:29.280 --> 00:11:33.800
So, these data sets can combine uh spans across many different settings.

s119
00:11:33.800 --> 00:11:42.400
And the agents are then able to interact with these different data sets that we have through a unified uh interface and through code execution.

s120
00:11:42.400 --> 00:11:47.000
In the DSGM environment, we also provide a unified infrastructure for the agents.

s121
00:11:47.000 --> 00:11:55.640
So, for example, the agents can actually spin up many different Docker containers to test their data science algorithms and actually run them in parallel.

s122
00:11:58.080 --> 00:12:10.360
So, in the process of actually creating the data sets and tasks for the data DSGM environment, so we initially actually wanted to incorporate some of the existing data science benchmarks that have been used to evaluate agents.

s123
00:12:10.360 --> 00:12:16.280
But we actually quickly realized that many of the existing widely-used benchmarks actually have many problems.

s124
00:12:16.280 --> 00:12:19.760
And one big problem is that they're actually very vulnerable to shortcuts.

s125
00:12:19.760 --> 00:12:26.640
By shortcut, I mean here is that uh down here what I'm showing are three different common popular data science benchmarks.

s126
00:12:26.640 --> 00:12:26.920
Right?

s127
00:12:26.920 --> 00:12:32.200
And the in green here we see shows like the performance of the agents on these benchmarks.

s128
00:12:32.200 --> 00:12:39.880
Uh but the red bar also shows how well they're able to the what fraction of the benchmark the agents can actually solve without actually using the data sets themselves.

s129
00:12:39.880 --> 00:12:40.040
Right?

s130
00:12:40.040 --> 00:12:47.320
So, just by reasoning or by, you know, uh doing other shortcuts without actually actually do working with the underlying data sets.

s131
00:12:47.320 --> 00:12:57.680
And across many of these different benchmarks, right, sometimes up to 20 to 50% of the tasks can be solved without actually looking at any of the underlying data.

s132
00:12:57.680 --> 00:13:02.600
Which I think is uh really a significant problem with many of the existing benchmarks.

s133
00:13:03.120 --> 00:13:12.440
So, to address that, we actually carefully curated at our own our own benchmarks, right, for both for scientific analysis and also for predictive modeling.

s134
00:13:12.440 --> 00:13:21.560
So, for scientific analysis and discovery, the way we did this is that we actually went through recently published papers and then carefully curated data and then also tasks from those papers.

s135
00:13:21.560 --> 00:13:26.040
And then we also had human scientists and experts to review each of those tasks.

s136
00:13:26.040 --> 00:13:34.720
And for predictive modeling, the way we did this is go through all the different Kaggle competitions to look for some of the recent Kaggle competitions that are still open

s137
00:13:34.720 --> 00:13:39.880
and and where also you have high quality data sets and also high quality evaluations.

s138
00:13:39.880 --> 00:13:48.440
Then we curated those into the DS Gym as a kind of task for evaluating how well models agents can actually build predictive models.

s139
00:13:48.440 --> 00:13:53.320
So all together in the DS Gym, we actually have created over a dozen different tasks.

s140
00:13:53.320 --> 00:14:01.080
They span across dozens of different scientific domains ranging from biology to physics to economics.

s141
00:14:01.080 --> 00:14:04.920
It also involves many different data types and data modalities.

s142
00:14:06.600 --> 00:14:12.320
So this actually makes it very easy for us to evaluate different models, both open and closed source models.

s143
00:14:12.320 --> 00:14:22.280
And one thing we found is that the existing models, even the frontier models, often are only still achieves like less than 50% accuracy performance on the DS Gym tasks.

s144
00:14:22.280 --> 00:14:22.400
Right?

s145
00:14:22.400 --> 00:14:26.280
So these are definitely not saturated benchmarks.

s146
00:14:26.280 --> 00:14:31.680
We can also use a DS Gym as sort of like a training factory to improve these open source models.

s147
00:14:31.680 --> 00:14:31.800
Right?

s148
00:14:31.800 --> 00:14:40.440
So one thing we did here is actually generate in the DS Gym actually the gym itself will actually create all these execution verified trajectories, which means that these are trajectories

s149
00:14:40.440 --> 00:14:47.680
generated by the agents that have been verified through the through through actually executing the code from the agents.

s150
00:14:47.680 --> 00:14:47.880
Right?

s151
00:14:47.880 --> 00:14:58.120
So by generating these execution verified trajectories, then we are able to like fine-tune sort of small open source models that actually now achieve sort of the they're sort of the best in class

s152
00:14:58.120 --> 00:15:01.400
open source models in terms of solving these kind of data science tasks.

s153
00:15:01.400 --> 00:15:01.600
Right?

s154
00:15:01.600 --> 00:15:07.560
And these models are small enough that you can actually run them locally on your laptops and your computers.

s155
00:15:08.440 --> 00:15:11.560
So just to summarize the this part was the data science gym.

s156
00:15:11.560 --> 00:15:11.720
Right?

s157
00:15:11.720 --> 00:15:21.480
So we with DS Gym, we created this unified execution layer so people can actually run and all these different tasks across dozens of different tasks across many different domains.

s158
00:15:21.480 --> 00:15:29.040
We have carefully verified that there are no shortcuts in these tasks, which has been sort of a common challenge with existing data science benchmarks.

s159
00:15:29.040 --> 00:15:39.320
And we also enable in the DSG and a way to generate synthetic data, so you can easily use that to improve and to train your own data science agents.

s160
00:15:39.680 --> 00:15:50.080
So, just to summarize the presentation, um I think the main takeaway here is that I think we're in seeing this interesting progression as in terms of how we build different AI systems.

s161
00:15:50.080 --> 00:15:50.280
Right.

s162
00:15:50.280 --> 00:15:56.320
So, then in the past, people have been building these AI systems mostly by designing individual models or individual tools.

s163
00:15:56.320 --> 00:16:02.320
And currently, there's a lot of focus on creating designing agents or harnesses and workflows around agents.

s164
00:16:02.320 --> 00:16:13.160
But what our research shows is that I think we're already moving towards the next stage, where you're not then trying to design workflows or specific or specific agents, what we really want to do is to design environments,

s165
00:16:13.160 --> 00:16:20.960
which is a set of infrastructure and incentives that in that motivates the agents that you solve more and more challenging problems.

s166
00:16:20.960 --> 00:16:31.839
And with appropriate designs these environments can actually unlock much more creativity and collective intelligence from the agents that's that's limited by the existing workflows.

s167
00:16:31.839 --> 00:16:36.800
And here are some of the references for the papers that we published that describes these in more detail.

s168
00:16:36.800 --> 00:16:38.886
So, thank you very much.

s169
00:16:38.886 --> 00:16:40.886
[applause]

s170
00:16:51.986 --> 00:16:53.986
[music]
