Training Taste — Thais Castello Branco, Taste Labs

AI Engineer · 15 min · 187 sentences · from YouTube's caption track

Each timecode opens YouTube at the start of that sentence. Line anchors (#s42) are the cue ids in the WebVTT, and every line carries its start and end seconds. All transcripts has every talk, and the whole corpus as one file.

  1. 00:01[music]
  2. 00:12Test.
  3. 00:13Okay, amazing.
  4. 00:15It's great to meet everyone.
  5. 00:16I'm Taís.
  6. 00:16I'm the founder of Taste Labs.
  7. 00:19Uh for those of you who don't know us, we came out of stealth a few weeks ago and our whole mission is basically how do we end AI slop?
  8. 00:25I that's my personal enemy.
  9. 00:27Um and so we really believe that to solve this problem of slop, we have to like decode subjective domains.
  10. 00:34Uh there's been so much effort being put into getting models and agents amazing at things like coding and math.
  11. 00:39Uh and it's time that we put all that same effort into making them great at things like design uh and writing.
  12. 00:44And so design is this first pillar that we're starting with and it's been it's been incredibly exciting.
  13. 00:49Um We work primarily in two ways.
  14. 00:51So we work a lot with the frontier labs on how do we evaluate their models, understand where they're breaking, understand what could be better about them, and then construct the right either post-training data or RL environments to basically fix that problem.
  15. 01:04And part of this is like how do you take something as fuzzy and large as design and break it down to a level that you can identify what is best solved through each method.
  16. 01:13What are elements of design that are almost like once you kind of boil down the problem becomes so specific that they almost become deterministic.
  17. 01:19So for example, uh if you're trying to train a model to be good at selecting color palettes or have contrast or alignment, those are things that if you define the problem and the context in a specific enough way,
  18. 01:29uh you you can get to an answer that's like pretty objective or that at least most experts would agree to.
  19. 01:33But maybe other things like uh aesthetics, you naturally will see this expert disagreement.
  20. 01:38And so then you want to lean on to things that are closer to to data.
  21. 01:41So anyway, we spend a lot of time thinking about all those problems.
  22. 01:44Uh but on the other side is also without even touching the model layer, right?
  23. 01:47How do we actually help agents and app layer companies produce better things?
  24. 01:52And there's a lot that goes into that, right?
  25. 01:53You have these different sets of problems at the application layer because you're using an off-the-shelf model that tends to collapse in terms uh uh of style tends to collapse to the mean.
  26. 02:01So, how do we force that creativity back to the system?
  27. 02:04How do we avoid these patterns of slop, which we'll talk about a lot today?
  28. 02:07Uh how do you understand like user preferences or brand preferences preference so that you can uh maintain endurance to that style?
  29. 02:15Uh so, there's lots of things that are actually need to be solved as context or judgment or verification at the app layer, which is why we kind of work across both.
  30. 02:25Maybe I'll start with more of a a philosophical question of like how how do you define something that is great?
  31. 02:30Like how do you define greatness?
  32. 02:32And for something like math, it's easier, right?
  33. 02:34Because there's kind of one objective answer, and uh great is the same as correct.
  34. 02:39But then for something like writing or design, it's much harder, right?
  35. 02:43Like how do you define what's like a great tweet or what's a great art piece or what's a great website?
  36. 02:49Um I don't know what's the last time that you interacted with a poem or walked into a coffee shop and for some reason it kind of like hit different and it felt
  37. 02:56very special.
  38. 02:58Uh but probably it's a combination of things that it it felt very unique.
  39. 03:01It felt almost a little different.
  40. 03:02It kind of called your attention.
  41. 03:03Uh it felt like there it was made with a lot of care and attention to to detail and craft, and it almost had the sense of of like authenticity.
  42. 03:11Um and I think that's a lot of what AI is missing today.
  43. 03:13It's like how do we take uh things that are not necessarily average, right?
  44. 03:16How do we produce things that are purposely like out of distribution?
  45. 03:19Um and slop is kind of the opposite of that, right?
  46. 03:22I think it is hard to define what is great sometimes, but I think it's pretty pretty easy to define what is slop in the sense that most people would agree.
  47. 03:28I think the sense of like repetition of kind of soullessness is something that all of us feel right now when using AI, and I think it's quite magical, by the way, that
  48. 03:36AI has gotten to a point that any human on the planet that is not even a designer, that is not an engineer, can click a button and suddenly make an entire PowerPoint or make a website or make a web app.
  49. 03:46That's pretty cool.
  50. 03:47But it comes with consequences, right?
  51. 03:49Uh it comes with consequences of suddenly now the cost of generation is basically going to zero.
  52. 03:54Uh and But the average person hasn't necessarily honed their taste.
  53. 03:58Like I does think about the amount of effort and work that a designer puts in throughout their life to like build up their taste, right?
  54. 04:04Like there's all this process of like getting exposed to many things and learning to like spot patterns and learning to develop a point of view and like kind of do things in a in a courageous way that maybe are a little bit against the norm.
  55. 04:15Learning what not to do and how to like have restraint and that's very hard.
  56. 04:19Like the average person doesn't necessarily have the the time nor the skills to go and develop taste in everything, let's say in design.
  57. 04:26And so um I think it would be a bad case scenario for us to just like be like, "Okay, the way to fix slop is for everyone to have taste."
  58. 04:32cuz I don't think that's necessarily realistic.
  59. 04:35Um I think how do we how can we understand this better so that we can make even for the average person the ability to create something great and to understand maybe their own taste
  60. 04:43um easy more more more easy.
  61. 04:45So that's that's a lot of what we're we're focusing on.
  62. 04:49Um So yeah, I mean this phenomenon of slop, by the way, is not new.
  63. 04:52Uh if you were in the internet uh as social media emerged, you probably saw a lot of slop before that.
  64. 04:58But I do think that AI has been this kind of like accelerating force, right?
  65. 05:01Of like being able to create things very easily uh with a click of a button and that like thoughtlessness around it.
  66. 05:05And there's kind of these three characteristics that I I would say repeat in slop.
  67. 05:10Uh so A, repetition.
  68. 05:11So you start seeing the same thing many many many times.
  69. 05:15Um the second is lack of fit, which I actually think is is very related.
  70. 05:18So fit is kind of this ability for something to feel correct for a specific context, right?
  71. 05:23For a specific moment in time, for a specific person.
  72. 05:25Uh but suddenly if you have repetition and let's say one person asks for a website for their pet shop and the other one asks for a website for their finance firm and somehow those designs converge and look the same.
  73. 05:37That's quite odd, right?
  74. 05:38Like if that was in if you were actually crafting that with care, that wouldn't you wouldn't converge necessarily on those things.
  75. 05:44And so this lack of fit and lack of understanding of context is actually huge problem that like leads to slop.
  76. 05:49Um and the third is maybe low intent, which is probably a mix of Yeah, you're going to have a bunch of people prompting really quickly and maybe just wanting to one shot something.
  77. 05:56But I think there's actually this like intent interpretation piece that's missing in the systems that we're building.
  78. 06:01Like how can you help your user, right?
  79. 06:02Like how can you help them better understand the intent that they have um so that you can add more color and add more context on onto what you're trying to create.
  80. 06:12Okay.
  81. 06:12And I I'm a big believer by the way that you in order to fix something first have to measure it and you first have to understand it.
  82. 06:19I think that's exactly why we're so focused on like how do we uh turn these domains into something a bit more verifiable so that we can attach a measure to it.
  83. 06:27So uh you'll you'll go on a little bit of a research journey with me here now, but we basically wanted to figure out can we measure slop?
  84. 06:34Like can we actually measure this quantitatively and spot this and what does that like look like?
  85. 06:41So we analyzed over 2 million websites from the past like 10 years kind of like way back machine style to try to understand all the trends across like design,
  86. 06:49how is the internet changing, uh how are how is like design changing over time?
  87. 06:53And two things were interesting.
  88. 06:55And we also met by the way then kind of synthetically generated a set of uh design websites so we could kind of like compare like how does human-made sites compare to AI-generated
  89. 07:06ones?
  90. 07:07And there were a few things that were interesting.
  91. 07:08So one was that you already kind of saw a a bit of like a collapse uh on the internet before even AI.
  92. 07:15So you saw kind of the internet becoming more homogeneous, using more similar color palettes, using more similar layouts, uh which is probably a function of more uh I would say this like kind of trends spreading more and more and more quickly, let's say.
  93. 07:27Uh but with AI I think you saw this repetition happening a lot more and being almost more like um identified kind of regardless of context.
  94. 07:34So even in completely different buckets you saw patterns that were very similar.
  95. 07:37So we we built this I I call this probes, but basically we uh we did two things.
  96. 07:42So we did this like pattern mining on all this data to understand like what are features that we can extract from all these sites.
  97. 07:47What are all these characteristics that we can make more objective, right?
  98. 07:50Colors, typography, layout, audience.
  99. 07:53Like, how can we like distill this down into things that become almost like uh structured?
  100. 07:57And then how do we uh train up these like probes?
  101. 07:59So, think of these as like baby classifiers.
  102. 08:01Like, how do we uh train the ability to spot this one characteristic?
  103. 08:05And for all these slop sites, we identify we started identifying like what are the probes that basically mean the site is very likely to be AI slop.
  104. 08:14Um and especially when you start combining them use and you see the frequency of multiple of these happening at once, it became very likely that you could actually like measure
  105. 08:23uh and predict slop.
  106. 08:24And we saw a a super high basically ability to do that prediction, which was really cool to see.
  107. 08:29This performed better, by the way, than like most LLM as a judge methods of like asking an LLM to like judge if that uh is like great human quality versus like AI-generated
  108. 08:37slop.
  109. 08:38Uh so, that was pretty cool to see.
  110. 08:39And I think kind of shows this pattern that we see in AI really being uh an actually quantitative thing that we can see in slop, uh which I find really cool.
  111. 08:49But, obviously, we don't want to stop there, right?
  112. 08:50We don't want to just measure slop.
  113. 08:52We want to also solve it.
  114. 08:53And so, um there's a few I I think I mentioned this before, but like the as the cost of production basically goes to zero, I think the thing that becomes
  115. 09:02expensive and matters more than ever is judgment.
  116. 09:05Um I don't even want to use the word taste here.
  117. 09:07Is judgment.
  118. 09:08I think it's this ability to discern what's right.
  119. 09:10It's this ability to break down a problem so that you can actually understand it and create solutions for it.
  120. 09:15And so, yes, there's the side of judgment that is human judgment that I actually think is more valuable than ever.
  121. 09:19But, there's also the side of like how do we build the right tools and systems to like fix pieces of this problem, right?
  122. 09:27So, yeah, how do we how do we fight slop, my my enemy?
  123. 09:30Um And, by the way, I think there's there's a lot of conversation going around how do you fight slop at the model layer?
  124. 09:37Like, how do we make models better?
  125. 09:39How do we make models have a higher bar?
  126. 09:40Which don't get me wrong, it has to be solved and we're working very hard to solve that, too.
  127. 09:44But I actually think this problem of inference time is equally, if not even more important.
  128. 09:49Because that's actually when you interact with the end user.
  129. 09:52And this kind of back and forth of how do you understand this context and intent happens at the moment of inference time.
  130. 09:57So, I don't think that we can ignore and just make models better and not solve this, otherwise slop will keep existing.
  131. 10:02Um so, maybe breaking down a few of those pieces and kind of um a few of the ways that we've thought about solving this or a few solutions that we built to solve this.
  132. 10:10But I think, for example, for something like repetition, one of the things that we're working on is I I've nicknamed it, I don't know if that's going to be the official name, but like the creativity API.
  133. 10:17How can we create a system that almost becomes an inspiration machine for your agent?
  134. 10:22So that it can produce something that's actually out of distribution instead of something that is in that same average and kind of mean that we're seeing happen with like the slop sites.
  135. 10:30Um so, this is one of the ways that practically, if we can intentionally produce something that's out of distribution, you can improve this like overall uh quality.
  136. 10:38And by the way, I I don't think that this can be something just like randomness.
  137. 10:43So, it's not just about like turning up the temperature of the model and and kind of fingers crossed hoping for the best.
  138. 10:47I think it's much more like how do we understand um even like what are rules or expectations in specific domains?
  139. 10:53Like let's say that you asked for a slide deck for for the pitch of your startup.
  140. 10:57Like what is a what does a good pitch deck look like?
  141. 11:00And then how do you almost like intentionally break rules uh to create things that are more creative, right?
  142. 11:05Because usually creativity isn't like randomness, isn't doing something that completely feels off for that situation.
  143. 11:11It's like you intentionally maybe diverge on a couple of things while maintaining kind of um adherence to to expectations of that category, let's say for for others.
  144. 11:20So, that's one of the things we're working on.
  145. 11:22The second one of this problem of fit, I think um it's interesting, but brands, as probably a lot of you who are designers know, takes so much effort to create great brands.
  146. 11:31Like great brands are the work of dozens of designers putting in a lot of like craft and thought and care.
  147. 11:39And so we've almost like already pre-done the work of defining what is great for that specific company and then we're not using it well.
  148. 11:44So this like brand endurance actually think is a huge problem and one of the things that can very more easily let's say like raise that bar of quality.
  149. 11:51So I'll I'll touch on an example on this one specifically.
  150. 11:54And then same with like intent and judgment I think the baby classifiers was a good example.
  151. 11:58Um like how it how we can actually like use this to be even become a gate for slop and and not let your agent ship slop.
  152. 12:06But so the brand API is the first product that we're releasing to to the public.
  153. 12:09This is already in in beta testing with a bunch of our design partners.
  154. 12:13And essentially what it does is it can take let's say a brand URL and extract this into like very specific components that are good for an agent to follow.
  155. 12:22So basically how do we turn something as fuzzy as a brand into something so structured that it becomes easy to for your agent to follow that but also for you to judge against it, right?
  156. 12:30Because I think the piece that we can't forget here is this judgment and verification.
  157. 12:35So yes, this goes and helps your agent to produce something better.
  158. 12:38But how can we also add a way for you to judge okay, is the agent actually staying on track?
  159. 12:43Is it actually performing well to adhere to this brand or how is it failing or where is it failing?
  160. 12:47So this is the first flow I would say that we we're seeing that is really helping to improve quality.
  161. 12:53And what's cool is of course we're talking here about an example of a brand that already exists.
  162. 12:56But let's say you have an agent or you have an app and the person that is using your app actually doesn't have a brand.
  163. 13:03Let's say they're an average consumer.
  164. 13:04Can we actually One of the things that we're creating is basically like a repository, like an index of brands of pre almost like pre-created brand system so that if they want something that feels dreamy,
  165. 13:14why not retrieve a dreamy brand system that already has been thought out to be cohesive instead of doing like a generative approach the moment of that might end up not so great or might end up again in those pillars of slop.
  166. 13:26And I want to show you a real example of this in action.
  167. 13:28So, um there's this company that I think is awesome called the General Intelligence Company of New York.
  168. 13:32They have a sick website, you guys should check it out.
  169. 13:34Um but basically, if you ask Cloud Design to create a slide deck uh in their branding, the the middle one is basically what it comes up with.
  170. 13:41So, the one on the left is is the original brand.
  171. 13:44Uh this is the kind of the the default.
  172. 13:46And if you kind of use this extraction actually in the process, it creates something that's way more high fidelity with the original.
  173. 13:52Um and that even like in the details, I would say like feels right.
  174. 13:56So, this is just to show an example of it in in action.
  175. 13:59Um but yeah, I think we I think all of us would agree that like human human taste and kind of the peak of human craft is always going to be like deeply valuable.
  176. 14:11And that right now, I think the challenge is we are almost even not earning the right to debate this like how can we have uh models like reach this like pinnacle of taste.
  177. 14:21I don't think it's about that at all.
  178. 14:22It's like how do we first just like raise the bar.
  179. 14:25Like the bar is currently, I would say, on the ground.
  180. 14:27And so, I think all of this work that we're putting into like how do we decompose a problem and how do we measure it is exactly so that we can at least like improve this bar
  181. 14:36um of quality.
  182. 14:36And I think we have to start with that.
  183. 14:41That's it.
  184. 14:41Uh thank you very much for for the time.
  185. 14:43Uh this is this is awesome.
  186. 14:46[applause]
  187. 15:03[music]