WEBVTT

NOTE Sentence-level transcript of https://www.youtube.com/watch?v=AhQpRalYlyg

NOTE One cue per sentence. Cue ids are the line anchors on /transcripts/AhQpRalYlyg.html. A cue ends where the next begins, or 2 s after its last word.

s1
00:00:01.309 --> 00:00:03.309
[music]

s2
00:00:13.880 --> 00:00:15.360
Okay, good morning folks.

s3
00:00:15.360 --> 00:00:17.200
Thank you for showing up.

s4
00:00:17.200 --> 00:00:17.760
I'm Nidhi.

s5
00:00:17.760 --> 00:00:20.320
I'm a product person at Google DeepMind.

s6
00:00:20.320 --> 00:00:24.200
And today I'll be talking about multimodal collaborative agents.

s7
00:00:24.200 --> 00:00:36.920
Typically, these are agents that work with fuzzy intent intent when the intent when it's not clear the user hasn't gotten the right keywords to specify and when they come in for a shopping intent, how can you have the agents

s8
00:00:36.920 --> 00:00:44.160
still guide the user towards their goal with high agency execution, proactive elicitation, and a lot of hand-holding.

s9
00:00:44.160 --> 00:00:57.120
The framework that we'll be discussing today are grounded in shopping or commerce because I wanted to show you a few patterns that are much easier to see in commerce, but they're quite applicable in other consumer verticals as well, finance, education,

s10
00:00:57.120 --> 00:00:59.200
or whatever you guys work on.

s11
00:00:59.200 --> 00:01:05.840
Um Yeah, so before we even start into the frameworks, I wanted to discuss about why this is an existing problem.

s12
00:01:05.840 --> 00:01:12.080
Currently, a lot of the agents that we have act more like a wrapper to the search bar.

s13
00:01:12.080 --> 00:01:19.840
They assume that the user has a well-defined intent, has the right keywords, already knows what they're looking for.

s14
00:01:19.840 --> 00:01:24.440
And so when they come in, they just have the right vocabulary to interact with the agent.

s15
00:01:24.440 --> 00:01:28.000
However, there is quite a huge articulation gap.

s16
00:01:28.000 --> 00:01:31.640
When the users come in, rarely they have their intent well formed.

s17
00:01:31.640 --> 00:01:37.760
Rather, they kind of have a fuzzy feeling or a vibe when they are kind of looking to shop.

s18
00:01:37.760 --> 00:01:43.120
So, the agent needs to play quite a huge role in hand-holding them.

s19
00:01:43.120 --> 00:01:50.200
They need to work with the user in first understanding their preferences or even eliciting these preferences more proactively.

s20
00:01:50.200 --> 00:02:03.400
And then helping the user kind of show them different possibilities of what could be possible that they can shop and then eventually move towards kind of recommendations that work with these constraints that they have in mind.

s21
00:02:03.400 --> 00:02:15.720
So, what we're going to be discussing is this kind of a flywheel or a loop that goes from very, very fuzzy intent all the way towards kind of achieving the user's goal.

s22
00:02:15.959 --> 00:02:18.880
So, this is how the shopping loop looks like right now.

s23
00:02:18.880 --> 00:02:25.240
The very first thing that the agent prepares for when the user comes in is what we call the discovery phase.

s24
00:02:25.240 --> 00:02:31.040
This is where the where the agent takes all the different contextual information that they have about the user.

s25
00:02:31.040 --> 00:02:35.880
This could be past conversations, what the user has specified in their query.

s26
00:02:35.880 --> 00:02:41.760
It could be information present in their personal context, or even the references that the user have provided.

s27
00:02:41.760 --> 00:02:50.320
And then comes up with a collaborative strategy, and we'll be going to the details of this, but it comes up with a collaborative strategy on what more does the agent need to elicitate

s28
00:02:50.320 --> 00:02:57.560
back from the user in order to help get the intent in a better shape, to get clarity on what exactly the user is looking for.

s29
00:02:57.560 --> 00:03:01.480
And then it moves to the second phase, which is what we call the research phase.

s30
00:03:01.480 --> 00:03:03.480
Which is again a two-step process.

s31
00:03:03.480 --> 00:03:08.760
First is where it learns what is the best way to elicitate this preference.

s32
00:03:08.760 --> 00:03:18.480
So, for example, a lot of times the text the text base elicitation may not be the best way to get the preferences from the user because sometimes a user don't know what they're looking for.

s33
00:03:18.480 --> 00:03:23.200
So, then how can the agent start being more creative in terms of elicitating these preferences?

s34
00:03:23.200 --> 00:03:30.640
Can they start using some kind of visual references or visual visual inspiration boards for grounding and speaking a common language with the user?

s35
00:03:30.640 --> 00:03:39.120
And then the second phase to this is also coming up going into the background and doing the heavy lifting for the user, so taking the burden away from the user to describing

s36
00:03:39.120 --> 00:03:51.000
what they want, and rather going into the background and doing all the comparisons, trade-offs, summarization of all the information that they're looking for, and coming back with the right set of potential options for the user.

s37
00:03:51.000 --> 00:04:03.680
And then the third phase, once um the user and the agent is ready to go into some kind of a is ready to go into the last phase, which is where um the agent is ready to give out a response.

s38
00:04:03.680 --> 00:04:09.360
This is where the agent needs to start adapting the response in a way that is most useful for the user.

s39
00:04:09.360 --> 00:04:13.480
So, typically a lot of systems fail here where they just give out a text-heavy response.

s40
00:04:13.480 --> 00:04:17.519
What the agent should rather be doing is adapting to the query that the user had in mind.

s41
00:04:17.519 --> 00:04:20.440
So, should it be using um some kind of a bulleted list?

s42
00:04:20.440 --> 00:04:22.240
Should it be using comparison tables?

s43
00:04:22.240 --> 00:04:24.800
Should it be using um visual boards?

s44
00:04:24.800 --> 00:04:37.919
So, this is where the there's adaptive response happening as well, where the uh agent starts to develop uh more more of a uh smarter uh presentation skills uh for the user to really find the answer that they're looking for.

s45
00:04:37.919 --> 00:04:43.400
So, we'll be diving into all these topics as I uh kind of work through the presentation.

s46
00:04:43.680 --> 00:04:46.200
So, first and foremost, we're looking at discovery.

s47
00:04:46.200 --> 00:04:52.840
Like I was mentioning, this is where uh the the the agent needs to remember what exactly matters.

s48
00:04:52.840 --> 00:04:56.840
Um you the agent starts to look at the context across a bunch of different signals.

s49
00:04:56.840 --> 00:04:58.720
It starts to look at past conversations.

s50
00:04:58.720 --> 00:05:03.080
It looks at some of the reference images or reference links that the user might have provided.

s51
00:05:03.080 --> 00:05:07.840
It also starts to look at personal context and starts to build out a working state.

s52
00:05:07.840 --> 00:05:15.320
So, if you look at the sample code that we have here, like there is there is a goal uh one of the let's let's work with this query where the user is trying to redo their

s53
00:05:15.320 --> 00:05:17.560
living room with a certain budget in mind.

s54
00:05:17.560 --> 00:05:23.280
And some of the um some of the things that the agent develops as part of the working state is the session history.

s55
00:05:23.280 --> 00:05:25.080
It has a user context.

s56
00:05:25.080 --> 00:05:30.560
It also kind of extracts out the hard constraint that the user might have provided in the query.

s57
00:05:30.560 --> 00:05:34.360
But, things start getting getting interesting when we come to the softer constraints.

s58
00:05:34.360 --> 00:05:43.280
So, this is where the user may not be able to describe what they're looking for and may might have provided like a reference image on, you know, an inspiration that they had in mind

s59
00:05:43.280 --> 00:05:50.120
or some kind of a layout design that they really liked and is trying to get to the agent in terms of uh this is what speaks more to me.

s60
00:05:50.120 --> 00:06:01.040
So, this is where the agent starts to be more proactive and pulls out some of the salient signals from the from the reference images and starts to develop a mental model of what the user might really be looking to get at.

s61
00:06:01.040 --> 00:06:09.440
So, here the user is uh so the agent has specified uh has identified that maybe the style that is working out well for them is of a certain kind.

s62
00:06:09.440 --> 00:06:20.640
It also uh is also is also starting to work towards a confidence score on like how confident it is in terms of pulling out some of this information present in the multimodal inputs provided.

s63
00:06:20.640 --> 00:06:31.240
And then the last thing uh that happens as part of developing this working state is also figuring out what are the variables that the agent needs to pull out almost in real time because these variables

s64
00:06:31.240 --> 00:06:34.919
uh this variables do affect how the results will be displayed back to the user.

s65
00:06:34.919 --> 00:06:44.280
This could be variables that need to be refreshed in real time like uh inventory because if if what you're providing back to the user is stale information, then it's kind of a moot point.

s66
00:06:44.280 --> 00:06:51.600
So, these are the uh these are the variables that you want to refresh in real time and is and becomes a part of the agent's working state.

s67
00:06:51.600 --> 00:07:05.640
Um few ways that you can evaluate this state uh the way we have developed our auto raters, we we do make sure that all facts are retained meaning that whatever was mentioned in the part of uh whatever was mentioned in part of the context

s68
00:07:05.640 --> 00:07:09.160
is properly represented in the working state.

s69
00:07:09.160 --> 00:07:19.919
We also capture if the confidence collaboration was within a certain error bound because uh the agent needs to be able to confidently pull out these signals from the input.

s70
00:07:19.919 --> 00:07:24.720
We also focus on getting out um the counterfactual sensitivity uh counterfactual sensitivity.

s71
00:07:24.720 --> 00:07:35.960
So, we do this by flipping some parts of the queries and making sure that when the query changes, the underlying constraints pulled out by the agent the those also change and the ones that are not relevant stay the same.

s72
00:07:35.960 --> 00:07:40.080
So, we kind of want to measure the sensitivity both ways.

s73
00:07:40.080 --> 00:07:49.400
And then the second part that happens in the discovery phase as well is once the agent knows what they have what they already know about the user what more do they need to know about the user.

s74
00:07:49.400 --> 00:07:55.000
So what we call this what we call this is kind of discovering the intent gap.

s75
00:07:55.000 --> 00:08:06.840
There's a lot of times that there are a lot of unknown variables before the agent can provide the best answer and amongst all these unknown variables the agent doesn't need to go and find all the unknown variables up front.

s76
00:08:06.840 --> 00:08:22.160
So what I mean by that is in this unknown in in our working example some of the unknown variables that the agent said they need to find out more about is maybe knowing what the room width would be for the user or even kind of working on improving the confidence for the style before they can recommend back

s77
00:08:22.160 --> 00:08:23.160
the results.

s78
00:08:23.160 --> 00:08:35.159
Now once these unknown variables are figured out the agent also needs to work on a collaborative strategy compare all the different moves possible and then figure out what is that one unknown variable that it should prioritize

s79
00:08:35.159 --> 00:08:38.159
such that it has the maximal information gain at that point.

s80
00:08:38.159 --> 00:08:54.320
So in this case I mean one could argue that maybe finding out the room width is the best next move for the agent because if the products that the the agent is recommending doesn't fit into the room then again it's a moot point and that is one variable that is going to meaningfully change the

s81
00:08:54.320 --> 00:09:02.320
the direction of the conversation and that's what the agent works on in this step kind of figuring out what's the best next thing to ask to the user and what's

s82
00:09:02.320 --> 00:09:05.040
why is that the best next thing to ask as well.

s83
00:09:05.040 --> 00:09:19.600
Once it has okay so I'll go ahead into the auto data section on like why how do we evaluate this collaborative strategy we first focus on making sure that the agent is able to identify all the different blockers that are needed to be answered before

s84
00:09:19.600 --> 00:09:22.560
the agent can come back with meaningful responses.

s85
00:09:22.560 --> 00:09:37.000
We also work towards making sure that the agent is um optimal in trying to get some of these responses so we also don't want to have the agent constantly going into the loop and continually continuously asking these questions so over asking is definitely something we flag.

s86
00:09:37.000 --> 00:09:39.440
We also work on question utility.

s87
00:09:39.440 --> 00:09:42.920
So, again, how optimally is the question being asked?

s88
00:09:42.920 --> 00:09:48.080
Is the question indeed useful to get the right or elicit the right preference from the user?

s89
00:09:48.080 --> 00:09:50.480
And many more.

s90
00:09:51.160 --> 00:09:52.280
Okay.

s91
00:09:52.280 --> 00:09:53.000
Moving on.

s92
00:09:53.000 --> 00:09:57.640
The The second part that I was mentioning is the multimodal elicitation.

s93
00:09:57.640 --> 00:10:01.040
This is where the research phase happens.

s94
00:10:01.040 --> 00:10:13.400
So, play along, but like let's say the the agent has asked about the room width, the user has provided the response, and then the agent goes to the next step, which is where the agent is asking about

s95
00:10:13.400 --> 00:10:17.760
the next constraint that it needs to know about is what was the style preference of the user.

s96
00:10:17.760 --> 00:10:31.120
Now, the first first thing that the agent needs to do is kind of form this temporary bridge between the constraint itself and how that maps back to the the product catalog and the ontology in your knowledge knowledge database.

s97
00:10:31.120 --> 00:10:38.520
And this is going to be important because when you start retrieving these products, you want to have a way to map these constraints back to your knowledge base.

s98
00:10:38.520 --> 00:10:48.320
So, this almost happens in real time where we map the known constraint or the constraint that the agent is exploring back to the knowledge database.

s99
00:10:48.320 --> 00:10:57.920
We also work towards um figuring out We also work on like the agent knowing what is the best way to get the response for this constraint as well.

s100
00:10:57.920 --> 00:11:07.920
So, in this case, the agent has decided that maybe since this is this kind of a subjective constraint, the textual the actual elicitation is not the best way to do this.

s101
00:11:07.920 --> 00:11:17.560
So, one of the ways that the agent thinks this could this could be elicited back from the user is using some kind of a visual preference board.

s102
00:11:17.560 --> 00:11:23.360
So, the agent then goes back to determining what is the best form of options to show to the user.

s103
00:11:23.360 --> 00:11:34.000
This could be a combination of figuring out from the existing constraints, the past conversation, and then the temporary mapping that you've created from the constraints back to your product ontology.

s104
00:11:34.000 --> 00:11:46.400
So, in this case, the the agent thinks that maybe coming up with a few styles that are most similar to what the user had provided as reference image could be a good way to start thinking or guiding the user towards a common language on what

s105
00:11:46.400 --> 00:11:49.720
could be uh something that the user is interested in.

s106
00:11:49.720 --> 00:11:59.320
And then the the then the agent also goes into kind of um observing the space of what kind of reactions the user is giving.

s107
00:11:59.320 --> 00:12:08.560
So, it could start looking at these micro signals of if there was a hover or a click in a certain direction and starts updating its its confidence model on what kind of signals

s108
00:12:08.560 --> 00:12:14.920
uh what kind of signals can be used to im- improve the confidence in like what could be the style preference for the user.

s109
00:12:14.920 --> 00:12:23.080
Um some of the auto raters we use here, we do look at how efficient the how efficient the agent is in discovering hidden preferences.

s110
00:12:23.080 --> 00:12:32.280
So, typically, we would use a user simulator, feed it with some constraints, and then we'll see how how efficient the agent was in kind of eliciting some of these constraints.

s111
00:12:32.280 --> 00:12:33.960
We also focus on turn efficiency.

s112
00:12:33.960 --> 00:12:42.040
So, how efficient was the user in how many turns did it take for the user to be able to to be able to uncover all these hidden preferences.

s113
00:12:42.040 --> 00:12:46.320
Ideally, we don't want the user to go into this loop and keep asking the same questions again and again and again.

s114
00:12:46.320 --> 00:12:56.600
Or also, we don't want to go into the loop of asking some um some questions which may not give you the best uh which may not give you the best information required to proceed the conversation.

s115
00:12:56.600 --> 00:12:59.040
And then we also look at forma- format selection accuracy.

s116
00:12:59.040 --> 00:13:04.360
So, we typically also look for if the agent is asking the right right question in the right format.

s117
00:13:04.360 --> 00:13:11.200
So, for example, if the if the question was something that was easily statable, the right format could be a textual elicitation.

s118
00:13:11.200 --> 00:13:22.560
But if if this was more of a fuzzy question where the user is clearly having an articulation gap and is not able to describe the preference, maybe the best way to do this is speak a common language and come up with some

s119
00:13:22.560 --> 00:13:28.640
visual anchor points for the user to say what speaks more to them.

s120
00:13:29.480 --> 00:13:39.720
Um Yeah, so then the last step in this process, once the So, to recap like basically the Now the agent knows exactly what they what the user is looking for.

s121
00:13:39.720 --> 00:13:49.320
The agent has identified a collaboration strategy, has figured out all the preferences for the user, and knows what is going to be most optimal in terms of um the different priorities for the user.

s122
00:13:49.320 --> 00:13:56.400
The last step in this process is to also use the model intelligence to figure out what is the best way to provide this response back to the user.

s123
00:13:56.400 --> 00:14:06.920
So, in this working example, we found out that the agent knows what the style preference is, has figured out that the user is looking to buy products under a certain budget,

s124
00:14:06.920 --> 00:14:21.200
and knows what are the different dimensions along which it needs to find uh these products for because through the conversation they figured out some of the different um different metadata information that is going to be relevant to surface when they're giving out this product information back to the user.

s125
00:14:21.200 --> 00:14:28.400
So, the one one important step that happens at this this stage is figuring out what's the best way to surface back this response.

s126
00:14:28.400 --> 00:14:39.600
So, for example, like if the user was looking for a particular policy or review information about a particular product, maybe the best way to do this is to go give out a summary or a bulleted list.

s127
00:14:39.600 --> 00:14:50.600
But if the if the user was looking more towards comparing two different products, maybe the best response is to give out a trade-off table or a comparison table across the different axis that the user cares about.

s128
00:14:50.600 --> 00:15:05.800
Um and in case that in our example where the user was looking more towards kind of style inspiration or ideas on how they can kind of redo their room, then maybe the best way is to give out some visual references and product inspiration photos on like how What are the different options and possibilities

s129
00:15:05.800 --> 00:15:07.440
uh for the user to uncover.

s130
00:15:07.440 --> 00:15:14.440
So, some other ways that we focus on evaluating uh this stage is focusing on the format accuracy.

s131
00:15:14.440 --> 00:15:22.400
So, we want to make sure that the response format is kind of optimal for the user query so that they they find out exactly what they're looking for.

s132
00:15:22.400 --> 00:15:33.040
Uh ideally the information that they're looking for should not be buried in the response, but should be easy for the user to spot so they can come into the next stage in the intent journey, which is to basically buy the product.

s133
00:15:33.040 --> 00:15:44.920
We also look at data fidelity, which is to make sure that the model is not hallucinating and it's really capturing the information in the correct format in the you know, the the information is just accurate and is captured across

s134
00:15:44.920 --> 00:15:45.920
the response.

s135
00:15:45.920 --> 00:15:48.040
We also look at user actionability.

s136
00:15:48.040 --> 00:15:59.480
So, this is ensuring that the response format is such that the user is very confident and and commits to the next action, which is like I said, the action to purchase the product.

s137
00:15:59.480 --> 00:16:06.600
So, just to recap so far, what we have is um you want to design the product.

s138
00:16:06.600 --> 00:16:11.680
Um you want to you want to design the product such that you're prepared to accept vibes, like I say.

s139
00:16:11.680 --> 00:16:14.440
So, that is to say that users will come with fuzzy intent.

s140
00:16:14.440 --> 00:16:16.600
Users will not have a well-defined goal.

s141
00:16:16.600 --> 00:16:21.520
So, you want to make sure that your system is able to work through queries that are not clean.

s142
00:16:21.520 --> 00:16:29.440
Um the second takeaway is you want to focus on showing and asking rather than always asking with textual textual with textual elicitations.

s143
00:16:29.440 --> 00:16:33.080
Again, like visuals and comparisons do reveal preferences much much faster.

s144
00:16:33.080 --> 00:16:37.520
It allows you the agent and the user to speak a common language.

s145
00:16:37.520 --> 00:16:39.800
Third one I would say is shape the answer.

s146
00:16:39.800 --> 00:16:47.520
So, do focus on making sure that the presentation format is ideal for the user being able to find the right information.

s147
00:16:47.520 --> 00:16:52.960
Um the way you have the model response structure is also very much part of the intelligence.

s148
00:16:52.960 --> 00:16:56.440
And then the last one is make sure that you grade the loop.

s149
00:16:56.440 --> 00:17:00.520
You have the right auto rater set up on every step of the process.

s150
00:17:00.520 --> 00:17:05.880
Um and honestly, developing these auto raters is a is is almost like an evolving system.

s151
00:17:05.880 --> 00:17:13.959
It starts very simple, but as and when you start the the system starts evolving, you want the auto raters to kind of gradually grow with your system and start

s152
00:17:13.959 --> 00:17:17.319
um yeah, it which just gradually grow with your system.

s153
00:17:17.319 --> 00:17:19.040
So, that's all I had.

s154
00:17:19.040 --> 00:17:26.839
Um I can take a few questions, but hopefully the learnings we shared are useful for whatever you folks are building.

s155
00:17:26.839 --> 00:17:28.839
Yeah.

s156
00:18:12.080 --> 00:18:12.680
Yeah, great question.

s157
00:18:12.680 --> 00:18:19.680
So, the question is about what should the how should the ontology be structured on the merchant side so that it's fair both for the agent and the merchants.

s158
00:18:19.680 --> 00:18:31.200
So, yes, we do take a lot of advantage on the domain expertise of the merchant on as to what they're trying to sell and like we do work towards creating So, you remember how I was mentioning about the bridge between

s159
00:18:31.200 --> 00:18:36.720
the constraints that the user might have specified and then what the agent kind of understands.

s160
00:18:36.720 --> 00:18:44.960
That is where we do expect a lot of intelligence to flow from the merchant side where the ontology on how that constraint could map to the different metadata that the agent has.

s161
00:18:44.960 --> 00:18:47.000
Sorry, the that the merchant has maps in.

s162
00:18:47.000 --> 00:18:57.240
So, yes, we do partner a lot and then there's also like the UCP stuff that we launched recently which allows all the merchants to kind of start speaking the common language with the agent as well.

s163
00:19:20.560 --> 00:19:30.840
Yeah, so ideally, I mean, honestly right now we focus on making sure all of this flows back to the agent and the agent makes the decisions because you want to build like a horizontal common layer across all the different merchants.

s164
00:19:30.840 --> 00:19:37.400
So, and also like it should be a seamless experience for the for the user who's interacting with our apps.

s165
00:19:37.400 --> 00:19:41.680
So, right now the response format is very much part of the agent's intelligence.

s166
00:19:41.680 --> 00:19:45.720
It is not something that the merchant gets to decide.

s167
00:19:45.720 --> 00:19:46.840
Yeah.

s168
00:19:46.840 --> 00:19:47.560
Yes.

s169
00:19:47.560 --> 00:19:53.160
I'm just curious your opinion what happens when your user is not a user anymore, it's an agent.

s170
00:19:53.160 --> 00:19:54.920
That's a great question.

s171
00:19:54.920 --> 00:20:02.120
We're in the early stages of building this out, but I think I mean there could be a case where um Yeah.

s172
00:20:02.120 --> 00:20:02.880
Yeah, yeah, exactly.

s173
00:20:02.880 --> 00:20:07.320
So, yeah, I would I would expect like an MCP to be the interface between the two for sure.

s174
00:20:07.320 --> 00:20:12.240
Honestly, we haven't gotten to a point where we have agents interacting with our agents just yet.

s175
00:20:12.240 --> 00:20:21.080
Also, like what we've realized at least from our user studies is users really like to be more involved in the process of choosing or even exploring the different possibilities.

s176
00:20:21.080 --> 00:20:29.520
So, during the upper funnel journeys where users is looking more towards discovery, inspiration, that is where they would rather be interacting with the system than with their agent.

s177
00:20:29.520 --> 00:20:40.960
I think where the agent typically comes in or even where what we've heard is like the was the lower end of the journey where they're just looking to compare or negotiate or compare prices across different merchants, but very much up there in the funnel,

s178
00:20:40.960 --> 00:20:43.520
it's the users who kind of interact more with our systems.

s179
00:20:43.520 --> 00:20:45.560
So, yeah.

s180
00:20:46.080 --> 00:20:51.640
Yeah, I think yeah, I can take questions outside, but thank you folks for coming.

s181
00:21:05.474 --> 00:21:07.474
[music]
