1
00:00:09,090 --> 00:00:14,690
What is up everyone? 
Hope you were jamming to that 

2
00:00:14,690 --> 00:00:18,210
music as much as I was. 
I can see the chat is already 

3
00:00:18,210 --> 00:00:22,330
getting off to a good start. 
I'm excited because we've got a 

4
00:00:22,330 --> 00:00:28,520
packed house today to talk all 
about good old state of agentic 

5
00:00:28,520 --> 00:00:31,480
retrieval. 
There's so much that we're going

6
00:00:31,480 --> 00:00:35,080
to get into and we've got so 
many incredible folks here 

7
00:00:35,400 --> 00:00:37,920
today. 
But I want to set the scene 

8
00:00:38,200 --> 00:00:41,720
before we bring out the guests 
of honor. 

9
00:00:41,960 --> 00:00:45,920
And I want to talk to you about 
why this session feels so 

10
00:00:45,920 --> 00:00:49,280
important. 
And it is probably why you 

11
00:00:49,320 --> 00:00:52,520
actually came. 
You know, there's so many 

12
00:00:52,520 --> 00:00:57,160
different challenges and changes
that we've had with retrieval 

13
00:00:57,400 --> 00:01:04,360
over the past two years. 
I have seen such an evolution. 

14
00:01:04,400 --> 00:01:08,960
And so I thought, why don't we 
get together some people who 

15
00:01:08,960 --> 00:01:12,920
have been knee deep in this 
whole retrieval game to talk to 

16
00:01:12,920 --> 00:01:16,560
us about what it was like, what 
did we have back in the day and 

17
00:01:16,560 --> 00:01:19,280
what do we have now and what are
ways that we can optimize it so 

18
00:01:19,280 --> 00:01:21,320
that we are finding success 
right now. 

19
00:01:21,960 --> 00:01:26,400
So without further ado, I'm 
going to bring on the team from 

20
00:01:26,400 --> 00:01:29,360
Quadrant. 
We've got the Devrel folks 

21
00:01:30,080 --> 00:01:35,280
coming out. 
First up is my man Dylan. 

22
00:01:36,400 --> 00:01:39,640
Where are you at it? 
Hey Dylan, how you doing? 

23
00:01:40,320 --> 00:01:47,120
We've got Jenny, we've got Eva, 
Neil and Andre. 

24
00:01:47,600 --> 00:01:51,640
I think this is the most people 
that we've ever had on a round 

25
00:01:51,640 --> 00:01:56,280
table. 
So let's just kick it off, start

26
00:01:56,280 --> 00:01:59,920
strong with some hard hitting 
questions. 

27
00:02:00,320 --> 00:02:06,320
I'm going to ask what has 
changed in the last two years? 

28
00:02:06,320 --> 00:02:10,320
I I've seen so much. 
What do we need to focus on? 

29
00:02:11,640 --> 00:02:14,400
Yeah, sure. 
Yeah, I can take that. 

30
00:02:14,600 --> 00:02:17,960
Hi everyone, Neil here. 
Yeah, thanks for joining the 

31
00:02:17,960 --> 00:02:19,480
session and thanks for the 
intro. 

32
00:02:19,480 --> 00:02:23,600
Demetrios. 
There's so much happening in 

33
00:02:23,640 --> 00:02:28,040
Retrieval and Surge. 
I, before I joined Quadrant, I 

34
00:02:28,040 --> 00:02:30,400
didn't fully appreciate the, the
field of search. 

35
00:02:30,520 --> 00:02:34,360
And we kind of just expect 
search to work. 

36
00:02:34,400 --> 00:02:38,480
And we definitely notice when 
search doesn't work well and we 

37
00:02:38,480 --> 00:02:44,280
notice it through just getting 
like terrible results and, or it

38
00:02:44,280 --> 00:02:45,560
takes a long time to get 
results. 

39
00:02:45,560 --> 00:02:48,000
We can't find what we're looking
for. 

40
00:02:48,000 --> 00:02:52,160
And the more data that we 
capture and from the real world,

41
00:02:52,160 --> 00:02:56,720
the more we need to get that 
information, get, get the data, 

42
00:02:56,720 --> 00:02:59,800
the right data we need, get the 
information out of our, the data

43
00:02:59,800 --> 00:03:01,280
we're storing. 
So what the back to the 

44
00:03:01,280 --> 00:03:04,720
question, what's changed? 
One of the biggest things has 

45
00:03:04,720 --> 00:03:10,880
changed is agents and the topic 
of today's round table agents, a

46
00:03:10,880 --> 00:03:16,040
lot of them don't really search 
well, I would say, in my 

47
00:03:16,040 --> 00:03:22,920
opinion, there's a lot of they, 
they kind of search like a, like

48
00:03:22,960 --> 00:03:26,240
a, a beginner would just search 
by just like kind of throwing a 

49
00:03:26,240 --> 00:03:31,080
lot of things into the, into the
hat and just like trying to like

50
00:03:31,080 --> 00:03:34,520
pull, pull, pull results. 
And they go through lots of 

51
00:03:34,520 --> 00:03:39,760
turns and they burn a lot of 
tokens and they try to find this

52
00:03:39,760 --> 00:03:42,520
context, but they're not using 
all the tools available for 

53
00:03:42,520 --> 00:03:44,400
search and this. 
And if you don't have the right 

54
00:03:44,400 --> 00:03:48,920
context in your AI, that's going
to really make your AI less 

55
00:03:48,920 --> 00:03:52,880
efficient, less effective. 
It's going to prevent you from 

56
00:03:52,880 --> 00:03:56,360
reaching your goals. 
So agents and, and one of the 

57
00:03:56,360 --> 00:03:59,040
things that agents are doing 
that's also interesting is like 

58
00:03:59,040 --> 00:04:02,920
a human might search like once 
or twice a minute or so. 

59
00:04:02,920 --> 00:04:05,080
Agents are searching thousands 
of times per minute. 

60
00:04:05,080 --> 00:04:07,000
So they're just going boom, 
boom, boom, boom, boom hitting 

61
00:04:07,000 --> 00:04:09,960
it. 
And they're not able to evaluate

62
00:04:10,280 --> 00:04:13,200
those search results. 
So same way human can, because 

63
00:04:14,400 --> 00:04:17,279
humans have a lot of more 
judgement that they can apply. 

64
00:04:17,279 --> 00:04:20,000
So yeah, I'll stop there. 
But those are those are some 

65
00:04:20,000 --> 00:04:23,120
things that have really changed.
Dude, you're preaching to the 

66
00:04:23,120 --> 00:04:27,600
choir on the different ways that
agents search and then the ways 

67
00:04:27,600 --> 00:04:32,040
that they interpret that data. 
And so I, I have some questions 

68
00:04:32,040 --> 00:04:36,680
here about like, hey, now that 
we've got agents and these 

69
00:04:36,680 --> 00:04:40,880
agents are figuring things out, 
they're kind of, we've all seen 

70
00:04:40,880 --> 00:04:42,760
it. 
They will stumble through 

71
00:04:42,760 --> 00:04:46,480
different problem statements 
where they'll try and figure it 

72
00:04:46,480 --> 00:04:51,800
out and they don't necessarily 
go about it the most efficient 

73
00:04:51,800 --> 00:04:56,760
way possible. 
I wonder if you all have found 

74
00:04:56,760 --> 00:05:02,400
tricks on nudging the agents to 
be like, Oh, yeah, by the way, 

75
00:05:02,520 --> 00:05:07,320
like there's obviously the go 
look in that file trick, but 

76
00:05:07,720 --> 00:05:10,600
there potentially could be 
plenty more that you've seen. 

77
00:05:10,600 --> 00:05:14,600
Because this at the end of the 
day, is a search and retrieval 

78
00:05:14,600 --> 00:05:17,640
problem. 
And the faster that you can get 

79
00:05:17,640 --> 00:05:21,280
those agents, that knowledge, 
the faster they can hopefully do

80
00:05:21,280 --> 00:05:26,320
what you've asked them to do. 
Yeah, I'm I'm actually going to 

81
00:05:26,320 --> 00:05:28,840
take jump in here again, but I'm
going to pass the Dylan on this.

82
00:05:28,840 --> 00:05:34,280
So Dylan, get ready. 
So just to set like a little 

83
00:05:34,280 --> 00:05:39,320
more context there is, you know,
it's not just the efficiency, 

84
00:05:39,320 --> 00:05:41,680
but it's the effectiveness of 
search like finding the most 

85
00:05:42,280 --> 00:05:44,600
relevant response, relevant 
results. 

86
00:05:45,120 --> 00:05:49,480
And I think that it goes into 
evaluation. 

87
00:05:49,480 --> 00:05:54,160
How do you evaluate that the 
results are good enough to use? 

88
00:05:54,400 --> 00:05:57,200
Should you try again? 
Should you change your query? 

89
00:05:57,680 --> 00:06:01,600
And if you should try again, 
when do you stop? 

90
00:06:01,600 --> 00:06:04,960
When like how how long does the 
agent go before it stops? 

91
00:06:05,440 --> 00:06:08,360
What tools can it use available 
to it? 

92
00:06:08,360 --> 00:06:14,920
And Dylan's actually done self 
correcting agent loops that he 

93
00:06:14,920 --> 00:06:16,840
can talk a little bit about. 
So I'm going to hand to you to 

94
00:06:16,840 --> 00:06:18,400
talk a little bit about what you
found. 

95
00:06:19,320 --> 00:06:20,720
Yeah, absolutely. 
Thank you. 

96
00:06:21,200 --> 00:06:25,000
So my first points that I'm 
going to talk about is that 

97
00:06:25,040 --> 00:06:29,160
Neil, you just mentioned 
evaluations and I feel like a 

98
00:06:29,160 --> 00:06:33,360
lot of people are familiar with 
like the general, you know, 

99
00:06:33,440 --> 00:06:37,800
evaluation frameworks and evals 
that you can run on an agent. 

100
00:06:38,240 --> 00:06:42,080
But also in the information 
retrieval space, we have a lot 

101
00:06:42,080 --> 00:06:46,520
of our own like evals and 
metrics like NDCJMR that are 

102
00:06:46,520 --> 00:06:51,480
like really relevant for, you 
know, search pipelines and that 

103
00:06:51,480 --> 00:06:56,560
that we use extensively. 
But going into your what, what 

104
00:06:56,560 --> 00:06:58,840
you mentioned about like self 
evaluating agents. 

105
00:06:59,680 --> 00:07:03,360
So, you know, there's a lot of 
techniques for for that. 

106
00:07:03,360 --> 00:07:07,720
You know, the most obvious one 
is to spin up like a small agent

107
00:07:07,720 --> 00:07:10,640
and say, hey, was that result 
relevance? 

108
00:07:10,640 --> 00:07:14,320
Did it contain, you know, all 
the information that the user 

109
00:07:14,320 --> 00:07:17,720
was looking for? 
But the problem here is that, 

110
00:07:17,720 --> 00:07:21,680
you know, you add another like 
LLM roundtrip so that it is 

111
00:07:21,680 --> 00:07:25,920
extra latency, extra cost, Ian, 
and it makes your pipeline 

112
00:07:25,960 --> 00:07:29,920
slower, more complex. 
And so something that we've been

113
00:07:29,920 --> 00:07:36,680
working on was computing 
statistical signals that, you 

114
00:07:36,680 --> 00:07:40,920
know, very cheap signals based 
on the, the results. 

115
00:07:41,440 --> 00:07:45,000
And so, so for example, you 
know, when when you retrieve, 

116
00:07:45,880 --> 00:07:48,920
you know, using like dense or 
sparse embeddings, basic 

117
00:07:48,920 --> 00:07:53,080
basically you get like a score 
as a result. 

118
00:07:53,080 --> 00:07:57,320
So for like dense embeddings, it
is between, you know, like 0 and

119
00:07:57,320 --> 00:07:59,440
one which one being like the 
most confidence. 

120
00:07:59,840 --> 00:08:03,760
And you know, you return like 
the top K results, so the top 

121
00:08:03,760 --> 00:08:08,480
five or the top ten. 
And basically you can do some 

122
00:08:08,480 --> 00:08:13,560
statistic based on like the 
spread of those scores or how 

123
00:08:14,160 --> 00:08:17,800
the difference between the the 
top one and the top second 

124
00:08:17,800 --> 00:08:21,640
results. 
And basically from here, try to 

125
00:08:21,640 --> 00:08:28,160
create a programmatic, you know,
signals that will that that will

126
00:08:28,160 --> 00:08:32,120
be able to tell you if the 
quality of your retrieval was 

127
00:08:32,120 --> 00:08:37,280
good or not without virtually 
any adding any compute to your 

128
00:08:37,280 --> 00:08:40,280
pipeline. 
And then once you you have that 

129
00:08:40,280 --> 00:08:43,679
that signals that signal that 
tells you, you know, the the 

130
00:08:43,679 --> 00:08:48,040
retrieval was good or not, then 
you can basically route your 

131
00:08:48,040 --> 00:08:50,800
retrieval agents to like 
different paths. 

132
00:08:51,320 --> 00:08:55,240
And you can choose, for example,
to do like more expensive 

133
00:08:55,240 --> 00:08:58,880
retrieval method, like a cross 
encoder or late interaction 

134
00:08:58,880 --> 00:09:01,920
model. 
Or if the result was not good 

135
00:09:01,920 --> 00:09:05,280
enough or not like not good at 
all at all. 

136
00:09:05,280 --> 00:09:08,480
It's like, you know, there was 
like a huge divergent in in 

137
00:09:08,480 --> 00:09:10,680
between like all the scores that
you retrieved. 

138
00:09:10,880 --> 00:09:13,720
This is when you can invoke an 
additional LLM. 

139
00:09:14,120 --> 00:09:19,480
So the in order testings, this 
was like really a good way to 

140
00:09:19,560 --> 00:09:22,520
tell apart like good from bad 
retrieval in a way that is 

141
00:09:22,520 --> 00:09:27,200
really cheap and allows you to 
spend more money only only when 

142
00:09:27,960 --> 00:09:34,120
it was necessary. 
And then we also have some 

143
00:09:34,160 --> 00:09:36,840
something on that topic on 
skills. 

144
00:09:36,840 --> 00:09:40,960
So I'll pass it over to Jenny. 
I think it's super funny. 

145
00:09:40,960 --> 00:09:44,920
We're working like, you know, 
actually a multi agent chain out

146
00:09:44,920 --> 00:09:48,120
there because we're passing 
tasks to each other and 

147
00:09:48,120 --> 00:09:51,400
Demetrius, you're the evaluator 
of how the pipeline is going, 

148
00:09:51,520 --> 00:09:55,840
right? 
So we're just emulating search 

149
00:09:55,840 --> 00:09:58,800
and informational retrieval from
our heads in the going. 

150
00:09:58,800 --> 00:10:01,760
And yeah, like the little 
addition as I think everybody 

151
00:10:01,760 --> 00:10:04,200
knows about the genetic skills 
also, right? 

152
00:10:05,720 --> 00:10:07,840
So they're also applicable in 
search. 

153
00:10:07,840 --> 00:10:13,760
The problem is agents need to 
know how to use search as humans

154
00:10:13,760 --> 00:10:16,480
actually sometimes also do it 
bad because they also don't know

155
00:10:16,480 --> 00:10:18,840
how to search properly for 
something, but they do it 

156
00:10:18,840 --> 00:10:21,200
differently. 
With agents. 

157
00:10:21,200 --> 00:10:24,520
There is a good solution of 
teach them how to approach it is

158
00:10:24,520 --> 00:10:29,720
to show them skills. 
For example, we designed some 

159
00:10:29,720 --> 00:10:34,440
bucket of search engineer skills
which we're trying to maintain 

160
00:10:34,440 --> 00:10:39,560
in quadrant which explained our 
agent which approaches our 

161
00:10:39,880 --> 00:10:44,560
vector search engine on what to 
use and how to combine it to get

162
00:10:44,560 --> 00:10:48,160
different outcomes. 
For example, if you have a 

163
00:10:48,160 --> 00:10:52,400
problem that your search doesn't
show for you as an agent, 

164
00:10:52,400 --> 00:10:55,320
relevant results at the top and 
you feel like there is something

165
00:10:55,320 --> 00:10:59,200
more to it, then there is a 
specific markdown part of the 

166
00:10:59,200 --> 00:11:02,360
skill which tells, Hey, you 
should do this, this, this with 

167
00:11:02,360 --> 00:11:07,760
our, with our APIs, I think we 
can share the link later. 

168
00:11:08,200 --> 00:11:11,280
Or I could try to share the 
screen, but I'm so afraid to 

169
00:11:11,280 --> 00:11:13,760
share. 
You know, it's the hardest 

170
00:11:13,760 --> 00:11:15,200
thing. 
That's why like I'm senior 

171
00:11:15,200 --> 00:11:18,480
developer relations, but I never
managed to share a screen 

172
00:11:18,480 --> 00:11:20,080
without sharing my private 
chats. 

173
00:11:20,080 --> 00:11:23,600
So let's see, worst case we're 
just dumping the link and 

174
00:11:23,600 --> 00:11:28,240
showing how the skill evolvement
for the search pipelines is also

175
00:11:28,240 --> 00:11:32,400
helping our agents to use our 
search infrastructure properly. 

176
00:11:34,000 --> 00:11:36,080
Classic. 
OK, I've got some follow up 

177
00:11:36,080 --> 00:11:40,040
questions and there are some 
awesome questions that are 

178
00:11:40,040 --> 00:11:44,640
coming through in the chat, but 
ever I wanted to give you some 

179
00:11:44,640 --> 00:11:47,920
time in case there's anything 
you wanted to tag on here. 

180
00:11:50,120 --> 00:11:52,800
Sure. 
So Jenny mentioned a little bit 

181
00:11:52,800 --> 00:11:57,680
of how we can teach agents how 
to know search. 

182
00:11:57,680 --> 00:12:03,160
So essentially when we added the
agents, we moved from a very 

183
00:12:03,160 --> 00:12:06,880
static system of RAG to 
something much more dynamic. 

184
00:12:06,960 --> 00:12:10,000
But actually the search 
primitives that are underlying 

185
00:12:10,000 --> 00:12:13,960
remain the same. 
So one research paper that came 

186
00:12:13,960 --> 00:12:17,760
quite recently last month that I
found really interesting on the 

187
00:12:17,760 --> 00:12:22,600
topic is called Sera, which is 
the Super intelligence retrieval

188
00:12:22,600 --> 00:12:26,200
agent. 
And the whole idea there is to 

189
00:12:27,480 --> 00:12:30,840
achieving super intelligence, 
which is when the multi round 

190
00:12:30,840 --> 00:12:34,520
search is compressed into a 
single corpus retrieval. 

191
00:12:34,520 --> 00:12:40,080
So we've essentially avoided the
latency that Dylan was talking 

192
00:12:40,080 --> 00:12:45,640
about the the latency lags. 
And what they've done that was 

193
00:12:45,920 --> 00:12:49,720
quite relevant to to our today's
topic is they figured out how to

194
00:12:49,840 --> 00:12:52,720
use the LLM to enrich the 
documents. 

195
00:12:52,720 --> 00:12:58,240
So have this iterative loop 
whenever the search vocabulary 

196
00:12:58,240 --> 00:13:00,400
was missing. 
So they looked at the queries 

197
00:13:00,400 --> 00:13:04,760
that were generated, evaluated 
what queries that were missing, 

198
00:13:04,760 --> 00:13:09,000
and essentially predicted 
omitted vocabulary and added to 

199
00:13:09,000 --> 00:13:13,400
it before they passed it on to 
more traditional search 

200
00:13:13,400 --> 00:13:17,760
paradigms like BM25. 
OK. 

201
00:13:17,760 --> 00:13:22,720
There's a lot to unpack with all
three of you saying this stuff 

202
00:13:22,720 --> 00:13:28,640
because basically, let me just 
say what I caught real fast and 

203
00:13:28,640 --> 00:13:31,640
then you can correct me if I'm 
wrong there. 

204
00:13:32,200 --> 00:13:36,560
Dylan, You're talking about 
different ways that you can use 

205
00:13:36,560 --> 00:13:41,120
statistics to leverage and get 
better results with the 

206
00:13:41,120 --> 00:13:45,800
retrieval and with the agents 
recognizing what to grab and 

207
00:13:45,800 --> 00:13:48,440
what to retrieve and what is 
used. 

208
00:13:53,360 --> 00:13:54,800
Tell me what? 
Yeah, tell me what? 

209
00:13:55,000 --> 00:13:57,720
Sorry, yeah. 
So basically it was like a cheap

210
00:13:57,720 --> 00:14:03,400
way to get early signs into bad 
or good retrieval without, you 

211
00:14:03,960 --> 00:14:07,920
know, adding any latency or cost
to your pipeline and then routes

212
00:14:08,360 --> 00:14:12,280
based on those signals. 
All right. 

213
00:14:12,280 --> 00:14:18,120
Then from there, Jenny, you 
mentioned using skills and how 

214
00:14:18,120 --> 00:14:21,160
you all at Quadrant have created
a bunch of skills. 

215
00:14:21,160 --> 00:14:22,760
Are those open source? 
I didn't catch that one. 

216
00:14:22,760 --> 00:14:25,640
That was the quick. 
Everything that we do, 

217
00:14:26,320 --> 00:14:29,200
everything ish that we do is 
open source. 

218
00:14:29,200 --> 00:14:34,600
So, yeah, obviously. 
Drop that into the chat here and

219
00:14:34,600 --> 00:14:40,160
I can relay it to the chat or 
put it in wherever, whatever 

220
00:14:40,160 --> 00:14:42,960
chat you want. 
But we definitely I want to see 

221
00:14:42,960 --> 00:14:46,800
all of those skills. 
And lastly, Eva, you were saying

222
00:14:46,800 --> 00:14:51,960
that there is a paper on the 
Super intelligent retrieval 

223
00:14:51,960 --> 00:14:58,000
agents and the way that they're 
able to leverage and I didn't 

224
00:14:58,120 --> 00:14:59,920
catch now. 
Now my memory is starting to 

225
00:14:59,920 --> 00:15:03,120
fade trying to hold all these 
things in my hand head. 

226
00:15:03,640 --> 00:15:07,000
The Super intelligent retrieval 
agents were doing what now? 

227
00:15:08,600 --> 00:15:11,680
So it's just a framework of how 
do we know that we've build an 

228
00:15:11,680 --> 00:15:16,240
agent that essentially does 
successful retrieval? 

229
00:15:17,240 --> 00:15:18,960
Incredible. 
OK, cool. 

230
00:15:18,960 --> 00:15:25,000
So before I hit some of these 
questions from the chat and and 

231
00:15:25,120 --> 00:15:29,560
I'm just like looking at the 
chat about big question that's 

232
00:15:29,560 --> 00:15:31,480
coming up is like, hey, what 
about ground truth? 

233
00:15:32,120 --> 00:15:36,680
How do we figure that out? 
What is cause to know if the 

234
00:15:36,680 --> 00:15:40,040
retrieval is correct, right? 
We also need to know what the 

235
00:15:40,760 --> 00:15:43,760
the actual thing is, the ground 
truth that we're looking for. 

236
00:15:44,240 --> 00:15:50,560
Does anyone have opinions there?
I see some things taken I. 

237
00:15:54,920 --> 00:16:01,160
Would say, I would say a proper 
bunch and each of us has like I 

238
00:16:01,160 --> 00:16:05,640
would say different sides of it.
And also about what ground truth

239
00:16:05,640 --> 00:16:12,680
is, I think our problem is that 
we got the agentic systems and 

240
00:16:12,680 --> 00:16:15,600
we got super excited and we're 
like, are they solving search 

241
00:16:15,600 --> 00:16:17,560
fully? 
But the problem is that 

242
00:16:17,560 --> 00:16:20,920
evaluations weren't fully solved
before with the classical 

243
00:16:20,920 --> 00:16:22,960
search. 
And now with agentic search, it 

244
00:16:22,960 --> 00:16:26,200
gets even more tricky because 
you're like, is the ground truth

245
00:16:26,200 --> 00:16:29,600
that the agent successfully 
fulfilled the whole task and 

246
00:16:29,600 --> 00:16:33,680
arrived to the right point? 
Or is there is a ground truth on

247
00:16:33,680 --> 00:16:38,280
the age point of the iteration 
of what it does in the process 

248
00:16:38,400 --> 00:16:40,200
and there are several approaches
to it. 

249
00:16:40,640 --> 00:16:44,920
There is ground truth in 
retrieval as probably people 

250
00:16:44,920 --> 00:16:50,320
know can be created from some 
golden data sets per SE. 

251
00:16:50,320 --> 00:16:55,160
So some from the data that you 
already have and you know that 

252
00:16:55,160 --> 00:16:59,640
this question should get this 
answer from your data set. 

253
00:17:00,040 --> 00:17:05,720
Usually in production it's not 
so easy to get it, but LLMS got 

254
00:17:05,720 --> 00:17:09,319
pretty good at helping to 
generate the synthetical data 

255
00:17:09,319 --> 00:17:13,599
sets with golden truth. 
The one tip and opinion that I 

256
00:17:13,599 --> 00:17:15,560
have as the person who 
previously worked in 

257
00:17:15,560 --> 00:17:19,760
crowdsourcing as a devil of 
crowdsourcing approach, so where

258
00:17:19,760 --> 00:17:23,280
the humans were gathering these 
data sets is that don't 

259
00:17:23,720 --> 00:17:28,000
immediately expect dumping your 
protectional data to LLM that it

260
00:17:28,000 --> 00:17:31,720
will create your ideal golden 
set evaluation. 

261
00:17:31,720 --> 00:17:35,680
And creating the golden truth 
still is the work where you need

262
00:17:35,680 --> 00:17:38,920
to be in the dialogue like 
you're in the dialogue 

263
00:17:38,920 --> 00:17:42,560
developing a cold project. 
So the relevance notion, the 

264
00:17:42,560 --> 00:17:46,800
ground truth notion still comes 
from you as a domain expert and 

265
00:17:46,800 --> 00:17:51,960
that can be like LLM could be 
used as a tool to develop the 

266
00:17:51,960 --> 00:17:55,560
ground truth data set that you 
can then inject in your 

267
00:17:55,560 --> 00:17:59,240
evaluations. 
But when it comes, I know I have

268
00:17:59,240 --> 00:18:02,960
too much to say 1 little thing, 
one little 11 little thing. 

269
00:18:03,880 --> 00:18:08,880
But when it comes to the whole 
thing as the whole pipeline to 

270
00:18:08,880 --> 00:18:14,120
be successful for you like 
solving the task, there are very

271
00:18:14,120 --> 00:18:18,560
different approaches on how to 
teach model to do the correct 

272
00:18:18,560 --> 00:18:22,800
search trajectories which goes a
little bit in more into 

273
00:18:22,800 --> 00:18:25,720
reinforcement learning domain. 
And I feel like there is a lot 

274
00:18:25,720 --> 00:18:29,240
of emerging there teaching 
agents to do the right decisions

275
00:18:29,240 --> 00:18:32,680
in search. 
And there are the ground truth 

276
00:18:32,760 --> 00:18:38,720
is the result being correct 
given the task, given the search

277
00:18:38,720 --> 00:18:42,120
task, the search input. 
Maybe your audience knows Ralph 

278
00:18:42,120 --> 00:18:44,680
root loops. 
So they're also to some extent 

279
00:18:44,680 --> 00:18:50,160
applicable to the search. 
But Dylan here is coming from a 

280
00:18:50,160 --> 00:18:54,800
rise and he was doing 
evaluations as as the bread and 

281
00:18:54,800 --> 00:19:00,920
butter so. 
So, Dylan, this dovetails nicely

282
00:19:00,920 --> 00:19:04,080
into one of the questions that's
in the chat. 

283
00:19:05,200 --> 00:19:10,520
Someone was asking about how 
teams are versioning and 

284
00:19:10,520 --> 00:19:13,040
evaluating retrieval pipelines 
in production. 

285
00:19:14,680 --> 00:19:19,160
Yeah, absolutely. 
So you know, something that I 

286
00:19:19,320 --> 00:19:24,480
want to preface for preface with
first is that we are a vector 

287
00:19:24,480 --> 00:19:27,960
search engine. 
So you know, we only provide 

288
00:19:27,960 --> 00:19:30,680
only one of the cogs in the 
machine. 

289
00:19:31,120 --> 00:19:34,160
And you know, this is really by 
purpose. 

290
00:19:34,160 --> 00:19:38,280
We don't want to be selling like
or providing like a retrieval 

291
00:19:38,280 --> 00:19:40,440
agent. 
We really want to stay very 

292
00:19:41,480 --> 00:19:45,360
close to the metal and provide 
like the, the best retrieval 

293
00:19:45,360 --> 00:19:48,880
engine possible. 
But when it comes to those kinds

294
00:19:48,880 --> 00:19:53,160
of like evaluations and 
versioning, I will usually 

295
00:19:53,160 --> 00:19:56,800
suggest to use like an 
evaluation framework or 

296
00:19:56,800 --> 00:19:59,400
platform. 
Of course, you know, I come from

297
00:19:59,400 --> 00:20:03,480
our IAIS, so I'm, I'm, I might, 
my opinion might be biased 

298
00:20:03,480 --> 00:20:09,520
there, but there's a bunch of 
free evaluation tooling that can

299
00:20:09,560 --> 00:20:13,160
help you do that that versioning
there. 

300
00:20:15,760 --> 00:20:19,840
Awesome. 
And I think Andre, you had some 

301
00:20:19,840 --> 00:20:23,480
other thoughts about this too. 
Where are you? 

302
00:20:23,840 --> 00:20:26,360
I got to change the view to see.
There he is. 

303
00:20:26,360 --> 00:20:28,400
There we go. 
Yeah, Andre's been doing. 

304
00:20:29,960 --> 00:20:34,640
Space on the screen. 
So I, I would like to say that 

305
00:20:34,840 --> 00:20:38,640
from my experience, what can 
also help if we jump back for 

306
00:20:38,640 --> 00:20:43,720
the, for the golden set is when 
you kinda have some kind of 

307
00:20:43,760 --> 00:20:49,000
assumptions based on the human, 
on some human manual labor work.

308
00:20:49,080 --> 00:20:54,120
What can be good for your golden
set so that the LLM has some 

309
00:20:54,120 --> 00:20:57,200
kind of criterias that it can 
evaluate about. 

310
00:20:57,840 --> 00:21:00,960
And I would like to say that 
it's like an iterative process 

311
00:21:00,960 --> 00:21:05,040
in which you try to improve your
LLM as a judge. 

312
00:21:05,040 --> 00:21:08,960
Yeah, time from time so that the
golden set that it is creating 

313
00:21:08,960 --> 00:21:11,520
and the evaluation in general 
becomes better. 

314
00:21:11,520 --> 00:21:16,080
So it's a long process. 
It's iterative process that just

315
00:21:16,320 --> 00:21:20,040
becomes better as you train it 
more, yeah. 

316
00:21:21,880 --> 00:21:25,760
Awesome. 
Well, we've got more questions 

317
00:21:25,760 --> 00:21:29,800
in the chat. 
I see the let me just let me 

318
00:21:29,800 --> 00:21:35,040
just grab this. 
So Eva missing search terms are 

319
00:21:35,040 --> 00:21:38,880
Grafton dynamically. 
How does that handle prompt 

320
00:21:38,880 --> 00:21:41,520
injection shenanigans, if at 
all? 

321
00:21:44,640 --> 00:21:49,320
This is a good question so in 
the paper and feel free to share

322
00:21:49,320 --> 00:21:52,160
the link in the chat. 
Hopefully you have it. 

323
00:21:52,680 --> 00:21:56,800
I I haven't found the 
cybersecurity aspect of it, but 

324
00:21:56,840 --> 00:22:02,080
I think this would be like a 
really nice segue to agentic 

325
00:22:02,080 --> 00:22:06,080
harnesses. 
And as you design your system, 

326
00:22:07,360 --> 00:22:10,320
retrieval will not solve 
everything. 

327
00:22:10,320 --> 00:22:14,920
You still need to add tools onto
your agent that would make it 

328
00:22:14,920 --> 00:22:17,520
more secure. 
But I think this is a really 

329
00:22:17,520 --> 00:22:21,800
important design principle to to
consider because prompt 

330
00:22:21,800 --> 00:22:26,960
injection is definitely a very 
serious serious threats. 

331
00:22:26,960 --> 00:22:32,600
And I think especially recently 
we had open claw and or miss an 

332
00:22:32,640 --> 00:22:35,800
agent that came out. 
And with open claw there was a 

333
00:22:35,800 --> 00:22:40,720
lot of hype, but also we saw 
when we did not properly sandbox

334
00:22:40,720 --> 00:22:44,120
it that it was able to just go 
completely wild and access all 

335
00:22:44,120 --> 00:22:48,760
of our data. 
So yeah, I think that is not 

336
00:22:48,760 --> 00:22:52,240
specifically mentioned in that 
paper, but consideration that 

337
00:22:52,800 --> 00:22:56,120
every single good agentic 
retrieval system should have, 

338
00:22:56,640 --> 00:23:01,280
yes. 
So let me continue with a few 

339
00:23:01,280 --> 00:23:05,000
more questions from the chat and
then we will go on to the next 

340
00:23:05,280 --> 00:23:10,600
topic that I had in my notes. 
There is my biggest pain point 

341
00:23:10,600 --> 00:23:15,720
from John is he's saying biggest
pain point I've had with rag 

342
00:23:15,720 --> 00:23:20,920
stacks is defining edge and 
entities dynamically with LLM 

343
00:23:21,320 --> 00:23:24,800
analysts. 
Can you speak to different 

344
00:23:24,800 --> 00:23:28,960
approaches to that problem? 
And I'll just throw that to the 

345
00:23:28,960 --> 00:23:31,520
crowd. 
Anybody have strong thoughts? 

346
00:23:35,520 --> 00:23:38,960
I agree because I've heard it a 
lot on the conferences, but I 

347
00:23:38,960 --> 00:23:44,440
believe it's more of the graph 
ontology construction problem. 

348
00:23:45,440 --> 00:23:50,840
And we as the vector search 
engine don't do the ontology in 

349
00:23:50,840 --> 00:23:54,240
construction. 
We usually combine because in 

350
00:23:54,240 --> 00:23:57,880
many, many domains it's a very 
good complementary thing. 

351
00:23:57,880 --> 00:24:01,200
For example, in memory of 
agents, it makes sense to mix 

352
00:24:01,200 --> 00:24:06,440
both graph and vector search. 
That's how many agentic memory 

353
00:24:06,440 --> 00:24:09,360
providers do, for example, 
Cogni. 

354
00:24:10,680 --> 00:24:15,800
But I've heard that list from 
practitioners on the conferences

355
00:24:15,800 --> 00:24:18,240
and don't quote on me. 
I would ask Neo 4 J guys on 

356
00:24:18,240 --> 00:24:24,880
that, that if you atomize the 
task of ontology instruction in 

357
00:24:24,880 --> 00:24:29,120
the sense you don't try to build
the whole ontology of the 

358
00:24:29,120 --> 00:24:35,080
domain, but you do it bit by bit
in the more atomized setting 

359
00:24:35,080 --> 00:24:39,080
that actually performs much 
better, like the outcome is much

360
00:24:39,080 --> 00:24:43,840
better. 
But I also know that Andrey 

361
00:24:43,840 --> 00:24:47,600
before worked with ontology. 
So maybe you have like a little 

362
00:24:47,600 --> 00:24:51,960
point of view on that. 
Yeah. 

363
00:24:52,040 --> 00:24:57,240
So can you repeat the question 
again so that I I have the 

364
00:24:57,240 --> 00:25:00,440
latest? 
The biggest pain point I've had 

365
00:25:00,440 --> 00:25:05,280
with Rag stacks is defining edge
and entities dynamically with 

366
00:25:05,280 --> 00:25:10,240
LLM and analysis. 
Do you have any approaches on 

367
00:25:10,240 --> 00:25:16,560
that? 
Oh, I would say that, yeah, the 

368
00:25:16,560 --> 00:25:21,280
problem is that these are kind 
as, as Jane, you already said, 

369
00:25:21,280 --> 00:25:25,320
some kind of ontologies that you
try to provide to the vector 

370
00:25:25,560 --> 00:25:30,480
search and to the semantics. 
So you you you kinda try to put 

371
00:25:30,480 --> 00:25:33,600
the explicit semantics into the 
vector semantics. 

372
00:25:33,640 --> 00:25:39,680
And I think these are laying 
down two different layers. 

373
00:25:39,680 --> 00:25:45,440
Yeah, 1 is application layer and
another one is, is more of 

374
00:25:45,600 --> 00:25:47,840
implementation right of the 
technology. 

375
00:25:48,200 --> 00:25:55,760
So I would say right now what 
you can do is basically just 

376
00:25:56,440 --> 00:26:00,360
like the the current state of 
art we can say is that there are

377
00:26:00,360 --> 00:26:08,680
lots of GitHub profiles and 
GitHub projects that are trying 

378
00:26:08,680 --> 00:26:15,600
to combine new 4G approach here 
with explicit entities and also 

379
00:26:15,600 --> 00:26:19,040
nodes with something that is 
vector based. 

380
00:26:19,520 --> 00:26:26,000
And there are some successes 
there because or what does the 

381
00:26:26,000 --> 00:26:28,400
graph tell you? 
It can tell you the 

382
00:26:28,400 --> 00:26:32,400
relationship, can tell you the 
dependencies and in some sense 

383
00:26:32,400 --> 00:26:35,400
it can give you additional 
context that semantics not 

384
00:26:35,400 --> 00:26:38,040
always can get, you know, 
because like you can have 

385
00:26:38,320 --> 00:26:45,160
something that is 2000 from 2001
and something from 2014 and you 

386
00:26:45,160 --> 00:26:51,320
want something specifically that
is from 2014 to be more aligned 

387
00:26:51,560 --> 00:26:54,600
on top. 
You can you can fix it of 

388
00:26:54,600 --> 00:27:00,080
course, with some custom, some 
custom scoring, some custom 

389
00:27:01,520 --> 00:27:04,600
relevance metrics that you 
provide by yourself. 

390
00:27:04,600 --> 00:27:09,280
Maybe again, it depends on the 
domain, but I would say there 

391
00:27:09,280 --> 00:27:13,880
are these approaches that they 
basically try to combine vector 

392
00:27:13,880 --> 00:27:16,560
and graph in one stack. 
Yeah. 

393
00:27:17,640 --> 00:27:23,320
Or you can try to encode in some
way information into the red. 

394
00:27:23,320 --> 00:27:27,400
But yeah, it's quite hard, yeah.
Awesome. 

395
00:27:27,680 --> 00:27:30,720
Well, let's keep it cruising 
because I want to get into 

396
00:27:30,720 --> 00:27:36,040
memory and it feels like that is
a perfect segue into memory and 

397
00:27:36,040 --> 00:27:41,200
how folks are approaching this, 
how you all have seen the best 

398
00:27:41,200 --> 00:27:44,840
in the business do it. 
And especially like I know there

399
00:27:44,840 --> 00:27:50,360
is a lot of talk about long term
memory and the architectures 

400
00:27:50,360 --> 00:27:53,600
that you have for those. 
And also if they are that 

401
00:27:53,600 --> 00:27:59,160
valuable because sometimes I 
will not necessarily want 

402
00:27:59,160 --> 00:28:01,440
something to be remembered, but 
it gets remembered. 

403
00:28:01,440 --> 00:28:06,040
And then for every single 
question that I'm asking my 

404
00:28:06,040 --> 00:28:09,200
agent, and I'll give you a 
concrete example of this, I 

405
00:28:09,200 --> 00:28:14,520
told, I told the chatbot that I 
was vegetarian. 

406
00:28:15,080 --> 00:28:18,720
And now for things that have 
nothing to do with food, it 

407
00:28:18,720 --> 00:28:21,920
says, well, given your vegan 
lifestyle. 

408
00:28:22,200 --> 00:28:25,360
And I'm like, I asked you a 
question about my taxes. 

409
00:28:25,680 --> 00:28:29,600
Get the fuck out of here. 
Why would you need to reference 

410
00:28:29,760 --> 00:28:32,080
my vegan lifestyle? 
So anyway, that's just a little 

411
00:28:32,080 --> 00:28:34,480
bit of a tangent. 
Maybe we can talk long term 

412
00:28:34,480 --> 00:28:37,400
memory or we can also. 
I'm going to throw different 

413
00:28:37,400 --> 00:28:39,680
things out there like episodic 
memory. 

414
00:28:39,800 --> 00:28:41,600
Do you want to go event based 
verse? 

415
00:28:41,600 --> 00:28:44,400
Semantic memory? 
How do you do the factual verse 

416
00:28:44,400 --> 00:28:47,280
conceptual? 
Who wants to take this one? 

417
00:28:47,560 --> 00:28:49,760
I'll throw it. 
Throw it up there and let 

418
00:28:49,760 --> 00:28:59,000
anybody rise to the challenge. 
Let me try to jump in because I 

419
00:28:59,000 --> 00:29:02,880
relate a lot to this part about 
the vegan taxes. 

420
00:29:03,400 --> 00:29:07,440
I think it's the classical, you 
know, LLMS, like please don't 

421
00:29:07,440 --> 00:29:09,800
think about elephants. 
I think Dylan taught me that. 

422
00:29:10,000 --> 00:29:12,840
And then immediately whatever 
you do, it's going to recall 

423
00:29:13,360 --> 00:29:15,880
this elephants. 
That's why it's also remember 

424
00:29:15,880 --> 00:29:19,080
when you do anything around 
memory is also to think about 

425
00:29:19,080 --> 00:29:22,400
forgetting as a conception. 
But it's a very hard one, an 

426
00:29:22,400 --> 00:29:24,480
interesting one. 
And there are like several tool 

427
00:29:24,480 --> 00:29:29,480
links on how to do that. 
So basically whatever I wanted 

428
00:29:29,480 --> 00:29:34,880
to say is very quick is that I 
think people out there probably 

429
00:29:34,880 --> 00:29:40,080
tried Hermes or open claw agents
because who didn't? 

430
00:29:40,080 --> 00:29:42,600
It was my first Wow effect, to 
be honest. 

431
00:29:43,640 --> 00:29:48,000
And we recently recently been 
like day before yesterday, 

432
00:29:48,000 --> 00:29:51,880
actually, it's, you know, hype 
driven development a little bit 

433
00:29:52,160 --> 00:29:55,880
released a plug in which allows 
to store episodic and semantic 

434
00:29:55,880 --> 00:29:59,240
memory of a Hermes agent in 
quadrant. 

435
00:30:00,120 --> 00:30:05,200
And I don't have a super sick 
demo or anything, but I chat 

436
00:30:05,200 --> 00:30:09,080
with my RMS agent a little bit 
and some of the memories of me 

437
00:30:09,080 --> 00:30:12,160
trying to prepare for this 
envelopes meetups got actually 

438
00:30:12,160 --> 00:30:16,240
started in my quadrant classer. 
I can show it but before also 

439
00:30:16,240 --> 00:30:20,800
like a quick note in case 
somebody just this not very well

440
00:30:20,800 --> 00:30:22,760
versed with the concept of 
different memories. 

441
00:30:23,040 --> 00:30:25,440
There is a working one which is 
the context window. 

442
00:30:25,440 --> 00:30:29,920
The chat you're in there is 
semantic or factual memory. 

443
00:30:30,080 --> 00:30:32,520
It's the one that something is 
true about you. 

444
00:30:32,520 --> 00:30:35,760
For example, that the Metris is 
vegan and likes to do taxes. 

445
00:30:36,800 --> 00:30:39,720
Do you? 
No, OK, doesn't like to do taxes

446
00:30:39,760 --> 00:30:44,720
then happens to all of us. 
Then there is a procedure, 

447
00:30:44,720 --> 00:30:47,600
procedural one. 
It's actually more like skills. 

448
00:30:47,600 --> 00:30:50,520
So what to do in order to 
achieve something. 

449
00:30:50,920 --> 00:30:53,640
And the fourth one would be 
episodic one. 

450
00:30:53,640 --> 00:30:56,280
This perfect thing for the 
vector search. 

451
00:30:56,280 --> 00:30:59,920
Actually, it's all the recalls 
of what you have discussed in 

452
00:30:59,920 --> 00:31:05,000
the past and that it could help 
your agent to understand how to 

453
00:31:05,000 --> 00:31:08,040
deal with the stuff better 
considering all its private 

454
00:31:08,800 --> 00:31:15,960
previous knowledge. 
So I wired my Army's agent to my

455
00:31:15,960 --> 00:31:20,560
cordon cluster specifically for 
that to show you. 

456
00:31:20,560 --> 00:31:23,640
Let me try to demonstrate it. 
It's going to be a very lame 

457
00:31:23,640 --> 00:31:27,520
demo, but we have a cooler 1 so 
don't get don't get discouraged.

458
00:31:29,240 --> 00:31:33,080
This is nice because people were
asking for demos in the chat, 

459
00:31:33,080 --> 00:31:38,320
and so I think I just asked in 
our background chat if anybody 

460
00:31:38,320 --> 00:31:41,680
else has one. 
Dylan, you have one too, huh? 

461
00:31:41,800 --> 00:31:44,160
That we can throw later on. 
All right, cool. 

462
00:31:44,480 --> 00:31:46,640
Yeah, I'll follow up with them 
as well. 

463
00:31:47,040 --> 00:31:50,600
OK, I'm I'm going to tell you, 
you're going to love Dylan's one

464
00:31:50,600 --> 00:31:53,560
and you're going to maybe like 
mine. 

465
00:31:53,880 --> 00:31:58,440
What am I currently showing? 
Is it my Slack or is it like UI 

466
00:31:58,440 --> 00:32:02,560
of collections in quadrant? 
Hold on, I got to check it out. 

467
00:32:03,160 --> 00:32:09,080
I see this no us so yeah, 
quadrant you're. 

468
00:32:09,080 --> 00:32:10,520
Good. 
I'm good. 

469
00:32:11,000 --> 00:32:14,320
OK. 
So as you see, this is a very 

470
00:32:14,320 --> 00:32:18,240
impressive memory of 37 points 
approximately in quadrant. 

471
00:32:18,520 --> 00:32:23,800
But basically what happens, all 
of the facts that I have in the 

472
00:32:23,800 --> 00:32:26,480
conversation with my Hermes 
agent, which I'm not going to 

473
00:32:26,480 --> 00:32:29,760
obviously share because I ask 
embarrassing questions all the 

474
00:32:29,760 --> 00:32:32,640
time, that's why I wipe the 
memory. 

475
00:32:32,880 --> 00:32:38,080
So all the facts and all the 
turns in the conversations are 

476
00:32:38,080 --> 00:32:42,480
saved with the different 
metadata of the sessions that 

477
00:32:42,480 --> 00:32:46,400
they happened in. 
And that will help agent recall 

478
00:32:46,400 --> 00:32:48,560
some similar information. 
For example. 

479
00:32:48,560 --> 00:32:53,200
I'm a big fan of this GE Chepa 
approach which Jan Dikun is 

480
00:32:53,200 --> 00:32:56,120
recently saying that it's the 
next big thing for embeddings. 

481
00:32:56,520 --> 00:32:59,440
So agent can find similar 
memories. 

482
00:32:59,440 --> 00:33:02,920
You can see that with dense 
vector search for example about 

483
00:33:02,920 --> 00:33:06,640
this method you will be able to 
find some other memories which 

484
00:33:06,640 --> 00:33:12,560
are kinda about the same fact. 
And this plug in you can also 

485
00:33:12,560 --> 00:33:18,840
visualize kinda the memories. 
There is not so much to see yet 

486
00:33:18,960 --> 00:33:21,240
because there is not too much 
memories. 

487
00:33:21,640 --> 00:33:25,240
But we can see that for example 
with the semantical embeddings, 

488
00:33:25,240 --> 00:33:29,200
there are similar memories about
me preparing some smart facts 

489
00:33:29,200 --> 00:33:33,440
about the Transformers attention
window will being very important

490
00:33:33,440 --> 00:33:36,920
for context windows being very 
important for vector search. 

491
00:33:36,920 --> 00:33:41,800
And that's why vector search now
won't be dead as much as RAC can

492
00:33:41,800 --> 00:33:47,680
become dead at some point. 
And if I would start chatting 

493
00:33:47,680 --> 00:33:53,200
now with my Hermes agent, for 
example, that I will add the 

494
00:33:53,200 --> 00:33:58,920
fact that well that Demetrius is
vegan. 

495
00:33:58,920 --> 00:34:02,760
So now all of my inputs are also
going to be polluted about. 

496
00:34:02,760 --> 00:34:04,720
This. 
No, it was the other one that 

497
00:34:04,720 --> 00:34:06,880
was true. 
I like doing taxes. 

498
00:34:07,200 --> 00:34:13,520
And Demetrius likes to do. 
I'm not showing you my telegram 

499
00:34:13,520 --> 00:34:16,400
where I'm chatting with Ernies, 
but that's what's happening. 

500
00:34:16,760 --> 00:34:21,800
So technically, if the double 
goals are nice to me, we should 

501
00:34:21,800 --> 00:34:25,000
see that the points are going to
get updated at some point. 

502
00:34:25,760 --> 00:34:28,920
But if they are not, the demo 
gods are just not nice to me and

503
00:34:28,920 --> 00:34:31,600
you will see the nicer demo. 
We can come back. 

504
00:34:31,960 --> 00:34:35,880
We can come back because I think
my RMS agent obviously decided 

505
00:34:35,880 --> 00:34:38,520
to sleep in the moment I decided
to demonstrate something. 

506
00:34:38,520 --> 00:34:43,320
But the general conception is 
that basically you have your 

507
00:34:43,320 --> 00:34:48,280
memory bank because information 
about you that you do now will 

508
00:34:48,280 --> 00:34:52,480
grow and grow and grow. 
And some point MD files won't be

509
00:34:52,480 --> 00:34:54,800
enough to recall all of the 
information. 

510
00:34:54,800 --> 00:34:58,080
And then you need some, well 
index which will organize this 

511
00:34:58,080 --> 00:35:02,840
memories, let them forget and 
let them be surfaced in the 

512
00:35:03,080 --> 00:35:05,520
right moments. 
And I think this is very 

513
00:35:05,520 --> 00:35:06,960
important. 
Where are we going? 

514
00:35:06,960 --> 00:35:09,480
Because the information, as Neil
said at the very beginning, 

515
00:35:09,480 --> 00:35:13,800
grows and grows and grows. 
Demo gods absolutely obliterated

516
00:35:13,800 --> 00:35:16,000
me. 
But I know that Dylan has a very

517
00:35:16,000 --> 00:35:19,720
cool 1 so everybody can forget 
this. 

518
00:35:20,960 --> 00:35:24,520
And oh, go ahead, sorry, we 
finished it. 

519
00:35:24,520 --> 00:35:28,480
I. 
Think we've got, well, I wanted 

520
00:35:28,480 --> 00:35:34,000
to talk for a minute about 
forgetting and that whole thing 

521
00:35:34,000 --> 00:35:35,680
because I know that can be very 
difficult. 

522
00:35:35,680 --> 00:35:39,160
And there's also a question in 
the chat that I want to bring up

523
00:35:39,800 --> 00:35:43,160
about like vector databases 
still being the preferred choice

524
00:35:43,160 --> 00:35:47,400
for long term memory. 
So maybe Eva, can you talk 

525
00:35:47,400 --> 00:35:49,760
forgetting real fast? 
And then if you have any 

526
00:35:49,760 --> 00:35:53,800
thoughts about vector databases,
I can imagine I know which way 

527
00:35:53,800 --> 00:35:54,800
you're going to follow on that 
one. 

528
00:35:55,440 --> 00:35:59,920
But let's let's go. 
Hopefully you won't forget that 

529
00:35:59,920 --> 00:36:04,480
one. 
So yeah, I mean, Jenny showed 

530
00:36:04,720 --> 00:36:09,040
what's like the benefit of of 
memory and dealing of all of 

531
00:36:09,040 --> 00:36:11,840
that. 
But I think memory is also 

532
00:36:12,160 --> 00:36:14,240
polluting a lot our context 
window. 

533
00:36:14,240 --> 00:36:18,160
So forgetting is actually a 
massive and really interesting 

534
00:36:18,200 --> 00:36:23,920
topic of how do we even decide. 
Being vegetarian might not be 

535
00:36:23,920 --> 00:36:27,920
relevant when you fill your 
taxes, but on your next session 

536
00:36:27,920 --> 00:36:31,880
you might be trying to optimize 
your diet and that information 

537
00:36:31,880 --> 00:36:37,440
will be very important. 
So this is the world of decay 

538
00:36:37,440 --> 00:36:39,840
functions and also relevance 
feedback. 

539
00:36:39,840 --> 00:36:46,720
So agents have this nice thing 
that you can iteratively tell 

540
00:36:46,720 --> 00:36:51,000
them what exactly in that 
particular search query is 

541
00:36:51,120 --> 00:36:55,880
relevant and you can boost that.
So this is the ability where I 

542
00:36:55,880 --> 00:36:59,000
think vector search actually 
really shines because you can 

543
00:36:59,760 --> 00:37:03,800
figure out for that particular 
session, what information should

544
00:37:03,800 --> 00:37:06,840
I include in my decay function 
and just forget. 

545
00:37:06,840 --> 00:37:09,920
And that can be based on 
temporary time stamp or that 

546
00:37:09,920 --> 00:37:13,520
could be filtered by keywords, 
etc. 

547
00:37:14,360 --> 00:37:18,680
And then there's another thing, 
which is some memories actually 

548
00:37:18,680 --> 00:37:22,320
overtime that you're putting 
into the sections into our 

549
00:37:22,320 --> 00:37:26,800
sessions, my duplicates. 
So that's another really 

550
00:37:26,800 --> 00:37:30,880
important thing is in order not 
to completely jam packed your 

551
00:37:30,880 --> 00:37:34,560
context window and actually get 
out of your engine agent what 

552
00:37:34,560 --> 00:37:38,560
you want, you can use vector 
search to de duplicate those 

553
00:37:38,560 --> 00:37:41,000
memories. 
So we have very clean context 

554
00:37:41,000 --> 00:37:44,240
window and next time you're 
filing your taxes, it's just 

555
00:37:44,240 --> 00:37:49,320
that and you can boost it and 
decay based on certain filters 

556
00:37:49,320 --> 00:37:52,080
and scores the information 
that's not relevant for the 

557
00:37:52,080 --> 00:37:56,920
session. 
That's awesome, Neil. 

558
00:37:56,920 --> 00:37:58,880
I feel like you had some things 
to say. 

559
00:37:59,120 --> 00:38:01,000
Yeah, I haven't. 
Talked to you in a while. 

560
00:38:01,400 --> 00:38:04,880
Yeah, I don't have anything 
insightful. 

561
00:38:04,880 --> 00:38:09,600
I just wanted to hand over to 
Dylan with a little bit of 

562
00:38:09,600 --> 00:38:12,360
context. 
We're talking about memory and 

563
00:38:12,360 --> 00:38:17,760
forgetting and search. 
And we, we weren't planning on 

564
00:38:17,800 --> 00:38:21,280
showing this demo on, but since 
there were requests for demos, 

565
00:38:21,280 --> 00:38:24,480
we'll go ahead and show it. 
This is a demo demo. 

566
00:38:24,480 --> 00:38:27,960
Dylan's gonna show on actually 
on device search. 

567
00:38:28,400 --> 00:38:30,920
And so when we think about 
memory and forgetting an AI, 

568
00:38:30,920 --> 00:38:35,440
like physical AI is becoming a 
really interesting area and how 

569
00:38:35,640 --> 00:38:39,560
agents and AI on physical 
devices need to be able to 

570
00:38:39,600 --> 00:38:44,640
remember certain contexts, be 
able to use that context, and 

571
00:38:44,640 --> 00:38:46,640
then how humans can interact 
with that. 

572
00:38:46,800 --> 00:38:50,280
There's also something really 
cool that Dylan can can show. 

573
00:38:50,280 --> 00:38:52,600
So yeah, Dylan handing over to 
you. 

574
00:38:53,400 --> 00:38:54,560
Absolutely. 
Thank you. 

575
00:38:54,600 --> 00:38:58,120
And yeah, so you know, we're 
kind of like steering away from 

576
00:38:58,120 --> 00:39:01,800
like the agency topic. 
But this is the good thing about

577
00:39:01,800 --> 00:39:05,560
vector search is that it's not 
only limited to to rag and 

578
00:39:05,560 --> 00:39:08,000
Agentic. 
And so we're today we're so I'm 

579
00:39:08,000 --> 00:39:11,400
going to show showcase your 
project that I built for a 

580
00:39:11,520 --> 00:39:16,160
robotic use case. 
So can you guys see my screen? 

581
00:39:16,160 --> 00:39:20,480
OK. 
Not yet Hold. 

582
00:39:20,560 --> 00:39:23,360
On Let me do my Job, there we 
go. 

583
00:39:23,600 --> 00:39:27,440
All right, perfect. 
So this is like just like the 

584
00:39:27,560 --> 00:39:31,240
the GitHub project. 
It is completely public and free

585
00:39:31,240 --> 00:39:33,560
and open source. 
So if you guys want you to run 

586
00:39:33,560 --> 00:39:36,680
this yourself, so today we're 
going to run it on like a 

587
00:39:36,680 --> 00:39:40,600
prerecorded video, but it works 
with like any, any kind of 

588
00:39:40,600 --> 00:39:42,720
input. 
You can connect your your camera

589
00:39:42,720 --> 00:39:46,480
and it will stop working live. 
Can you make it a little bigger?

590
00:39:47,120 --> 00:39:51,440
Yes, absolutely. 
There we go. 

591
00:39:51,440 --> 00:39:55,400
Now we're getting there. 
All right, so basically we're 

592
00:39:55,400 --> 00:39:59,280
starting with zero that agents 
knows nothing about the world. 

593
00:39:59,280 --> 00:40:03,600
It doesn't have like a database 
of like labels of what items 

594
00:40:03,600 --> 00:40:06,720
looks like, of what is a chair, 
what is the floor lamp, what is 

595
00:40:06,720 --> 00:40:09,360
a coffee table. 
And so we have like three kinds 

596
00:40:09,360 --> 00:40:12,760
of models that run in parallel. 
So it's all all running on 

597
00:40:12,760 --> 00:40:14,960
device, it's all running 
locally. 

598
00:40:16,040 --> 00:40:17,960
You know, if I turn off my 
network right now, you would 

599
00:40:17,960 --> 00:40:21,360
lose me, but I would not lose 
that product. 

600
00:40:21,720 --> 00:40:24,840
And so basically what what's 
happening here is that we have 

601
00:40:24,840 --> 00:40:29,280
first a an image recognition, an
object recognition model that is

602
00:40:29,280 --> 00:40:34,240
called yellow, and then a second
model that is an image to text 

603
00:40:34,240 --> 00:40:36,720
model. 
And this model basically creates

604
00:40:36,760 --> 00:40:41,600
a label and also a description 
for each item that it sees. 

605
00:40:42,480 --> 00:40:46,400
And then what we can do is that 
basically we can do semantic 

606
00:40:46,400 --> 00:40:52,280
search on every item that the 
robot has seen before. 

607
00:40:52,640 --> 00:40:56,000
And so, you know, it started 
with like a completely clean 

608
00:40:56,000 --> 00:40:59,960
memory when I, you know, when I 
opened up that that app, that 

609
00:40:59,960 --> 00:41:02,360
robot has never seen anything 
before. 

610
00:41:02,520 --> 00:41:06,920
And now it's able to recognize 
every single objects and 

611
00:41:06,920 --> 00:41:10,560
memorize them and then recall 
when and where is so those 

612
00:41:10,560 --> 00:41:12,840
objects. 
So, you know, the applications 

613
00:41:12,840 --> 00:41:17,240
for robotics are pretty pretty 
much endless. 

614
00:41:17,600 --> 00:41:22,480
And here basically you can so 
you know, this graph is a 2D 

615
00:41:22,480 --> 00:41:25,560
representation of the actual 
embedding space. 

616
00:41:25,880 --> 00:41:30,080
And so you can see the, the 
robot, the robot's brain being 

617
00:41:30,080 --> 00:41:34,560
built in real time and you can 
see all the memory of basically 

618
00:41:34,840 --> 00:41:37,120
all the concepts that are close 
together. 

619
00:41:37,120 --> 00:41:40,400
So you see like all the hallways
are grouped together here. 

620
00:41:40,400 --> 00:41:44,120
We can see like there's like the
dining table and chairs that is 

621
00:41:44,120 --> 00:41:47,200
grouped together here. 
You know, we can see that we, we

622
00:41:47,200 --> 00:41:51,400
went into the bathroom so that 
those memories are pretty far 

623
00:41:51,400 --> 00:41:56,200
away in the embedding space. 
And then you can really recall, 

624
00:41:56,640 --> 00:41:59,920
you know, everything that you 
has that you have seen before 

625
00:42:00,880 --> 00:42:05,040
and the the robot will be 
basically to able to recall 

626
00:42:05,040 --> 00:42:07,320
every single mirror that he has 
seen. 

627
00:42:08,680 --> 00:42:10,360
And yeah. 
And so you know, this is just 

628
00:42:10,360 --> 00:42:15,560
to, so you know, we're, we're 
not selling embedded models, 

629
00:42:15,560 --> 00:42:20,520
we're not selling vision models.
We're really just that memory 

630
00:42:20,520 --> 00:42:21,720
layer. 
We're just. 

631
00:42:21,920 --> 00:42:25,880
We have a model that creates 
embeddings based on those image 

632
00:42:25,920 --> 00:42:31,280
and text, and we allow for very 
quick search and recall on those

633
00:42:31,280 --> 00:42:34,520
memories. 
That's so awesome. 

634
00:42:36,520 --> 00:42:38,760
Whose house is that? 
Is that your house? 

635
00:42:39,720 --> 00:42:44,200
Oh, I wish. 
I was going to say that you 

636
00:42:44,200 --> 00:42:48,080
negotiated a pretty good salary 
if that is your house. 

637
00:42:48,560 --> 00:42:51,520
No, that that was just, you 
know, like on online tour from 

638
00:42:51,520 --> 00:42:53,640
like a leasing agent or a 
realtor. 

639
00:42:54,800 --> 00:42:56,960
Oh that's awesome. 
That is so cool. 

640
00:42:56,960 --> 00:43:01,200
The fact that it's on device too
with these the small models is 

641
00:43:01,400 --> 00:43:03,680
really impressive. 
Yep. 

642
00:43:03,680 --> 00:43:08,200
So, so you know, this was like a
way to demo like Quadrant Edge, 

643
00:43:08,200 --> 00:43:13,040
which is the embedded version of
quadrants and this version can 

644
00:43:13,040 --> 00:43:15,920
run on a Raspberry Pi. 
I'm not even like the top of the

645
00:43:15,920 --> 00:43:18,720
line US1. 
I have one form like five years 

646
00:43:18,720 --> 00:43:22,800
ago that runs it just fine. 
Awesome. 

647
00:43:23,440 --> 00:43:27,880
Now there was a question that I 
said in the chat I was going to 

648
00:43:28,040 --> 00:43:29,960
ask and I totally skipped over 
it. 

649
00:43:30,520 --> 00:43:34,640
I told folks so we're I want to 
get to it. 

650
00:43:34,640 --> 00:43:38,000
It's asking about vector 
databases still being the 

651
00:43:38,000 --> 00:43:42,520
preferred choice for long term 
memory or our knowledge graphs 

652
00:43:42,520 --> 00:43:45,800
and structured memory stores 
gaining traction. 

653
00:43:50,880 --> 00:43:57,760
I, I would say from what I'm 
saying, it's not an either orbit

654
00:43:57,760 --> 00:44:01,160
of both. 
We're, we're they, they both 

655
00:44:01,160 --> 00:44:05,440
have their uses in different 
contexts. 

656
00:44:05,440 --> 00:44:10,240
But you know, like dense, dense 
vectors, dense embeddings are 

657
00:44:10,240 --> 00:44:13,640
going to give you a lot of like 
semantic meaning, but like 

658
00:44:13,800 --> 00:44:17,200
vector stores are not really the
best for like relationships 

659
00:44:17,200 --> 00:44:20,720
between those different 
entities. 

660
00:44:21,080 --> 00:44:26,040
And so using like we, we have a 
great partnership with Neo 4J 

661
00:44:27,440 --> 00:44:30,240
and there's a video I put in the
private chat. 

662
00:44:30,240 --> 00:44:33,880
Maybe you can share Demetrios, 
but like graph rag and using 

663
00:44:33,880 --> 00:44:37,960
knowledge graphs with vector 
stores together can make a 

664
00:44:37,960 --> 00:44:40,560
really good substrate for memory
overall. 

665
00:44:40,880 --> 00:44:44,960
And just to kind of like shout 
out a couple other companies 

666
00:44:44,960 --> 00:44:48,520
doing this is like Cogney and 
Men 0. 

667
00:44:49,080 --> 00:44:56,160
Cogney has they use both like 
transactional data or well, they

668
00:44:56,160 --> 00:45:00,120
use like transactional data, 
they use vector data stores, 

669
00:45:00,120 --> 00:45:02,800
then they use knowledge graphs 
all together. 

670
00:45:02,800 --> 00:45:08,000
So you can check those out too. 
Nice. 

671
00:45:08,560 --> 00:45:15,200
A question came through about 
the demo and it was mainly about

672
00:45:15,200 --> 00:45:17,880
how detailed all of these 
captions are. 

673
00:45:17,920 --> 00:45:21,800
Like are you getting things down
to the level of types of 

674
00:45:21,800 --> 00:45:26,080
material and style or is it just
coffee table? 

675
00:45:29,280 --> 00:45:32,720
For Dylan. 
Sorry, can you repeat that 

676
00:45:32,720 --> 00:45:34,600
question? 
You were looking at the chat 

677
00:45:34,600 --> 00:45:36,280
huh? 
You were you were busy or you 

678
00:45:36,280 --> 00:45:39,160
were re watching the demo 
thinking damn this is a really 

679
00:45:39,160 --> 00:45:43,760
good product that I made How? 
How? 

680
00:45:44,200 --> 00:45:48,280
How do you know me so well? 
That's well played. 

681
00:45:48,560 --> 00:45:54,600
So, so basically how much detail
do you get from those different 

682
00:45:55,080 --> 00:45:58,480
captions or whatever it's being 
labeled? 

683
00:45:59,560 --> 00:46:04,280
Is it going down to the like hey
this is a postmodern coffee 

684
00:46:04,280 --> 00:46:08,200
table made out of glass or is it
just like coffee table? 

685
00:46:08,800 --> 00:46:12,600
Yeah. 
So it really depends on like you

686
00:46:12,600 --> 00:46:17,160
know, the image to text model 
that you you want to use. 

687
00:46:17,640 --> 00:46:22,720
So I use SIGLIP 2, if I remember
correctly, which is a very like,

688
00:46:23,360 --> 00:46:27,120
you know, basic like open source
model that can that can run on 

689
00:46:27,120 --> 00:46:30,920
your on your device. 
And so that model will usual. 

690
00:46:30,960 --> 00:46:36,480
And I did not try to ask for 
much in depth definition. 

691
00:46:36,840 --> 00:46:39,520
You know, I was just looking for
coffee table table here. 

692
00:46:39,760 --> 00:46:42,840
So maybe that model has the 
capability to have a more in 

693
00:46:42,840 --> 00:46:46,440
depth definition. 
But if not this one, I'm sure 

694
00:46:46,440 --> 00:46:50,520
that, you know, image to text 
models are very good these days.

695
00:46:50,760 --> 00:46:55,280
And if you want a really like 
detailed explain the issue of or

696
00:46:55,280 --> 00:46:58,680
like description of the object, 
I'm sure this is something that 

697
00:46:58,680 --> 00:47:02,000
can be done. 
Nice. 

698
00:47:02,400 --> 00:47:08,920
So we've got, I want to bring us
back a little bit to the topic 

699
00:47:09,280 --> 00:47:12,680
du jour, which is a genetic 
retrieval. 

700
00:47:13,000 --> 00:47:17,160
We were going hard on that and 
the chat was loving it. 

701
00:47:17,160 --> 00:47:20,600
And then we also, the chat was 
asking for some demos and we got

702
00:47:20,600 --> 00:47:25,200
some cool stuff. 
But there's a lot that we can 

703
00:47:25,200 --> 00:47:27,640
talk about still with agentic 
retrieval. 

704
00:47:28,240 --> 00:47:34,080
And so maybe we can veer in that
direction. 

705
00:47:34,160 --> 00:47:38,080
Neil, I feel like you have some 
things you wanted to say. 

706
00:47:39,720 --> 00:47:41,560
Yeah, sure. 
I can. 

707
00:47:41,920 --> 00:47:44,840
You know, we got like roughly 10
minutes left. 

708
00:47:44,840 --> 00:47:52,200
And so as we center back on 
agentic retrieval, some general 

709
00:47:52,200 --> 00:47:56,360
things that I think the audience
should, you know, some ground 

710
00:47:56,360 --> 00:47:58,960
setting, ground levelling things
that I think the audience should

711
00:47:58,960 --> 00:48:05,400
all be aware of with vector 
search is that even whether it's

712
00:48:05,440 --> 00:48:09,080
pertaining to retrieve a genetic
retrieval or just any retrieval,

713
00:48:09,400 --> 00:48:12,840
it's not only semantic search. 
And this goes back to the Neo 4 

714
00:48:12,840 --> 00:48:15,000
J and graph Rag kind of question
too. 

715
00:48:15,000 --> 00:48:21,000
But vector search is really 
vectors and embeddings are going

716
00:48:21,000 --> 00:48:23,840
to be the highest density format
in which you can store 

717
00:48:23,840 --> 00:48:27,280
information. 
You're taking in dense 

718
00:48:27,280 --> 00:48:29,640
embeddings, you're taking tons 
of unstructured data. 

719
00:48:29,640 --> 00:48:33,480
And yes, you're storing that 
into a fixed dimensional vector,

720
00:48:33,840 --> 00:48:35,600
but there's also sparse 
embeddings. 

721
00:48:35,600 --> 00:48:37,400
There's also, there's different 
embedding models. 

722
00:48:37,560 --> 00:48:42,040
All vectors are is a format and 
you as you use different 

723
00:48:42,040 --> 00:48:47,600
embedding models for images and 
videos and memory conversations 

724
00:48:47,600 --> 00:48:50,800
and all this like embeddings are
just a way, they're a vehicle 

725
00:48:50,800 --> 00:48:55,240
for this. 
And we can do keyword search and

726
00:48:55,240 --> 00:48:59,240
lexical search and we can do all
these different exotic things or

727
00:48:59,240 --> 00:49:03,360
maybe exotic strong words, just 
advanced skilled ways to find 

728
00:49:03,360 --> 00:49:08,160
the needles in the haystack. 
And that becomes really 

729
00:49:08,160 --> 00:49:12,760
important for agentic retrieval 
because just centering back on 

730
00:49:12,760 --> 00:49:18,600
agentic retrieval, you want to 
have access to information in 

731
00:49:18,600 --> 00:49:22,080
the most efficient way. 
And we, we have looked at file 

732
00:49:22,080 --> 00:49:24,920
search and other approaches and 
there are places for file 

733
00:49:24,920 --> 00:49:26,880
search. 
There is not necessarily you 

734
00:49:26,880 --> 00:49:28,840
must use vector search for 
everything. 

735
00:49:30,160 --> 00:49:32,920
You must set up an embedding 
model and you must set up your 

736
00:49:32,920 --> 00:49:34,200
chunky. 
You must do that for everything.

737
00:49:34,200 --> 00:49:39,480
It may not make sense in certain
scenarios, but when you as we 

738
00:49:39,480 --> 00:49:43,640
start collecting more data in 
our, our agents become more 

739
00:49:43,640 --> 00:49:48,320
sophisticated. 
Are their search tools become 

740
00:49:48,320 --> 00:49:51,840
better? 
I think that vector search has a

741
00:49:51,840 --> 00:49:54,680
really strong place there. 
And I also like intentionally 

742
00:49:54,680 --> 00:49:57,080
use that term vector search 
instead of vector database 

743
00:49:57,080 --> 00:50:04,120
because you want a system that 
optimizes for not just storing 

744
00:50:04,160 --> 00:50:09,520
and not just storing and having 
like a blind retrieval method to

745
00:50:09,600 --> 00:50:12,560
get your vectors out. 
You want to actually have 

746
00:50:12,560 --> 00:50:17,040
something that is focused on the
retrieval side of the equation, 

747
00:50:17,040 --> 00:50:20,040
the search side of the equation.
So like that's why I quadrant we

748
00:50:20,040 --> 00:50:21,680
consider ourselves a vector 
search engine. 

749
00:50:22,600 --> 00:50:24,960
But yeah, I'll stop kind of 
rambling there. 

750
00:50:24,960 --> 00:50:28,600
I don't know if anyone else 
wants to add thoughts to where 

751
00:50:28,600 --> 00:50:31,960
they see vector search fitting 
into agentic retrieval overall. 

752
00:50:35,200 --> 00:50:41,120
I can jump in on the scale. 
So Neil mentioned how like we're

753
00:50:41,120 --> 00:50:45,360
operating at millions of 
billions of of vectors scale and

754
00:50:45,360 --> 00:50:47,840
search will certainly get much, 
much bigger. 

755
00:50:48,240 --> 00:50:52,280
So early search primitives like 
the BM20 fives of of the world 

756
00:50:52,280 --> 00:50:57,480
are already 5 orders of 
magnitude smaller in terms of 

757
00:50:57,960 --> 00:51:01,880
memory that they consume on the 
machine with the latest one. 

758
00:51:01,880 --> 00:51:05,840
And another research that I 
wanted to share with you comes 

759
00:51:05,840 --> 00:51:09,040
from the group research group 
called said one. 

760
00:51:10,200 --> 00:51:13,520
They have a really interesting 
approach and they think that 

761
00:51:13,640 --> 00:51:18,040
reinforcement learning for 
identic retrieval is the way to 

762
00:51:18,040 --> 00:51:25,120
get to that next 5 order of 
magnitude efficient identic 

763
00:51:25,120 --> 00:51:27,440
retrieval. 
When the context window gets 

764
00:51:27,560 --> 00:51:30,320
bigger, when the number of 
vectors that we're have to parse

765
00:51:30,320 --> 00:51:32,480
through, it gets much, much 
bigger. 

766
00:51:32,960 --> 00:51:37,760
As I'll share the link so you 
can read into their approach and

767
00:51:37,760 --> 00:51:40,480
why reinforce reinforcement 
learning exactly. 

768
00:51:41,320 --> 00:51:43,720
Yeah, I know, Jenny, you had 
mentioned that earlier too 

769
00:51:43,720 --> 00:51:50,000
maybe, but can you talk a little
bit more to that because I like 

770
00:51:50,000 --> 00:51:52,840
this idea. 
We were talking about RL 2 for 

771
00:51:52,840 --> 00:51:58,360
evals and being able to create 
these environments, which are 

772
00:51:59,600 --> 00:52:03,440
super trendy these days, I guess
you could say. 

773
00:52:04,480 --> 00:52:07,120
And so you have these 
environments and then you can 

774
00:52:07,120 --> 00:52:10,600
eval them. 
But the RL for retrieval, what 

775
00:52:10,600 --> 00:52:13,120
is that even like? 
Break that down for me a little 

776
00:52:13,120 --> 00:52:14,200
more. 
Sure. 

777
00:52:14,200 --> 00:52:20,280
So what I got out of that paper 
is that the TLDR is basically 

778
00:52:20,280 --> 00:52:24,440
they're using a multi turn RL 
and it's a mixture of synthetic 

779
00:52:24,440 --> 00:52:27,720
and real questions. 
So Jenny mentioned before the 

780
00:52:27,800 --> 00:52:32,240
golden data set and the 
reinforcement learning approach 

781
00:52:32,240 --> 00:52:37,000
would be not that we have like 
human generated golden data set,

782
00:52:37,000 --> 00:52:42,720
but it is the environment where 
you learn, you interact with 

783
00:52:42,760 --> 00:52:47,640
over time with some real 
questions that resemble that 

784
00:52:47,640 --> 00:52:49,840
golden data set. 
And some of them are completely 

785
00:52:50,760 --> 00:52:54,240
AI generated. 
And there needs to be a reward 

786
00:52:54,240 --> 00:52:59,160
system designed in place that 
rewards every single time that 

787
00:52:59,160 --> 00:53:03,320
our agentic system hits the 
correct answer and punishes it 

788
00:53:03,320 --> 00:53:06,680
whenever the learning step was 
not made correctly. 

789
00:53:06,920 --> 00:53:11,960
So the paper outlines how you 
can design those, but I, I won't

790
00:53:11,960 --> 00:53:13,960
be diving too much in the detail
there. 

791
00:53:14,280 --> 00:53:20,680
So that's the, that's the high 
level overview of what exactly 

792
00:53:20,680 --> 00:53:25,480
they're doing there. 
Yeah, that is so cool to see and

793
00:53:25,600 --> 00:53:27,440
I feel like there's so much 
potential there. 

794
00:53:28,360 --> 00:53:32,640
The thing that I always wonder, 
and I also see that folks in the

795
00:53:32,640 --> 00:53:37,600
chat are asking this, it's more 
along the lines of like, hey, 

796
00:53:37,600 --> 00:53:41,920
I've got this agent. 
How do I either set up a 

797
00:53:41,920 --> 00:53:50,200
framework or encourage it to do 
different kinds of searches or 

798
00:53:50,200 --> 00:53:52,320
search techniques at different 
times? 

799
00:53:52,760 --> 00:53:56,760
So I want to do some kind of 
semantic search when there's 

800
00:53:56,800 --> 00:54:00,680
XYZ, like is there a framework? 
Is there a way that is it a 

801
00:54:00,680 --> 00:54:03,640
skill that we were talking about
earlier that you can say like, 

802
00:54:03,640 --> 00:54:07,200
hey, agent, here's whenever you 
need to retrieve something, 

803
00:54:07,200 --> 00:54:08,840
here's what you should go 
through. 

804
00:54:09,320 --> 00:54:12,280
And that feels like probably the
easiest. 

805
00:54:12,280 --> 00:54:14,480
When I see Jenny, you're shaking
your head. 

806
00:54:14,640 --> 00:54:17,920
You might have some thoughts. 
No, because I think it's 

807
00:54:18,160 --> 00:54:20,840
exactly, you answer the 
question, there are two ways, 

808
00:54:20,840 --> 00:54:23,880
right. 
There is like the way zero shot 

809
00:54:23,880 --> 00:54:28,160
way you teach it procedurally, 
which is like the set of skills,

810
00:54:28,160 --> 00:54:31,680
the set of search skills and it 
has its pros because it's easy 

811
00:54:31,680 --> 00:54:35,680
enough and you can adopt it 
easily enough because it's just 

812
00:54:35,680 --> 00:54:38,360
writing. 
I mean it's not just, but it's 

813
00:54:38,360 --> 00:54:41,200
composing the set of 
instructions and we kind of know

814
00:54:41,200 --> 00:54:44,680
how to approach that. 
And the second branch that we 

815
00:54:44,680 --> 00:54:49,360
are seeing now emerging is to 
actively teaching agents to 

816
00:54:49,360 --> 00:54:51,280
search through this 
reinforcement learning 

817
00:54:51,280 --> 00:54:56,240
environments. 
I don't see now easy ways of any

818
00:54:56,240 --> 00:55:01,040
practitioner to do that, but I 
feel like that might be the next

819
00:55:01,040 --> 00:55:03,240
thing in the future because 
everything becomes more 

820
00:55:03,240 --> 00:55:07,480
accessible now, right. 
So maybe we will get our Aral 

821
00:55:07,800 --> 00:55:10,400
gyms. 
Do you see my biceps? 

822
00:55:10,400 --> 00:55:14,240
Yeah, I'm saying the Aral Gym, 
but we will get them soon where 

823
00:55:14,240 --> 00:55:19,080
you can actually give the set of
what you want to achieve and the

824
00:55:19,080 --> 00:55:23,480
set of the search tools and your
problem and it will converge to 

825
00:55:23,480 --> 00:55:27,240
the like the small search agent 
will be able to converge to a 

826
00:55:27,240 --> 00:55:32,160
path which actually is, well 
what you want to see in your 

827
00:55:32,160 --> 00:55:36,160
specific pipelines. 
We're going to stay and watch 

828
00:55:36,160 --> 00:55:42,440
that because we represent kind 
of the default layer that humans

829
00:55:42,440 --> 00:55:45,680
or agents build upon. 
I wouldn't discard humans steal 

830
00:55:45,680 --> 00:55:47,960
from all of this picture, by the
way, because I'm human. 

831
00:55:49,040 --> 00:55:52,120
So I still, I still sometimes 
want to search stuff myself. 

832
00:55:52,400 --> 00:55:56,680
So I think it's just important 
to have the tooling which will 

833
00:55:57,000 --> 00:56:02,040
be usable by all the categories 
of the search users. 

834
00:56:02,400 --> 00:56:04,840
I think it was a tautology, but 
you get me I guess. 

835
00:56:05,280 --> 00:56:10,120
And I think what we were also 
trying to say that vector search

836
00:56:10,120 --> 00:56:13,720
is also so much more than just 
retrieving you a semantically 

837
00:56:13,720 --> 00:56:17,560
similar facts. 
I feel like we kind of fell into

838
00:56:17,560 --> 00:56:22,880
the strap of the rag cage where 
you just think about it as the 

839
00:56:22,880 --> 00:56:26,200
machine that spits out you the 
text chunk based on your 

840
00:56:26,200 --> 00:56:29,120
question. 
But I think there are other 

841
00:56:29,120 --> 00:56:34,360
emerging interesting parts where
it could be used for anomaly 

842
00:56:34,360 --> 00:56:38,400
detection, data analysis at 
scale, image to text search, 

843
00:56:38,400 --> 00:56:43,080
videos to whatever audios, so 
all of that stuff. 

844
00:56:43,080 --> 00:56:46,640
And at some point agents will be
also able to use it 3D. 

845
00:56:46,640 --> 00:56:49,560
And I'm really looking forward 
to see how they're going to 

846
00:56:49,560 --> 00:56:51,840
approach this part of vector 
search. 

847
00:56:52,920 --> 00:56:58,480
Well said. 
I think that is a perfect way to

848
00:56:58,480 --> 00:57:01,600
wrap it up. 
I know that there are still 

849
00:57:01,600 --> 00:57:04,880
folks in the chat that are 
stoked and asking questions. 

850
00:57:05,320 --> 00:57:10,280
I will just mentioned that 
everyone who is in the chat and 

851
00:57:10,280 --> 00:57:15,280
watching this right now, on your
left hand sidebar there is a 

852
00:57:16,760 --> 00:57:18,920
button that you can press that 
says match. 

853
00:57:19,280 --> 00:57:24,480
If you go there and more than 
two people do it, you will 

854
00:57:24,480 --> 00:57:28,360
randomly get put together with 
somebody else that is watching 

855
00:57:28,360 --> 00:57:31,120
this. 
So you can meet someone that is 

856
00:57:31,480 --> 00:57:38,200
just about as crazy on retrieval
as we are, and you can talk to 

857
00:57:38,200 --> 00:57:40,200
somebody that is watching this 
too. 

858
00:57:40,440 --> 00:57:43,520
So it's a great way to bond with
the rest of the community if you

859
00:57:43,520 --> 00:57:46,280
want to stick around for more 
time. 

860
00:57:46,480 --> 00:57:50,640
But for this session, I think we
are going to wrap. 

861
00:57:50,640 --> 00:57:54,360
I know that there are some great
places the Quadrant team hangs 

862
00:57:54,360 --> 00:57:57,480
out. 
You all have an awesome Discord,

863
00:57:57,720 --> 00:58:00,400
so I'll drop a link for that in 
the chat. 

864
00:58:00,400 --> 00:58:07,160
And then of course, if you 
weren't already bought into 

865
00:58:07,840 --> 00:58:11,200
everything that quadrants do in 
and following these good folks 

866
00:58:11,200 --> 00:58:15,320
that are here with us today, you
should definitely do that. 

867
00:58:15,320 --> 00:58:21,120
Go and give Quadrant a star if 
you haven't already on GitHub 

868
00:58:21,680 --> 00:58:27,160
and stay up to date with what 
all of their amazing Devrel is 

869
00:58:27,160 --> 00:58:29,520
doing. 
Follow them on LinkedIn and X 

870
00:58:29,520 --> 00:58:33,520
and all the places. 
So thanks everyone for this 

871
00:58:33,760 --> 00:58:37,400
excellent session. 
I will bid you farewell. 

872
00:58:37,440 --> 00:58:41,080
And for those that want to hang 
around, click the match button 

873
00:58:41,640 --> 00:58:45,320
on the left hand sidebar and 
I'll see you all later. 

874
00:58:46,160 --> 00:58:47,440
Thank you. 
Thanks everyone. 

875
00:58:47,760 --> 00:58:48,760
Thank you. 
Thanks to meet you. 

876
00:58:49,320 --> 00:58:49,400
Bye.
