1
00:00:00,040 --> 00:00:03,560
With agents becoming more 
powerful by the day, do you even

2
00:00:03,560 --> 00:00:06,920
need a vector database? 
On episode 79 of Tool Use, we're

3
00:00:06,920 --> 00:00:09,320
joined by Arjun Patel. 
He's the Senior Developer 

4
00:00:09,320 --> 00:00:11,920
Advocate at Pine Cone, which is 
a vector database for AI 

5
00:00:11,920 --> 00:00:14,640
applications. 
We're going to discuss what use 

6
00:00:14,640 --> 00:00:18,520
cases are best suited for vector
databases, strategies to improve

7
00:00:18,520 --> 00:00:21,800
the results from your vector 
database, and discuss the full 

8
00:00:21,800 --> 00:00:25,480
range from low level details up 
to no code implementation. 

9
00:00:25,880 --> 00:00:28,480
So please enjoy this 
conversation with Arjun Patel. 

10
00:00:28,560 --> 00:00:31,880
So a vector database is a 
specialized data store that 

11
00:00:31,880 --> 00:00:36,520
allows you to quickly index and 
query over information called 

12
00:00:36,520 --> 00:00:39,120
vectors. 
Vectors can be described many 

13
00:00:39,120 --> 00:00:41,120
different ways, but for our 
purposes we can just. 

14
00:00:41,120 --> 00:00:43,880
We can think of them as 
meaningful lists of numbers. 

15
00:00:43,920 --> 00:00:47,680
And it turns out that you can 
encode things that are 

16
00:00:47,680 --> 00:00:52,240
unstructured, like images, text,
audio, video, sometimes even 

17
00:00:52,240 --> 00:00:56,320
tabular or data or PDFs, into 
these vectors that correspond to

18
00:00:56,320 --> 00:00:57,760
meaning. 
And you can search over those 

19
00:00:57,760 --> 00:01:00,680
vectors really quickly by using 
other ones. 

20
00:01:01,000 --> 00:01:03,920
So the trick is, if you have 
hundreds of thousands of these 

21
00:01:03,920 --> 00:01:06,520
vectors, millions, 10s of 
millions, how can you do that 

22
00:01:06,520 --> 00:01:08,760
quickly and at scale? 
That's what a vector database is

23
00:01:08,760 --> 00:01:10,760
for. 
And could you give some concrete

24
00:01:10,760 --> 00:01:13,480
examples what gets unlocked when
we use a vector database? 

25
00:01:13,480 --> 00:01:15,360
What is possible that wasn't 
possible prior? 

26
00:01:15,400 --> 00:01:17,440
Yeah. 
So any any operation that 

27
00:01:17,440 --> 00:01:20,160
requires searching over 
unstructured data at scale 

28
00:01:20,160 --> 00:01:23,000
becomes possible with a vector 
database, provided that you can 

29
00:01:23,000 --> 00:01:24,600
convert those things into 
vectors. 

30
00:01:24,920 --> 00:01:29,240
So for example, if you wanted to
search across hundreds of 

31
00:01:29,240 --> 00:01:32,680
thousands of images with other 
images to find the most similar 

32
00:01:32,680 --> 00:01:35,480
ones, that operation is possible
with the vector database and 

33
00:01:35,480 --> 00:01:37,440
embedding model. 
But let's say you're embedding 

34
00:01:37,440 --> 00:01:41,480
model also handles text input at
the same time. 

35
00:01:41,680 --> 00:01:44,720
Now you've unlocked the ability 
to do a sort of multimodal 

36
00:01:44,720 --> 00:01:47,400
search where you can pass in 
text queries and get back images

37
00:01:47,520 --> 00:01:50,600
really quickly in that scale. 
And because the vector database 

38
00:01:50,600 --> 00:01:53,560
has this nice property of being 
able to index this data very 

39
00:01:53,560 --> 00:01:56,160
efficiently, the response times 
you get are going to be very, 

40
00:01:56,160 --> 00:01:59,400
very quick, like under seconds, 
under milliseconds, right? 

41
00:01:59,640 --> 00:02:02,520
So you're able to do these types
of fast operations much more 

42
00:02:02,520 --> 00:02:04,560
quickly than you would if you 
did a brute first search. 

43
00:02:04,720 --> 00:02:07,760
And just specifically around 
Pine Cone, I've heard about PG 

44
00:02:07,800 --> 00:02:10,520
vector and I think even SQL Lite
has embeddings. 

45
00:02:10,800 --> 00:02:13,040
What what are the ranges in a 
vector database? 

46
00:02:13,040 --> 00:02:15,040
Like what's the what makes 
something premier versus 

47
00:02:15,040 --> 00:02:16,320
something that's more entry 
level? 

48
00:02:16,400 --> 00:02:18,800
Absolutely. 
So Pine Cone is a great option 

49
00:02:18,800 --> 00:02:21,800
for anybody who's trying to do 
really high throughput 

50
00:02:21,800 --> 00:02:24,840
operations or really high scale.
But it's also wonderful for the 

51
00:02:24,840 --> 00:02:27,360
starting developer, right? 
All you have to do with Pine 

52
00:02:27,360 --> 00:02:29,920
Cone, we manage everything for 
you, so you can just interact 

53
00:02:29,920 --> 00:02:33,320
with us through the APIs. 
Nice thing about Pine Cone is 

54
00:02:33,320 --> 00:02:35,440
that it's very specialized in 
doing this type of vector 

55
00:02:35,440 --> 00:02:38,200
search. 
So you could use other services 

56
00:02:38,200 --> 00:02:39,880
that offer vector search 
capabilities. 

57
00:02:39,880 --> 00:02:42,520
But if you're really wanting 
something that's performing and 

58
00:02:42,520 --> 00:02:46,360
fast and focused on being able 
to return these items quickly, 

59
00:02:46,360 --> 00:02:48,680
you want to reach for Pine Cone,
and that's where we specialize 

60
00:02:48,680 --> 00:02:54,080
and work well at. 
Is there any prep work or or or 

61
00:02:54,080 --> 00:02:56,160
things like that that people 
need to do before they go to a 

62
00:02:56,160 --> 00:02:58,160
vector database? 
You need to structure their data

63
00:02:58,160 --> 00:03:01,600
certain way or prepare to to 
query the database differently. 

64
00:03:01,760 --> 00:03:04,080
What's the the prerequisites for
this? 

65
00:03:04,360 --> 00:03:06,800
Absolutely. 
So traditionally, and a lot of 

66
00:03:06,800 --> 00:03:08,720
this will be in the context of 
text data because when you start

67
00:03:08,720 --> 00:03:10,840
thinking about other modalities,
things get really hairy really 

68
00:03:10,840 --> 00:03:14,200
quickly. 
But for text data, the main the 

69
00:03:14,240 --> 00:03:16,760
thing that you usually need to 
do is turn it into vectors 

70
00:03:16,760 --> 00:03:18,800
before they can get indexed into
a vector database. 

71
00:03:18,960 --> 00:03:22,320
With Pine Cone, we have our own 
embedding models you can use 

72
00:03:22,960 --> 00:03:24,640
through something we call Pine 
Cone Inference. 

73
00:03:24,640 --> 00:03:27,520
So you can interact with the 
database entirely in text. 

74
00:03:27,800 --> 00:03:31,240
But the second order issue 
becomes how much text are you 

75
00:03:31,240 --> 00:03:34,600
allowed to represent in a fixed 
size vector, and how can you 

76
00:03:34,600 --> 00:03:38,600
search across those text vectors
in order to get the best, most 

77
00:03:38,600 --> 00:03:40,560
appropriate response? 
So most people have to do 

78
00:03:40,560 --> 00:03:44,240
something called chunking, which
means you take your input text 

79
00:03:44,240 --> 00:03:47,640
and you break it up into the 
literal chunks that are big 

80
00:03:47,640 --> 00:03:51,040
enough to fit in size the 
context window of the embedding 

81
00:03:51,040 --> 00:03:53,800
model that you're using. 
That's usually the extent of the

82
00:03:53,800 --> 00:03:57,240
pre processing you have to do 
for a simple chunk based search.

83
00:03:57,480 --> 00:04:01,480
So you can imagine if we had a 
bunch of sentences from the 

84
00:04:01,480 --> 00:04:04,520
titles of product reviews that 
we were embedding inside a 

85
00:04:04,520 --> 00:04:07,600
database and you wanted to do 
some sort of similarity search 

86
00:04:07,600 --> 00:04:11,760
where you're looking for certain
product qualities that satisfy 

87
00:04:11,760 --> 00:04:14,160
some conditions. 
You could pass in that query 

88
00:04:14,200 --> 00:04:17,880
embedded, embed all the titles 
of all the product reviews and 

89
00:04:17,880 --> 00:04:20,320
have the most similar ones 
returned inside the vector 

90
00:04:20,320 --> 00:04:22,640
database and you'd be chunking 
at the sentence level. 

91
00:04:22,800 --> 00:04:25,000
But if you had a Wikipedia 
article, you might want to chunk

92
00:04:25,000 --> 00:04:27,680
at the paragraph level because 
each paragraph encodes some sort

93
00:04:27,680 --> 00:04:30,720
of meeting. 
However, at Pine Cone, we have a

94
00:04:30,720 --> 00:04:32,920
product called Pine Cone 
Assistant now where you can just

95
00:04:32,920 --> 00:04:35,000
throw your PD FS in and it'll do
the chunking for you. 

96
00:04:35,000 --> 00:04:37,600
So there's lots of different 
options to use the underlying 

97
00:04:37,600 --> 00:04:40,280
vector database, but 
traditionally you you got to do 

98
00:04:40,280 --> 00:04:42,520
some chunking. 
I've heard different techniques 

99
00:04:42,520 --> 00:04:46,080
around chunking strategies. 
We had Kirk Marple on a while 

100
00:04:46,080 --> 00:04:48,480
ago and he talked about wanting 
your chunks to have a little bit

101
00:04:48,480 --> 00:04:50,840
of overlap with each other just 
to make sure you're not doing a 

102
00:04:50,840 --> 00:04:53,480
hard stop to to kind of bucket 
the context. 

103
00:04:53,800 --> 00:04:55,800
Do you have any insights there? 
Even though I know Pine Cone 

104
00:04:55,800 --> 00:04:57,320
Assistant could help with that. 
Just for people get an 

105
00:04:57,320 --> 00:04:59,040
understanding, what should they 
think about when they do their 

106
00:04:59,040 --> 00:04:59,920
chunking? 
Yeah. 

107
00:04:59,920 --> 00:05:03,400
So this is one of the most 
frequent questions we get asked 

108
00:05:03,400 --> 00:05:06,040
to the point where like we have 
an article on our website that's

109
00:05:06,040 --> 00:05:08,600
like one of the most viewed 
articles on our website that's 

110
00:05:08,600 --> 00:05:12,920
about chunking strategies. 
And the the hard truth is that 

111
00:05:12,920 --> 00:05:16,200
it really depends on the type of
application that you're building

112
00:05:16,200 --> 00:05:18,160
and the type of search that you 
want to benefit from. 

113
00:05:18,160 --> 00:05:20,440
But there are some heuristics 
that are pretty helpful. 

114
00:05:21,000 --> 00:05:24,240
What I normally tell people is 
to try the simplest chunking 

115
00:05:24,240 --> 00:05:27,960
strategy that you can implement 
on whatever data that you're 

116
00:05:27,960 --> 00:05:32,560
working on and hit your data set
with some queries and just see 

117
00:05:32,560 --> 00:05:36,000
how it feels querying the data 
and getting the responses back. 

118
00:05:36,000 --> 00:05:39,680
And as soon as you start looking
at how your data looks like, 

119
00:05:39,680 --> 00:05:42,160
once you've applied some sort of
chunking strategy, you will 

120
00:05:42,160 --> 00:05:44,720
quickly start to see some of the
holes and issues. 

121
00:05:44,960 --> 00:05:48,680
For example, one of the common 
techniques that people implement

122
00:05:48,760 --> 00:05:52,120
to fix a OK chunking strategy is
something called chunk 

123
00:05:52,120 --> 00:05:54,680
expansion. 
So if you have a bunch of chunks

124
00:05:54,680 --> 00:05:57,120
in a document, let's go back to 
our Wikipedia example, you have 

125
00:05:57,120 --> 00:05:59,400
a bunch of paragraphs that 
you're chunking. 

126
00:05:59,640 --> 00:06:02,360
Sometimes the paragraph that 
you've embedded might not 

127
00:06:02,360 --> 00:06:05,160
contain the entirety of the 
argument that you're trying to 

128
00:06:05,160 --> 00:06:07,960
retrieve given some query. 
So it might just be a helpful 

129
00:06:07,960 --> 00:06:10,960
heuristic for you to look at 
where that paragraph is relative

130
00:06:10,960 --> 00:06:14,360
to the rest of the document, and
then pull out the additional 

131
00:06:14,360 --> 00:06:17,600
chunks above and below that 
chunk in order to return an 

132
00:06:17,600 --> 00:06:20,920
entire kind of document there. 
Padgrown Assistant does that. 

133
00:06:20,920 --> 00:06:23,120
You could also do that yourself 
by including chunk numbers and 

134
00:06:23,120 --> 00:06:26,640
then just retrieving the -1 
chunk and the +1 chunk. 

135
00:06:26,640 --> 00:06:31,400
And that can kind of remediate 
any sort of like not optimal 

136
00:06:31,400 --> 00:06:33,920
chunking strategy. 
So I would say like start at 

137
00:06:33,920 --> 00:06:36,320
like a meaningful unit of 
information that you feel like 

138
00:06:36,320 --> 00:06:38,840
is relevant for you. 
So for Wikipedia articles like 

139
00:06:38,840 --> 00:06:41,480
paragraphs are pretty good. 
Maybe you have a really 

140
00:06:41,480 --> 00:06:44,200
structure like textbook and 
you're like, OK, these chapters 

141
00:06:44,200 --> 00:06:47,280
and sub chapters should be like 
really the minimum amount of 

142
00:06:47,280 --> 00:06:49,920
information you look at. 
Or sometimes people do entire 

143
00:06:49,920 --> 00:06:51,480
pages. 
If you're dealing with like 

144
00:06:51,480 --> 00:06:53,800
multimodal PD FS and like all 
right, each page will be my 

145
00:06:53,800 --> 00:06:55,840
chunk. 
Trying something is better than 

146
00:06:55,840 --> 00:06:58,440
nothing, and trying to 
understand how far you can go 

147
00:06:58,440 --> 00:07:00,840
with a simple method is better 
than applying something really 

148
00:07:00,840 --> 00:07:03,360
complicated, spending a lot of 
time, and then not getting good 

149
00:07:03,360 --> 00:07:04,400
results. 
Yeah. 

150
00:07:04,400 --> 00:07:06,840
And when you mentioned 
potentially doing a full page, 

151
00:07:06,840 --> 00:07:08,880
is there an upper bound to how 
much you put in is? 

152
00:07:08,880 --> 00:07:10,320
Does it depend on the vector 
side? 

153
00:07:10,960 --> 00:07:12,680
So that's a really good 
question. 

154
00:07:12,680 --> 00:07:15,800
It's really dependent on the 
type of embedding model that 

155
00:07:15,800 --> 00:07:17,560
you're using, what it's context 
window is. 

156
00:07:17,560 --> 00:07:20,000
But like with chunking, there's 
some heuristics that you can 

157
00:07:20,000 --> 00:07:23,280
follow that make it easy to kind
of do so. 

158
00:07:23,640 --> 00:07:26,680
It used to be the case that 
people would say that you could 

159
00:07:26,680 --> 00:07:29,800
fill up the entirety of the 
chunking window, the context 

160
00:07:29,800 --> 00:07:32,960
window of an embedding model. 
So if an embedding models like 

161
00:07:33,120 --> 00:07:37,480
context windows like 1024 tokens
or 2048, you could just count 

162
00:07:37,720 --> 00:07:39,760
the number you should tokenize 
your text, fill up to that 

163
00:07:39,760 --> 00:07:41,720
amount, and just throw it inside
the embedding model. 

164
00:07:41,720 --> 00:07:44,680
But there's some anecdotal 
evidence that says that if you 

165
00:07:44,680 --> 00:07:49,000
do this for embedding models 
that have particularly large 

166
00:07:49,000 --> 00:07:53,560
context windows, so greater than
2048, so on so forth, you might 

167
00:07:53,560 --> 00:07:58,120
start a washing out or averaging
over the information that exists

168
00:07:58,120 --> 00:08:00,600
in that chunk. 
And therefore it no longer 

169
00:08:00,600 --> 00:08:04,200
becomes relevant to any query 
than any other chunk would be. 

170
00:08:04,600 --> 00:08:09,480
So in that regime, you probably 
want to only fill up like half 

171
00:08:09,680 --> 00:08:12,960
to maybe 2/3 of a context window
and kind of go from there. 

172
00:08:13,520 --> 00:08:14,960
So that's that's kind of what I 
would say. 

173
00:08:14,960 --> 00:08:16,880
And it also depends on the 
complexity of your text. 

174
00:08:16,880 --> 00:08:19,880
So if your text is very 
information dense, then you want

175
00:08:19,880 --> 00:08:22,280
to have smaller chunks that 
allow you to retrieve that 

176
00:08:22,280 --> 00:08:24,760
information kind of 
independently of one another. 

177
00:08:24,760 --> 00:08:27,000
But if your text is not 
information dense, then it's OK 

178
00:08:27,000 --> 00:08:29,000
to kind of pack more information
in there. 

179
00:08:29,120 --> 00:08:31,120
But yeah, it depends on your 
embedding model. 

180
00:08:31,120 --> 00:08:35,000
It depends on your depends on 
the amount of information you're

181
00:08:35,000 --> 00:08:38,440
passing, but the vibes kind of 
tell me that. 

182
00:08:38,440 --> 00:08:41,440
Like as embedding models get 
really good at getting more and 

183
00:08:41,440 --> 00:08:44,120
more information in, that 
doesn't necessarily scale the 

184
00:08:44,120 --> 00:08:46,880
same way as how that information
is getting represented. 

185
00:08:47,160 --> 00:08:50,360
Similar to how if you have a 
context window for a model 

186
00:08:50,360 --> 00:08:53,880
that's like 200 K, 1,000,000 
tokens, even if you jam all that

187
00:08:53,880 --> 00:08:56,320
up, you're not going to get that
great of a response out. 

188
00:08:56,680 --> 00:09:00,040
Yeah, we very traditionally keep
it under 50% if not lower just 

189
00:09:00,040 --> 00:09:02,360
to make sure we get the better 
results, especially with coding 

190
00:09:02,360 --> 00:09:04,320
models. 
Just on the note of embedding 

191
00:09:04,320 --> 00:09:06,960
models, I'm, I'm admittedly not 
very well versed in it. 

192
00:09:06,960 --> 00:09:09,160
When you talk about embedding 
models being better than others 

193
00:09:09,600 --> 00:09:14,000
is better just in terms of 
having because I, I OK, backing 

194
00:09:14,000 --> 00:09:16,080
up a bit. 
My idea is it takes a certain 

195
00:09:16,080 --> 00:09:18,480
amount of information and puts 
into a number of dimensions with

196
00:09:18,480 --> 00:09:21,320
a value for each dimension. 
When an embedding model is 

197
00:09:21,320 --> 00:09:24,800
better than others, does it just
mean that number like the vector

198
00:09:24,800 --> 00:09:27,520
has more precision? 
Does it just have a better 

199
00:09:27,520 --> 00:09:28,880
representation of the 
information you give it? 

200
00:09:28,880 --> 00:09:31,000
Like what makes 1 embedding 
model better than another? 

201
00:09:31,640 --> 00:09:34,080
Man, this is like one of my 
favorite topics. 

202
00:09:34,160 --> 00:09:37,160
And this is like something that 
I've thought a lot about for 

203
00:09:37,160 --> 00:09:40,560
many years. 
So one of the one courses I took

204
00:09:40,560 --> 00:09:45,080
in college on AI machine 
learning, because there weren't 

205
00:09:45,160 --> 00:09:48,040
that many offers at the time I 
was there, was on vector 

206
00:09:48,040 --> 00:09:49,640
embedding. 
So this is like the first kind 

207
00:09:49,640 --> 00:09:51,040
of thing that I was introduced 
to. 

208
00:09:51,040 --> 00:09:54,160
And all the techniques there are
going to sound almost alien to 

209
00:09:54,160 --> 00:09:56,440
like anybody who's working with 
vector embeddings now, because 

210
00:09:56,440 --> 00:09:59,960
we're talking about techniques 
that work much differently than 

211
00:09:59,960 --> 00:10:01,360
the embedding models that do 
now. 

212
00:10:01,360 --> 00:10:05,160
For example, we could be talking
about bag of word like 

213
00:10:05,160 --> 00:10:08,600
techniques which rely on the 
frequency of how often a word 

214
00:10:08,600 --> 00:10:11,800
occurs in a text and apply some 
sort of waiting algorithm to 

215
00:10:11,800 --> 00:10:13,920
that frequency in order to 
determine it's importance or 

216
00:10:13,920 --> 00:10:15,840
meaning. 
Whereas now we have this 

217
00:10:15,840 --> 00:10:18,800
completely unsupervised method 
that just produces a list of 

218
00:10:18,800 --> 00:10:20,760
numbers and it's not necessarily
interpretable inside those 

219
00:10:20,760 --> 00:10:22,800
numbers. 
So I'm gonna try to give you a 

220
00:10:22,800 --> 00:10:26,000
brief overview of how these 
things are supposed to work and 

221
00:10:26,000 --> 00:10:28,440
then we can talk about what 
makes different things better 

222
00:10:28,440 --> 00:10:32,040
and more interesting. 
So at first when you had 

223
00:10:32,040 --> 00:10:35,480
embedding models, you had very 
simple heuristics that didn't 

224
00:10:35,480 --> 00:10:38,040
deal so much with the 
relationship between the words 

225
00:10:38,040 --> 00:10:40,200
in a given text. 
So if you think about searching 

226
00:10:40,200 --> 00:10:43,960
on Google before like, I don't 
know, 20/18/2017 2016 like 

227
00:10:43,960 --> 00:10:46,560
before they switched to a 
transformer kind of model. 

228
00:10:46,880 --> 00:10:50,600
When you search something, you 
type in as many keywords as you 

229
00:10:50,600 --> 00:10:53,680
can before you get a text that 
kind of maps with that keyword, 

230
00:10:53,680 --> 00:10:56,800
because you're basically just 
trying to hit all of those words

231
00:10:56,800 --> 00:10:59,320
that are inside the thing that 
you're relevant and I think that

232
00:10:59,320 --> 00:11:00,480
you're interested in looking 
for. 

233
00:11:00,880 --> 00:11:03,960
But the trick is, is that these 
words actually have meaning 

234
00:11:03,960 --> 00:11:06,320
associated with them. 
And the meaning of the words 

235
00:11:06,520 --> 00:11:09,640
depend on how they occur inside 
of a sentence. 

236
00:11:10,080 --> 00:11:13,160
So you can think about like 
homophones, like I went to the 

237
00:11:13,160 --> 00:11:16,960
park or I'm going to park, like 
those mean different things in 

238
00:11:16,960 --> 00:11:19,440
different contexts because of 
the way that they kind of relate

239
00:11:19,480 --> 00:11:21,720
to one another. 
So there's this idea called the 

240
00:11:21,720 --> 00:11:26,120
distributional hypothesis, which
claims that words can kind of, 

241
00:11:26,200 --> 00:11:28,640
you can kind of understand what 
words mean based on how they 

242
00:11:28,640 --> 00:11:30,680
cluster around the words that 
are around them. 

243
00:11:30,880 --> 00:11:33,240
So if you look at a word and you
look at the words next to them, 

244
00:11:33,720 --> 00:11:36,240
so there's something about the 
words that Co locate with that 

245
00:11:36,240 --> 00:11:38,680
center word that can give you 
information about what that word

246
00:11:38,680 --> 00:11:40,520
means. 
And it turns out that if you 

247
00:11:40,520 --> 00:11:45,040
train a model to say, hey, 
here's the 1st 4 words in a 

248
00:11:45,040 --> 00:11:46,840
sentence, can you tell me the 
fifth word? 

249
00:11:47,080 --> 00:11:49,400
Or here's a sentence that has 10
words. 

250
00:11:49,640 --> 00:11:52,480
I'm going to throw away 3 or 4 
of them randomly. 

251
00:11:52,720 --> 00:11:55,120
And I want you to guess what 
words should kind of go in 

252
00:11:55,120 --> 00:11:56,840
there. 
If you train a model with that 

253
00:11:56,840 --> 00:11:59,360
objective, it turns out that you
can learn a lot about the 

254
00:11:59,360 --> 00:12:02,800
meaning of words from that task.
So if you think of like Mad Libs

255
00:12:02,800 --> 00:12:04,320
or you can think of just like 
charades, right? 

256
00:12:04,320 --> 00:12:07,040
Just like guessing words and 
trying to complete those things.

257
00:12:07,480 --> 00:12:10,560
So we figured out how to train 
models that can follow that sort

258
00:12:10,560 --> 00:12:13,560
of learning objective. 
And the representations that 

259
00:12:13,560 --> 00:12:16,480
these models learned in order to
achieve that objective turned 

260
00:12:16,480 --> 00:12:18,760
out to be very good at 
representing the underlying 

261
00:12:18,760 --> 00:12:20,480
concept. 
So this is like the second wave 

262
00:12:20,480 --> 00:12:23,440
of like embedding models being 
able to represent what these 

263
00:12:23,440 --> 00:12:27,280
words kind of mean, right? 
And so the differentiation here 

264
00:12:27,280 --> 00:12:30,800
is the loss function. 
So how is the model learning 

265
00:12:30,800 --> 00:12:33,400
this information to generate 
this representation? 

266
00:12:33,680 --> 00:12:36,560
Is it getting better at guessing
the words that are there and 

267
00:12:36,560 --> 00:12:38,440
then also the data that's being 
trained on? 

268
00:12:38,440 --> 00:12:42,040
So the Wikipedia, the Internet, 
whatever corpuses you're going 

269
00:12:42,040 --> 00:12:44,640
on are gonna affect like this 
distribution of words that are 

270
00:12:44,640 --> 00:12:48,240
kind of occurring, right? 
And it turns out that you could 

271
00:12:48,240 --> 00:12:51,560
abandoned this idea of working 
only in English. 

272
00:12:51,840 --> 00:12:56,120
You could train a model that 
predicts on a sub word level or 

273
00:12:56,120 --> 00:12:59,280
token level, literally looking 
at like how often certain 

274
00:12:59,280 --> 00:13:02,800
characters are Co occurring and 
train on tons and tons of 

275
00:13:02,800 --> 00:13:07,480
languages and learn something 
fundamental about how tokens Co 

276
00:13:07,480 --> 00:13:09,560
occur, which inform the meaning 
of those words. 

277
00:13:09,560 --> 00:13:12,800
So you start getting surprising 
research results like, oh, we 

278
00:13:12,800 --> 00:13:15,440
were able to kind of train on 
the entire corpus of 

279
00:13:15,440 --> 00:13:17,520
multilingual data. 
And now we have a way of kind of

280
00:13:17,520 --> 00:13:20,720
searching across languages, even
though we haven't necessarily 

281
00:13:20,720 --> 00:13:22,840
cracked a deterministic way of 
translating across those 

282
00:13:22,840 --> 00:13:25,320
languages, right? 
There's something inherent about

283
00:13:25,320 --> 00:13:28,400
how these words are coexisting 
that informs information there. 

284
00:13:28,640 --> 00:13:33,800
So going down to this sub token 
level, token level allowed us to

285
00:13:33,800 --> 00:13:36,520
kind of capture more of this 
generalized information, right? 

286
00:13:37,000 --> 00:13:40,560
The last kind of way you can get
more power out of embedding 

287
00:13:40,560 --> 00:13:44,040
models or build embedding models
that are cleverer is gonna be 

288
00:13:44,040 --> 00:13:46,360
your data set distribution. 
So the previous thing I 

289
00:13:46,360 --> 00:13:49,200
described was like loss 
functions and then like clever 

290
00:13:49,200 --> 00:13:52,000
ways of preprocessing the data, 
looking at the sub token level. 

291
00:13:52,520 --> 00:13:54,320
And the last thing is data set 
distribution. 

292
00:13:54,320 --> 00:13:56,280
So this is where you get 
specialized embedding models 

293
00:13:56,280 --> 00:13:58,400
from. 
So if you want to say, do legal 

294
00:13:58,400 --> 00:14:02,120
search really, really well and 
it matters to you how to 

295
00:14:02,120 --> 00:14:05,480
represent this domain specific 
terminology, then you probably 

296
00:14:05,480 --> 00:14:08,800
need an embedding model that has
been fine-tuned on that data in 

297
00:14:08,800 --> 00:14:11,080
order to really understand 
what's going on there. 

298
00:14:11,080 --> 00:14:14,080
Because for the same reason that
words mean different things in 

299
00:14:14,080 --> 00:14:16,280
different context, there are 
certain words that have 

300
00:14:16,280 --> 00:14:19,200
extremely powerful meanings in 
specific context that these 

301
00:14:19,200 --> 00:14:21,520
generalized embedding models 
will never have access to. 

302
00:14:21,800 --> 00:14:24,240
So if you're working at a large 
organization, you have tons of 

303
00:14:24,240 --> 00:14:27,640
internal data, maybe there's a 
concept or an idea that you 

304
00:14:27,640 --> 00:14:30,840
refer to that means something 
way different than it would out 

305
00:14:30,840 --> 00:14:32,320
on the Internet or out in the 
wild. 

306
00:14:32,640 --> 00:14:35,240
So changing the data 
distribution of how these 

307
00:14:35,240 --> 00:14:37,480
embedding models are trained on 
are gonna really affect their 

308
00:14:37,480 --> 00:14:39,320
ability to retrieve this 
information. 

309
00:14:39,400 --> 00:14:42,960
And there's also this idea of 
fine tuning for information 

310
00:14:42,960 --> 00:14:46,680
retrieval. 
So not only do we, so before I 

311
00:14:46,680 --> 00:14:49,880
describe tasks of like, oh, 
guessing words like, oh, like we

312
00:14:49,880 --> 00:14:53,560
can like complete a sentence or 
we can fill in like what's 

313
00:14:53,560 --> 00:14:56,360
called masked sentences. 
Like, that's cute, right? 

314
00:14:56,360 --> 00:14:58,800
But we, we care about actually 
finding information. 

315
00:14:58,800 --> 00:15:01,560
So why don't we train on finding
information? 

316
00:15:01,560 --> 00:15:04,520
So you'll have data sets that 
are like here's the input query 

317
00:15:04,720 --> 00:15:07,840
or real question someone asked, 
and here's like all the 

318
00:15:07,840 --> 00:15:11,320
documents that are actually 
relevant to that input query. 

319
00:15:11,640 --> 00:15:14,880
Can we produce a model that 
takes that input query, creates 

320
00:15:14,880 --> 00:15:17,400
an embedding, embeds these 
documents, creates another 

321
00:15:17,400 --> 00:15:21,080
embedding such that the relevant
documents cluster closer in 

322
00:15:21,080 --> 00:15:23,520
vector space than the unrelevant
ones? 

323
00:15:23,560 --> 00:15:26,920
So now you're actually learning 
the signal that is important for

324
00:15:26,920 --> 00:15:29,280
the test that we're describing, 
which is semantic search. 

325
00:15:29,600 --> 00:15:32,160
And that's kind of like this 
latter class of embedding models

326
00:15:32,160 --> 00:15:35,160
that are getting really good at 
representing user questions and 

327
00:15:35,160 --> 00:15:37,760
kind of returning things back. 
So that was a little long 

328
00:15:37,760 --> 00:15:40,720
winded, but that's kind of like 
why you have to care about 

329
00:15:40,720 --> 00:15:42,440
embedding models and what they 
kind of do. 

330
00:15:42,640 --> 00:15:45,640
And the ones that we host have 
these properties of like, not 

331
00:15:45,640 --> 00:15:47,880
only are they good at 
generalizing, not only are they 

332
00:15:47,880 --> 00:15:51,120
good at blah, blah, they're also
really good at doing information

333
00:15:51,120 --> 00:15:53,240
retrieval, which is why that 
we've kind of selected them to 

334
00:15:53,240 --> 00:15:55,000
be on our platform. 
OK, very cool. 

335
00:15:55,000 --> 00:15:58,840
Appreciate that explanation. 
I'm going to take a swing at my 

336
00:15:58,920 --> 00:16:01,600
like just explain my mental lot 
of what semantic similarity is 

337
00:16:01,600 --> 00:16:04,440
and and you just tell me how off
I am because I had trouble. 

338
00:16:04,560 --> 00:16:06,440
I had trouble just like 
picturing things in, you know, 

339
00:16:06,640 --> 00:16:08,960
1024 dimensions. 
I just can't really comprehend 

340
00:16:08,960 --> 00:16:10,160
that. 
But when someone broke it down 

341
00:16:10,160 --> 00:16:13,040
and just brought down to 2 two 
basic dimensions, 2 arbitrary 

342
00:16:13,040 --> 00:16:15,720
points in in space, the semantic
similar is kind of like the 

343
00:16:15,760 --> 00:16:18,240
angle between them. 
And if you want to have 

344
00:16:18,240 --> 00:16:21,160
something like with a smaller 
angle, it's more similar and a 

345
00:16:21,320 --> 00:16:23,960
bigger angle is less similar. 
So that's why semantic 

346
00:16:23,960 --> 00:16:26,560
similarity kind of gauges like 
what your threshold is for 

347
00:16:26,560 --> 00:16:28,040
difference. 
Is that okay? 

348
00:16:28,040 --> 00:16:30,160
Am I kind of close? 
Right, right, right. 

349
00:16:30,160 --> 00:16:33,600
That is one measure of semantic 
similarity. 

350
00:16:33,600 --> 00:16:36,560
And again, it depends on the 
embedding model and what 

351
00:16:36,560 --> 00:16:38,840
distribution of angles you're 
embedding model produces. 

352
00:16:38,840 --> 00:16:39,920
But that's like a detailed, 
right? 

353
00:16:39,920 --> 00:16:45,400
The overall idea is the closer 
in angle you are to this input 

354
00:16:45,400 --> 00:16:49,320
vector or query vector, the more
semantically similar and 

355
00:16:49,320 --> 00:16:51,680
relevant you could be to 
whatever question that you're 

356
00:16:51,680 --> 00:16:53,520
asking. 
And it's crazy, like when you 

357
00:16:53,520 --> 00:16:55,720
look at some of these word plots
that some of these embedding 

358
00:16:55,720 --> 00:16:58,080
models produce, you'll get 
things that make a lot of sense.

359
00:16:58,080 --> 00:17:02,520
Like you'll have the embedding 
for dog, and it'll be next to 

360
00:17:02,520 --> 00:17:05,280
like things that dogs do or 
things that dogs like. 

361
00:17:05,280 --> 00:17:07,440
And you'll have the embedding 
for a cat, and we'll be next to 

362
00:17:07,440 --> 00:17:09,280
the things that cats do or cats 
like. 

363
00:17:09,280 --> 00:17:11,960
But hot dog won't be near any of
those things because it means a 

364
00:17:11,960 --> 00:17:13,839
completely different thing. 
And you're sitting there like, 

365
00:17:13,960 --> 00:17:15,720
how does this work? 
Like, how does this make any 

366
00:17:15,720 --> 00:17:17,480
sense? 
And it just, it just turns out 

367
00:17:17,480 --> 00:17:19,119
it just does. 
Like we can just learn 

368
00:17:19,119 --> 00:17:21,960
information from how these words
are used in context. 

369
00:17:21,960 --> 00:17:24,359
That's pretty wild. 
But that that explanation is 

370
00:17:24,359 --> 00:17:25,480
more or less correct. 
Yeah. 

371
00:17:25,599 --> 00:17:29,680
But in terms of trying to fine 
tune a system so I I have my 

372
00:17:29,920 --> 00:17:33,080
vector database up, I have a few
queries and I'm not totally in 

373
00:17:33,080 --> 00:17:34,480
love with the results I'm 
getting. 

374
00:17:34,720 --> 00:17:37,800
Is semantic similarity the one 
knob that I turn? 

375
00:17:37,800 --> 00:17:40,240
Or what else would I try to 
adjust to improve the quality of

376
00:17:40,240 --> 00:17:41,200
my results? 
Right. 

377
00:17:41,200 --> 00:17:45,520
So having a good embedding model
is one piece of the puzzle of 

378
00:17:45,520 --> 00:17:47,680
building a performance semantic 
search system. 

379
00:17:47,680 --> 00:17:51,280
And semantic search is like a 
good layer of abstraction to 

380
00:17:51,280 --> 00:17:54,480
kind of talk about all kinds of 
applications. 

381
00:17:54,480 --> 00:17:57,320
You can build with Pine Cone or 
with other vector databases. 

382
00:17:57,720 --> 00:18:00,000
O we're talking about 
recommendation systems, we're 

383
00:18:00,000 --> 00:18:02,800
talking about like rag chat 
bots, all that stuff. 

384
00:18:02,800 --> 00:18:05,080
It's basically doing some sort 
of search under the hood. 

385
00:18:05,080 --> 00:18:06,680
So this is kind of like one 
layer, right? 

386
00:18:07,120 --> 00:18:11,240
And the embedding model will 
control how much fidelity you 

387
00:18:11,240 --> 00:18:14,080
can initially get with like your
query and like all of your 

388
00:18:14,080 --> 00:18:15,600
documents that are being 
represented. 

389
00:18:15,800 --> 00:18:18,440
But there's still other 
engineering decisions or even 

390
00:18:18,440 --> 00:18:21,640
other models that you can kind 
of apply to the system in order 

391
00:18:21,640 --> 00:18:23,960
to improve the quality of the 
results. 

392
00:18:23,960 --> 00:18:26,480
But beyond like what the 
embedding model is giving you, 

393
00:18:26,480 --> 00:18:28,520
it's kind of like one thing you 
can do. 

394
00:18:28,520 --> 00:18:31,680
So a common recommendation we 
give at Pine Cone is to apply a 

395
00:18:31,680 --> 00:18:34,520
reranker to your results. 
So we talked about embedding 

396
00:18:34,520 --> 00:18:37,480
models which kind of like allow 
you to represent queries and 

397
00:18:37,480 --> 00:18:41,720
documents in vector space with 
respect to their meaning using 

398
00:18:41,720 --> 00:18:45,400
vectors. 
Rerankers do that kind of, but 

399
00:18:45,400 --> 00:18:48,200
instead of outputting and 
embedding, they output a score. 

400
00:18:48,480 --> 00:18:52,720
So the there's two trade-offs to
think about with these types of 

401
00:18:53,040 --> 00:18:57,680
models. 
So ideally what happens is you 

402
00:18:57,680 --> 00:18:59,840
have a query, you embed that 
query and you have a bunch of 

403
00:18:59,840 --> 00:19:01,480
documents, you embed those 
documents. 

404
00:19:01,800 --> 00:19:05,040
But when you embed those things,
you're trying to extract the 

405
00:19:05,080 --> 00:19:08,440
meaning or the information 
inside the query and each 

406
00:19:08,440 --> 00:19:11,680
document independently, they 
don't know about one another. 

407
00:19:12,080 --> 00:19:15,360
And that's nice because it means
you can embed all of your 

408
00:19:15,360 --> 00:19:17,960
documents independently of 
embedding your query. 

409
00:19:17,960 --> 00:19:20,600
So you can pre index your vector
database without having to 

410
00:19:20,600 --> 00:19:23,600
rebuild your index every single 
time as you get a new query. 

411
00:19:23,600 --> 00:19:25,560
And if you had to think about 
them in context. 

412
00:19:25,920 --> 00:19:28,920
So that makes that type of 
system very performant. 

413
00:19:28,920 --> 00:19:31,480
And this is a property of like 
any any vector database. 

414
00:19:32,040 --> 00:19:35,560
And the compromise though is 
that there are situations in 

415
00:19:35,560 --> 00:19:41,400
which you'd want to adjust the 
influence a a the idea of 

416
00:19:42,720 --> 00:19:46,440
relevance or similarity has 
based on the query and the 

417
00:19:46,440 --> 00:19:49,560
document at the same time. 
So what a reranker does is that 

418
00:19:49,560 --> 00:19:54,280
it inputs a query and a document
at the same time with context of

419
00:19:54,280 --> 00:19:58,240
one another, and it outputs a 
score between zero and one. 

420
00:19:58,280 --> 00:20:02,600
Usually that score correlates to
not only the semantic 

421
00:20:02,600 --> 00:20:06,760
similarity, but the relevance of
the document to the query. 

422
00:20:06,880 --> 00:20:09,320
O It's possible to have 
documents that are semantically 

423
00:20:09,320 --> 00:20:12,400
similar but aren't relevant. 
It could be about things that 

424
00:20:12,720 --> 00:20:15,360
are close in meaning to the 
thing that you're trying to ask,

425
00:20:15,520 --> 00:20:17,720
but they're not the thing that 
you're actually looking for. 

426
00:20:18,000 --> 00:20:21,240
So we consider this as sort of 
refining step after doing an 

427
00:20:21,240 --> 00:20:23,560
initial semantic search. 
So you do a search, you have a 

428
00:20:23,560 --> 00:20:27,400
million documents, you reduce it
down to like 100 or 200 

429
00:20:27,400 --> 00:20:29,920
documents, and then you apply a 
reranker and you reduce it down 

430
00:20:29,920 --> 00:20:33,240
even further to the top, like 10
or 15 documents and you return 

431
00:20:33,240 --> 00:20:35,840
that back. 
So that allows you to increase 

432
00:20:35,840 --> 00:20:38,880
the quality of the already good 
results to something even better

433
00:20:38,880 --> 00:20:40,320
without having to re index 
everything. 

434
00:20:40,960 --> 00:20:43,640
So that's like one thing you can
do, and there's tons of other 

435
00:20:43,640 --> 00:20:46,320
tools in the toolbox to make 
search better, but that's like 

436
00:20:46,320 --> 00:20:50,000
one thing we recommend. 
OK, and just so I'm making sure 

437
00:20:50,000 --> 00:20:54,600
I understand the rear maker 
concept so it won't impact the 

438
00:20:54,600 --> 00:20:58,040
embeddings at any weight. 
It's just like a filter after a 

439
00:20:58,040 --> 00:21:01,440
query returns a number of chunks
or a number of embeddings I 

440
00:21:01,440 --> 00:21:04,120
should say. 
It is a, it is a scoring 

441
00:21:04,120 --> 00:21:06,880
algorithm. 
So it you, you could say it 

442
00:21:06,880 --> 00:21:09,760
filters for like the most 
relevant documents, but 

443
00:21:09,760 --> 00:21:13,240
mechanically it's rescoring 
everything, giving it a new 

444
00:21:13,240 --> 00:21:15,040
score. 
And then you order by that score

445
00:21:15,040 --> 00:21:17,280
and just return everything past 
the cut off. 

446
00:21:17,880 --> 00:21:21,080
That's why I like to refer to it
as refinement, because you get 

447
00:21:21,080 --> 00:21:23,680
like 100 documents and then 
you're saying, OK, I want to 

448
00:21:24,000 --> 00:21:27,640
further refine the quality of 
these 100 documents in the 

449
00:21:27,640 --> 00:21:30,120
context of this query. 
So then you reorder those 

450
00:21:30,120 --> 00:21:33,160
documents, and then you just 
take the top 15 most relevant 

451
00:21:33,160 --> 00:21:34,880
ones and return that back to the
user. 

452
00:21:35,160 --> 00:21:37,080
Right. 
And are rerankers, the type of 

453
00:21:37,080 --> 00:21:40,120
thing that can be fine-tuned 
based people, you know, they see

454
00:21:40,120 --> 00:21:42,680
the performance like I actually 
want a number 12 or 13, let's 

455
00:21:42,680 --> 00:21:44,600
readjust it so that gets higher,
higher ranking. 

456
00:21:45,000 --> 00:21:47,840
You could. 
There are limits to how 

457
00:21:47,840 --> 00:21:52,520
effective fine tuning is for 
like the developer. 

458
00:21:52,520 --> 00:21:55,680
It depends on how much data you 
already have and if your data 

459
00:21:55,680 --> 00:21:57,360
are labeled right. 
That's not easy. 

460
00:21:57,440 --> 00:22:01,080
Need to fine tune. 
Usually when you want to fine 

461
00:22:01,080 --> 00:22:04,480
tune something like information 
retrieval or a re ranker, you 

462
00:22:04,480 --> 00:22:07,040
need data sets that are like 
here's the query and here are 

463
00:22:07,040 --> 00:22:09,680
the relevant documents. 
And most people don't have that 

464
00:22:09,680 --> 00:22:13,560
because they don't have a system
in production yet or they don't 

465
00:22:13,560 --> 00:22:15,280
have it in the format that they 
actually need. 

466
00:22:15,280 --> 00:22:18,880
So you might have like data 
about, I don't know, things that

467
00:22:18,880 --> 00:22:21,000
users are clicking on, but you 
might not know if that's 

468
00:22:21,000 --> 00:22:22,280
actually relevant to what they 
need. 

469
00:22:22,280 --> 00:22:23,640
It might just be a good 
heuristic. 

470
00:22:23,640 --> 00:22:28,360
So it is possible to fine tune 
any of these models, it just 

471
00:22:28,360 --> 00:22:31,880
depends on how specialized your 
domain is and whether you have 

472
00:22:31,880 --> 00:22:33,600
the data to support that sort of
operation. 

473
00:22:33,720 --> 00:22:38,280
And is there anything such as 
metadata or I'm thinking like a 

474
00:22:38,280 --> 00:22:42,360
tagging system to put on 
embeddings to kind of highlight 

475
00:22:42,360 --> 00:22:43,440
them? 
Maybe it's from like a certain 

476
00:22:43,440 --> 00:22:45,480
book or certain chapter or 
something so you can instead of 

477
00:22:45,680 --> 00:22:47,440
going across everything, be like
I definitely want from this 

478
00:22:47,440 --> 00:22:49,440
area. 
Reduce the list that way. 

479
00:22:49,720 --> 00:22:52,280
Yeah. 
So this is another common thing 

480
00:22:52,280 --> 00:22:55,000
that people do with vector 
databases and on Pine Cone. 

481
00:22:55,200 --> 00:22:58,000
So we have two main systems of 
letting you. 

482
00:22:58,000 --> 00:23:00,280
And this is where like the 
engineering concerns of like 

483
00:23:00,280 --> 00:23:02,560
building a search engine really 
come into play. 

484
00:23:03,160 --> 00:23:06,840
So you can imagine that if you 
were some large like 

485
00:23:07,120 --> 00:23:10,840
organization or maybe you were 
just like a, a developer that's 

486
00:23:10,840 --> 00:23:13,400
really good at like getting data
into whatever system that you're

487
00:23:13,400 --> 00:23:15,960
kind of making, building your 
own Wikipedia or whatever, 

488
00:23:15,960 --> 00:23:17,680
right? 
Indexing all of your notes, you 

489
00:23:17,680 --> 00:23:20,200
could have like a million 
documents or 100 million 

490
00:23:20,200 --> 00:23:21,880
documents or something insane, 
right? 

491
00:23:22,280 --> 00:23:26,320
And even if we have performant 
indexing algorithms, you might 

492
00:23:26,320 --> 00:23:30,040
not want to do a search that 
like does something to all of 

493
00:23:30,040 --> 00:23:33,440
those operations, right? 
You want to reduce the scope of 

494
00:23:33,440 --> 00:23:35,760
the search somewhat if you have 
some information at the 

495
00:23:35,760 --> 00:23:38,680
beginning, and there's two ways 
to do this with Pine Cone. 

496
00:23:38,680 --> 00:23:40,760
One is metadata filtering, which
is what you describe. 

497
00:23:40,760 --> 00:23:44,280
So every time you upsert A 
vector, which is our terminology

498
00:23:44,280 --> 00:23:48,080
for putting data into your 
vector database, you can attach 

499
00:23:48,080 --> 00:23:53,800
metadata to that vector, and you
can also filter by that before 

500
00:23:53,800 --> 00:23:56,760
you're doing the vector search. 
So that means, hey, everything 

501
00:23:56,760 --> 00:23:59,880
that satisfies this, pull that 
and then do the vector search 

502
00:23:59,880 --> 00:24:01,840
over that information. 
So you're still getting the 

503
00:24:01,840 --> 00:24:05,880
benefits of a really performant 
vector search, but you're not 

504
00:24:05,880 --> 00:24:08,360
scanning like the entirety of 
your index because you don't 

505
00:24:08,360 --> 00:24:11,000
need to, because you already 
specified the filter in advance.

506
00:24:11,320 --> 00:24:13,760
So you could do this with 
content, you could do it with 

507
00:24:13,760 --> 00:24:16,920
languages if you had multiple 
languages in a vector database 

508
00:24:17,120 --> 00:24:19,200
or sections of a document, any 
anything. 

509
00:24:19,760 --> 00:24:22,720
And the second, a layer of 
abstraction, which is more 

510
00:24:22,720 --> 00:24:24,880
strict is namespaces. 
Excuse me? 

511
00:24:25,160 --> 00:24:29,080
So namespaces are like 
partitions on your index that 

512
00:24:29,080 --> 00:24:32,840
allow you to enable a multi 
tenant architecture for your 

513
00:24:32,840 --> 00:24:35,200
users. 
So let's say you have a ton of 

514
00:24:35,200 --> 00:24:38,600
users and you don't want users 
to query each other's data, but 

515
00:24:38,600 --> 00:24:41,320
you want them all to live on the
same index so that you can 

516
00:24:41,320 --> 00:24:44,560
benefit from pine cones like 
auto scaling infrastructure. 

517
00:24:44,560 --> 00:24:46,560
We scale independently of 
queries and rights. 

518
00:24:46,560 --> 00:24:48,920
And you don't necessarily want 
everyone to have like their own 

519
00:24:48,920 --> 00:24:50,720
server or whatever. 
You kinda want them to all live 

520
00:24:50,720 --> 00:24:53,760
on the same infrastructure, so 
their latency is like quite low.

521
00:24:54,120 --> 00:24:58,240
Namespaces let you specify that 
level of partition in your 

522
00:24:58,240 --> 00:25:00,880
database so that you're just 
querying inside there. 

523
00:25:00,880 --> 00:25:02,960
So each user could have their 
own namespace and then you're 

524
00:25:02,960 --> 00:25:05,320
just querying inside their 
partition of the vector 

525
00:25:05,320 --> 00:25:08,680
database. 
Are there any other free filters

526
00:25:08,680 --> 00:25:11,120
or things that people should 
consider to help making sure 

527
00:25:11,120 --> 00:25:13,240
that the results they retrieve 
are accurate? 

528
00:25:13,320 --> 00:25:16,320
So metadata filtering is the big
one. 

529
00:25:16,320 --> 00:25:19,640
And then they're sort of like, 
and of course, metadata 

530
00:25:19,640 --> 00:25:23,080
filtering can be used for 
content type, any sort of 

531
00:25:23,080 --> 00:25:26,320
operation that allows you to, if
you have this information in 

532
00:25:26,320 --> 00:25:29,120
advance, what data you're kind 
of putting into Pine Cone and 

533
00:25:29,120 --> 00:25:32,960
how you can sort of tag it. 
But another thing you could do 

534
00:25:32,960 --> 00:25:36,800
is have separate indexes and 
then use some sort of routing 

535
00:25:36,800 --> 00:25:40,120
mechanism in order to push 
people toward one index or 

536
00:25:40,120 --> 00:25:43,280
another. 
So, for example, you can embed 

537
00:25:43,280 --> 00:25:46,440
data in different ways, which 
allow for different sorts of 

538
00:25:46,440 --> 00:25:49,120
benefits. 
Like you can embed information 

539
00:25:49,120 --> 00:25:52,120
with a dense embedding model, 
which kind of exposes more 

540
00:25:52,120 --> 00:25:55,400
semantic meaningful concepts out
of your data. 

541
00:25:55,720 --> 00:25:59,280
Or you can use a keyword or 
sparse embedding model, which 

542
00:25:59,280 --> 00:26:01,920
really focuses on the words that
are being used and their 

543
00:26:01,920 --> 00:26:04,880
meanings inside that context. 
And there might be situations in

544
00:26:04,880 --> 00:26:07,720
which you want to route a user 
toward one type of search method

545
00:26:07,720 --> 00:26:10,200
versus the other. 
And that way you're not hitting 

546
00:26:10,200 --> 00:26:12,240
the entirety of the index at the
same time, but you're just 

547
00:26:12,240 --> 00:26:14,640
hitting the specific kind of 
search method that they need. 

548
00:26:14,880 --> 00:26:18,200
That's something that somebody 
could implement like on top of 

549
00:26:18,200 --> 00:26:20,800
Pine Cone to route queries. 
We call that query routing. 

550
00:26:20,880 --> 00:26:23,160
That's another way of kind of 
like filtering results down. 

551
00:26:23,160 --> 00:26:26,120
Or if you, if you think you have
an agent and it has access to 

552
00:26:26,120 --> 00:26:29,160
many different databases, you 
don't want it to hit every 

553
00:26:29,160 --> 00:26:30,800
single database on every single 
query. 

554
00:26:30,800 --> 00:26:33,920
You wanted to know about what 
information is in inside each 

555
00:26:33,920 --> 00:26:37,520
database and then only use the 
data from the relevant 

556
00:26:37,520 --> 00:26:39,560
databases. 
It kind of meta there, right? 

557
00:26:39,560 --> 00:26:42,040
You're doing a relevant database
search before you do a relevant 

558
00:26:42,040 --> 00:26:46,480
data search, and that can reduce
your search space as well, so be

559
00:26:46,480 --> 00:26:48,560
beyond that. 
It's more engineering concerns 

560
00:26:48,560 --> 00:26:51,720
before you get to Pine Cone, but
metadata filtering is the really

561
00:26:51,720 --> 00:26:53,880
big one. 
How well does a hybrid 

562
00:26:53,880 --> 00:26:56,080
architecture or or general 
advice you have? 

563
00:26:56,120 --> 00:26:59,080
If someone wants to use sequel 
for a portion of the information

564
00:26:59,080 --> 00:27:01,600
in vector for another portion, 
how should I think about 

565
00:27:01,600 --> 00:27:03,920
dividing it? 
How should they orchestrate it? 

566
00:27:03,920 --> 00:27:07,320
Is it just simply having two 
simultaneous ones or or what? 

567
00:27:07,320 --> 00:27:08,600
What have you seen there? 
Yeah. 

568
00:27:08,600 --> 00:27:12,680
So this is a problem that a lot 
of like developers run into, 

569
00:27:12,720 --> 00:27:15,720
especially when they already 
have one database and they want 

570
00:27:15,720 --> 00:27:19,480
to kind of figure out another. 
So sometimes people will wanna 

571
00:27:19,480 --> 00:27:23,280
do like tabular structured 
operations on their vector data 

572
00:27:23,280 --> 00:27:26,680
or they'll have their tabular 
structured operations, but they 

573
00:27:26,680 --> 00:27:29,560
wanna do a vector search and 
they're like, well, what are 

574
00:27:29,560 --> 00:27:32,920
these heuristics I need to think
about and which data stores best

575
00:27:32,920 --> 00:27:35,360
for what? 
So my advice would be is to 

576
00:27:35,360 --> 00:27:38,400
think about the type of search 
that you want to implement and 

577
00:27:38,400 --> 00:27:41,880
what field or fields of your 
data that you're already 

578
00:27:41,880 --> 00:27:44,480
collecting are gonna be 
important for that type of 

579
00:27:44,480 --> 00:27:46,520
search. 
That is the information that you

580
00:27:46,520 --> 00:27:49,200
want to put inside Pine Cone 
because you can only embed like 

581
00:27:49,200 --> 00:27:52,240
one kind of concept in each of 
the indexes that you're kind of 

582
00:27:52,240 --> 00:27:54,400
working with. 
So an example would be like, 

583
00:27:54,440 --> 00:27:57,480
let's say you have product 
reviews and you have all this 

584
00:27:57,480 --> 00:28:00,080
information about the products, 
like what it is, how much it 

585
00:28:00,080 --> 00:28:01,760
costs, so on and so forth. 
That's like kind of your 

586
00:28:01,760 --> 00:28:03,800
structured data. 
And then you have the product 

587
00:28:03,800 --> 00:28:05,640
reviews which map back to the 
products. 

588
00:28:05,880 --> 00:28:07,920
And the product reviews are 
completely unstructured. 

589
00:28:07,920 --> 00:28:09,640
People can write like whatever 
they want. 

590
00:28:10,000 --> 00:28:13,200
So you might want to store 
information in Pine Cone being 

591
00:28:13,200 --> 00:28:15,560
the product reviews. 
Maybe you have another index for

592
00:28:15,560 --> 00:28:17,920
just the titles, or you put the 
titles and the reviews in the 

593
00:28:17,920 --> 00:28:21,040
same index and then use metadata
filters to identify one from the

594
00:28:21,040 --> 00:28:24,080
other, whatever you want to do. 
And you have a unique ID that 

595
00:28:24,080 --> 00:28:26,560
maps back to your structured 
database if you need to pull 

596
00:28:26,560 --> 00:28:28,440
this other information that's 
relevant. 

597
00:28:28,800 --> 00:28:32,200
And that way you're 
concentrating this like very 

598
00:28:32,320 --> 00:28:35,360
specific kind of data that 
requires this vector search 

599
00:28:35,360 --> 00:28:37,800
inside Pine Cone, which is what 
Pine Cone is really good for. 

600
00:28:37,800 --> 00:28:40,080
And then whenever you need to 
run your structured queries, you

601
00:28:40,080 --> 00:28:41,600
can do that with your other 
databases. 

602
00:28:41,840 --> 00:28:43,960
It's really tempting for 
developers to kind of like 

603
00:28:43,960 --> 00:28:47,560
flatten their existing databases
and then put that all inside 

604
00:28:47,560 --> 00:28:49,760
Pine code and then try to do 
searches over that. 

605
00:28:49,760 --> 00:28:53,040
So like you can imagine if we 
had a product description and 

606
00:28:53,040 --> 00:28:55,760
all these attributes and all 
this stuff, you could put it 

607
00:28:55,760 --> 00:28:58,160
into one big text chunk and then
embed that whole thing. 

608
00:28:58,360 --> 00:29:02,120
That ends up doing the sort of 
averaging of information that we

609
00:29:02,120 --> 00:29:04,760
talked about earlier where 
there's like too much stuff in 

610
00:29:04,760 --> 00:29:07,560
each chunk and that's not what 
you want to do. 

611
00:29:07,560 --> 00:29:10,560
You want to figure out the exact
thing you want to help people 

612
00:29:10,560 --> 00:29:12,680
find and build the search for 
that thing. 

613
00:29:12,840 --> 00:29:16,400
OK, so for, for concrete 
example, in my world, yeah, at 

614
00:29:16,400 --> 00:29:20,200
Box 1 Ventures, if I have 
information on a company, the 

615
00:29:20,320 --> 00:29:23,400
location, the headcount, the 
founding year, all of that still

616
00:29:23,400 --> 00:29:26,440
sits in sequel because it's very
direct, easy to query, but a 

617
00:29:26,440 --> 00:29:29,160
description of what the company 
does, a description of the 

618
00:29:29,160 --> 00:29:30,800
science that led to their 
discovery. 

619
00:29:30,800 --> 00:29:33,240
What not things where people 
would say, hey, like I need 

620
00:29:33,240 --> 00:29:35,880
information on certain types of 
therapeutic companies that would

621
00:29:35,880 --> 00:29:38,800
be more appropriate for a pine 
cone or vector database because 

622
00:29:38,800 --> 00:29:41,440
it's not going to be as binary. 
Give me this one field. 

623
00:29:42,000 --> 00:29:43,680
Exactly. 
And maybe there are certain 

624
00:29:43,680 --> 00:29:46,760
operations that you do along 
with that that are just useful 

625
00:29:46,760 --> 00:29:48,840
to have inside Pine Cone as 
metadata. 

626
00:29:48,840 --> 00:29:51,760
So maybe the industry, you've 
already pulled this information 

627
00:29:51,760 --> 00:29:54,960
out from the company. 
So it just makes sense to be 

628
00:29:54,960 --> 00:29:58,160
able to segment your searches 
inside Pine Cone by industry. 

629
00:29:58,160 --> 00:30:00,200
So if you care about medical 
applications and then you're 

630
00:30:00,200 --> 00:30:03,440
searching inside the subsection 
of medical companies about the 

631
00:30:03,440 --> 00:30:06,440
specific innovation that they're
exposing, you can see how that 

632
00:30:06,440 --> 00:30:08,280
would work. 
But what's interesting about 

633
00:30:08,280 --> 00:30:11,280
this conversation is that it 
reveals A trend that has changed

634
00:30:11,360 --> 00:30:14,440
over the past five or six years.
And something that I've seen a 

635
00:30:14,440 --> 00:30:18,240
lot of because when I started 
developing as a data scientist, 

636
00:30:18,680 --> 00:30:21,240
like the initial answer would 
have been to build models to 

637
00:30:21,240 --> 00:30:24,640
kind of pull information out of 
these structured descriptions in

638
00:30:24,640 --> 00:30:27,440
order to populate the structured
database and then do queries 

639
00:30:27,440 --> 00:30:29,080
over those, right? 
Like build a model that can 

640
00:30:29,080 --> 00:30:32,680
predict the industry of a 
company or try to extract like 

641
00:30:32,680 --> 00:30:34,200
the thing that they're building,
all that sort of stuff. 

642
00:30:34,560 --> 00:30:37,600
Now we have these powerful ways 
of kind of abstracting that 

643
00:30:37,600 --> 00:30:40,080
search out so we don't have to 
do this upfront engineering 

644
00:30:40,080 --> 00:30:41,920
work, right? 
It's kind of amazing. 

645
00:30:41,920 --> 00:30:45,440
Like we can offload this type of
querying to this agent or the 

646
00:30:45,760 --> 00:30:47,880
representation of the 
information to this embedding 

647
00:30:47,880 --> 00:30:50,320
model so that we can find this 
information faster. 

648
00:30:50,320 --> 00:30:53,600
That's the key value there. 
And in one thing that I've been 

649
00:30:53,600 --> 00:30:57,040
exploring with is just having a,
let's just say my Obsidian 

650
00:30:57,040 --> 00:30:58,560
vault. 
So a clutch and markdown files 

651
00:30:58,760 --> 00:31:01,280
let Claude code loose on it. 
It'll figure out what it needs 

652
00:31:01,280 --> 00:31:02,960
and it'll bring back that 
information to a natural 

653
00:31:02,960 --> 00:31:05,000
language query similar to how a 
vector database would without 

654
00:31:05,000 --> 00:31:07,440
all the steps in the middle. 
What are the pros and cons to 

655
00:31:07,440 --> 00:31:08,680
that approach? 
Because like you mentioned, the 

656
00:31:08,680 --> 00:31:11,000
speed is very fast with vector 
database and having Claude go 

657
00:31:11,000 --> 00:31:12,160
through which things is very 
slow. 

658
00:31:12,160 --> 00:31:14,320
But what else should people 
consider of those two 

659
00:31:14,320 --> 00:31:15,520
approaches? 
Yeah. 

660
00:31:15,520 --> 00:31:18,520
So this is like a very, this is 
like very interesting argument, 

661
00:31:18,560 --> 00:31:20,800
right? 
And I think that this depends, 

662
00:31:21,360 --> 00:31:24,480
obviously there's a lot of like 
qualifiers on the exact type of 

663
00:31:24,480 --> 00:31:26,960
search that you're doing, what 
your tolerance is for the 

664
00:31:26,960 --> 00:31:29,480
relevance of the results, how 
long you're willing to wait. 

665
00:31:29,840 --> 00:31:33,280
And I think it's really tempting
for for people to draw the 

666
00:31:33,280 --> 00:31:37,840
conclusion that as long as I 
start a task and then have the 

667
00:31:37,840 --> 00:31:39,920
results of the task come back, 
that's OK. 

668
00:31:39,920 --> 00:31:41,760
That's like kind of good enough.
And there are a lot of 

669
00:31:41,760 --> 00:31:44,120
situations where that's 
completely fine, right? 

670
00:31:44,120 --> 00:31:48,600
Like maybe Mike, you don't need 
like the most low latent search 

671
00:31:48,600 --> 00:31:51,240
for your upsetting docs because 
you're gonna ask a lot of 

672
00:31:51,240 --> 00:31:52,560
questions. 
You're gonna walk away, then 

673
00:31:52,560 --> 00:31:54,120
you'll come back and you'll kind
of have the answer. 

674
00:31:54,120 --> 00:31:57,040
And that's like, OK, right? 
That's completely fine, but 

675
00:31:57,240 --> 00:31:59,840
maybe you're realizing that, 
hey, you know what Claude is 

676
00:31:59,840 --> 00:32:04,320
doing a lot of queries and a lot
of text searches and a lot of 

677
00:32:04,320 --> 00:32:05,720
tokens are kind of being used 
up. 

678
00:32:05,720 --> 00:32:08,280
Maybe this is a problem. 
Maybe I need a way to kind of 

679
00:32:09,080 --> 00:32:12,160
quickly identify which documents
are relevant, and then once I do

680
00:32:12,160 --> 00:32:14,560
that, cloud can do grep search 
inside there. 

681
00:32:14,560 --> 00:32:17,560
So it doesn't necessarily have 
to be mutually exclusive at all.

682
00:32:17,640 --> 00:32:20,200
Of course, there are engineering
concerns like spinning, like 

683
00:32:20,400 --> 00:32:22,680
making sure you upsert all of 
your data and keeping that In 

684
00:32:22,680 --> 00:32:25,640
Sync. 
But we find that semantic search

685
00:32:25,640 --> 00:32:28,560
can be a lot more flexible than 
just doing keyword searches 

686
00:32:28,560 --> 00:32:29,960
repeatedly. 
And of course, you're going to 

687
00:32:29,960 --> 00:32:32,440
save a lot of tokens if you're 
getting the results back the 

688
00:32:32,440 --> 00:32:34,960
first few times. 
So yeah, that's kind of like 

689
00:32:34,960 --> 00:32:37,160
what I would say. 
It's completely fine to just use

690
00:32:37,160 --> 00:32:40,640
what you think is helpful, but I
think if you look closely, 

691
00:32:40,640 --> 00:32:43,600
especially at how Claude is 
doing searches, it is a brute 

692
00:32:43,680 --> 00:32:45,960
force method. 
It doesn't know about what data 

693
00:32:45,960 --> 00:32:49,160
you have, so it's just trying to
find things as you could, and 

694
00:32:49,320 --> 00:32:52,520
it's kind of fast at it, so it 
can do it quickly, but you have 

695
00:32:52,520 --> 00:32:55,360
to think about the consequences 
of doing so many searches. 

696
00:32:55,360 --> 00:32:57,800
You're going to fill up your 
context window really quickly, 

697
00:32:57,800 --> 00:32:59,800
and at some point you have to 
decide if that's worth it over a

698
00:32:59,800 --> 00:33:02,800
really large database. 
Yeah, And similar type mentioned

699
00:33:02,800 --> 00:33:05,680
earlier, if it's just grepping 
through a file for dog and 

700
00:33:05,680 --> 00:33:08,880
there's something about, you 
know, a leash, it's not 

701
00:33:08,880 --> 00:33:09,920
necessarily going to come up 
with it. 

702
00:33:09,920 --> 00:33:12,240
So I might miss it completely. 
Just on the note of concrete 

703
00:33:12,240 --> 00:33:15,520
examples, having worked at Pine 
Code where some cool unique use 

704
00:33:15,520 --> 00:33:18,120
cases you've seen where people 
might not traditionally think a 

705
00:33:18,120 --> 00:33:20,480
vector database is the right 
solution, but it turned out to 

706
00:33:20,480 --> 00:33:22,360
be great for them. 
Yeah. 

707
00:33:22,360 --> 00:33:26,640
So I think that a lot of people 
think about pine cone and vector

708
00:33:26,640 --> 00:33:30,360
databases for retrieval 
augmented generation, which is 

709
00:33:30,360 --> 00:33:33,200
like, hey, you have a chat bot, 
you upload some documents and 

710
00:33:33,200 --> 00:33:36,800
then it searches over the stuff 
and you get answers back. 

711
00:33:37,040 --> 00:33:40,400
And I think that's like fine. 
But there are a lot of really 

712
00:33:40,440 --> 00:33:43,760
interesting, like industry 
specific applications that are 

713
00:33:43,760 --> 00:33:47,800
particularly powerful that rely 
more on this idea of being able 

714
00:33:47,800 --> 00:33:50,720
to do semantic search really 
quickly rather than like, you 

715
00:33:50,720 --> 00:33:53,680
know, being able to generate 
like the appropriate or correct 

716
00:33:53,680 --> 00:33:56,720
response. 
So one of and and outside of 

717
00:33:56,720 --> 00:33:58,920
that, I think one of the 
interesting applications of 

718
00:33:58,920 --> 00:34:00,920
semantic search is 
recommendation systems. 

719
00:34:01,280 --> 00:34:04,440
So what people might not realize
or think about is that 

720
00:34:04,440 --> 00:34:07,320
recommendation systems are a 
form of search where you have an

721
00:34:07,320 --> 00:34:10,679
input query, which might be 
users preferences or their 

722
00:34:10,679 --> 00:34:14,840
previous shopping history or 
whatever might have you, and you

723
00:34:14,840 --> 00:34:17,560
have a bunch of information 
that's been embedded that can 

724
00:34:18,000 --> 00:34:20,159
quickly be things that are 
relevant to them. 

725
00:34:20,679 --> 00:34:24,440
And the ability to kind of make 
that sort of operation possible 

726
00:34:24,440 --> 00:34:27,480
very quickly is super cool. 
And I think that building 

727
00:34:27,480 --> 00:34:30,320
recommendation systems on top of
Pine Cone is like very, very 

728
00:34:30,320 --> 00:34:33,440
underrated. 
I have a fun example of spinning

729
00:34:33,440 --> 00:34:36,560
up kind of a quick simple one 
that I can show you in a couple 

730
00:34:36,560 --> 00:34:38,159
of minutes. 
Another thing that I think is 

731
00:34:38,159 --> 00:34:42,679
really cool personally, that's 
kind of outside of a super funky

732
00:34:42,679 --> 00:34:46,440
enterprise application is the 
ability to encode data that 

733
00:34:46,440 --> 00:34:49,120
normally you wouldn't think 
would be easy to encode or map 

734
00:34:49,120 --> 00:34:50,480
up. 
So I'm talking about like 

735
00:34:50,480 --> 00:34:53,199
multimodal search or 
multilingual or cross lingual 

736
00:34:53,199 --> 00:34:54,239
search. 
One of the first things I 

737
00:34:54,239 --> 00:34:56,760
learned at Pine Cone is that 
there is such a thing called 

738
00:34:56,760 --> 00:34:59,440
multilingual embedding models. 
And these are embedding models 

739
00:34:59,440 --> 00:35:02,680
that have somehow figured out 
how to represent like hundreds 

740
00:35:02,680 --> 00:35:04,520
of languages in the same 
embedding space. 

741
00:35:04,760 --> 00:35:07,560
So if you ask a question in 
English, you'll get the correct 

742
00:35:07,560 --> 00:35:09,560
responses back in other 
languages. 

743
00:35:09,840 --> 00:35:12,960
And there's not a clear reason 
as to how this translation is 

744
00:35:12,960 --> 00:35:14,200
actually occurring under the 
hood. 

745
00:35:14,200 --> 00:35:16,880
It's just the property of how 
the languages are being used, 

746
00:35:16,880 --> 00:35:19,640
like across the world. 
So you can do crazy things like 

747
00:35:19,720 --> 00:35:22,880
just do translation without 
having to actually translate 

748
00:35:22,880 --> 00:35:25,200
because you're using an 
embedding model that finds the 

749
00:35:25,200 --> 00:35:28,760
most relevant information to the
given input query purely off the

750
00:35:28,760 --> 00:35:31,480
property of the semantic 
similarity search. 

751
00:35:31,480 --> 00:35:34,800
And once I built that demo, I 
was like, oh, there's something 

752
00:35:34,800 --> 00:35:38,440
very interesting here. 
And that extends to any, any 

753
00:35:38,440 --> 00:35:40,240
modality. 
So you can do text to image, 

754
00:35:40,240 --> 00:35:42,880
image to text, video to video, 
audio to video. 

755
00:35:43,080 --> 00:35:45,400
If you can represent it in the 
same embedding space, then you 

756
00:35:45,400 --> 00:35:47,760
can search over it and it gets 
things get really crazy really 

757
00:35:47,760 --> 00:35:49,800
fast. 
Well, you, you mentioned demo 

758
00:35:49,920 --> 00:35:52,480
and we talked before that you 
have something set up with the 

759
00:35:52,480 --> 00:35:54,880
cloud code which is getting 
hotter than ever, especially 

760
00:35:54,880 --> 00:35:57,040
with Four 6 coming out today. 
So would you mind giving us a 

761
00:35:57,040 --> 00:35:58,960
little demo? 
Yeah, absolutely. 

762
00:35:58,960 --> 00:36:02,000
So a little a little preface. 
So at Pine Cone, we think really

763
00:36:02,000 --> 00:36:06,120
hard about how people are using 
US and their agentic IDs like 

764
00:36:06,120 --> 00:36:09,600
Cursor, Cloud Code, Copilot, 
Anti Gravity, so on and so 

765
00:36:09,600 --> 00:36:11,480
forth. 
And so we've been developing a 

766
00:36:11,480 --> 00:36:15,120
lot of content at Pine Cone to 
help people be successful with 

767
00:36:15,120 --> 00:36:16,920
our products inside those 
environments. 

768
00:36:17,040 --> 00:36:19,880
One of the things we developed 
was a plug in for Cloud Code. 

769
00:36:20,120 --> 00:36:23,040
You can find this plug in on the
Cloud Code official marketplace.

770
00:36:23,240 --> 00:36:26,320
Chances are if you have Cloud 
Code, it's already on the 

771
00:36:26,320 --> 00:36:28,080
marketplace, is already 
installed in your session. 

772
00:36:28,200 --> 00:36:31,360
All you have to do is hit slash 
plugin, install Pine Cone, and 

773
00:36:31,360 --> 00:36:34,640
you can start using our plugin. 
So I'll walk you through some 

774
00:36:34,640 --> 00:36:37,720
demos on Semantic Search and 
Assistant if we have time. 

775
00:36:37,720 --> 00:36:40,400
It'll be super fun. 
Let me go ahead and share my 

776
00:36:40,400 --> 00:36:44,160
screen and get started. 
So I'm inside VS Code right now 

777
00:36:44,160 --> 00:36:47,480
and I've fired up Claude here. 
I've already have our plugin 

778
00:36:47,480 --> 00:36:49,200
installed and I'll show you what
that looks like. 

779
00:36:49,480 --> 00:36:53,120
And I've also have claimed an 
API key from my Pine Cone 

780
00:36:53,120 --> 00:36:55,200
account and I've exported it in 
my environment. 

781
00:36:55,200 --> 00:36:57,760
Before I go here, I'm not going 
to show you that for obvious 

782
00:36:57,760 --> 00:37:01,040
reasons, but be aware that 
you'll have to export your API 

783
00:37:01,040 --> 00:37:03,120
key. 
You can get one for free at Pine

784
00:37:03,120 --> 00:37:04,480
cone dot IO. 
Cool. 

785
00:37:04,480 --> 00:37:07,160
So what I'm going to do is I'm 
going to find the help command 

786
00:37:07,160 --> 00:37:09,280
from Pine Cone. 
So I'm trying to type help here.

787
00:37:09,280 --> 00:37:12,280
And you can see we have the 
slash command which explains how

788
00:37:12,280 --> 00:37:14,080
to use Pine Cone inside Cloud 
Code. 

789
00:37:14,080 --> 00:37:15,400
So I'm going to go ahead and hit
that. 

790
00:37:15,840 --> 00:37:18,960
And what it's going to do is 
it's going to check that an API 

791
00:37:18,960 --> 00:37:21,680
key has been set and that we can
list our indexes. 

792
00:37:22,080 --> 00:37:24,480
So we can see that we've invoked
the help command. 

793
00:37:24,920 --> 00:37:26,880
We have some features we have 
access to. 

794
00:37:27,080 --> 00:37:29,240
We have a bunch of Pine Cone 
assisted commands. 

795
00:37:29,240 --> 00:37:31,120
I'll go into that in a few 
seconds. 

796
00:37:31,320 --> 00:37:34,200
We have access to our Pine Cone 
MCP server, which means we can 

797
00:37:34,200 --> 00:37:38,280
create indexes, do semantic 
search, list our indexes, all 

798
00:37:38,280 --> 00:37:41,440
that sort of fancy fun stuff, 
some slash commands which let us

799
00:37:41,440 --> 00:37:44,760
explore our data, and we have 
some setup instructions. 

800
00:37:45,440 --> 00:37:48,760
So it looks like we were able to
successfully list our indexes. 

801
00:37:48,760 --> 00:37:50,520
Here. 
I'll use Control O to kind of 

802
00:37:50,520 --> 00:37:55,920
show you what that looks like. 
This is an output from our MCP 

803
00:37:55,920 --> 00:37:58,960
server, which shows us the 
indexes that we have access to. 

804
00:37:59,280 --> 00:38:02,080
So it looks like we're all set 
up and we can start exploring 

805
00:38:02,080 --> 00:38:05,240
our indices. 
So I have some fun indices that 

806
00:38:05,240 --> 00:38:07,120
I've set up for us to kind of 
look through. 

807
00:38:07,160 --> 00:38:09,600
I have a few here. 
I have one that's just like a 

808
00:38:09,600 --> 00:38:14,200
simple sentence search. 
I have one that indexes 50,000 

809
00:38:14,200 --> 00:38:17,160
Wikipedia pages of North 
American birds with two 

810
00:38:17,160 --> 00:38:19,320
different methods. 
So we can kind of understand 

811
00:38:19,320 --> 00:38:23,200
like how semantic search works 
in one domain and how sparse 

812
00:38:23,200 --> 00:38:27,400
search works in another domain. 
And I also have a recommendation

813
00:38:27,400 --> 00:38:32,120
system example that kind of 
shows us how someone could do 

814
00:38:32,120 --> 00:38:35,600
recommendation systems with food
as as kind of the domain. 

815
00:38:35,840 --> 00:38:39,720
So what I will do is I will 
start with this first a simple 

816
00:38:39,720 --> 00:38:42,240
semantic search example so we 
can get our bearings and then we

817
00:38:42,240 --> 00:38:45,400
can do the bird search one or 
the food one, depending on what 

818
00:38:45,400 --> 00:38:47,960
you're interested in, Mike. 
So we got we got to do bird is 

819
00:38:47,960 --> 00:38:49,960
search for a puffin. 
OK, cool. 

820
00:38:49,960 --> 00:38:50,760
Awesome. 
Hell yeah. 

821
00:38:51,080 --> 00:38:53,120
Mike Bird, of course. 
My gosh, how could? 

822
00:38:53,120 --> 00:38:56,200
I it all works out. 
Even not notice. 

823
00:38:56,200 --> 00:39:00,320
So let me go ahead and query the
semantic search index that I'm 

824
00:39:00,320 --> 00:39:02,640
describing. 
So this is an index of English 

825
00:39:02,640 --> 00:39:07,040
sentences that are using that 
are about going to the park and 

826
00:39:07,040 --> 00:39:09,640
parking your car. 
So a very simple example of 

827
00:39:09,640 --> 00:39:12,840
semantic search is just trying 
to disambiguate the fact that 

828
00:39:12,840 --> 00:39:14,760
park is being used in two 
different contexts. 

829
00:39:14,760 --> 00:39:17,680
So I'm going to go ahead and 
say, OK, I want to query my 

830
00:39:17,960 --> 00:39:24,160
semantic search index and I want
to say where do I have to park? 

831
00:39:24,960 --> 00:39:29,320
Right as I'm running the slash 
command, Claude is going to 

832
00:39:29,320 --> 00:39:31,920
figure out what index I'm kind 
of working with. 

833
00:39:32,040 --> 00:39:35,240
It's looking what parameters it 
needs to kind of pass in, and 

834
00:39:35,240 --> 00:39:36,800
it's gonna go ahead and do that 
search. 

835
00:39:36,800 --> 00:39:40,040
This is all MCP calls. 
They're being enabled by Claude.

836
00:39:40,360 --> 00:39:42,480
And now we're getting some 
outputs back and we can see 

837
00:39:42,480 --> 00:39:45,040
what's kind of going on. 
So I said, where do I have to 

838
00:39:45,040 --> 00:39:47,200
park? 
And I got a bunch of queries 

839
00:39:47,200 --> 00:39:50,600
that are coming back that are 
about parking my car, right? 

840
00:39:50,920 --> 00:39:54,480
And this is interesting because 
we're not at any point doing a 

841
00:39:55,640 --> 00:39:58,640
mapping over the query that I'm 
looking at and the text that's 

842
00:39:58,640 --> 00:40:00,280
coming back. 
This is just the meaning of the 

843
00:40:00,280 --> 00:40:02,720
sentences. 
And I can demonstrate that by 

844
00:40:02,720 --> 00:40:04,120
asking a different kind of 
query. 

845
00:40:04,120 --> 00:40:07,720
So instead of where do I have to
park, I can say something like I

846
00:40:07,720 --> 00:40:11,800
want to play and I can see what 
sort of sentences kind of have 

847
00:40:12,600 --> 00:40:15,680
kind of come out of that. 
So I can say I want to play at 

848
00:40:15,680 --> 00:40:18,800
the park and the same type of 
operation will happen again. 

849
00:40:18,800 --> 00:40:20,960
It's going to search the same 
index because it's saw that I'm 

850
00:40:21,040 --> 00:40:22,880
kind of passing in the same 
query again. 

851
00:40:23,040 --> 00:40:25,840
And it'll generate a list of the
things that kind of come back. 

852
00:40:25,840 --> 00:40:28,440
So I went to the park to play 
tennis, I'm playing in the park,

853
00:40:28,440 --> 00:40:30,520
so on so forth. 
And we didn't get any of the 

854
00:40:30,520 --> 00:40:33,280
other sentences that I showed 
you prior of like where can I 

855
00:40:33,280 --> 00:40:36,800
park, blah, blah, blah. 
So this is like a very like one 

856
00:40:36,800 --> 00:40:39,240
O one example. 
I like to show people to like 

857
00:40:39,240 --> 00:40:43,000
really nail this point. 
At no point are we actually like

858
00:40:43,160 --> 00:40:46,080
looking at the words that we're 
using inside the documents or 

859
00:40:46,080 --> 00:40:48,240
the query. 
This is all embedding models, 

860
00:40:48,240 --> 00:40:51,280
like all the way down, which is 
crazy to think about. 

861
00:40:52,400 --> 00:40:55,240
So now that I've shown you that,
Are you ready for like this bird

862
00:40:55,240 --> 00:40:58,080
kind of search example? 
Yeah, absolutely awesome. 

863
00:40:58,080 --> 00:41:02,520
So I actually built an 
application that basically shows

864
00:41:02,520 --> 00:41:07,000
users how to index information 
in different ways in order to 

865
00:41:07,000 --> 00:41:09,640
expose different properties of 
semantic search. 

866
00:41:09,920 --> 00:41:13,160
So we've been talking a lot 
about like representing concepts

867
00:41:13,160 --> 00:41:16,400
and ideas, and a lot of that 
comes from using what's called a

868
00:41:16,400 --> 00:41:19,360
dense embedding model. 
These are embedding models that 

869
00:41:19,360 --> 00:41:22,400
take in documents and produce 
embeddings based on the entirety

870
00:41:22,400 --> 00:41:25,200
of the meeting in the document, 
and less so on the words that 

871
00:41:25,200 --> 00:41:28,680
are being used. 
Sparse embedding models place a 

872
00:41:28,680 --> 00:41:32,520
greater emphasis on the words 
that are being used rather than 

873
00:41:32,520 --> 00:41:35,800
the entirety of the meaning of 
the document, but they reweight 

874
00:41:35,880 --> 00:41:38,800
the influence of the words based
on the other things that are 

875
00:41:38,800 --> 00:41:40,480
kind of occurring inside that 
document. 

876
00:41:40,480 --> 00:41:44,120
So why does this matter, right? 
You can imagine situations where

877
00:41:44,120 --> 00:41:47,160
you'd want to do a sort of 
keyword search, but you also 

878
00:41:47,160 --> 00:41:50,360
want to preserve some of the 
meaning of the query that you're

879
00:41:50,360 --> 00:41:53,880
passing in in order to adjust, 
like what words kind of 

880
00:41:53,880 --> 00:41:55,480
happened. 
We were doing a little bit of 

881
00:41:55,480 --> 00:41:58,720
that prior where I was saying, 
hey, I want to go play at the 

882
00:41:58,720 --> 00:42:00,600
park. 
So we want to find sentences 

883
00:42:00,600 --> 00:42:04,320
that include the word park, 
which is relevant for us, but we

884
00:42:04,320 --> 00:42:07,360
care about how the word park is 
being used in that context. 

885
00:42:07,360 --> 00:42:10,080
So we still need an embedding 
model to inform the 

886
00:42:10,080 --> 00:42:13,320
representation of that word and 
change it because of the context

887
00:42:13,320 --> 00:42:16,720
being shown there. 
So I have an example of having 

888
00:42:16,720 --> 00:42:20,240
indexed the same data with these
two different methods, and I'll 

889
00:42:20,240 --> 00:42:23,160
show you the differences in the 
data based on those querying 

890
00:42:23,160 --> 00:42:26,480
methods. 
So what I'm gonna say is, hey, 

891
00:42:26,480 --> 00:42:32,360
Claude, can you list my indexes 
that are named Ferd? 

892
00:42:33,400 --> 00:42:37,120
And those are going to be my 
bird dense index, my bird sparse

893
00:42:37,120 --> 00:42:38,960
index. 
So I'm going to go ahead and 

894
00:42:38,960 --> 00:42:42,120
have Claude list those indexes 
and show me which ones they are 

895
00:42:42,120 --> 00:42:43,720
before we can actually query 
them. 

896
00:42:44,200 --> 00:42:47,520
OK, great. 
So I have a sparse bird index 

897
00:42:47,760 --> 00:42:50,800
and a dense bird index, and I 
have a Bestbatch 25 one, which 

898
00:42:50,800 --> 00:42:53,240
is just regular keyword search. 
We won't talk about that one 

899
00:42:53,240 --> 00:42:56,120
here because we can't easily 
query that one with the way all 

900
00:42:56,120 --> 00:42:58,080
this is set up. 
So we're going to run our 

901
00:42:58,080 --> 00:43:00,760
experiments on the sparse bird 
search and the dense bird 

902
00:43:00,760 --> 00:43:02,920
search. 
So I have some prompts here. 

903
00:43:03,400 --> 00:43:06,840
This one is large predatory bird
that hunts small mammals at 

904
00:43:06,840 --> 00:43:09,800
night. 
OK, before I send this out, what

905
00:43:09,800 --> 00:43:12,760
kinds of data do you think are 
going to return if we were to 

906
00:43:12,760 --> 00:43:16,640
ask this in like Wikipedia 
search for example, or in a very

907
00:43:16,640 --> 00:43:19,640
keyword based way like what do 
you what do you think we're 

908
00:43:19,640 --> 00:43:22,240
going to see like in the return 
information? 

909
00:43:22,320 --> 00:43:25,680
Sorry, like an owl. 
Yeah, it could be like an owl. 

910
00:43:25,680 --> 00:43:29,880
It could be probably a paragraph
about an owl describing its 

911
00:43:29,880 --> 00:43:32,360
hunting tactics. 
Right, Right. 

912
00:43:33,000 --> 00:43:35,320
But what's interesting about 
this sentence is that we're 

913
00:43:35,320 --> 00:43:38,200
never describing the name of the
bird, right? 

914
00:43:38,200 --> 00:43:41,120
We're describing the property 
and the idea of the bird. 

915
00:43:41,400 --> 00:43:44,960
So my hunch is that this will 
probably work better with a 

916
00:43:44,960 --> 00:43:47,200
dense search over a sparse 
search because we don't really 

917
00:43:47,200 --> 00:43:50,240
care about like, specific 
terminology or information. 

918
00:43:50,400 --> 00:43:53,160
There's this idea of this bird 
that we're trying to look for. 

919
00:43:53,360 --> 00:43:54,760
So I'm going to say, hey, 
Claude. 

920
00:43:55,480 --> 00:43:59,920
Hey, Claude, I'm going to, I'm 
going to use my query slash 

921
00:43:59,920 --> 00:44:01,360
commands. 
So I'm going to go down and hit 

922
00:44:01,360 --> 00:44:07,200
tab and I'm going to say, hey 
Claude, query this sentence in 

923
00:44:07,200 --> 00:44:14,120
both sparse and dense indexes 
and return the results. 

924
00:44:14,840 --> 00:44:17,920
And then I'll paste the sentence
and we'll just compare the 

925
00:44:17,920 --> 00:44:21,440
results together and see which 
one is better, one or the other.

926
00:44:22,000 --> 00:44:24,760
So we can see it's making a 
bunch of MCP calls and then 

927
00:44:24,760 --> 00:44:27,200
it'll format both results for 
us. 

928
00:44:27,360 --> 00:44:31,560
It's searching both of these 
things and it's inputting that 

929
00:44:31,560 --> 00:44:34,440
query that we care about. 
And now it's it's honking. 

930
00:44:34,440 --> 00:44:38,080
That's so fun. 
Pretty odd point there. 

931
00:44:39,120 --> 00:44:41,320
And then we'll just see what 
kind of pops up here. 

932
00:44:41,840 --> 00:44:45,360
OK, awesome. 
So let's look at the search 

933
00:44:45,360 --> 00:44:48,200
results here. 
So for the 1st result we did 

934
00:44:48,400 --> 00:44:52,160
using pine Cone sparse, we got 
the Costa Rican pygmy Owl and 

935
00:44:52,160 --> 00:44:54,680
the details that we got is that 
it forges day and night and 

936
00:44:54,680 --> 00:44:57,560
hunts from a low perch. 
The diet includes birds, small 

937
00:44:57,560 --> 00:44:59,400
mammals, vertebrae, and 
arthropods. 

938
00:44:59,840 --> 00:45:03,520
And then with the dense results,
we got something that 

939
00:45:03,520 --> 00:45:05,800
specifically preys on 
hummingbirds, which is 

940
00:45:05,800 --> 00:45:08,480
interesting because we didn't 
specify hummingbirds at all, 

941
00:45:08,640 --> 00:45:11,560
right? 
And it's also a bird of prey. 

942
00:45:11,560 --> 00:45:15,320
We got the Western screech owl, 
which is maybe kind of large, 

943
00:45:15,320 --> 00:45:18,600
and it kind of eats these other 
things like mice, rats, flying 

944
00:45:18,600 --> 00:45:20,160
squirrels, bats, so on, so 
forth. 

945
00:45:20,440 --> 00:45:22,560
We're not seeing the raw 
information. 

946
00:45:22,560 --> 00:45:24,880
I think Claude is helping us 
here by kind of cleaning up the 

947
00:45:24,880 --> 00:45:26,840
return results that are being 
generated. 

948
00:45:27,120 --> 00:45:31,080
But you can see that we're kind 
of extracting different ideas or

949
00:45:31,080 --> 00:45:33,280
properties because of what's 
going on here. 

950
00:45:33,280 --> 00:45:36,640
And let me go ahead and just 
like show you the results. 

951
00:45:37,480 --> 00:45:42,800
So let's look at our sparse 
search here. 

952
00:45:44,000 --> 00:45:47,680
So we got the Costa Rican pygmy 
owl forges both day and night. 

953
00:45:47,960 --> 00:45:50,960
So we got knight in our query 
here, and we got Knight kind of 

954
00:45:50,960 --> 00:45:52,840
returned here. 
It hunts from a low perch and it

955
00:45:52,840 --> 00:45:55,600
takes prey. 
It's dye has not been defined in

956
00:45:55,600 --> 00:45:58,040
detail, but it's include birds, 
small mammals, and other 

957
00:45:58,040 --> 00:45:59,680
verbrae. 
So we have another hit here on 

958
00:45:59,680 --> 00:46:01,840
small mammals, which is why we 
got this result back. 

959
00:46:02,280 --> 00:46:04,320
Now let's look at our dense 
search. 

960
00:46:04,880 --> 00:46:06,680
Let's see the 1st result that's 
coming back. 

961
00:46:06,960 --> 00:46:10,920
It's diurnal, it's found 
primarily on humid forests, and 

962
00:46:10,920 --> 00:46:12,440
it's known to prey on 
hummingbirds. 

963
00:46:12,440 --> 00:46:15,080
So you can immediately see the 
difference in the quality of the

964
00:46:15,080 --> 00:46:16,680
results, right? 
So here. 

965
00:46:16,720 --> 00:46:19,240
We got sorry even describes it 
as a small bird rather. 

966
00:46:19,240 --> 00:46:24,040
Than exactly exactly, which is 
probably just something that is 

967
00:46:24,040 --> 00:46:27,120
not as relevant, right? 
Like it probably got confused by

968
00:46:27,120 --> 00:46:30,080
this idea of like praying on 
hummingbirds versus being small.

969
00:46:30,080 --> 00:46:32,520
There's like too many attributes
that we're kind of searching 

970
00:46:32,520 --> 00:46:36,360
here, but it's still really 
interesting that we got a result

971
00:46:36,360 --> 00:46:39,640
that's like, hey, this is eating
like really small hopping 

972
00:46:39,680 --> 00:46:41,800
hummingbirds, right? 
There's like something like 

973
00:46:41,800 --> 00:46:45,480
valuable or interesting there. 
This second one is probably way 

974
00:46:45,480 --> 00:46:47,680
more relevant. 
So these nocturnal birds wait on

975
00:46:47,680 --> 00:46:50,320
purchase to Snoop down. 
They catch insects in flight, 

976
00:46:50,320 --> 00:46:53,680
not mammals, but they do eat 
mice, rats, flank squirrels and 

977
00:46:53,680 --> 00:46:56,800
bats, other birds, and they're 
opportunistic predators. 

978
00:46:56,800 --> 00:46:59,760
So this result would probably be
a lot more relevant than some of

979
00:46:59,760 --> 00:47:01,640
the other things that we got 
with the other model. 

980
00:47:02,320 --> 00:47:04,640
OK. 
You did say you wanted to search

981
00:47:04,800 --> 00:47:07,600
information about puffins. 
So I wonder if you wanted to 

982
00:47:07,600 --> 00:47:10,520
take a stab at forming a query 
and seeing which index gives you

983
00:47:10,520 --> 00:47:12,440
better results here. 
Is there anything you want to 

984
00:47:12,440 --> 00:47:14,320
know about puffins? 
And then we could kind of like 

985
00:47:14,320 --> 00:47:16,280
go from there. 
Migratory Habits. 

986
00:47:16,720 --> 00:47:19,320
OK. 
So what are some migratory 

987
00:47:19,320 --> 00:47:24,280
habits of puffins, right? 
And are there puffins in North 

988
00:47:24,280 --> 00:47:25,800
America? 
I hope there are actually. 

989
00:47:26,680 --> 00:47:28,680
Different ones on the West and 
the east so. 

990
00:47:28,680 --> 00:47:29,800
Oh, really? 
Wow. 

991
00:47:29,800 --> 00:47:32,240
OK, we got, we got an expert. 
OK, awesome. 

992
00:47:32,240 --> 00:47:36,440
So I'm thinking that this would 
probably perform better in our 

993
00:47:36,440 --> 00:47:39,720
sparse search because we're 
really keying in on this 

994
00:47:39,720 --> 00:47:43,240
terminology of migratory habits.
But we might get some sentences 

995
00:47:43,240 --> 00:47:45,280
that are relevant as well. 
So let let's go ahead and 

996
00:47:45,280 --> 00:47:49,840
inspect the results ourselves. 
So here's our sparse search. 

997
00:47:50,280 --> 00:47:52,320
First thing we got was the 
Ellen's Hummingbird. 

998
00:47:52,320 --> 00:47:54,800
All right, that's not that's not
what we're looking for. 

999
00:47:55,880 --> 00:47:59,440
And it looks like we're getting 
migratory patterns back. 

1000
00:47:59,440 --> 00:48:03,560
So this probably over indexed on
this idea of migratory patterns 

1001
00:48:03,560 --> 00:48:05,960
and not on the actual bird that 
we're interested in. 

1002
00:48:05,960 --> 00:48:10,120
However, the second block that 
we get back has something about 

1003
00:48:10,120 --> 00:48:14,200
the Atlantic puffin, which is 
pretty awesome, but it's more 

1004
00:48:14,200 --> 00:48:17,280
about the characteristics of 
that puffin rather than the 

1005
00:48:17,480 --> 00:48:19,000
results that we were interested 
in. 

1006
00:48:20,320 --> 00:48:22,760
Let's go back here. 
OK, great. 

1007
00:48:22,760 --> 00:48:25,680
So those are our first two kind 
of like sparse results, right? 

1008
00:48:25,680 --> 00:48:30,280
We got another puffin kind of 
section here, probably another 

1009
00:48:30,280 --> 00:48:33,440
part of the same article, and 
then we got tufted puffins back 

1010
00:48:33,440 --> 00:48:35,560
here. 
So it looks like we got 

1011
00:48:35,560 --> 00:48:37,960
information about puffins, but 
not necessarily about the 

1012
00:48:37,960 --> 00:48:39,920
migratory habits. 
So that could just be a data set

1013
00:48:39,920 --> 00:48:42,800
issue, like we might not have 
good information there. 

1014
00:48:43,360 --> 00:48:46,880
In the dense search. 
We got information about puffins

1015
00:48:46,880 --> 00:48:48,320
right away, which is really 
cool. 

1016
00:48:48,320 --> 00:48:49,920
It looks like we have the 
Atlantic puffin. 

1017
00:48:50,280 --> 00:48:53,560
We got the doc horn puffin, 
which we didn't see in the 

1018
00:48:53,560 --> 00:48:56,480
opposite results, we got the 
Atlantic puffin again, but the 

1019
00:48:56,480 --> 00:48:59,560
sentence is more about like 
their habits that they kind of 

1020
00:48:59,560 --> 00:49:01,880
do in danger. 
And then we got one sentence 

1021
00:49:01,880 --> 00:49:06,000
about migratory patterns from a 
different bird, and then the 

1022
00:49:06,000 --> 00:49:08,960
tufted puffins, which also are 
not really talking about their 

1023
00:49:08,960 --> 00:49:10,640
migratory patterns. 
So we learned something 

1024
00:49:10,640 --> 00:49:12,600
interesting here, right? 
Like maybe there's something 

1025
00:49:12,600 --> 00:49:15,920
about the way we're chunking our
data that's just not capturing 

1026
00:49:15,920 --> 00:49:19,240
this migratory information at 
the same time as it's capturing 

1027
00:49:19,240 --> 00:49:21,280
information about the animal 
that we're interested in. 

1028
00:49:21,720 --> 00:49:24,800
But we were able to kind of see 
the different results that each 

1029
00:49:24,800 --> 00:49:28,000
search method produced and maybe
make actionable choices based on

1030
00:49:28,000 --> 00:49:30,240
that. 
So, OK, I know that we're not 

1031
00:49:30,240 --> 00:49:32,840
supporting migratory queries 
that well, so maybe we'll have 

1032
00:49:32,840 --> 00:49:36,280
to think about how we want to 
chunk around that specific idea 

1033
00:49:36,280 --> 00:49:39,360
in general, especially if a lot 
of our bird search users are 

1034
00:49:39,360 --> 00:49:41,160
trying to find that sort of 
information. 

1035
00:49:42,240 --> 00:49:43,800
But yeah, that's, that's kind of
the idea. 

1036
00:49:43,800 --> 00:49:46,120
That's that's kind of how the 
plug in works. 

1037
00:49:46,120 --> 00:49:47,960
I have. 
I have a few more things to show

1038
00:49:48,000 --> 00:49:49,840
you, but I know we're kind of 
running up on time, so I don't 

1039
00:49:49,840 --> 00:49:52,840
want to take up too much of. 
It yeah, yeah, no happy to to 

1040
00:49:52,880 --> 00:49:54,440
keep going with this. 
OK, cool. 

1041
00:49:54,440 --> 00:49:57,600
Awesome. 
So that's kind of our search 

1042
00:49:57,840 --> 00:50:00,360
demos there. 
There's another functionality of

1043
00:50:00,360 --> 00:50:03,840
Pine Cone of our Pine Cone plug 
in that people might be 

1044
00:50:03,840 --> 00:50:07,200
interested in if they want to do
the whole retrieval augmented 

1045
00:50:07,200 --> 00:50:10,040
generation pipeline without 
having to build it entirely 

1046
00:50:10,040 --> 00:50:12,400
themselves. 
So we have a product called Pine

1047
00:50:12,400 --> 00:50:15,280
Cone Assistant, which sits on 
top of our Pine Cone database, 

1048
00:50:15,280 --> 00:50:17,680
which is like our trusted 
knowledge kind of infrastructure

1049
00:50:17,680 --> 00:50:20,160
layer. 
And the idea is a lot of people 

1050
00:50:20,160 --> 00:50:23,080
want to chunk their data, but 
they don't know how. 

1051
00:50:23,120 --> 00:50:27,080
Or a lot of people have their 
data in PDFs or text files or 

1052
00:50:27,080 --> 00:50:29,520
all that sort of stuff, and they
just want to start asking 

1053
00:50:29,520 --> 00:50:32,680
questions over that data. 
So we created a way for you to 

1054
00:50:32,680 --> 00:50:37,000
instantly chat with your 
documents, upload them, and then

1055
00:50:37,000 --> 00:50:41,240
get generated answers off of 
them along with citations and 

1056
00:50:41,240 --> 00:50:43,960
references back to the documents
that you're kind of looking at. 

1057
00:50:44,160 --> 00:50:46,280
So you don't have to build this 
whole pipeline yourself. 

1058
00:50:46,280 --> 00:50:49,600
You can just say, hey, here's 
all my documents, go put them 

1059
00:50:49,600 --> 00:50:52,360
inside Pine Cone and then let me
query over that data. 

1060
00:50:52,360 --> 00:50:55,960
But we didn't stop there. 
We also handled the query 

1061
00:50:55,960 --> 00:50:57,760
planning and orchestration for 
you. 

1062
00:50:57,760 --> 00:51:00,600
So it's not just a simple 
semantic search over the 

1063
00:51:00,600 --> 00:51:04,120
information that you care about.
We do a lot of sub queries that 

1064
00:51:04,120 --> 00:51:05,920
breakdown the question that 
you're asking. 

1065
00:51:05,920 --> 00:51:08,840
We search all of those results, 
we re rank them, we represent 

1066
00:51:09,240 --> 00:51:13,040
the results back to a language 
model like GPT or Claude, and 

1067
00:51:13,040 --> 00:51:14,840
then we generate a response for 
you. 

1068
00:51:14,920 --> 00:51:18,320
And then you can even see how 
the model generated the response

1069
00:51:18,320 --> 00:51:20,080
based on the data that it 
searched over. 

1070
00:51:20,640 --> 00:51:23,360
Now there's a lot of ways to 
make a pine cone assistant. 

1071
00:51:23,360 --> 00:51:26,880
You can make one inside of our 
console, but we were able to, I 

1072
00:51:26,880 --> 00:51:30,400
was able to kind of create a no 
code experience inside the 

1073
00:51:30,400 --> 00:51:32,120
plugin. 
So I can just tell Claude to 

1074
00:51:32,120 --> 00:51:34,680
make my assistant for me upload 
these documents and start 

1075
00:51:34,680 --> 00:51:36,160
chatting with it. 
So that's what we're going to 

1076
00:51:36,160 --> 00:51:40,200
try to do here. 
So I have this folder here 

1077
00:51:40,960 --> 00:51:44,480
called Assistant Documents and 
I've uploaded, I've put in a 

1078
00:51:44,480 --> 00:51:48,400
bunch of PDFs in this folder. 
They're all about different 

1079
00:51:48,400 --> 00:51:52,240
kinds of AI research. 
So 1 is about how people write 

1080
00:51:52,240 --> 00:51:54,360
agent dot MD files. 
If you're familiar. 

1081
00:51:54,360 --> 00:51:56,800
They kind of help control the 
generation of the code that 

1082
00:51:56,800 --> 00:52:00,160
you're creating. 1 is a research
recent research paper. 

1083
00:52:00,160 --> 00:52:03,200
Oops, don't want to open that 
inside VS Code. 1 is a research,

1084
00:52:03,320 --> 00:52:08,400
a recent research paper that 
discusses skill formation over 

1085
00:52:08,480 --> 00:52:10,920
people that are using code 
generating models that Anthropic

1086
00:52:10,920 --> 00:52:14,040
kind of produced. 
And then I have, and then I have

1087
00:52:14,160 --> 00:52:17,720
research papers and articles 
that we wrote, articles that we 

1088
00:52:17,720 --> 00:52:21,120
wrote and research papers that 
we gathered relating to sparse 

1089
00:52:21,120 --> 00:52:23,760
and dense search. 
So I can say that my goal is, 

1090
00:52:23,760 --> 00:52:26,120
hey, I want to combine all of 
these sources and I want to 

1091
00:52:26,120 --> 00:52:30,240
think about how I can write like
an agents dot MD or a skill dot 

1092
00:52:30,240 --> 00:52:33,440
MD file to teach developers how 
to do sparse or dense search. 

1093
00:52:33,440 --> 00:52:35,720
Like, here's all this 
information, synthesize an 

1094
00:52:35,720 --> 00:52:38,080
answer for me. 
So what I'm going to do is I'm 

1095
00:52:38,080 --> 00:52:42,720
going to 1st ask Claude to 
create an assistant and I'm 

1096
00:52:42,720 --> 00:52:46,200
going to ask it to upload all 
this data from this document 

1097
00:52:46,200 --> 00:52:47,440
here. 
So I'm going to say, hey, 

1098
00:52:47,440 --> 00:52:53,680
Claude, using the assistant 
skill, can you please create an 

1099
00:52:53,680 --> 00:52:57,600
assistant with the I'm going to 
add my folder. 

1100
00:52:57,800 --> 00:53:01,920
Oops. 
Typing, Typing is hard docs 

1101
00:53:04,040 --> 00:53:07,320
folder, and we'll go ahead and 
see what happens. 

1102
00:53:07,320 --> 00:53:10,800
So ideally, we've created a 
skill and a bunch of slash 

1103
00:53:10,800 --> 00:53:13,680
commands kind of make it easy 
for Cloud to create assistance 

1104
00:53:13,680 --> 00:53:15,480
to chat with them, upload 
documents. 

1105
00:53:16,200 --> 00:53:18,400
And you can see that it found 
our Pine Cone skill. 

1106
00:53:18,720 --> 00:53:21,440
And what this skill's gonna do 
is that it's gonna create, 

1107
00:53:21,440 --> 00:53:25,080
manage, and interact with Pine 
Cone assistance, so on and so 

1108
00:53:25,080 --> 00:53:26,440
forth. 
So I'm gonna go ahead and click 

1109
00:53:26,440 --> 00:53:28,120
yes. 
It's gonna load information 

1110
00:53:28,120 --> 00:53:31,040
about what it's able to do, and 
now it's going to create our 

1111
00:53:31,040 --> 00:53:34,520
assistant and it'll probably ask
me for a name. 

1112
00:53:34,920 --> 00:53:36,680
It created the name for me. 
That's fine. 

1113
00:53:36,720 --> 00:53:39,040
We'll call it the tool Use 
Podcast Stocks Assistant. 

1114
00:53:39,040 --> 00:53:41,520
All right, that's cool. 
And now it's going to go ahead 

1115
00:53:41,520 --> 00:53:44,680
and upload the document. 
So what's happening under the 

1116
00:53:44,680 --> 00:53:49,000
hood here is that inside of our 
plugin, we've packaged the skill

1117
00:53:49,120 --> 00:53:53,280
slash commands and a bunch of 
scripts that Claude can use in 

1118
00:53:53,280 --> 00:53:57,160
order to say, upload these files
to our assistant or query them. 

1119
00:53:57,480 --> 00:54:00,960
And this property is really nice
because we don't want Claude to 

1120
00:54:00,960 --> 00:54:04,400
generate the code to upsert the 
data every single time because 

1121
00:54:04,400 --> 00:54:06,120
that's a lot of tokens that are 
being generated. 

1122
00:54:06,120 --> 00:54:09,360
So instead, we're giving access 
to tools that all it has to do 

1123
00:54:09,360 --> 00:54:12,040
is generate the call over so 
that you're preserving a lot of 

1124
00:54:12,040 --> 00:54:15,040
tokens and the script can do a 
lot of the hard work. 

1125
00:54:15,560 --> 00:54:18,640
This will probably take maybe 
like 3040 seconds because 

1126
00:54:18,640 --> 00:54:20,320
there's a lot of different 
documents that are kind of 

1127
00:54:20,320 --> 00:54:24,240
uploading while that's going. 
I can show you where they're 

1128
00:54:24,240 --> 00:54:27,520
being uploaded inside the 
console if you'd like to see 

1129
00:54:28,360 --> 00:54:30,520
Yeah, perfect. 
I'm inside the pine cone console

1130
00:54:30,520 --> 00:54:33,360
and I can hit tool use podcast 
docs and I'll look inside my 

1131
00:54:33,360 --> 00:54:36,280
folder and you can see that we 
have all of our files kind of 

1132
00:54:36,280 --> 00:54:39,640
uploaded inside of our assistant
here, which is really awesome. 

1133
00:54:40,040 --> 00:54:42,720
So that's great. 
Our assistant was made. 

1134
00:54:42,720 --> 00:54:45,360
Our documents are here. 
We can start querying over them 

1135
00:54:45,360 --> 00:54:47,440
and see what's what. 
Awesome. 

1136
00:54:47,440 --> 00:54:51,000
So we're here and now we can 
kind of ask this question that 

1137
00:54:51,000 --> 00:54:55,200
we're thinking about. 
So I pre wrote the prompt 

1138
00:54:55,200 --> 00:54:59,520
because I find that it's always 
helpful to have that information

1139
00:54:59,520 --> 00:55:03,840
kind of there so I can copy 
paste, especially when I'm doing

1140
00:55:04,400 --> 00:55:07,000
live demos. 
So basically what this is gonna 

1141
00:55:07,000 --> 00:55:10,640
say is, hey, what can you tell 
me about skill formation with AI

1142
00:55:10,920 --> 00:55:14,000
and how does that impact agent 
dot MD file creation? 

1143
00:55:14,320 --> 00:55:17,800
What I need to do is create an 
agent dot MD file and skill 

1144
00:55:17,800 --> 00:55:20,080
files that help users make 
better semantic search 

1145
00:55:20,080 --> 00:55:23,600
pipelines, especially ones that 
deal with sparse or multilingual

1146
00:55:23,600 --> 00:55:26,040
data, which are the two 
embedding models that we kind of

1147
00:55:26,360 --> 00:55:28,360
one or two of the three 
embedding models that we 

1148
00:55:28,360 --> 00:55:31,520
support. 
So I'm going to say pine cone, 

1149
00:55:31,560 --> 00:55:34,200
whoops. 
I'm going to say slash pine cone

1150
00:55:35,400 --> 00:55:39,480
assistant chat, and I'm going to
say this query and I'm going to 

1151
00:55:39,480 --> 00:55:42,240
hit enter. 
And you can see that whenever 

1152
00:55:42,240 --> 00:55:45,640
I'm using some of these slash 
commands, there's options to 

1153
00:55:45,640 --> 00:55:48,000
kind of fill in the arguments. 
But I'm just going ahead and 

1154
00:55:48,000 --> 00:55:51,040
putting in my query and I'm 
letting Claude kind of figure 

1155
00:55:51,040 --> 00:55:54,200
out what information you should 
put in each field based on the 

1156
00:55:54,200 --> 00:55:57,320
context of the chat. 
And that's a really powerful way

1157
00:55:57,320 --> 00:56:00,240
of using Claude code because it 
has context to your previous 

1158
00:56:00,240 --> 00:56:03,040
conversation and your workspace 
and the data that we've kind of 

1159
00:56:03,280 --> 00:56:07,200
included with the pine cone plug
in, we can kind of let it figure

1160
00:56:07,200 --> 00:56:10,840
out all that sort of stuff. 
So we can see that I think we're

1161
00:56:10,840 --> 00:56:14,280
getting a response back here. 
Oh yeah, this is super cool. 

1162
00:56:14,640 --> 00:56:17,640
So we're seeing this backward 
because I have to kind of scroll

1163
00:56:17,640 --> 00:56:21,400
up, but I'll show you what the 
response is so we can see the 

1164
00:56:21,400 --> 00:56:23,640
assistant response. 
Skill formation with AI and its 

1165
00:56:23,640 --> 00:56:26,160
implications can be instituted 
through three aspects, the AI 

1166
00:56:26,160 --> 00:56:28,400
and skill development role 
structure, context files. 

1167
00:56:29,800 --> 00:56:32,320
I'm gonna see if. 
Yeah, right. 

1168
00:56:32,320 --> 00:56:34,360
So we have some information here
to complement the agent's 

1169
00:56:34,360 --> 00:56:36,360
identity file, create skill 
files that guide users in 

1170
00:56:36,360 --> 00:56:38,720
billing, semantic search 
pipelines, which can include 

1171
00:56:38,720 --> 00:56:40,600
tutorials on blah, blah, blah, 
blah, blah. 

1172
00:56:40,920 --> 00:56:45,240
And we get page level citations 
for this type of information as 

1173
00:56:45,240 --> 00:56:46,720
well. 
So if we were inside the 

1174
00:56:46,720 --> 00:56:51,320
console, we could actually click
on these pages and we would be 

1175
00:56:51,320 --> 00:56:55,480
able to see inside the console 
like where this information is 

1176
00:56:55,480 --> 00:56:57,960
getting pulled from. 
And so that's super cool. 

1177
00:56:57,960 --> 00:57:00,680
We were just able to do rag with
our documents just like 

1178
00:57:00,680 --> 00:57:04,000
out-of-the-box using cloud code 
without having to write any code

1179
00:57:04,000 --> 00:57:06,640
ourselves, like purely on the 
information that's kind of 

1180
00:57:06,640 --> 00:57:10,480
there. 
But that's the kind of benefit 

1181
00:57:10,480 --> 00:57:12,960
of making using the pine cone 
plug in there. 

1182
00:57:12,960 --> 00:57:14,440
That's the extra thing I wanted 
to show you. 

1183
00:57:14,920 --> 00:57:17,400
Yeah, very cool. 
It's making it so much more 

1184
00:57:17,400 --> 00:57:21,560
accessible and removing a lot of
the questions I know I had that 

1185
00:57:21,560 --> 00:57:23,760
might have been intimidating me 
from getting into using vector 

1186
00:57:23,760 --> 00:57:25,080
databases. 
If I can just have a quick chat 

1187
00:57:25,080 --> 00:57:27,000
with Claude, fire it up, get 
some documents up there, 

1188
00:57:27,600 --> 00:57:29,280
seamless. 
Exactly. 

1189
00:57:29,280 --> 00:57:32,920
And very often people could take
this directly into production 

1190
00:57:32,920 --> 00:57:36,040
because you could interact with 
Pine Cone assistance via an API.

1191
00:57:36,040 --> 00:57:38,320
So if you want to create your 
own web interface, you can just 

1192
00:57:38,560 --> 00:57:40,800
have it pipe into our Pine Cone 
assistant. 

1193
00:57:41,040 --> 00:57:43,800
Or maybe you're like, hey, I 
want apoc really quick. 

1194
00:57:43,800 --> 00:57:46,280
Like I just want to know what it
would be like to like implement 

1195
00:57:46,360 --> 00:57:49,560
an agentic search or like really
simple rag over my documents, 

1196
00:57:50,000 --> 00:57:52,320
throw it into assistant, see 
that it works really well, and 

1197
00:57:52,320 --> 00:57:54,840
then learn how to build it 
yourself using our plugin and 

1198
00:57:54,840 --> 00:57:57,680
the pine cone vector database. 
So it's nice, it has a nice 

1199
00:57:57,680 --> 00:58:00,160
optionality and it's also super 
fun for these kinds of demos. 

1200
00:58:00,160 --> 00:58:03,400
Like we can learn a lot about 
our data really quickly this 

1201
00:58:03,400 --> 00:58:04,240
way. 
Absolutely. 

1202
00:58:04,240 --> 00:58:05,360
This was awesome. 
Arjud. 

1203
00:58:05,360 --> 00:58:08,360
Really appreciate coming on to 
teach us about vector databases 

1204
00:58:08,360 --> 00:58:10,280
and Pine cone. 
Before I let you go, is there 

1205
00:58:10,280 --> 00:58:11,280
anything you'd like the audience
to? 

1206
00:58:11,640 --> 00:58:14,360
Know hey, you can always find me
on LinkedIn or you can find me 

1207
00:58:14,360 --> 00:58:17,040
inside the pine cone discord. 
Just search my name, Arjun 

1208
00:58:17,040 --> 00:58:19,200
Patel, a pine cone. 
You'll probably find me. 

1209
00:58:19,280 --> 00:58:22,240
Feel free to shoot me a message 
or interact with me there or 

1210
00:58:22,240 --> 00:58:24,320
inside the disk inside our Pine 
cone Discord. 

1211
00:58:24,480 --> 00:58:25,680
It was awesome hanging out with 
you, Mike. 

1212
00:58:25,680 --> 00:58:27,640
This was so much fun. 
Thank you for listening to my 

1213
00:58:27,640 --> 00:58:30,720
conversation with Arjun Patel. 
I really wanted to make sure 

1214
00:58:30,720 --> 00:58:34,600
that both you and I understood 
the need for vector databases in

1215
00:58:34,600 --> 00:58:37,240
this point in time. 
They were all the rage before 

1216
00:58:37,240 --> 00:58:40,120
and kind of fell out of fashion,
but there's still a lot of 

1217
00:58:40,120 --> 00:58:43,560
viable use cases for it. 
So I do hope you experiment with

1218
00:58:43,560 --> 00:58:46,280
it, play around, try different 
chunking strategies, see how it 

1219
00:58:46,280 --> 00:58:47,920
works for you. 
Because there might be a 

1220
00:58:47,920 --> 00:58:50,440
recommendation system that you 
didn't know you could build, 

1221
00:58:50,440 --> 00:58:52,000
which is enabled by this 
technology. 

1222
00:58:52,200 --> 00:58:55,160
Or maybe just your little chat 
bot can get better performance 

1223
00:58:55,160 --> 00:58:57,240
if you get a little more 
meticulous in the way you 

1224
00:58:57,240 --> 00:58:58,960
structure it. 
I hope you enjoy this 

1225
00:58:58,960 --> 00:59:00,520
conversation and I'll see you 
next week.

