1
00:00:06,240 --> 00:00:09,280
Hello and welcome to MongoDB 
Podcast Live. 

2
00:00:09,280 --> 00:00:11,360
I'm Shay McAllister. 
I'm one of the leads on the 

3
00:00:11,360 --> 00:00:13,600
Developer Relations team here at
MongoDB. 

4
00:00:14,000 --> 00:00:17,760
Imagine dealing with 11,000,000 
land documents, historical 

5
00:00:17,760 --> 00:00:21,840
records, contracts, leases, all 
scattered across different 

6
00:00:21,840 --> 00:00:25,640
formats and different systems. 
The challenge is making sense of

7
00:00:25,640 --> 00:00:29,680
all that unstructured data 
efficiently without sinking of 

8
00:00:30,040 --> 00:00:32,600
millions of dollars into manual 
processing. 

9
00:00:32,759 --> 00:00:36,480
That's where AI and large 
language models and MongoDB come

10
00:00:36,480 --> 00:00:39,080
into play. 
So let's get started and to 

11
00:00:39,080 --> 00:00:42,040
discuss all of this. 
I'm really pleased to welcome 

12
00:00:42,080 --> 00:00:47,800
Alex and Andrew from oxy.com 
onto the show and to unpack how 

13
00:00:47,800 --> 00:00:51,440
they did this and why they chose
Mongo DB and what it means to 

14
00:00:51,440 --> 00:00:54,560
the future of AI powered 
document management. 

15
00:00:54,640 --> 00:00:56,800
Alex, Andrew, you're both very 
welcome to the show. 

16
00:00:56,800 --> 00:00:58,840
It's great to have you. 
Thanks Shane for having us all. 

17
00:00:58,920 --> 00:01:01,120
Thanks to you and Mongo DB for 
having us in. 

18
00:01:01,320 --> 00:01:03,320
No, it's great. 
It's great to have both of you 

19
00:01:03,320 --> 00:01:04,120
join. 
It's great. 

20
00:01:04,120 --> 00:01:06,280
As I said, when we were 
preparing to have you both in 

21
00:01:06,280 --> 00:01:09,520
the same room, it's it's always 
much easier to deal with two 

22
00:01:09,520 --> 00:01:12,840
guests in that. 
And before we get started, I 

23
00:01:12,840 --> 00:01:15,120
love that we have both of you 
here. 

24
00:01:15,280 --> 00:01:18,760
You know, could you pick Andrew 
or Alex to go first to talk a 

25
00:01:18,760 --> 00:01:21,880
little bit about your career 
path to date, you know, what you

26
00:01:21,880 --> 00:01:24,720
studied, how you got here and 
how you ended up at Oxy? 

27
00:01:24,840 --> 00:01:28,320
Yeah, I actually started at Oxy 
probably 10 years. 

28
00:01:28,320 --> 00:01:31,680
Ago just about. 11 years ago, 
undergrad at the 

29
00:01:31,680 --> 00:01:34,600
telemachineering at the 
University of Texas at Austate. 

30
00:01:34,760 --> 00:01:39,560
Pretty early on in my career, I,
I went more towards data and 

31
00:01:39,560 --> 00:01:43,040
less towards, it's kind of the 
traditional petroleum reservoir 

32
00:01:43,040 --> 00:01:45,000
production engineer of oil and 
gaps. 

33
00:01:45,000 --> 00:01:48,720
And so I, I kind of managed data
for most of my career. 

34
00:01:48,800 --> 00:01:52,800
About six or seven years ago, I 
did a master's in computer 

35
00:01:52,800 --> 00:01:55,320
science because I felt like that
was really where my interest 

36
00:01:55,320 --> 00:01:58,720
was. 
And since then it's been kind of

37
00:01:58,720 --> 00:02:01,360
a journey. 
I've done a little bit of more 

38
00:02:01,360 --> 00:02:04,200
like embedded systems, 
microcontroller stuff. 

39
00:02:04,320 --> 00:02:07,560
And then a couple years ago I 
joined Alex's team and Alex runs

40
00:02:07,560 --> 00:02:10,080
the AI team here at Oxy. 
I'll let him take over. 

41
00:02:10,600 --> 00:02:12,520
Sure. 
So much like Andrew, I'm kind of

42
00:02:12,520 --> 00:02:16,560
the confused engineer slash data
scientist archetype where yeah, 

43
00:02:16,560 --> 00:02:19,400
major petroleum engineering at 
University of Tulsa and then 

44
00:02:19,480 --> 00:02:23,280
ended up working for Oxy in 2017
and increasingly got more of 

45
00:02:23,280 --> 00:02:26,560
these data focus, AI focus role.
So, yeah, now the team that 

46
00:02:26,560 --> 00:02:29,840
Andrew and I work on, it's it's 
an OPS focused AI development 

47
00:02:29,840 --> 00:02:32,800
team and Oxy and we get to have 
a lot of fun with with 

48
00:02:32,800 --> 00:02:35,720
applications mainly on the oily 
gas like kind of production side

49
00:02:35,720 --> 00:02:38,360
of things, but occasionally 
dabbling with other departments 

50
00:02:38,360 --> 00:02:41,840
like like legal or land and this
case that we're sharing today. 

51
00:02:41,960 --> 00:02:43,720
So yeah, excited to share the 
work. 

52
00:02:43,720 --> 00:02:46,840
I think, you know, the AI space 
people disagree about whether 

53
00:02:46,840 --> 00:02:49,400
it's over hyped, under hyped, 
but I'll just say this, it's a 

54
00:02:49,400 --> 00:02:51,240
lot of fun to work in. 
So we've been, we've been having

55
00:02:51,240 --> 00:02:52,200
a good time. 
Last year or? 

56
00:02:52,560 --> 00:02:54,040
Two excellent. 
Thank you, Alex. 

57
00:02:54,040 --> 00:02:57,280
And I suppose look, before we 
dive into that and the problem 

58
00:02:57,280 --> 00:02:59,400
that you're solving and the 
reason why you're you're on the 

59
00:02:59,400 --> 00:03:02,640
show here today. 
Like I suppose people consider 

60
00:03:02,760 --> 00:03:06,200
the oil and gas industries, as 
you know, Heavy Industries, 

61
00:03:06,200 --> 00:03:09,840
large machineries, large 
refineries, distribution, all of

62
00:03:09,840 --> 00:03:12,680
that sort of stuff. 
They wouldn't particularly maybe

63
00:03:12,680 --> 00:03:16,320
think of it as using data, 
etcetera, etcetera, bar where to

64
00:03:16,320 --> 00:03:19,680
drill or what to do or minerals 
or geological surveys, etcetera.

65
00:03:19,920 --> 00:03:23,240
Tell us a little bit about the 
Oxy as a company and what it 

66
00:03:23,240 --> 00:03:26,800
does and, and I suppose the 
industry as a whole and, and how

67
00:03:27,000 --> 00:03:29,880
you know it's using digital 
these days before we dive into 

68
00:03:29,880 --> 00:03:31,560
the specific solution that 
you've built. 

69
00:03:31,800 --> 00:03:34,200
Maybe I'll take a quick pass and
then if you feel on what I miss,

70
00:03:34,200 --> 00:03:36,960
but I'll just say I mean, oil 
and gas has always been 

71
00:03:37,120 --> 00:03:38,400
technology's been really 
important. 

72
00:03:38,400 --> 00:03:40,280
Oil and gas especially you look 
at, you know, the history of 

73
00:03:40,520 --> 00:03:43,800
seismic data processing and some
of the original large, large 

74
00:03:43,800 --> 00:03:47,280
scale data processing projects 
were were usually seismic focus.

75
00:03:47,280 --> 00:03:51,280
So OXY in particular has taken a
really aggressive stance towards

76
00:03:51,440 --> 00:03:54,400
innovation and AI in general. 
So there's a couple fronts. 1 is

77
00:03:54,520 --> 00:03:58,040
OXY has I think the most 
aggressive direct air capture 

78
00:03:58,040 --> 00:04:00,280
effort of any, any oil and gas 
company right now. 

79
00:04:00,280 --> 00:04:02,200
So obviously that's out of scope
for our talk today. 

80
00:04:02,200 --> 00:04:04,680
But I mean, if you check on the 
website oxy.com, you'll see some

81
00:04:04,800 --> 00:04:07,800
really phenomenal work happening
in the technology space there. 

82
00:04:07,960 --> 00:04:11,480
And then second, with AI, Oxy 
has been really, really with, in

83
00:04:11,480 --> 00:04:14,760
terms of its, its aggressiveness
with applying AII think our 

84
00:04:14,760 --> 00:04:16,560
team, you know, some of the work
we're going to share today as 

85
00:04:16,560 --> 00:04:18,120
part of that. 
And obviously there's a lot of 

86
00:04:18,120 --> 00:04:20,160
other work besides that, that, 
you know, we'd love to talk 

87
00:04:20,160 --> 00:04:22,400
about another time. 
But I, I think Oxy and, and 

88
00:04:22,400 --> 00:04:25,200
certainly I, I agree that the AI
is going to be one of the key 

89
00:04:25,200 --> 00:04:26,880
differentiators. 
When you look which oil 

90
00:04:26,880 --> 00:04:29,320
companies are going to be the 
most successful three to five 

91
00:04:29,320 --> 00:04:30,680
years from now? 
I think it's going to be the 

92
00:04:30,680 --> 00:04:33,320
ones that were most successful 
in applying this to, to 

93
00:04:33,640 --> 00:04:36,160
streamlining the work we do and 
proving the work we do, 

94
00:04:36,280 --> 00:04:38,800
increasing safety, reducing our 
cost per barrel. 

95
00:04:38,800 --> 00:04:40,440
So so I think Oxy's got the 
right. 

96
00:04:40,560 --> 00:04:42,960
Attitude towards this and I 
think that kind of bears out in 

97
00:04:42,960 --> 00:04:45,960
the work that's going on here. 
Yeah, I think Alex pretty well 

98
00:04:45,960 --> 00:04:47,960
covered it. 
Just more on the data front, 

99
00:04:48,040 --> 00:04:51,480
there's a lot of data we gather 
aside from the seismic together,

100
00:04:51,480 --> 00:04:54,480
tons of SCADA data. 
So every single well is equipped

101
00:04:54,480 --> 00:04:58,320
with, I don't know, probably on 
average 5 to 10 sensors that are

102
00:04:58,320 --> 00:05:01,880
all gathering data in real time.
Facilities has orders of 

103
00:05:01,880 --> 00:05:05,120
magnitude more sensors than that
gathering data in real time. 

104
00:05:05,400 --> 00:05:08,560
So yeah, I, I think the, the 
amount of information being 

105
00:05:08,560 --> 00:05:10,240
captured is often 
underestimated. 

106
00:05:10,240 --> 00:05:13,880
And then if you get into 
drilling, it's like just massive

107
00:05:13,880 --> 00:05:16,360
amounts of data for any given 
well that you drill. 

108
00:05:16,360 --> 00:05:18,840
So yeah, oil and gas, I don't 
think they're a stranger. 

109
00:05:19,000 --> 00:05:20,800
To. 
Large pond use of data. 

110
00:05:20,800 --> 00:05:24,080
I think they've traditionally 
been viewed as maybe definitely 

111
00:05:24,080 --> 00:05:27,040
not on the, you know, the 
frontier tech, but there's a lot

112
00:05:27,040 --> 00:05:28,760
they do that's that might 
surprise you. 

113
00:05:28,960 --> 00:05:30,520
Excellent. 
Then look, I know we're going to

114
00:05:30,520 --> 00:05:34,040
delve into that as well too. 
And thanks for everybody who has

115
00:05:34,040 --> 00:05:36,920
joined us on the stream as well.
We see from folks from 

116
00:05:36,920 --> 00:05:38,960
Minneapolis and Pakistan 
etcetera. 

117
00:05:38,960 --> 00:05:42,480
And Jeff gives you a shout out 
there, Alex and Jeff Schmidt who

118
00:05:42,480 --> 00:05:44,960
obviously you know, so look, 
it's great to have everybody 

119
00:05:44,960 --> 00:05:48,240
join us as well too. 
Size and scale of Oxy as a 

120
00:05:48,240 --> 00:05:51,800
company, how big is the company 
and how big is is kind of the AI

121
00:05:51,800 --> 00:05:53,120
team that you guys are working 
on? 

122
00:05:53,600 --> 00:05:55,880
Sure. 
So I mean, First off, he's been 

123
00:05:55,880 --> 00:05:59,560
around for over 100 years now. 
I've got quite a quite a history

124
00:05:59,560 --> 00:06:02,400
of international operations. 
I couldn't even count how many 

125
00:06:02,400 --> 00:06:05,200
countries I think today we 
operate in maybe 5-5 or so. 

126
00:06:05,360 --> 00:06:07,880
Don't, don't quote me on that, 
roughly 5 or so countries. 

127
00:06:07,880 --> 00:06:11,280
But in terms of scale, what 
15,000 employees or they're 

128
00:06:11,280 --> 00:06:13,160
about? 
It's a large. 

129
00:06:13,200 --> 00:06:16,600
Fortune 100 company significant,
significant oil and gas 

130
00:06:16,600 --> 00:06:19,640
presence, especially these days 
in the Permian Basin, Gulf of 

131
00:06:19,640 --> 00:06:22,320
Mexico and and and DJ Basin in 
the Rockies. 

132
00:06:22,320 --> 00:06:24,280
So yeah, Andrew, anything? 
Yeah. 

133
00:06:24,400 --> 00:06:26,960
As far as production, if you're 
familiar with it, we're about 

134
00:06:26,960 --> 00:06:31,920
1.4 to 1.5 million Poe per day. 
Are we the biggest Permian 

135
00:06:31,920 --> 00:06:33,960
producer? 
Yeah, we're either close to or 

136
00:06:33,960 --> 00:06:36,520
or the biggest first. 
Or second Permian operator. 

137
00:06:36,600 --> 00:06:38,400
OK, OK. 
Yeah, we do a lot. 

138
00:06:38,640 --> 00:06:41,880
On shore US, I'll say one thing 
about Oxy, I think for those 

139
00:06:41,880 --> 00:06:44,280
that are joining that are 
outside of oil and gas, I think 

140
00:06:44,280 --> 00:06:46,600
the oil companies we know about 
are the ones that we buy our 

141
00:06:46,600 --> 00:06:48,960
gasoline from. 
And so a lot of the times Oxy 

142
00:06:48,960 --> 00:06:50,360
doesn't I mean Oxy doesn't do 
that, right. 

143
00:06:50,360 --> 00:06:53,240
There's no, no downstream 
providing selling the customers.

144
00:06:53,240 --> 00:06:56,720
So I think that kind of reduces 
our our our public presence for 

145
00:06:56,720 --> 00:06:58,720
those outside of the industry 
unless they happen to attend an 

146
00:06:58,720 --> 00:07:01,600
Astros game, in which case you 
can't miss the Oxy logo on the 

147
00:07:01,640 --> 00:07:03,720
field there. 
But but yeah, that's that's a 

148
00:07:03,720 --> 00:07:06,720
little bit Oxy. 
So Oxy's big enough to be a team

149
00:07:06,720 --> 00:07:08,640
sponsor then. 
So that's a good way, a good 

150
00:07:08,640 --> 00:07:10,680
bracket to put things into as 
well, too. 

151
00:07:10,960 --> 00:07:13,920
And I appreciate you, Alex, not 
courting controversy there and 

152
00:07:13,920 --> 00:07:15,920
calling it the Gulf of America. 
You called it the. 

153
00:07:15,920 --> 00:07:17,880
Gulf. 
Oh goodness, it's all good. 

154
00:07:18,000 --> 00:07:20,120
That aside, I didn't even think 
about it. 

155
00:07:20,320 --> 00:07:21,920
Perfect. 
And listen to the others that 

156
00:07:21,920 --> 00:07:25,040
are joining us from from Spain 
and from India as well too. 

157
00:07:25,040 --> 00:07:26,640
It's great to have you all on 
board. 

158
00:07:26,880 --> 00:07:29,240
So look, at the beginning. 
I said, you know, in the 

159
00:07:29,240 --> 00:07:33,240
introduction I talked about 
11,000,000 documents, et cetera.

160
00:07:33,400 --> 00:07:36,840
You walk us through kind of the 
challenge that you are faced. 

161
00:07:36,840 --> 00:07:39,680
When did this start? 
When, you know, everybody kind 

162
00:07:39,680 --> 00:07:43,120
of considers, yeah, well, the 
public per SE considers, you 

163
00:07:43,120 --> 00:07:46,680
know, AI engine, AI and large 
language models having been born

164
00:07:46,680 --> 00:07:49,800
in November 22 when ChatGPT came
out to the public. 

165
00:07:49,800 --> 00:07:52,840
But obviously it's been around 
for a long, long time. 

166
00:07:52,840 --> 00:07:57,160
So tell me a little bit about 
the pre solution that you have. 

167
00:07:57,160 --> 00:08:00,960
What was what did that look like
for the folks at Oxy for all of 

168
00:08:00,960 --> 00:08:03,920
these documents they have before
we get into kind of your 

169
00:08:03,920 --> 00:08:06,960
approach to to remedy in that 
and make it easier for them? 

170
00:08:07,120 --> 00:08:10,120
Yeah, this might be a good 
chance to share kind of a visual

171
00:08:10,120 --> 00:08:13,120
because, you know, I come out on
podcasts and I talk to you about

172
00:08:13,240 --> 00:08:15,960
our document problems and you 
might be thinking of something 

173
00:08:15,960 --> 00:08:17,640
different than what I'm actually
talking about. 

174
00:08:17,640 --> 00:08:20,080
So you might be thinking of it's
kind of document. 

175
00:08:20,160 --> 00:08:21,560
Right. 
Yeah, yes, yeah. 

176
00:08:21,720 --> 00:08:24,280
We we we know what documents are
here, but they're not the 

177
00:08:24,280 --> 00:08:25,960
documents here. 
Right. 

178
00:08:26,000 --> 00:08:28,640
So you know, the documents we're
here to talk about is actually, 

179
00:08:28,800 --> 00:08:32,400
it's this kind of document. 
So it's a scan, a paper 

180
00:08:32,400 --> 00:08:34,559
document. 
So like I said, Oxy's been 

181
00:08:34,559 --> 00:08:37,240
around for over 100 years. 
So we've accumulated quite a lot

182
00:08:37,240 --> 00:08:40,480
of documents, paper documents. 
And some of them look like this.

183
00:08:40,480 --> 00:08:42,600
This one's pretty, pretty easy 
on the eyes. 

184
00:08:42,600 --> 00:08:44,480
It's it's just a classic lease 
agreement. 

185
00:08:44,480 --> 00:08:48,360
Then I might show you this one a
bit harder to read, a bit crusty

186
00:08:48,360 --> 00:08:50,920
looking maybe. 
And then I could show you some 

187
00:08:50,920 --> 00:08:54,480
other examples like this or this
getting a little bit harder, 

188
00:08:54,480 --> 00:08:57,000
right? 
Or this or this or this. 

189
00:08:57,000 --> 00:08:59,080
One's one of my favorites. 
I like how the hand actually 

190
00:08:59,080 --> 00:09:02,720
made it into the scan. 
An important part of the scan? 

191
00:09:02,720 --> 00:09:03,880
Yes, definitely. 
Oh wow. 

192
00:09:04,040 --> 00:09:05,480
OK. 
And then you get the sense we 

193
00:09:05,480 --> 00:09:08,600
actually have a lot of these. 
So this is, this is 1 hall, one 

194
00:09:08,600 --> 00:09:11,680
kind of row of I think there's 
maybe 40 of them in this 

195
00:09:11,680 --> 00:09:13,960
particular file room. 
And there's a dozen or so far. 

196
00:09:14,040 --> 00:09:16,280
And then a fun, you know, that's
actually, that's me at the end 

197
00:09:16,280 --> 00:09:18,280
of the hall there. 
And I'm, I'm short, but I'm not 

198
00:09:18,280 --> 00:09:20,920
that sure that's, that's, it's a
lot of documents, suffice to 

199
00:09:20,920 --> 00:09:22,560
say. 
So kind of laying the 

200
00:09:22,560 --> 00:09:25,840
groundwork, obviously Oxy isn't 
unique in this. 

201
00:09:25,840 --> 00:09:29,440
Every oil and gas company that's
been around for decades or 100 

202
00:09:29,440 --> 00:09:32,000
years has vast rows of 
documents. 

203
00:09:32,160 --> 00:09:34,960
The ones that we're talking 
about today pertain to land, 

204
00:09:35,000 --> 00:09:37,920
which is the process of, of 
leasing minerals and 

205
00:09:37,920 --> 00:09:40,640
understanding the ownership of 
minerals and, and rights 

206
00:09:40,640 --> 00:09:43,040
throughout time. 
So land is not the only part of 

207
00:09:43,040 --> 00:09:45,320
the oil and gas industry that's 
very document heavy, but it's, 

208
00:09:45,320 --> 00:09:47,360
it's one of the very important 
parts. 

209
00:09:47,520 --> 00:09:51,000
And throughout the years, Oxy's 
undertaken a number of 

210
00:09:51,000 --> 00:09:54,360
digitization efforts. 
So that's scan coming in. 

211
00:09:54,480 --> 00:09:56,440
It's very rare that somebody's 
going and actually handling a 

212
00:09:56,440 --> 00:09:58,960
paper document these days. 
As you would hope these have 

213
00:09:58,960 --> 00:10:00,840
been digitized. 
But even once they're digitized,

214
00:10:00,840 --> 00:10:03,520
there's a lot of additional work
that needs to follow after that.

215
00:10:03,680 --> 00:10:07,160
And So what Andrew and I found 
out about last year, what was an

216
00:10:07,160 --> 00:10:10,320
effort to digitize, or rather 
they'd already been digitized, 

217
00:10:10,320 --> 00:10:13,800
but there was some additional 
work post digitization happening

218
00:10:14,200 --> 00:10:15,840
on one and a half million land 
documents. 

219
00:10:15,840 --> 00:10:18,640
So the idea here is, you know, 
we have these scanned in and we 

220
00:10:18,640 --> 00:10:21,520
know, OK, for this lease 
agreement, we have maybe 100 or 

221
00:10:21,520 --> 00:10:23,800
150 documents that pertain to 
it. 

222
00:10:23,920 --> 00:10:27,400
But if I'm working in the land 
department and I'm, I'm asked a 

223
00:10:27,400 --> 00:10:29,960
question about a lease, it's not
enough to go say, hey, here's 

224
00:10:29,960 --> 00:10:32,400
150 documents that govern this 
lease. 

225
00:10:32,440 --> 00:10:35,200
I, I would like to know which 
document is which type. 

226
00:10:35,360 --> 00:10:38,760
I would like to know some of the
metadata inside those documents 

227
00:10:38,760 --> 00:10:41,640
without having to open up each 
of those 150 documents. 

228
00:10:41,720 --> 00:10:44,400
So you can imagine it's. 
It's really useful to be able to

229
00:10:44,440 --> 00:10:46,800
categorize documents and extract
information out of them. 

230
00:10:47,240 --> 00:10:50,160
And I would imagine that that 
process originally even the 

231
00:10:50,160 --> 00:10:54,480
scanning obviously is, is labour
and time intensive and resource 

232
00:10:54,480 --> 00:10:57,000
intensive. 
But as you say, then having to 

233
00:10:57,000 --> 00:11:00,160
look these up etcetera, 
obviously again time and 

234
00:11:00,200 --> 00:11:03,200
resource intensive. 
You, you can imagine for, for 

235
00:11:03,240 --> 00:11:05,760
one and a half million documents
for someone to go in and 

236
00:11:05,760 --> 00:11:09,320
categorize it and then to 
extract maybe 5 or 6 pieces of 

237
00:11:09,320 --> 00:11:10,600
information out of that 
document. 

238
00:11:10,600 --> 00:11:11,600
That's going to take a long 
time. 

239
00:11:11,600 --> 00:11:15,240
And that's exactly the, the 
project that we walked into last

240
00:11:15,240 --> 00:11:18,960
year, Andrew and me and by the 
gentleman on the team, Eric, we 

241
00:11:18,960 --> 00:11:22,480
we learned of this effort to 
manually categorize and extract 

242
00:11:22,480 --> 00:11:24,640
information out of one and a 
half million documents that was 

243
00:11:24,640 --> 00:11:28,480
going to be done with a a small 
army of, of contract workers. 

244
00:11:29,160 --> 00:11:32,440
So yeah, I think 4040 people 
were scoped onto this project. 

245
00:11:32,440 --> 00:11:33,480
It was to take a year and a 
half. 

246
00:11:33,480 --> 00:11:36,160
And then so the, you know, you 
can, you can imagine the cost 

247
00:11:36,160 --> 00:11:38,280
associated with that. 
And there are also a lot of 

248
00:11:38,280 --> 00:11:41,160
tough decisions you have to make
when you're having an army of 

249
00:11:41,200 --> 00:11:45,000
non experts, contract workers 
doing a task like this. 

250
00:11:45,120 --> 00:11:47,000
The scope of the project has to 
be very limited. 

251
00:11:47,320 --> 00:11:49,640
For example, I mentioned 
categorization, you want to 

252
00:11:49,640 --> 00:11:51,560
categorize what type of 
documents these are. 

253
00:11:51,720 --> 00:11:54,840
When we're doing this manually 
and we're paying by the hour for

254
00:11:54,840 --> 00:11:57,440
the work to do this, we had 
massively limited, the land 

255
00:11:57,440 --> 00:12:00,080
department had massively limited
the scope of this effort to just

256
00:12:00,080 --> 00:12:02,840
six a different categories. 
So very broad stroke, broad 

257
00:12:02,840 --> 00:12:04,320
categories. 
OK. 

258
00:12:05,080 --> 00:12:07,040
And then when we talk about 
extracting some of the 

259
00:12:07,040 --> 00:12:10,200
information from inside these 
documents, similarly, they had 

260
00:12:10,200 --> 00:12:13,080
limited it to just the most 
basic facts and figures, dates, 

261
00:12:13,200 --> 00:12:16,400
things that were easy for a non 
expert to pick up and pull out 

262
00:12:16,400 --> 00:12:18,760
of these documents. 
But once it would nonetheless 

263
00:12:18,880 --> 00:12:21,360
help the land team as they're 
working with these documents 

264
00:12:21,360 --> 00:12:23,920
into the future. 
So, so not only were we dealing 

265
00:12:23,920 --> 00:12:27,680
with a very expensive project, a
very long duration project, 18 

266
00:12:27,680 --> 00:12:30,640
months, it was also going to be 
a very tight, narrow, limited 

267
00:12:30,640 --> 00:12:33,160
scope for the purpose of of cost
savings. 

268
00:12:33,400 --> 00:12:35,680
So that was kind of the scene 
where Andrew and I last year, 

269
00:12:35,680 --> 00:12:38,200
early last year were looking at 
each other and we said, man, 

270
00:12:38,360 --> 00:12:41,320
isn't this a perfect, perfect 
project for a large language 

271
00:12:41,320 --> 00:12:43,520
model? 
So timing was on your side 

272
00:12:43,520 --> 00:12:45,120
there. 
I suppose you're getting in at 

273
00:12:45,120 --> 00:12:48,480
the the ground roots of a 
project and seeing whether I 

274
00:12:48,480 --> 00:12:51,680
suppose and we'll talk a little 
bit about this obviously is your

275
00:12:51,680 --> 00:12:53,760
approach. 
How did that compare to the 

276
00:12:53,760 --> 00:12:57,280
approach that the pathway was 
already indicating lots of 

277
00:12:57,280 --> 00:13:02,040
people long time span and only 
actually constraint of only 

278
00:13:02,040 --> 00:13:04,720
getting some of that information
out of those documents, right? 

279
00:13:04,880 --> 00:13:08,080
Yeah, I think first, Andrew and 
I were no strangers to to 

280
00:13:08,080 --> 00:13:10,360
language models. 
We, we've done some work on them

281
00:13:10,480 --> 00:13:13,080
with them at Oxy already. 
I think all the work we had done

282
00:13:13,080 --> 00:13:15,320
before was obviously a much, 
much smaller scale. 

283
00:13:15,320 --> 00:13:18,320
So that was a bit what, what 
scared us about this not, not so

284
00:13:18,320 --> 00:13:20,120
much could we do it? 
We were pretty confident that 

285
00:13:20,120 --> 00:13:23,560
this is possible, but to do it 
at a larger scale and, and the 

286
00:13:23,560 --> 00:13:26,960
data management practices that 
that would require and just the 

287
00:13:26,960 --> 00:13:29,600
cloud resources that that would 
require, those were some of the 

288
00:13:29,600 --> 00:13:31,840
things that we were a little 
less confident in. 

289
00:13:32,000 --> 00:13:35,440
So, so we wanted to very quickly
de risk this and, and 

290
00:13:35,440 --> 00:13:36,880
understand, first of all, prove 
out. 

291
00:13:36,880 --> 00:13:38,840
Can we do this? 
Because you know, this other 

292
00:13:38,840 --> 00:13:41,000
team is going to take on this 
project, do this manually 

293
00:13:41,000 --> 00:13:43,200
contract this out. 
It's already been budgeted, you 

294
00:13:43,200 --> 00:13:45,280
know, the years underway. 
The project's about to start 

295
00:13:45,280 --> 00:13:48,280
just a month and a half away. 
So the question is, first of 

296
00:13:48,280 --> 00:13:49,840
all, can we prove that we can do
it? 

297
00:13:49,920 --> 00:13:53,120
Can we convince people, wait, 
stop what you're going to do, 

298
00:13:53,320 --> 00:13:55,720
let's do this instead? 
So that was our challenge. 

299
00:13:56,480 --> 00:13:58,960
Was that a proof of concept for 
you then? 

300
00:13:58,960 --> 00:14:02,400
Did you had that short window? 
You, you had four weeks, six 

301
00:14:02,400 --> 00:14:04,800
weeks to be able to show. 
That this would work to six 

302
00:14:04,800 --> 00:14:06,280
weeks. 
That's exactly right. 

303
00:14:06,280 --> 00:14:09,240
We wanted to do a proof of 
concept, so we took some of the 

304
00:14:09,320 --> 00:14:12,400
a subset of the documents, 1000 
of the documents and we set out 

305
00:14:12,400 --> 00:14:16,400
to to do again, there's two 
tasks here, classified and 

306
00:14:16,400 --> 00:14:19,360
extracting And just to back up, 
I just, you know, we looked at a

307
00:14:19,360 --> 00:14:21,480
lot of those documents. 
You get a sense that, you know, 

308
00:14:21,480 --> 00:14:23,920
some of these are easier to 
read, some of these are harder 

309
00:14:23,920 --> 00:14:26,440
to read, But one of the most 
important things we had to do 

310
00:14:26,440 --> 00:14:28,920
before we could even do the 
classification and extraction 

311
00:14:28,920 --> 00:14:33,120
was figure out OCR, optical 
character recognition and OCR. 

312
00:14:33,240 --> 00:14:36,480
You know, people don't think 
about it as AI, but it actually 

313
00:14:36,480 --> 00:14:39,160
somewhat is it's kind of a 
depending on how you define AI, 

314
00:14:39,160 --> 00:14:41,960
that is a form of AI and it's a 
really important to requisite 

315
00:14:41,960 --> 00:14:44,400
step here. 
A traditionally Oxy had used a 

316
00:14:44,400 --> 00:14:49,200
vendor Co fax or for our and you
can see with this particular 

317
00:14:49,200 --> 00:14:51,200
document, it actually it really 
struggled. 

318
00:14:51,200 --> 00:14:54,080
You can see Atlantic operating 
turned into Atlantic gyrating. 

319
00:14:54,560 --> 00:14:57,040
While it's kind of funny, we 
don't want that to make it into 

320
00:14:57,080 --> 00:14:59,200
our land database down the road.
Sure. 

321
00:14:59,440 --> 00:15:01,640
And you can see generally it 
struggles a lot with and writing

322
00:15:01,640 --> 00:15:03,920
here. 
So the first task we had to do 

323
00:15:04,240 --> 00:15:07,040
was benchmark a number of the 
different OCR providers. 

324
00:15:07,160 --> 00:15:09,840
And I think if anybody's 
experimented in this space 

325
00:15:09,840 --> 00:15:11,920
recently, it's actually changing
really rapidly. 

326
00:15:11,920 --> 00:15:15,440
I think a multimodal models are 
increasingly being used for OCR 

327
00:15:15,440 --> 00:15:18,800
and cloud vendors are improving 
quite rapidly their OCR 

328
00:15:18,800 --> 00:15:21,800
solutions. 
The open source space is also 

329
00:15:21,800 --> 00:15:24,600
improving rapidly, though we 
found a pretty big gap when we 

330
00:15:24,600 --> 00:15:27,640
went about benchmarking all the 
different OCR options, open 

331
00:15:27,640 --> 00:15:29,680
source and and through cloud 
providers. 

332
00:15:29,680 --> 00:15:33,000
We found the best results on our
documents with Azure Document 

333
00:15:33,000 --> 00:15:34,920
Intelligence. 
OK, that's what we ended up 

334
00:15:34,920 --> 00:15:36,480
using for this product. 
OK. 

335
00:15:36,920 --> 00:15:38,840
And I'll say real quick why 
that's so important. 

336
00:15:38,840 --> 00:15:41,680
If you imagine if you're the 
large language model, if we feed

337
00:15:41,680 --> 00:15:45,000
in that text in the top right 
here, it doesn't matter how 

338
00:15:45,000 --> 00:15:48,400
smart the language model is, 
it's not going to be able to fix

339
00:15:48,400 --> 00:15:49,880
what's what's broken on the 
input. 

340
00:15:50,080 --> 00:15:53,680
So it's very important to get 
this very boring OCR step. 

341
00:15:53,720 --> 00:15:55,560
It's important to get this right
for that reason. 

342
00:15:55,560 --> 00:15:59,360
Yeah, and it would, it would 
compound I suppose the, you 

343
00:15:59,360 --> 00:16:03,680
know, we accuse LLMS of being 
hallucinatory sometimes as well 

344
00:16:03,680 --> 00:16:06,280
too, but it's got fed the wrong 
information in the 1st place. 

345
00:16:06,280 --> 00:16:08,720
We don't really know where that 
came from. 

346
00:16:08,720 --> 00:16:11,920
And and we're kind of, I suppose
blaming the LLM as opposed to 

347
00:16:12,040 --> 00:16:16,360
actually the original OC or so. 
So that was a task in hand 

348
00:16:16,360 --> 00:16:20,040
anyway was to get, you know, 
very accurate, appropriate 

349
00:16:20,040 --> 00:16:23,640
proper OC or so that should have
would have taken a good chunk of

350
00:16:23,640 --> 00:16:26,440
your time at this kind of tiny 
six week window, right? 

351
00:16:26,640 --> 00:16:29,880
Yeah, unfortunately, as you said
it in this short window, we were

352
00:16:29,880 --> 00:16:30,920
lucky. 
We're again, we're dealing with 

353
00:16:30,920 --> 00:16:33,720
a small subset, just 1000 
documents instead of one and a 

354
00:16:33,720 --> 00:16:35,280
half million. 
We're just trying to prove the 

355
00:16:35,280 --> 00:16:37,520
point here. 
So we, we were able to get this 

356
00:16:37,520 --> 00:16:40,800
benchmarking done pretty quick 
and move on to the next task, 

357
00:16:40,800 --> 00:16:43,960
which was classification. 
We talked about classification. 

358
00:16:43,960 --> 00:16:47,160
I, I teased you earlier with the
scope being very narrow to six 

359
00:16:47,160 --> 00:16:50,040
categories. 
Well, not so lucky in our case. 

360
00:16:50,040 --> 00:16:53,280
I think once they realize, hey, 
AI can, can maybe do this. 

361
00:16:53,360 --> 00:16:56,000
And you said, well, instead of 
using the six categories that 

362
00:16:56,240 --> 00:16:58,360
you know, that the 40 
contractors would be asked to 

363
00:16:58,360 --> 00:17:00,480
do, let's, let's try expanding 
that scope a little bit. 

364
00:17:00,480 --> 00:17:03,120
Let's do 140 categories. 
This would be much more useful. 

365
00:17:03,320 --> 00:17:05,240
Let's 20X. 
Yeah, yeah. 

366
00:17:05,359 --> 00:17:07,800
So I don't know how we got 
hoodwinked into that, but 

367
00:17:07,800 --> 00:17:10,280
somehow that happened. 
You were up for a challenge, 

368
00:17:10,359 --> 00:17:12,839
Yeah. 
I will say we were really lucky 

369
00:17:12,920 --> 00:17:15,880
for this project to have for 
most of these categories we had 

370
00:17:15,880 --> 00:17:20,000
reliable hand labeled data from 
our existing legacy records. 

371
00:17:20,000 --> 00:17:22,599
So there's one and a half 
million new documents is what 

372
00:17:22,599 --> 00:17:24,599
we're going after. 
But for the older documents we 

373
00:17:24,599 --> 00:17:27,000
have a lot of of reliably hand 
labeled data. 

374
00:17:27,000 --> 00:17:30,280
So for those that do any kind of
machine learning in the in the 

375
00:17:30,280 --> 00:17:32,680
audience, you might kind of see 
where this is going that we can 

376
00:17:32,680 --> 00:17:35,480
use a supervised method for the 
classification here. 

377
00:17:35,560 --> 00:17:38,680
OK, OK. 
And so the approach we took is 

378
00:17:38,720 --> 00:17:42,120
we use an AI model for for 
getting an off the shelf model. 

379
00:17:42,360 --> 00:17:45,640
Azure Open AI or rather Open AI 
releases this model text 

380
00:17:45,640 --> 00:17:48,280
embedding 3 large, you can use 
it through Azure Open AI 

381
00:17:48,280 --> 00:17:51,080
service. 
And those embeddings are pretty 

382
00:17:51,320 --> 00:17:53,680
high dimensional. 
They're 3072 dimensions. 

383
00:17:53,800 --> 00:17:56,600
But, but once you have that high
dimensionality space, you can 

384
00:17:56,720 --> 00:17:59,160
here in this visual, I've 
reduced it so you can just kind 

385
00:17:59,160 --> 00:18:01,320
of see the difference between 
those documents, but you can 

386
00:18:01,320 --> 00:18:04,440
also train classifiers against 
those because those embeddings 

387
00:18:04,440 --> 00:18:06,840
are just numbers, right? 
Train a classifier. 

388
00:18:06,840 --> 00:18:10,200
So we just train a classifier 
supervised learning on these 

389
00:18:10,200 --> 00:18:12,320
embeddings plus some hand 
curated features. 

390
00:18:12,440 --> 00:18:14,840
And, and with that, we were able
to achieve really, really high 

391
00:18:14,840 --> 00:18:16,520
accuracy. 
We were able to achieve 

392
00:18:16,560 --> 00:18:19,280
depending on the exact amount of
categories and how difficult the

393
00:18:19,280 --> 00:18:23,120
documents were, we were we were 
on the order of 95 to 97% 

394
00:18:23,120 --> 00:18:25,200
accurate. 
Contrasting that with the the 

395
00:18:26,080 --> 00:18:30,960
non expert contract labeling was
on the order of 88 to 92% 

396
00:18:31,040 --> 00:18:34,440
accurate, but but again that's 
with a much more limited subset 

397
00:18:34,440 --> 00:18:36,400
of categories, so much easier 
problem. 

398
00:18:36,560 --> 00:18:38,480
So we were really, really 
pleased with with the 

399
00:18:38,880 --> 00:18:42,280
classification results. 
And in the same way that you, 

400
00:18:42,280 --> 00:18:45,360
sorry Alex, in the same way that
you had to, you know, check out 

401
00:18:45,360 --> 00:18:48,200
a few different OCR's to find a 
good one, did you have to do the

402
00:18:48,200 --> 00:18:51,680
same with regard to choosing the
embeddings models as well too? 

403
00:18:51,680 --> 00:18:53,560
Because there was. 
There's so many to choose from, 

404
00:18:53,760 --> 00:18:56,600
OK, definitely these days. 
And even a year ago, when you 

405
00:18:56,600 --> 00:18:58,600
were getting started with this, 
there was still quite a lot 

406
00:18:58,600 --> 00:19:01,000
knocking around. 
Yeah, I, I think when we were 

407
00:19:01,000 --> 00:19:04,320
first experimenting with this, 
we tried some, you know, I, I 

408
00:19:04,320 --> 00:19:06,760
don't remember when text 
embedding 3 large came out, but.

409
00:19:06,920 --> 00:19:08,920
Yeah, I think at the time 
embedding. 

410
00:19:08,920 --> 00:19:12,040
Three large had just come out. 
And so we now knew this was 

411
00:19:12,040 --> 00:19:14,400
state-of-the-art and it made it 
really easy on us. 

412
00:19:14,520 --> 00:19:16,240
We, we might have been in the 
middle of the, I don't remember,

413
00:19:16,240 --> 00:19:18,560
we might have been in the middle
of the POC when it came out, we 

414
00:19:18,680 --> 00:19:20,400
were using it. 
It's, it's one of those amazing 

415
00:19:20,400 --> 00:19:22,000
things about working in AI right
now. 

416
00:19:22,000 --> 00:19:23,640
It's like you're doing 
something, you're getting this 

417
00:19:23,760 --> 00:19:26,360
level of accuracy that next week
a new model comes out and 

418
00:19:26,480 --> 00:19:28,880
instantly with no effort bumps 
your accuracy up. 

419
00:19:28,880 --> 00:19:30,680
So it was like one of those 
moments where we're like, wow, 

420
00:19:30,840 --> 00:19:33,480
this is too easy. 
But, but not to say there wasn't

421
00:19:33,480 --> 00:19:35,920
any iteration to your point 
about trying different models, 

422
00:19:35,920 --> 00:19:37,840
yeah, we tried different 
embedding models, we tried 

423
00:19:37,840 --> 00:19:40,200
different classifiers. 
We did some experimentation. 

424
00:19:40,200 --> 00:19:43,400
But I I think by and large, 
certainly in the POC phase, most

425
00:19:43,400 --> 00:19:45,280
of this just worked quite well 
out-of-the-box. 

426
00:19:45,440 --> 00:19:48,840
Brilliant, Brilliant. 
So with the extracting being, 

427
00:19:49,080 --> 00:19:52,200
you know, on top of its game and
doing well, and with the 

428
00:19:52,200 --> 00:19:55,320
classifications now working, I 
mean, was that enough with those

429
00:19:55,320 --> 00:19:59,680
1000 documents in your POC to 
get the thumbs up for this 

430
00:19:59,680 --> 00:20:02,360
project as opposed to the 40 
humans in A room? 

431
00:20:02,560 --> 00:20:04,840
That's right. 
So just from the classification 

432
00:20:04,840 --> 00:20:07,720
results, there was growing 
confidence among the land team 

433
00:20:07,720 --> 00:20:09,840
that this was going to be the 
way to go, not just because it's

434
00:20:09,840 --> 00:20:13,120
more accurate, but because 
getting so many more categories 

435
00:20:13,120 --> 00:20:15,320
would add a lot more value to 
them down the road. 

436
00:20:15,680 --> 00:20:17,000
That makes sense of these 
documents. 

437
00:20:17,000 --> 00:20:19,680
So the last thing for us then 
was to just prove that we could 

438
00:20:19,680 --> 00:20:22,760
do the extractions. 
And with that, again, we took a 

439
00:20:22,760 --> 00:20:23,960
very, very simple approach 
approach. 

440
00:20:24,280 --> 00:20:27,280
We we used an out-of-the-box 
model to do these extractions. 

441
00:20:27,280 --> 00:20:30,000
Some of these were very simple 
extractions like things like 

442
00:20:30,160 --> 00:20:33,800
what I showed earlier, gates, 
counties, just very simple 

443
00:20:33,800 --> 00:20:36,920
fields that you could that an 
untrained person could catch. 

444
00:20:37,080 --> 00:20:39,400
Some of these were more 
difficult things like requiring 

445
00:20:39,400 --> 00:20:42,520
interpreting some minor simple 
interpretation of legal clauses 

446
00:20:42,600 --> 00:20:45,400
like a few clause or a cessation
of production clause in a 

447
00:20:45,400 --> 00:20:47,600
document. 
And actually that was where we 

448
00:20:47,600 --> 00:20:50,160
were a little unsure and we were
quite pleasantly surprised at 

449
00:20:50,160 --> 00:20:53,000
what the language models could 
do in terms of interpreting 

450
00:20:53,080 --> 00:20:56,640
paragraphs of and boiling it 
down to just a database friendly

451
00:20:56,640 --> 00:20:59,920
scalar answer that we could 
insert, which I guess we'll get 

452
00:20:59,920 --> 00:21:01,680
to in a minute. 
But yeah, again, off the shelf 

453
00:21:01,680 --> 00:21:05,640
model, we used GPT 4 O for this.
And again, this is more accurate

454
00:21:05,640 --> 00:21:08,120
than the the contractor level 
classifications. 

455
00:21:08,280 --> 00:21:10,440
Excellent. 
So, so just off the shelf, no 

456
00:21:10,440 --> 00:21:14,200
fine tuning, no nothing else 
with regard to the model, off 

457
00:21:14,200 --> 00:21:16,040
the shelf model. 
Off the shelf model. 

458
00:21:16,040 --> 00:21:19,000
No fine tuning, just a very very
elaborate prompt. 

459
00:21:19,000 --> 00:21:22,640
This was one of the early, early
versions of our prompt for for 

460
00:21:22,640 --> 00:21:26,240
doing some extractions and full 
credit to Eric Smith on our team

461
00:21:26,240 --> 00:21:28,520
was the person that developed 
these prompts. 

462
00:21:28,520 --> 00:21:32,600
He had worked as a as a land man
prior to his role on this AI 

463
00:21:32,600 --> 00:21:33,680
team. 
So he was actually quite 

464
00:21:33,680 --> 00:21:36,120
familiar with a lot of this and 
he was able to kind of translate

465
00:21:36,120 --> 00:21:39,360
from land to AI and back and 
forth to get these models 

466
00:21:39,360 --> 00:21:43,720
working the way we needed. 
OK, OK, so and the prompt text 

467
00:21:43,720 --> 00:21:45,960
is quite small. 
I'm sure most people can't read 

468
00:21:45,960 --> 00:21:48,080
what's up there. 
But this is this is where the 

469
00:21:48,080 --> 00:21:51,240
role of the clever prompt 
engineer comes in. 

470
00:21:51,240 --> 00:21:55,520
You know, being able to preface 
it, set the scene, ask the right

471
00:21:55,680 --> 00:21:59,520
type of question and explain how
you would like the result, 

472
00:21:59,520 --> 00:22:00,800
right? 
Absolutely. 

473
00:22:00,800 --> 00:22:03,480
And I think, you know, we could 
debate back and forth about if 

474
00:22:03,520 --> 00:22:06,240
the role of a prompt engineer is
a real job or not. 

475
00:22:06,240 --> 00:22:08,560
But I'll, I'll say for Eric, it 
sure was it's. 

476
00:22:08,560 --> 00:22:10,640
Another show, Alex. 
That's another show A. 

477
00:22:10,640 --> 00:22:13,640
Whole episode on that I imagine 
I'll say at least for the state 

478
00:22:13,640 --> 00:22:16,560
of the models they are today, 
there's you can always eke out 

479
00:22:16,560 --> 00:22:18,880
better performance with with 
better prompting. 

480
00:22:18,880 --> 00:22:22,000
So, so there's no doubt that 
it's a useful skill and I think 

481
00:22:22,200 --> 00:22:25,280
it increasingly drives what's 
interesting is it increasingly 

482
00:22:25,280 --> 00:22:28,600
drives more of the the work and 
the power into the hands of the 

483
00:22:28,600 --> 00:22:30,760
subject matter expert. 
I think there's a lot of fear 

484
00:22:30,760 --> 00:22:33,840
about, you know, AI is going to 
take away what what we do for 

485
00:22:33,840 --> 00:22:35,800
jobs. 
But in this case, it's the land 

486
00:22:35,800 --> 00:22:38,560
SM ES that they are most 
important because they're the 

487
00:22:38,560 --> 00:22:41,640
ones that you need to to craft 
this kind of prompt. 

488
00:22:41,640 --> 00:22:43,960
Or you know, some of the later 
versions that were far more 

489
00:22:44,040 --> 00:22:46,480
elaborate. 
Yeah, no, it's a very fairpoint 

490
00:22:46,480 --> 00:22:48,720
I think. 
And I deal a lot with the main 

491
00:22:48,720 --> 00:22:52,760
Mongo DB in helping the code 
assistance that developers use 

492
00:22:52,760 --> 00:22:55,560
day-to-day better at Mongo DB 
tasks. 

493
00:22:55,560 --> 00:22:58,080
And I can totally see that, 
that, you know, we do a lot of 

494
00:22:58,280 --> 00:23:01,800
working with partners to help 
fine tune and to train and, and 

495
00:23:01,800 --> 00:23:04,880
to give ground truths and do 
evaluation sets, etcetera as 

496
00:23:04,880 --> 00:23:07,640
well too. 
So, yeah, I'm, I fully agree 

497
00:23:07,720 --> 00:23:09,680
with you there. 
I don't think we're removing any

498
00:23:09,680 --> 00:23:11,120
jobs. 
I just think we're getting rid 

499
00:23:11,120 --> 00:23:14,320
of the, the, the kind of mundane
tasks, right? 

500
00:23:14,320 --> 00:23:16,360
And bringing everything up to a 
higher level. 

501
00:23:16,520 --> 00:23:18,400
And maybe I'm getting ahead of 
myself. 

502
00:23:18,400 --> 00:23:21,960
Where does Mongo DB fit into all
of this then Alex and Andrew? 

503
00:23:22,320 --> 00:23:24,360
Well, wasn't that awkward. 
We're on among the podcast, so 

504
00:23:24,360 --> 00:23:27,200
we probably mentioned Bobby DB 
yet I think that's where Andrew 

505
00:23:27,200 --> 00:23:29,600
is going to come in and maybe 
speak to it. 

506
00:23:29,600 --> 00:23:32,200
Perfect. 
So this just kind of closing out

507
00:23:32,200 --> 00:23:36,000
the story here for the POC 
phase, we obviously we had 

508
00:23:36,000 --> 00:23:39,000
really good results on this 
subset, this thousand documents 

509
00:23:39,000 --> 00:23:41,880
of the one half million. 
And so we were able to make the 

510
00:23:41,880 --> 00:23:44,520
decision between us, it and the 
land department. 

511
00:23:44,520 --> 00:23:46,680
We were able to make the 
decision that, hey, let's, let's

512
00:23:46,680 --> 00:23:50,040
ditch this 18 months manual 
expensive project. 

513
00:23:50,040 --> 00:23:53,640
Instead, let's try and do this 
with, with, with AI and, and now

514
00:23:53,640 --> 00:23:56,280
you know, you know, we came to, 
I came to Andrew, Eric and I 

515
00:23:56,280 --> 00:23:58,880
came to Andrew and said, Andrew,
can you help us scale this? 

516
00:23:58,920 --> 00:24:01,000
And he was like, what have you, 
What have you guys done? 

517
00:24:01,120 --> 00:24:03,160
But he'll talk a bit about how 
that scaling works. 

518
00:24:03,360 --> 00:24:06,240
I love how you gently you 
approach that topic. 

519
00:24:06,240 --> 00:24:09,680
Can you take our 1000 document 
sample and ramp it up to 

520
00:24:09,680 --> 00:24:11,400
11,000,000? 
Off you go, I mean. 

521
00:24:11,640 --> 00:24:13,160
How hard can it be? 
How hard? 

522
00:24:13,160 --> 00:24:15,560
Can it be? 
Yeah, Andrew, how hard was it? 

523
00:24:15,680 --> 00:24:19,360
Yeah, it turned out it was not 
too bad, very modest. 

524
00:24:19,680 --> 00:24:22,840
So yeah, the slide Alex is 
showing, I guess, I guess some 

525
00:24:22,840 --> 00:24:25,560
things I would know here. 
It's just originally for the 

526
00:24:25,560 --> 00:24:30,120
1000 document POC, we did go 
ahead and build out some 

527
00:24:30,120 --> 00:24:33,320
architecture of how we would do 
this at scale. 

528
00:24:33,440 --> 00:24:36,680
And that was part of the POC. 
It was like, hey, you don't just

529
00:24:36,680 --> 00:24:39,800
want to manually be running, you
know, a Python script locally 

530
00:24:39,920 --> 00:24:43,000
and that says that you could, 
you know, process millions of 

531
00:24:43,000 --> 00:24:46,280
documents. 
Obviously we were going to need 

532
00:24:46,440 --> 00:24:49,040
probably some compute and 
services to do this. 

533
00:24:49,040 --> 00:24:51,880
But yeah, to mention Mongo, I 
guess we actually started down 

534
00:24:51,880 --> 00:24:55,760
the Mongo path mainly because at
the time there was a very 

535
00:24:55,760 --> 00:25:00,160
limited amount of databases that
could handle embeddings, at 

536
00:25:00,160 --> 00:25:02,920
least properly. 
So I think the options were like

537
00:25:02,920 --> 00:25:06,400
Postgres, Mongo and then what's 
that one? 

538
00:25:06,640 --> 00:25:08,920
They were dedicated. 
Yeah, like dedicated. 

539
00:25:09,320 --> 00:25:12,480
Yeah, yeah, Pine Cone, there's 
some, but we wanted to go with 

540
00:25:12,480 --> 00:25:14,800
something more mainstream. 
And then between Postgres and 

541
00:25:14,800 --> 00:25:19,200
Mongo, we were, we were trying 
to move through the POC so 

542
00:25:19,200 --> 00:25:22,720
quickly that we really didn't 
want to be locked down to some 

543
00:25:22,720 --> 00:25:27,160
specific schema or design that 
we had worked out because we 

544
00:25:27,160 --> 00:25:28,520
didn't, we didn't know where it 
was gone. 

545
00:25:28,640 --> 00:25:32,200
You know, one day Land's asking 
for 140 classifications. 

546
00:25:32,200 --> 00:25:34,680
The next day, who knows what 
they're going to be asking for. 

547
00:25:34,800 --> 00:25:37,360
We might just, we might need to 
be changing things really 

548
00:25:37,360 --> 00:25:39,680
quickly. 
So we landed on Mongo just 

549
00:25:39,680 --> 00:25:41,160
wanted to be in the no sequel 
world. 

550
00:25:41,160 --> 00:25:46,480
And then we had talked to Peter 
and Robert over at Mongo DB and 

551
00:25:46,560 --> 00:25:50,040
they're kind of trying to, I 
guess warn them like, hey, you 

552
00:25:50,040 --> 00:25:52,440
know, this is a really intensive
process. 

553
00:25:52,440 --> 00:25:54,400
We're we're not sure if Mongo 
can handle it. 

554
00:25:54,560 --> 00:25:57,600
And it was like, yeah, this is 
listen, small potatoes, don't 

555
00:25:57,600 --> 00:25:58,880
worry about it. 
No problem. 

556
00:25:58,960 --> 00:26:02,280
And so sure enough, you know, 
Mongo at the thousand of course 

557
00:26:02,280 --> 00:26:05,040
it was fine. 
And early on we really were just

558
00:26:05,040 --> 00:26:08,040
doing inserts to Mongo. 
So what we would do is we would 

559
00:26:08,040 --> 00:26:12,520
take the data from Filenet, 
which is our CRM and just dump 

560
00:26:12,520 --> 00:26:15,560
that directly into S3. 
And this, this was pretty 

561
00:26:15,680 --> 00:26:18,880
simple, just Python script and 
get off one place, but in 

562
00:26:18,880 --> 00:26:21,360
another place. 
And then from there we would run

563
00:26:21,360 --> 00:26:24,080
things more in parallel, just 
through Lambda functions. 

564
00:26:24,080 --> 00:26:26,880
So we'd have the functions 
handling pretty much everything 

565
00:26:26,880 --> 00:26:29,640
Alex just talked about just kind
of steps. 

566
00:26:29,640 --> 00:26:34,880
So, you know, go out, re OCR the
PDF and we have good OCR data 

567
00:26:34,880 --> 00:26:37,640
now go out classify it. 
And now I have classifications. 

568
00:26:37,640 --> 00:26:41,040
All right, using classifications
and the PDF text data. 

569
00:26:41,040 --> 00:26:43,320
Let's go extract what we need to
from that. 

570
00:26:43,400 --> 00:26:46,840
And all the while we're kind of 
building up this document and 

571
00:26:46,840 --> 00:26:48,360
then we just dump it in the 
Mongo. 

572
00:26:48,560 --> 00:26:50,040
And that worked pretty well for 
a while. 

573
00:26:50,040 --> 00:26:51,680
For the 1000, that was no 
problem. 

574
00:26:51,680 --> 00:26:54,800
So, you know, if we're running 
let's say like 100 layout does 

575
00:26:54,960 --> 00:26:58,480
in parallel and each document 
maybe takes a minute to 

576
00:26:58,560 --> 00:27:01,360
classify, we're talking very 
little transactions. 

577
00:27:01,720 --> 00:27:03,400
Let me see if we go. 
Yeah. 

578
00:27:03,760 --> 00:27:08,120
So one issue we started running 
into though was we didn't really

579
00:27:08,200 --> 00:27:11,040
account for a lot of the edge 
cases that come with when you go

580
00:27:11,040 --> 00:27:14,760
from 1000 documents to 1,000,000
plus documents, you're 

581
00:27:14,760 --> 00:27:17,560
introduced to a whole lot of 
different types of documents 

582
00:27:17,600 --> 00:27:20,440
then you would have encountered.
And so we were running into 

583
00:27:20,440 --> 00:27:23,720
things like, OK, we, the biggest
document we saw early on was 

584
00:27:23,720 --> 00:27:26,240
maybe, I don't know, 50 
megabytes. 

585
00:27:26,400 --> 00:27:29,920
And now we're running into a 
document that's like 3 gigabytes

586
00:27:30,080 --> 00:27:33,000
that's going to break some stuff
if you're not ready to handle 

587
00:27:33,000 --> 00:27:35,760
it. 
And so again, like, I don't know

588
00:27:35,760 --> 00:27:39,360
if this is the best architecture
looking back on, it's probably 

589
00:27:39,360 --> 00:27:42,680
not. 
What we did was every time we 

590
00:27:42,680 --> 00:27:45,840
were kind of like timing out 
these lambdas from whatever it 

591
00:27:45,840 --> 00:27:47,720
might be. 
This document's really large. 

592
00:27:47,720 --> 00:27:50,160
It's going to take 20 minutes to
process. 

593
00:27:50,280 --> 00:27:52,360
Lambda's only going to give us 
15 minutes. 

594
00:27:52,520 --> 00:27:53,560
Yeah. 
What can we do? 

595
00:27:53,720 --> 00:27:57,120
We basically just start storing 
every step of the process in 

596
00:27:57,120 --> 00:27:59,360
Migo. 
So lots of reads, lots of 

597
00:27:59,440 --> 00:28:01,800
updates. 
OK, so we move from, you know, 

598
00:28:01,960 --> 00:28:05,960
one to five transactions a 
second to thousands of 

599
00:28:05,960 --> 00:28:08,720
transactions a second because 
we're just constantly hitting 

600
00:28:08,720 --> 00:28:11,440
Mongo for either getting a 
current state, writing a new 

601
00:28:11,440 --> 00:28:14,200
state, you know, finishing the 
dock, updating the whole dock. 

602
00:28:14,280 --> 00:28:17,680
And I mean, it sounds bad, but 
it's like I actually didn't 

603
00:28:17,840 --> 00:28:20,880
think about Mongo the whole time
we're doing it like nothing went

604
00:28:20,880 --> 00:28:22,760
wrong. 
You know, it's just like this, 

605
00:28:22,840 --> 00:28:26,000
this piece of in our text stack 
just kind of magically worked 

606
00:28:26,160 --> 00:28:28,760
and we didn't have to change the
configuration of it. 

607
00:28:28,880 --> 00:28:32,120
We didn't have to, you know, 
give Peter a call and ask for to

608
00:28:32,200 --> 00:28:33,960
come over and take a look at 
what we're doing. 

609
00:28:34,080 --> 00:28:35,880
It was really like the least of 
our problems. 

610
00:28:36,200 --> 00:28:39,120
And so OK, OK, Yeah, it wasn't a
problem at all, honestly. 

611
00:28:39,120 --> 00:28:41,480
So yeah. 
We'll quote you on that one. 

612
00:28:41,480 --> 00:28:45,240
Andrew, we need a quote from you
to say Mongo DB is not a problem

613
00:28:45,240 --> 00:28:47,720
at all, it just works. 
Or something to that effect. 

614
00:28:47,720 --> 00:28:50,080
Definitely, definitely. 
But I'm sure. 

615
00:28:50,080 --> 00:28:52,360
Our marketing team will reach 
out to you at that point. 

616
00:28:52,360 --> 00:28:55,840
So and and look, obviously you 
touched on the fact that you 

617
00:28:55,840 --> 00:28:59,040
didn't know the scheme and the 
flexible scheme of the document 

618
00:28:59,040 --> 00:29:01,960
model helped a lot. 
He touched on scale there as 

619
00:29:02,040 --> 00:29:05,280
well, too, et cetera. 
Was there any other, and you 

620
00:29:05,280 --> 00:29:08,360
touched on the various types of 
documents, small, large and 

621
00:29:08,560 --> 00:29:13,040
humongous documents as well too.
Any other specific challenges in

622
00:29:13,040 --> 00:29:17,040
that in this architecture here 
that you you might have seen if 

623
00:29:17,160 --> 00:29:20,560
at the POC the thousand 
documents but then surfaced up 

624
00:29:20,600 --> 00:29:22,920
other than the different size 
and scale of the documents that 

625
00:29:22,920 --> 00:29:25,320
you already mentioned, I think. 
We're talking about like token 

626
00:29:25,560 --> 00:29:27,480
surfer limitations. 
Yeah, yeah. 

627
00:29:27,520 --> 00:29:31,400
So like, again, you have to 
rewind by your two on this and 

628
00:29:31,400 --> 00:29:34,520
think about when the language 
models were really hitting cloud

629
00:29:34,520 --> 00:29:37,480
vendors in earnest. 
Everyone was trying to use no 

630
00:29:37,480 --> 00:29:41,760
one had enough GP us and so 
everyone was pretty locked down 

631
00:29:41,760 --> 00:29:45,400
on token allocations. 
And so at Oxy, what we did was 

632
00:29:45,560 --> 00:29:48,560
with Azure, we basically just 
had a shared resource stood up 

633
00:29:48,560 --> 00:29:51,360
where SM ES at the company who 
kind of knew what they were 

634
00:29:51,360 --> 00:29:53,640
doing with AI want to try some 
stuff. 

635
00:29:53,840 --> 00:29:57,400
They could get an API token to 
this shared resource and start 

636
00:29:57,400 --> 00:30:02,560
using AI models, which was fine.
You know, we, we want people 

637
00:30:02,560 --> 00:30:04,920
across ox in across domains here
to test stuff. 

638
00:30:04,920 --> 00:30:08,280
But what that caused for us is, 
yeah, we're trying to process 

639
00:30:08,280 --> 00:30:11,080
documents full throttle. 
We're trying to use every token 

640
00:30:11,080 --> 00:30:15,000
we got as they're coming up. 
And people would go run their 

641
00:30:15,000 --> 00:30:16,720
own experiments. 
Maybe they're doing something 

642
00:30:16,720 --> 00:30:19,560
else with documents that's going
to be very demanding or just 

643
00:30:19,560 --> 00:30:21,120
like they're doing their own 
POC. 

644
00:30:21,320 --> 00:30:23,640
And so we started running into 
all sorts of rate limiting 

645
00:30:23,640 --> 00:30:26,760
issues very randomly throughout 
the day, maybe three times a 

646
00:30:26,760 --> 00:30:30,960
day, we, you know, get 420 nines
for 20 minutes and then they'd 

647
00:30:30,960 --> 00:30:32,920
magically go away. 
And we realized what was 

648
00:30:32,920 --> 00:30:35,240
happening was other people are 
trying stuff. 

649
00:30:35,560 --> 00:30:39,080
And so we couldn't really just 
hard code like, you know, some, 

650
00:30:39,240 --> 00:30:42,280
some nice, like perfect amount 
of lambdas to be running to 

651
00:30:42,280 --> 00:30:44,400
where we never ran into token 
issues. 

652
00:30:44,400 --> 00:30:47,320
So the other thing we would do 
is just like if the Lambda 

653
00:30:47,320 --> 00:30:51,040
failed on calling an insured 
resource, that was totally fine.

654
00:30:51,040 --> 00:30:53,440
We weren't going to lose any 
progress because we're 

655
00:30:53,440 --> 00:30:56,680
destroying everything in Mongo 
as it was our state or our state

656
00:30:56,760 --> 00:30:58,240
machine. 
So yeah, that's kind of what we 

657
00:30:58,240 --> 00:30:59,800
did. 
And that worked really well. 

658
00:31:00,120 --> 00:31:03,880
And what I'll call out is like 
we're, we're AAI slash 

659
00:31:03,880 --> 00:31:06,280
engineering kind of team. 
We're not a software team, we're

660
00:31:06,280 --> 00:31:10,080
not a database team. 
So I, I think where Mongo really

661
00:31:10,120 --> 00:31:13,640
worked well for us here was it, 
it made it very easy to take 

662
00:31:13,640 --> 00:31:17,040
this transition from, OK, we 
have a, a working TOC. 

663
00:31:17,040 --> 00:31:19,040
How do we scale that up? 
And each step along the way, 

664
00:31:19,040 --> 00:31:21,240
you're finding new problems and 
you're having to bolt on new 

665
00:31:21,240 --> 00:31:24,400
solutions and new workarounds or
handle edge cases. 

666
00:31:24,560 --> 00:31:27,000
And just because of the 
flexibility and kind of 

667
00:31:27,000 --> 00:31:29,920
effortlessness of Mongo that 
that what Andrew was saying 

668
00:31:29,920 --> 00:31:32,880
before, it was very, very simple
for us to go through this 

669
00:31:32,880 --> 00:31:36,080
scaling process without having 
to do a total rewrite or start 

670
00:31:36,080 --> 00:31:39,320
from scratch or really do some 
major, major work. 

671
00:31:39,320 --> 00:31:40,880
It just made it very, very 
effortless. 

672
00:31:40,880 --> 00:31:43,880
So I think for us, you know, 
kind of from that perspective of

673
00:31:43,880 --> 00:31:47,080
an AI team that's trying to move
fast and change the fly, it was 

674
00:31:47,080 --> 00:31:49,720
the perfect, perfect thing for 
us to use in this case. 

675
00:31:49,840 --> 00:31:51,760
Excellent. 
Well, yeah, you're saying all 

676
00:31:51,760 --> 00:31:54,720
the good things about MongoDB. 
I haven't prompted you at all. 

677
00:31:54,720 --> 00:31:55,960
It's all. 
Yeah. 

678
00:31:55,960 --> 00:31:57,920
I'm not going to get a marketing
phone call after. 

679
00:31:58,080 --> 00:31:58,920
Yeah. 
Yeah, it's. 

680
00:31:58,920 --> 00:32:01,240
No, no, it's all good. 
And it's, you know, obviously 

681
00:32:01,240 --> 00:32:05,440
MongoDB is is cross cloud, but 
you're, you're combining Azure 

682
00:32:05,440 --> 00:32:07,600
and AWS technologies here as 
well. 

683
00:32:07,600 --> 00:32:11,840
Was there any in that kind of 
transform and loops and 

684
00:32:11,840 --> 00:32:15,280
everything else involved here 
any any issues in that space at 

685
00:32:15,280 --> 00:32:18,280
all or just as as with 
everything else was just working

686
00:32:18,280 --> 00:32:22,400
fine scaling up from that POC? 
Honestly the only challenges 

687
00:32:22,640 --> 00:32:27,280
here as far as cross cloud and 
stuff like that was probably 

688
00:32:27,360 --> 00:32:30,400
more internal to oxy. 
Like basically Mongo is a 

689
00:32:30,920 --> 00:32:35,320
marketplace app instead of, you 
know, just native US service. 

690
00:32:35,400 --> 00:32:37,920
And so we we just have to go 
through a couple things in 

691
00:32:37,920 --> 00:32:39,720
supply chain to kind of iron 
that. 

692
00:32:39,720 --> 00:32:43,640
Out and be able to use Mongo 
instincts, but that was very not

693
00:32:43,640 --> 00:32:47,240
a big deal easy for boosted AI 
work that we do that if you ever

694
00:32:47,240 --> 00:32:49,240
see cross cloud you kind of 
scratch your head you think 

695
00:32:49,240 --> 00:32:50,640
maybe that's going to cause some
trouble. 

696
00:32:50,640 --> 00:32:53,160
But generally if we're working 
with text like a lot of heavy 

697
00:32:53,160 --> 00:32:56,120
text data and most of these 
applications and and really the 

698
00:32:56,120 --> 00:32:58,200
amount of data you're sending 
back and forth since it's mostly

699
00:32:58,200 --> 00:33:00,680
text, it ends up being pretty 
small and and not causing a huge

700
00:33:00,680 --> 00:33:02,560
issue. 
There are certainly cases that 

701
00:33:02,560 --> 00:33:05,040
are like video or image driven 
where you might want to stay 

702
00:33:05,040 --> 00:33:08,600
better within the constraints of
11 cloud ecosystem, but for this

703
00:33:08,600 --> 00:33:11,000
it ended up not being an issue. 
OK, great. 

704
00:33:11,000 --> 00:33:13,240
And just for our audience 
joining us, if you've any 

705
00:33:13,240 --> 00:33:16,200
questions for Andrew or Alex or 
anything on the solution that 

706
00:33:16,200 --> 00:33:18,800
they've built, drop them in the 
chat and we'll try and take care

707
00:33:18,800 --> 00:33:20,800
of them as we as we talk through
a bit more. 

708
00:33:21,080 --> 00:33:24,560
So what stage are we at here now
or what state? 

709
00:33:24,560 --> 00:33:29,080
Like have you scanned and 
extracted and classified 

710
00:33:29,160 --> 00:33:32,960
everything that Oxy have now? 
Is that the case and and is it 

711
00:33:32,960 --> 00:33:35,800
the case that this architecture 
did you have up here is just 

712
00:33:35,800 --> 00:33:39,360
invoked as you sign a new 
agreement or lease somewhere? 

713
00:33:39,480 --> 00:33:41,200
So. 
Speaking to this one point. 5 

714
00:33:41,200 --> 00:33:44,600
billion that was like the 
original project scope that that

715
00:33:44,600 --> 00:33:47,600
finished. 
And do you want to talk to kind 

716
00:33:47,600 --> 00:33:52,000
of how this architecture kind of
helps us moving forward to to 

717
00:33:52,000 --> 00:33:54,240
other document projects? 
Yeah. 

718
00:33:54,240 --> 00:33:58,600
So I'd say for one thing is I 
think Alex mentioned early on 

719
00:33:58,680 --> 00:34:02,200
Land's not the only spot in Oxy 
that has lots of things like 

720
00:34:02,200 --> 00:34:05,480
PDF. 
Every company think of like HR, 

721
00:34:05,480 --> 00:34:09,280
legal, supply chain, you've got 
all these agreements, PDFs, 

722
00:34:09,280 --> 00:34:12,199
contracts, but not. 
So we, we actually did have a, a

723
00:34:12,199 --> 00:34:16,880
project with legal where they 
had some needs for OCR and 

724
00:34:16,880 --> 00:34:19,960
extracting, classifying many, 
many PDFs. 

725
00:34:20,120 --> 00:34:24,800
And they've gotten wind of what 
we did here with land and we 

726
00:34:24,800 --> 00:34:26,960
were approached by them and 
they're not to do the same 

727
00:34:26,960 --> 00:34:29,400
thing. 
You know, let's bring on a team 

728
00:34:29,400 --> 00:34:31,159
to help us process all these 
documents. 

729
00:34:31,239 --> 00:34:35,760
And we basically just dragged 
and dropped this thing onto the 

730
00:34:35,760 --> 00:34:39,080
legal documents and, you know, 
updated the prompts. 

731
00:34:39,239 --> 00:34:42,440
Couple things there, but like it
was a very low friction quick 

732
00:34:42,440 --> 00:34:44,600
solution. 
So where this might have taken, 

733
00:34:44,600 --> 00:34:47,320
I don't know, six months from 
start to finish, the legal one 

734
00:34:47,320 --> 00:34:48,760
took about six weeks. 
That was 2. 

735
00:34:49,000 --> 00:34:51,120
Two weeks, sorry, technically 5 
weeks, but one week. 

736
00:34:51,120 --> 00:34:53,280
Was waiting for feedback for the
legal team. 

737
00:34:53,560 --> 00:34:57,800
So we've seen these cases there.
And then on the to answer your 

738
00:34:57,800 --> 00:35:00,200
question more on like, you know,
what are we doing for all 

739
00:35:00,200 --> 00:35:03,600
documents across Oxy. 
I think this did churn up a lot 

740
00:35:03,600 --> 00:35:08,240
of interest in saying, OK, there
is a way we can take all these 

741
00:35:08,240 --> 00:35:11,640
documents we have in file net, 
which is a massive amount. 

742
00:35:11,760 --> 00:35:15,200
And maybe we can start the 
process of getting good OCR, 

743
00:35:15,320 --> 00:35:19,520
extracting useful metadata just 
on a very general scale of how 

744
00:35:19,520 --> 00:35:24,000
we can classify those documents 
and sort of rework these whole 

745
00:35:24,080 --> 00:35:28,320
document foundation into to 
something more up to date. 

746
00:35:28,400 --> 00:35:30,320
So I think that's being 
explored. 

747
00:35:30,320 --> 00:35:33,440
That's a much larger. 
Scope and I'll say we're being a

748
00:35:33,440 --> 00:35:37,880
little bit vague here on purpose
with what we're allowed to talk 

749
00:35:37,880 --> 00:35:40,880
to on some of these upcoming 
projects, but our ongoing 

750
00:35:40,880 --> 00:35:43,320
projects. 
But yeah, suffice to say I think

751
00:35:43,320 --> 00:35:46,800
the core capability here of of 
classifying and extracting 

752
00:35:46,800 --> 00:35:50,560
information from documents is 
super, super scalable to to many

753
00:35:50,560 --> 00:35:54,560
different applications. 
And I think where we see, I 

754
00:35:54,560 --> 00:35:57,960
think this everybody tries, I'm 
not trying to start a fight, but

755
00:35:58,040 --> 00:36:00,360
but I think everybody tries with
their AI project. 

756
00:36:00,360 --> 00:36:03,360
The first thing everybody tries 
is, is like a chat bot or a rag 

757
00:36:03,360 --> 00:36:05,920
based application. 
And everybody goes straight to 

758
00:36:05,920 --> 00:36:09,280
like the harder projects. 
And if I'm going to speak one 

759
00:36:09,280 --> 00:36:12,240
suggestion to anybody dabbling 
with AI and looking for internal

760
00:36:12,240 --> 00:36:15,640
applications, it is start 
simple, start with just 

761
00:36:15,640 --> 00:36:19,120
classification, just extraction,
like really, really easy. 

762
00:36:19,280 --> 00:36:23,240
And then go after those those 
rag or more complex applications

763
00:36:23,240 --> 00:36:24,760
next. 
And that's what we've done. 

764
00:36:24,760 --> 00:36:27,360
And I think doing it in that 
order worked out a lot better 

765
00:36:27,360 --> 00:36:30,240
for us than some of my peers who
I talked to that have tried to 

766
00:36:30,240 --> 00:36:31,960
do the reverse. 
I think a lot of people get 

767
00:36:31,960 --> 00:36:34,880
stuck on that rank project and 
and and spend a lot of waste a 

768
00:36:34,880 --> 00:36:37,120
lot of time. 
OK, OK, that makes sense. 

769
00:36:37,120 --> 00:36:40,640
So in this is like this is the 
extraction classification, but 

770
00:36:40,680 --> 00:36:43,240
tell me a little bit about that 
retrieval process. 

771
00:36:43,240 --> 00:36:46,760
So you you did or didn't do a 
rag approach to getting this 

772
00:36:46,760 --> 00:36:50,800
information back out again for 
whoever needs that is it is what

773
00:36:50,800 --> 00:36:53,880
does it look like to the user 
now to be able to access all of 

774
00:36:53,880 --> 00:36:55,960
these these land documents that 
you would have put through the 

775
00:36:55,960 --> 00:36:58,200
system? 
So with this original project, 

776
00:36:58,200 --> 00:37:01,040
the goal was to get that 
extracted data back into our 

777
00:37:01,040 --> 00:37:05,000
land information systems. 
QLS is the Quoruman system is, 

778
00:37:05,040 --> 00:37:07,840
is the main database we use or 
system we use behind that. 

779
00:37:07,840 --> 00:37:11,280
So once it's back in here, it's 
quite effortless for, for the 

780
00:37:11,680 --> 00:37:14,960
folks in the land department to 
access that data, do automated 

781
00:37:14,960 --> 00:37:17,880
analysis with that data. 
To your original question, which

782
00:37:17,880 --> 00:37:19,440
is what are we doing with with 
RAG? 

783
00:37:19,440 --> 00:37:21,960
Like what's the future for land 
in terms of are we doing 

784
00:37:21,960 --> 00:37:24,000
anything with RAG? 
And Andrew could speak briefly 

785
00:37:24,120 --> 00:37:26,880
to some some stuff with Atlas, I
guess. 

786
00:37:27,000 --> 00:37:28,160
Yeah, I don't know if you want. 
Yeah. 

787
00:37:28,160 --> 00:37:31,680
I mean, so I think like I was 
saying, we had very simple 

788
00:37:31,680 --> 00:37:35,440
objectives early on and now 
we're sort of exploring maybe 

789
00:37:35,440 --> 00:37:38,720
what you might call the more 
interesting use cases like or 

790
00:37:39,040 --> 00:37:41,400
talking to a document and 
getting some information out of 

791
00:37:41,400 --> 00:37:43,000
it. 
And so we do have a couple 

792
00:37:43,000 --> 00:37:46,000
applications for that. 
We've played around with Mongo's

793
00:37:46,000 --> 00:37:50,240
Atlas search index. 
I think Mongo just posted some 

794
00:37:50,240 --> 00:37:53,880
GitHub examples of just like 
basically AI and Mongo. 

795
00:37:53,880 --> 00:37:56,160
So some examples for people to 
follow along with. 

796
00:37:56,160 --> 00:37:58,640
So we'll probably test some of 
those out and see how they do on

797
00:37:58,640 --> 00:38:01,760
some of our documents. 
But yeah, we're, I guess we're 

798
00:38:02,080 --> 00:38:03,240
not sure what the future holds 
here. 

799
00:38:03,880 --> 00:38:05,080
We're going to try some of this 
about. 

800
00:38:05,920 --> 00:38:07,480
Good, good. 
And look, thank you for that 

801
00:38:07,480 --> 00:38:09,520
call out. 
Yes, we have, we have a repo, 

802
00:38:09,520 --> 00:38:13,560
public repo called the Gen. 
AI showcase that a lot of the AI

803
00:38:13,560 --> 00:38:16,760
Devrel folks and indeed a lot of
our engineers and product team 

804
00:38:16,760 --> 00:38:20,560
have been creating examples and 
demos of how to build. 

805
00:38:20,760 --> 00:38:23,760
Obviously you're very familiar 
with vector search, but you 

806
00:38:23,760 --> 00:38:26,560
know, people would know MongoDB 
as a database company back in 

807
00:38:26,560 --> 00:38:28,280
the day. 
And as you said, you know, there

808
00:38:28,280 --> 00:38:31,200
was, you know, two years ago 
there was dedicated vector 

809
00:38:31,360 --> 00:38:34,160
databases and you know, we 
brought vector search capability

810
00:38:34,160 --> 00:38:38,160
to Mongo DB back in June of 23. 
And and since then we've been 

811
00:38:38,280 --> 00:38:41,600
having great success being able 
to kind of participate in this 

812
00:38:41,600 --> 00:38:43,640
Gen. 
AI revolution that we see going 

813
00:38:43,640 --> 00:38:46,720
on at the moment. 
But it I just I and I was really

814
00:38:46,720 --> 00:38:49,920
good answer because I had 
assumed that the retrieval 

815
00:38:49,920 --> 00:38:52,440
portion of this bit was a rag 
type thing. 

816
00:38:52,440 --> 00:38:56,000
But exactly as you said, Alex 
and Andrew, so you it was the 

817
00:38:56,000 --> 00:38:59,080
extracting and the 
classifications that you then 

818
00:38:59,400 --> 00:39:03,120
was powering the existing system
that you already had in Oxy to 

819
00:39:03,120 --> 00:39:06,240
make everything work. 
How hard was that to integrate 

820
00:39:06,240 --> 00:39:09,560
with that old, I'm assuming an 
older existing system, right. 

821
00:39:09,800 --> 00:39:12,400
But it was just getting the data
in the shape that that system 

822
00:39:12,400 --> 00:39:15,200
needed automatically, yeah. 
It was extremely unpleasant. 

823
00:39:15,200 --> 00:39:18,720
A lot of these I think, yeah, 
anybody again, who's working 

824
00:39:18,720 --> 00:39:20,960
with a lot of AI products, a lot
of times the worst part is just 

825
00:39:21,120 --> 00:39:24,000
getting what you need out and 
back into these legacy, legacy 

826
00:39:24,000 --> 00:39:27,040
systems, which often times don't
have like nice API layers for 

827
00:39:27,040 --> 00:39:29,520
communicating and out of them. 
It's it's more awesome than 

828
00:39:29,520 --> 00:39:32,480
like, yeah, just just really 
clunky ways of getting data in 

829
00:39:32,480 --> 00:39:34,280
and out. 
Just to continue my rant, I 

830
00:39:34,280 --> 00:39:38,120
guess about against rag. 
I think the natural instinct, 

831
00:39:38,120 --> 00:39:40,760
what when looking at something 
like this would be to say like, 

832
00:39:40,760 --> 00:39:43,720
yeah, let's do a rag rag app 
where somebody can come and ask 

833
00:39:43,720 --> 00:39:47,160
like, oh, what's what's the how 
many days do I have before this 

834
00:39:47,160 --> 00:39:50,520
lease suffers from a cessation 
of production event and the 

835
00:39:50,520 --> 00:39:52,560
lease is cancelled. 
That would be a really, you 

836
00:39:52,560 --> 00:39:54,920
know, yeah, easy, you know, 
right example. 

837
00:39:54,920 --> 00:39:57,480
You kind of adjust it all. 
You have some filtering that's 

838
00:39:57,560 --> 00:40:00,680
like just just like, sure, space
and some that's more like vector

839
00:40:00,680 --> 00:40:03,480
search base. 
But that whole way of doing it 

840
00:40:03,600 --> 00:40:07,040
still requires a human in the 
loop to go to the chat bot and 

841
00:40:07,040 --> 00:40:09,960
ask the question. 
Much, much better if we have an 

842
00:40:09,960 --> 00:40:13,640
existing system, plus in this 
case, where we can populate all 

843
00:40:13,640 --> 00:40:17,120
the data properly and have some 
automated system working against

844
00:40:17,120 --> 00:40:20,720
that, which is effectively just 
a query or a scheduled task or 

845
00:40:20,720 --> 00:40:22,960
scheduled process where you 
don't have to have a human come 

846
00:40:22,960 --> 00:40:24,920
in and type that and remember to
do that. 

847
00:40:25,320 --> 00:40:28,440
I think, yeah, increasingly 
where we're trying, like if we 

848
00:40:28,440 --> 00:40:30,400
need a human in the loop, we'll 
keep the human in the loop. 

849
00:40:30,400 --> 00:40:33,440
But if we don't, let's let's 
just automate this and make it 

850
00:40:33,440 --> 00:40:36,440
as easy as possible rather than 
having a chat ball where people 

851
00:40:36,440 --> 00:40:38,800
have to go ask questions. 
OK, excellent. 

852
00:40:38,800 --> 00:40:41,880
Which is a nice segue into and 
you touched on this at the 

853
00:40:41,880 --> 00:40:46,000
beginning was you know this 
project after the POC and then 

854
00:40:46,000 --> 00:40:49,000
after you know putting it across
all the other you know kind of 

855
00:40:49,080 --> 00:40:52,000
taking it up to 1,000,000 
million and a half and beyond. 

856
00:40:52,160 --> 00:40:56,200
What was the business impact for
Oxy and what was the return on 

857
00:40:56,200 --> 00:40:58,360
investment? 
You had said that was going to 

858
00:40:58,360 --> 00:41:03,120
potentially be 40 people doing a
much smaller classification 

859
00:41:03,120 --> 00:41:05,000
project. 
Tell us a little bit about how 

860
00:41:05,000 --> 00:41:08,200
those numbers resonated within 
Oxy then, having done this 

861
00:41:08,200 --> 00:41:09,600
project yourselves. 
Yeah. 

862
00:41:09,760 --> 00:41:12,920
So like you said, yeah, a couple
million in savings by averting 

863
00:41:12,920 --> 00:41:16,720
the manual classification effort
got our answers 12 months sooner

864
00:41:16,720 --> 00:41:19,440
because instead of 18 months, 
this was a six month effort. 

865
00:41:19,440 --> 00:41:21,320
Wow, OK, wow. 
That's yeah. 

866
00:41:21,400 --> 00:41:23,600
Cost savings, time savings was 
big on this. 

867
00:41:23,600 --> 00:41:26,440
But but yeah, the larger thing 
is, is now that we're, you know,

868
00:41:26,440 --> 00:41:29,840
this year we're looking at a lot
of more sophisticated AI 

869
00:41:30,080 --> 00:41:33,400
projects and land. 
And I, I think what this first 

870
00:41:33,400 --> 00:41:37,160
project did is it really gave 
the, the wider company and land 

871
00:41:37,160 --> 00:41:39,800
department especially a glimpse 
of, of what this tech can do. 

872
00:41:39,800 --> 00:41:42,960
And it gave us to be honest, an 
idea of what, what this tech can

873
00:41:42,960 --> 00:41:45,000
do. 
And I think it's, it's really 

874
00:41:45,000 --> 00:41:48,080
given us the license to go in 
and look at any process that's 

875
00:41:48,080 --> 00:41:51,360
very repetitive, very manual, 
very unpleasant to do. 

876
00:41:51,360 --> 00:41:54,000
Today we're we're looking at 
automating large parts of those 

877
00:41:54,000 --> 00:41:57,480
projects and maybe maybe in 
eight months or so we'll have 

878
00:41:57,480 --> 00:42:00,120
something else we can share on 
the on the Next Iteration 

879
00:42:00,160 --> 00:42:02,640
podcast about some of that. 
But yeah, I think we started 

880
00:42:02,640 --> 00:42:05,320
simple and I think now land 
knows and we know what's 

881
00:42:05,320 --> 00:42:08,440
possible and there's a lot more,
a lot more getting done. 

882
00:42:08,600 --> 00:42:10,240
Yeah. 
And as you touched on earlier, 

883
00:42:10,400 --> 00:42:13,800
HR got wind of this and legal 
got wind of this. 

884
00:42:13,800 --> 00:42:16,920
And as you said, this, this 
architecture with little or no 

885
00:42:16,920 --> 00:42:20,440
changes was applicable to their 
use cases as well too. 

886
00:42:20,480 --> 00:42:24,520
So obviously this is, you know, 
spearheaded given the cost and 

887
00:42:24,520 --> 00:42:28,600
time saving potentially lots of 
other projects and more work for

888
00:42:28,600 --> 00:42:30,600
you guys, right? 
Yeah, for, for better or for. 

889
00:42:30,600 --> 00:42:33,240
Worse, right? 
Yeah, I had an old boss of mine 

890
00:42:33,240 --> 00:42:35,680
who would say the reward for 
good work is more work, 

891
00:42:35,680 --> 00:42:37,920
basically. 
And I think you 2 essentially 

892
00:42:37,920 --> 00:42:40,920
exemplify that, right? 
You, you did it, it was a third 

893
00:42:40,920 --> 00:42:44,800
of the time with only a few 
people, much, much less in cost.

894
00:42:44,800 --> 00:42:46,840
And now everybody else wants to 
use it, right? 

895
00:42:46,960 --> 00:42:48,560
That's right. 
Excellent. 

896
00:42:48,760 --> 00:42:52,560
So what advice given this 
project and I'm assuming given 

897
00:42:52,560 --> 00:42:56,800
our audience is pretty varied, 
what advice would you have? 

898
00:42:56,960 --> 00:43:00,240
You know, you went in this from 
Apoc saying, look, we got to 

899
00:43:00,240 --> 00:43:03,200
test the OCR, we got to test the
embeddings, we got to see, does 

900
00:43:03,200 --> 00:43:05,320
this work? 
What advice would you have for 

901
00:43:05,600 --> 00:43:09,040
companies or people in a similar
situation looking to to 

902
00:43:09,040 --> 00:43:11,920
modernize, in this instance, 
their document management 

903
00:43:11,920 --> 00:43:14,720
systems, but you know, modernize
their applications and their 

904
00:43:14,720 --> 00:43:16,880
database and using AI? 
That can start. 

905
00:43:17,040 --> 00:43:19,760
I I probably mentioned two 
things, one of which Alex 

906
00:43:19,760 --> 00:43:22,440
already talked on. 
What else talked about where 

907
00:43:22,600 --> 00:43:25,920
it's much better to just do 
something than have to kind of 

908
00:43:25,920 --> 00:43:29,320
have a hybrid solution of humans
still have to interact with it 

909
00:43:29,480 --> 00:43:32,160
and now it's doing it. 
But like when you, I guess when 

910
00:43:32,160 --> 00:43:35,560
you introduce tooling basically 
forces a human to learn 

911
00:43:35,720 --> 00:43:39,400
something else, that's something
usually faces some opposition. 

912
00:43:39,520 --> 00:43:43,600
Whereas if you introduce some 
tooling that instead says, hey, 

913
00:43:43,720 --> 00:43:47,240
you used to do this thing that 
you didn't enjoy, and now you 

914
00:43:47,240 --> 00:43:48,440
just don't have to do that 
anymore. 

915
00:43:48,480 --> 00:43:50,040
That's how much. 
People prefer that. 

916
00:43:50,360 --> 00:43:52,800
Yeah, like that's my ears, 
definitely. 

917
00:43:53,120 --> 00:43:55,960
Solution. 
You're not like, hey, check out 

918
00:43:55,960 --> 00:43:58,920
this training for this new front
end we go and it'll make your 

919
00:43:58,920 --> 00:44:01,480
life easier because like in that
person's head, they're just 

920
00:44:01,480 --> 00:44:03,440
thinking all this is one more 
training I have to do, one more 

921
00:44:03,440 --> 00:44:06,280
thing I have to learn something 
I still have to keep doing. 

922
00:44:06,400 --> 00:44:10,600
And so one just philosophy I 
think our team has taken is how 

923
00:44:10,600 --> 00:44:12,960
can we just cut processes all 
together? 

924
00:44:13,120 --> 00:44:16,840
You know, let's not think about 
like necessarily enhancing it 

925
00:44:17,000 --> 00:44:20,240
unless that enhancement means 
cutting workout completely. 

926
00:44:20,360 --> 00:44:22,480
So that's one thing. 
And then the other thing I've 

927
00:44:22,480 --> 00:44:25,760
seen on this team is I think we 
have a lot of but people who are

928
00:44:25,760 --> 00:44:27,480
willing to wear a lot of 
different hats. 

929
00:44:27,600 --> 00:44:31,720
And so we don't we don't have 
like, you know, Adba who does 

930
00:44:31,720 --> 00:44:34,800
all the manga work. 
And once the manga stuff feels 

931
00:44:34,800 --> 00:44:38,720
good, we have a cloud engineer 
who does only cloud stuff. 

932
00:44:38,720 --> 00:44:40,800
And then once that seems good, 
we have a dev OPS guy. 

933
00:44:40,800 --> 00:44:43,960
And so, you know, it's everyone 
is doing everything. 

934
00:44:44,000 --> 00:44:46,360
And yeah, different people have 
different strengths and that's 

935
00:44:46,360 --> 00:44:48,120
great to lean on that. 
But at the end of the day, I 

936
00:44:48,120 --> 00:44:53,120
think most of the team could 
probably do this whole project 

937
00:44:53,160 --> 00:44:57,000
start to finish. 
And and that makes that gives us

938
00:44:57,000 --> 00:45:00,560
a lot of flexibility on how 
quickly we can move stuff we can

939
00:45:00,560 --> 00:45:02,760
get done. 
And so, yeah, that's that's been

940
00:45:02,760 --> 00:45:04,000
something that's really helped 
us out. 

941
00:45:04,000 --> 00:45:07,200
It's just people are willing to 
figure out how to do stuff in a 

942
00:45:07,440 --> 00:45:10,240
fairly constrained environment 
and with Maple skills they 

943
00:45:10,240 --> 00:45:12,040
didn't have before, but they're 
willing to go learn. 

944
00:45:12,320 --> 00:45:14,320
Yeah, we were able to move 
really quickly on all this 

945
00:45:14,320 --> 00:45:15,320
stuff. 
Accent. 

946
00:45:15,440 --> 00:45:17,320
Yeah, I think that's really good
advice, yeah. 

947
00:45:17,480 --> 00:45:19,440
Yeah. 
The only other third like piece 

948
00:45:19,440 --> 00:45:23,240
of advice that I would add is is
going on what Andrew said is 

949
00:45:23,240 --> 00:45:25,080
just try things. 
I'm just experiment. 

950
00:45:25,080 --> 00:45:27,360
I think the only way you learn 
what these models can and can't 

951
00:45:27,360 --> 00:45:29,960
do is by experimenting with it. 
I think it's so new and it's 

952
00:45:29,960 --> 00:45:31,840
changing so fast. 
You have to try things. 

953
00:45:31,840 --> 00:45:34,600
I think we've had cases where 
some team comes with us, comes 

954
00:45:34,600 --> 00:45:36,840
to us with a problem and says 
hey, can't do this. 

955
00:45:36,840 --> 00:45:38,600
And we say, yeah, let's give it 
a try. 

956
00:45:38,600 --> 00:45:41,360
We try and we said, yeah, we 
it's kind of working, but I 

957
00:45:41,360 --> 00:45:43,960
think to get it to the accuracy 
you need, this might take six 

958
00:45:43,960 --> 00:45:46,840
weeks and then and then, you 
know, we kind of leave it and 

959
00:45:46,840 --> 00:45:48,880
then and then two weeks later a 
new model comes out. 

960
00:45:49,000 --> 00:45:52,640
We tried with a new model and it
just one shots it 100% accuracy,

961
00:45:52,640 --> 00:45:54,280
no extra work required on our 
side. 

962
00:45:54,280 --> 00:45:56,800
So the only way you'll learn 
what these things can and can't 

963
00:45:56,800 --> 00:45:59,800
do is just by experimenting, by 
trying it out and, and it's a 

964
00:45:59,800 --> 00:46:01,400
lot of fun. 
I think you just have to like, 

965
00:46:01,400 --> 00:46:04,000
like you kind of picked up on 
earlier, you have to de risk, 

966
00:46:04,000 --> 00:46:06,600
you have to do short POC's, you 
have to move fast. 

967
00:46:06,800 --> 00:46:09,200
And like Andrew said, that the 
team design needs to be such 

968
00:46:09,200 --> 00:46:12,320
that you're not having to have 
committees and multiple teams 

969
00:46:12,320 --> 00:46:14,680
meetings and scrum cycles and 
stuff. 

970
00:46:14,680 --> 00:46:17,160
None of that's going to work. 
fast-paced, fast-paced. 

971
00:46:17,280 --> 00:46:19,640
You just have to be able to move
quick a small team. 

972
00:46:19,840 --> 00:46:22,400
OK. 
And on that I suppose the new 

973
00:46:22,400 --> 00:46:25,680
and the fast and how things are 
changing so quickly, either of 

974
00:46:25,680 --> 00:46:28,920
you, what's what's the most 
exciting thing you're seeing now

975
00:46:28,920 --> 00:46:33,080
even outside of Oxy in terms of 
what AI is capable of doing? 

976
00:46:33,080 --> 00:46:36,160
Is there anything that has you 
really excited and potentially 

977
00:46:36,160 --> 00:46:39,200
you want to play around with 
even in your off hours, right, 

978
00:46:39,200 --> 00:46:41,280
not during the daytime perhaps 
as well too. 

979
00:46:41,280 --> 00:46:44,240
What's exciting you most? 
I'll say for me first, it's, 

980
00:46:44,240 --> 00:46:47,120
it's the reasoning models. 
I think we saw this the, you 

981
00:46:47,360 --> 00:46:49,800
know, we look at the big 
multimodal models, big, big 

982
00:46:49,960 --> 00:46:52,480
language models there, there 
were these, there was a scaling 

983
00:46:52,480 --> 00:46:56,000
paradigm that got us to GPT 3 
where we just had more and more 

984
00:46:56,000 --> 00:46:57,440
data. 
We scaled up free training. 

985
00:46:57,440 --> 00:47:00,240
We got these really huge and 
awesome models, but they weren't

986
00:47:00,320 --> 00:47:03,240
super applicable or useful yet. 
And then we saw the the 

987
00:47:03,560 --> 00:47:06,400
basically supervised fine tuning
feedback or our league feedback 

988
00:47:06,400 --> 00:47:09,480
that got us to GPT 3.5 ChatGPT. 
And that was really, really 

989
00:47:09,480 --> 00:47:11,360
exciting. 
And now we're kind of exiting 

990
00:47:11,360 --> 00:47:13,240
those two regimes. 
Like, you know, they're not, not

991
00:47:13,240 --> 00:47:15,480
to say they're fully exhausted, 
but we're kind of seeing that 

992
00:47:15,600 --> 00:47:17,800
signs of exhaustion of those 
scaling regimes. 

993
00:47:17,800 --> 00:47:20,320
And now we're moving to this, 
this reasoning based scaling, 

994
00:47:20,320 --> 00:47:23,640
which for any kind of verifiable
domain like like coding or 

995
00:47:23,640 --> 00:47:26,840
mathematics or some fundamental 
sciences, I, I think it's really

996
00:47:26,840 --> 00:47:29,200
quite promising what's going to 
be possible with these reasoning

997
00:47:29,200 --> 00:47:31,200
models. 
I think for us, where we had 

998
00:47:31,200 --> 00:47:34,880
seen some limitations on certain
types of tasks that required, 

999
00:47:35,040 --> 00:47:37,880
I'll give an example pertaining 
to land, which is a lot of the 

1000
00:47:37,880 --> 00:47:41,320
work we presented today was on, 
on single documents as a unit, 

1001
00:47:41,480 --> 00:47:42,960
right? 
We're having the language model 

1002
00:47:42,960 --> 00:47:46,520
classify and and extract on a 
single document at a time. 

1003
00:47:46,520 --> 00:47:49,400
And where we saw struggling in 
the past with previous models is

1004
00:47:49,400 --> 00:47:52,320
in cross document reasoning. 
And I'll say with the new 

1005
00:47:52,400 --> 00:47:55,760
reasoning models, we're seeing a
lot better results on cross 

1006
00:47:55,760 --> 00:47:57,840
document work. 
And I think that's going to be a

1007
00:47:57,840 --> 00:48:01,920
huge thing for for this specific
subject, but also for many, many

1008
00:48:01,920 --> 00:48:03,840
other topics. 
I think this, this, these 

1009
00:48:03,840 --> 00:48:06,440
reasoning models are going to be
a very, very impactful. 

1010
00:48:06,640 --> 00:48:08,480
Yeah, it's a good point. 
We've been playing around with 

1011
00:48:08,480 --> 00:48:11,200
say like the reasoning engine 
from Google etcetera as well 

1012
00:48:11,200 --> 00:48:13,040
too. 
And it's, yeah, I think that 

1013
00:48:13,040 --> 00:48:15,960
moves us into a new sphere 
because as you say, a lot of a 

1014
00:48:15,960 --> 00:48:18,080
lot of the applications are 
built, you know, they're they're

1015
00:48:18,080 --> 00:48:20,760
looking at a single force of 
truth, essentially. 

1016
00:48:20,960 --> 00:48:24,480
And in fact, as they grow and 
get bigger, is there anything we

1017
00:48:24,480 --> 00:48:27,760
hear agentic all the time now, 
anything in that space that 

1018
00:48:27,760 --> 00:48:29,560
excites you? 
Definitely trying to think of 

1019
00:48:29,560 --> 00:48:32,080
what I can say specifically. 
No worries. 

1020
00:48:32,160 --> 00:48:36,360
No, But yeah, what I'll say is 
for an agentic system to be 

1021
00:48:36,360 --> 00:48:39,840
successful, the standard for 
accuracy goes up a lot because 

1022
00:48:39,840 --> 00:48:42,200
generally with agentic systems, 
you're seeing multiple steps, 

1023
00:48:42,200 --> 00:48:43,520
right? 
So you'd have, you know, sit, 

1024
00:48:43,520 --> 00:48:46,040
say, instead of just an input 
output, you might instead have 

1025
00:48:46,040 --> 00:48:50,560
like a 5A5 step process where if
you're at 98% accuracy with each

1026
00:48:50,560 --> 00:48:54,240
of those steps, so you need to 
multiply .98 * .98 * .98, and 

1027
00:48:54,240 --> 00:48:56,840
then you're going to get maybe a
much worse accuracy on the other

1028
00:48:56,840 --> 00:49:00,600
side of that. 
So I'd say it's a much harder 

1029
00:49:00,720 --> 00:49:03,800
application to build an agentic 
system, but I think there's a 

1030
00:49:04,000 --> 00:49:06,240
ton of value there and we're 
seeing a ton of value there. 

1031
00:49:06,920 --> 00:49:08,880
I just want to start I guess 
it's all. 

1032
00:49:09,440 --> 00:49:11,280
Yeah, Yeah. 
I look, there's, there's a lot 

1033
00:49:11,280 --> 00:49:13,520
of ways to go. 
And then look, as you said, it's

1034
00:49:13,520 --> 00:49:15,640
moving so fast and everything's 
so new. 

1035
00:49:15,640 --> 00:49:18,480
I think you're going to see 
frameworks for Agentic come out 

1036
00:49:18,480 --> 00:49:21,680
to help you do that because you 
need to safeguards in place, 

1037
00:49:21,680 --> 00:49:24,920
right, to make sure it doesn't 
go off a Cliff at the other end 

1038
00:49:24,920 --> 00:49:28,280
if you've left it unmanned as it
were, to do its own reasoning 

1039
00:49:28,280 --> 00:49:30,640
and and make its way through 
number of stages, I suppose. 

1040
00:49:30,680 --> 00:49:32,160
Yeah. 
And I'll say kind of to that 

1041
00:49:32,160 --> 00:49:35,080
point, you know, in oil and gas,
we also need to be conscious 

1042
00:49:35,080 --> 00:49:37,640
about, you know, we have to be 
picky about our application. 

1043
00:49:37,960 --> 00:49:40,680
I think, I think there's there's
projects that we've looked at on

1044
00:49:40,680 --> 00:49:42,800
our team or we've been 
approached internally for, for 

1045
00:49:42,800 --> 00:49:45,320
projects where where we really 
need to have a human in the 

1046
00:49:45,320 --> 00:49:47,080
loop. 
It's not seem to have an AI 

1047
00:49:47,080 --> 00:49:50,480
system like something start to 
finish, you know, with land 

1048
00:49:50,480 --> 00:49:53,120
document classification. 
I think the risks are a lot more

1049
00:49:53,120 --> 00:49:54,600
limited. 
You know, maybe you're limited 

1050
00:49:54,600 --> 00:49:56,920
to financial risk, worst case 
scenario if you mess something 

1051
00:49:56,920 --> 00:49:58,120
up. 
But but yeah, certainly with 

1052
00:49:58,120 --> 00:50:00,800
things affecting equipment in 
the field or or safety 

1053
00:50:00,800 --> 00:50:03,400
operations or personnel 
decisions in the field, I think 

1054
00:50:03,400 --> 00:50:06,480
we rightfully are acting a lot 
more cautiously. 

1055
00:50:06,480 --> 00:50:08,480
And I'd encourage others to 
actually a lot more cautiously 

1056
00:50:08,480 --> 00:50:10,760
in those applications. 
Sure, sure. 

1057
00:50:11,040 --> 00:50:14,200
And in the same vein of 
everything moving new and fast, 

1058
00:50:14,200 --> 00:50:17,440
how Andrew, like how do you keep
up with what's going on in the 

1059
00:50:17,640 --> 00:50:20,680
AI space? 
For example, is it, is it blogs?

1060
00:50:20,680 --> 00:50:22,960
Is it news feeds? 
Is it podcasts? 

1061
00:50:23,160 --> 00:50:24,440
How do you keep on top of 
things? 

1062
00:50:24,680 --> 00:50:28,360
Yeah, For me, I mainly just 
listen to what people on my team

1063
00:50:28,360 --> 00:50:31,360
are saying, OK? 
I mean, we've got some guys who 

1064
00:50:31,360 --> 00:50:34,720
are actually just had a guy who 
is basically scanning the 

1065
00:50:35,040 --> 00:50:37,640
archive and the Columbia 
University. 

1066
00:50:38,320 --> 00:50:40,640
I just. 
Supposed to not I forget but but

1067
00:50:40,640 --> 00:50:42,440
in any case, yes, scanning for 
papers. 

1068
00:50:42,440 --> 00:50:45,280
And through papers, sending out 
any interesting papers every 

1069
00:50:45,280 --> 00:50:46,040
week. 
OK. 

1070
00:50:46,280 --> 00:50:49,600
So like he's passed in on? 
Top of the Alex stays on top. 

1071
00:50:49,600 --> 00:50:53,160
I'd say it's like it's reading 
papers and tech, you know, ML 

1072
00:50:53,160 --> 00:50:56,640
Twitter, which you have to be 
kind of ruthless about curating 

1073
00:50:56,640 --> 00:50:58,960
your your Twitter feed to keep 
it tech focused. 

1074
00:50:58,960 --> 00:51:02,840
But but that those are the two 
big sources for me is papers and

1075
00:51:02,840 --> 00:51:03,600
Twitter. 
Yeah. 

1076
00:51:03,600 --> 00:51:05,600
Because there's an awful lot 
going on. 

1077
00:51:05,600 --> 00:51:08,040
There's, as you said earlier, 
new things coming out all the 

1078
00:51:08,040 --> 00:51:09,480
time. 
There's a lot of leapfrogging, 

1079
00:51:09,520 --> 00:51:12,040
you know, you think, oh, this is
gonna be the solution we need. 

1080
00:51:12,040 --> 00:51:14,280
Then you see something else 
that's easier. 

1081
00:51:14,480 --> 00:51:17,520
Just just today frog 3 came out 
so now it's like, yeah, there 

1082
00:51:17,560 --> 00:51:19,520
goes my afternoon. 
I'm gonna have to be testing the

1083
00:51:19,520 --> 00:51:21,520
latest. 
Yeah, the new model that looks 

1084
00:51:21,520 --> 00:51:24,360
to be a bit better than 40 for 
for non reasoning model. 

1085
00:51:24,360 --> 00:51:25,920
It's it's probably the strongest
now. 

1086
00:51:25,960 --> 00:51:29,160
So yeah, it keeps it exciting. 
Just just, yeah, like, like you 

1087
00:51:29,160 --> 00:51:32,920
said, it's a lot of nights and 
weekends to to on top of what's 

1088
00:51:32,920 --> 00:51:35,200
changing in this fast baby. 
Excellent. 

1089
00:51:35,200 --> 00:51:37,800
Well, look, this has been 
fascinating and it's an amazing 

1090
00:51:37,800 --> 00:51:40,880
to see, I suppose. 
Look, I, I do this live stream 

1091
00:51:40,880 --> 00:51:43,480
an awful lot and a lot of the 
time we've got demos and we've 

1092
00:51:43,480 --> 00:51:46,920
got some Pocs etcetera. 
But to see it in action, to see 

1093
00:51:46,920 --> 00:51:50,640
it at the scale that you've 
managed to do this in Oxy to, to

1094
00:51:50,640 --> 00:51:55,440
see the proof of concept borne 
out into fruition, as in it was 

1095
00:51:55,440 --> 00:51:57,920
quicker. 
It didn't need 40 people and we 

1096
00:51:57,920 --> 00:51:59,320
did it. 
And everybody else wants to 

1097
00:51:59,320 --> 00:52:02,440
leverage our architecture going 
forward as well too. 

1098
00:52:02,480 --> 00:52:05,560
Is brilliant to see. 
And you know, you've explained 

1099
00:52:05,560 --> 00:52:08,560
your own processes internally 
really, really well. 

1100
00:52:08,720 --> 00:52:12,240
Any last, as we come to a close 
on the podcast, any last parting

1101
00:52:12,240 --> 00:52:15,680
comments for for our audience or
kind of words of wisdom or 

1102
00:52:15,680 --> 00:52:18,760
things you've learned or even 
looking back on this project, is

1103
00:52:18,760 --> 00:52:20,920
there anything that you would 
have done differently or 

1104
00:52:20,920 --> 00:52:24,120
anything that you really 
stumbled on or everything super 

1105
00:52:24,120 --> 00:52:26,800
smoothly and no problems Do it 
all again the same way. 

1106
00:52:27,000 --> 00:52:30,480
I think, I think the only thing 
I would say is one thing I've 

1107
00:52:30,480 --> 00:52:33,480
really felt especially just with
the app end of is language 

1108
00:52:33,480 --> 00:52:38,240
models has been having more of 
an understanding of a solution 

1109
00:52:38,240 --> 00:52:41,800
Architect has been, I think 
really helpful because it's, 

1110
00:52:42,040 --> 00:52:46,680
it's become so simple to dive 
deep into something you want to 

1111
00:52:46,680 --> 00:52:50,960
know how to do, but you do have 
to know what questions to ask 

1112
00:52:50,960 --> 00:52:53,160
kind of, I've found a lot of 
value. 

1113
00:52:53,400 --> 00:52:55,840
It's kind of kind of made me 
change my philosophy a little 

1114
00:52:55,840 --> 00:53:00,080
bit on learning going forward is
like, you know, do you spend the

1115
00:53:00,080 --> 00:53:03,960
time to become truly an expert 
at one thing or can you get a 

1116
00:53:03,960 --> 00:53:07,560
lot more value from sort of 
knowing a whole lot of different

1117
00:53:07,560 --> 00:53:08,960
things and how they work 
together? 

1118
00:53:09,040 --> 00:53:12,480
More of like a system design 
instead of, you know, deep SME, 

1119
00:53:12,480 --> 00:53:14,920
whatever it might be. 
And so I think there's a balance

1120
00:53:14,920 --> 00:53:16,520
there. 
But yeah, I've really felt 

1121
00:53:16,520 --> 00:53:20,440
lately that just just knowing 
how systems work and seeing how 

1122
00:53:20,440 --> 00:53:22,880
they should be built, it's 
become a lot easier to actually 

1123
00:53:22,880 --> 00:53:25,560
go build the system. 
But kind of understanding the 

1124
00:53:25,560 --> 00:53:28,800
edge cases, what to watch out 
for, where things go wrong, all 

1125
00:53:28,800 --> 00:53:30,200
that stuff. 
You really do have to have a 

1126
00:53:30,200 --> 00:53:32,880
good understanding out of where 
things will go very wrong. 

1127
00:53:32,960 --> 00:53:36,200
But yeah, building has become 
much simpler than it used to in 

1128
00:53:36,200 --> 00:53:38,480
the past, so I don't know. 
I'm still trying to figure that 

1129
00:53:38,480 --> 00:53:41,720
one out myself, but exploring. 
And Alex, anything to add to 

1130
00:53:41,720 --> 00:53:43,320
that? 
Yeah, I'll just, I'll just say 

1131
00:53:43,680 --> 00:53:48,160
there's one thing on my wish 
list for 2025 and I intend to do

1132
00:53:48,160 --> 00:53:50,400
a better job at this. 
I want to see more open source 

1133
00:53:50,400 --> 00:53:53,000
work and oil and gas. 
I think there's a few companies 

1134
00:53:53,200 --> 00:53:56,280
doing a really good job at this.
But that's my call to action for

1135
00:53:56,360 --> 00:53:59,240
for everybody listening and also
just for me personally, I think 

1136
00:53:59,240 --> 00:54:02,080
I think we need to share more. 
I think it's exciting to see 

1137
00:54:02,080 --> 00:54:04,720
what other people are helping, 
so I hope to see more of that. 

1138
00:54:04,720 --> 00:54:06,240
This year, Yeah. 
Look, I'm a big fan. 

1139
00:54:06,240 --> 00:54:09,120
It's, it's always, you know, 
great to to learn from other 

1140
00:54:09,120 --> 00:54:11,640
people and to see what they're 
building and to have that shared

1141
00:54:11,640 --> 00:54:13,920
experience and that sense of 
community. 

1142
00:54:13,920 --> 00:54:16,960
I suppose as well, too, that, 
you know, as, as you say, large 

1143
00:54:16,960 --> 00:54:19,280
companies, sometimes that's hard
to foster, Right. 

1144
00:54:19,280 --> 00:54:23,680
But I think, you know, everybody
is trying to solve not much the 

1145
00:54:23,680 --> 00:54:26,640
same problem, but, you know, 
architecturally much the same 

1146
00:54:26,640 --> 00:54:28,240
problem. 
And we could all learn from the 

1147
00:54:28,240 --> 00:54:29,280
collected. 
Yeah. 

1148
00:54:29,400 --> 00:54:32,000
Perfectly perfect. 
Well, look, this has been super 

1149
00:54:32,000 --> 00:54:34,000
informative. 
I've learned a lot, as I always 

1150
00:54:34,000 --> 00:54:35,600
do. 
It's the reason why I do these 

1151
00:54:35,600 --> 00:54:38,480
podcasts is from my own 
education as well too. 

1152
00:54:38,480 --> 00:54:41,320
But Andrew and Alex, you've made
it incredibly clear to follow 

1153
00:54:41,320 --> 00:54:44,000
the pathway. 
And, you know, as a really, as I

1154
00:54:44,000 --> 00:54:46,640
said earlier, a really good 
example of taking that, you 

1155
00:54:46,640 --> 00:54:49,920
know, understanding that 
inclination that I think we can 

1156
00:54:49,920 --> 00:54:53,680
do this in a different way. 
And then into the POC and then 

1157
00:54:53,680 --> 00:54:57,080
into the the actual production. 
And to take, you know, I suppose

1158
00:54:57,080 --> 00:55:00,840
that drudgery out of that 
classification and extracting 

1159
00:55:00,840 --> 00:55:04,080
and to help the, the collective 
team there and, and go, look, 

1160
00:55:04,240 --> 00:55:06,960
it's all in your system, work 
away as you always did, but it's

1161
00:55:06,960 --> 00:55:09,040
now just easier. 
So I think that's a brilliant, 

1162
00:55:09,040 --> 00:55:11,040
brilliant example. 
And very much, I think you 

1163
00:55:11,040 --> 00:55:13,200
alluded to, Alex, there might be
something in the future. 

1164
00:55:13,200 --> 00:55:15,600
So as soon as you build 
something new and different, 

1165
00:55:15,600 --> 00:55:18,480
hopefully with a bit of Mongo DB
in the background there as well 

1166
00:55:18,480 --> 00:55:19,880
too. 
We'll certainly get your back on

1167
00:55:19,880 --> 00:55:22,640
the podcast to show show it 
again as well too. 

1168
00:55:22,880 --> 00:55:25,720
But Alex, Andrew, thank you so 
much for your time. 

1169
00:55:25,720 --> 00:55:27,960
We do appreciate it. 
Thank you for everybody who 

1170
00:55:27,960 --> 00:55:30,080
joined us as well too, and this 
is great. 

1171
00:55:30,080 --> 00:55:33,080
And as I said at the intros, do 
keep up to date with what we're 

1172
00:55:33,080 --> 00:55:36,200
doing on these shows. 
By liking and subscribing on our

1173
00:55:36,200 --> 00:55:39,400
YouTube channel and following us
on LinkedIn, you will be able to

1174
00:55:39,400 --> 00:55:42,000
keep up to date with future 
episodes like this. 

1175
00:55:42,000 --> 00:55:45,000
And if you want to keep up to 
date with you know, our examples

1176
00:55:45,000 --> 00:55:47,640
that we mentioned earlier and 
other case studies, etcetera as 

1177
00:55:47,640 --> 00:55:50,480
well too, just go to 
developer.mongodb.com. 

1178
00:55:50,480 --> 00:55:53,800
You'll find everything that our 
Deborah team and others with 

1179
00:55:53,800 --> 00:55:57,960
inside MongoDB create and build.
And if you're not already using 

1180
00:55:57,960 --> 00:56:01,760
MongoDB, but you'd like to get 
started, we've a load of help. 

1181
00:56:01,760 --> 00:56:04,680
We've got a great forum, we've 
got a great community where all 

1182
00:56:04,680 --> 00:56:08,160
our users hang out and obviously
our engineers and our Devrel 

1183
00:56:08,160 --> 00:56:09,960
team and our product managers as
well too. 

1184
00:56:09,960 --> 00:56:13,560
So community at mongodb.com. 
That's, that's the end of my 

1185
00:56:13,560 --> 00:56:15,840
plugs. 
Any parting words, Andrew or 

1186
00:56:15,840 --> 00:56:16,960
Alex? 
Thanks for having us. 

1187
00:56:17,000 --> 00:56:17,920
Yeah. 
Thanks a lot, Shane. 

1188
00:56:18,080 --> 00:56:20,040
Spotify. 
No, listen, you made my job 

1189
00:56:20,040 --> 00:56:22,040
super easy. 
It's been a pleasure to host you

1190
00:56:22,040 --> 00:56:24,200
both. 
And yeah, I look forward to see 

1191
00:56:24,200 --> 00:56:27,040
what you build in the future. 
But for now, thank you so much. 

1192
00:56:27,040 --> 00:56:28,800
It's been a pleasure. 
Take care of you, everybody. 

1193
00:56:28,800 --> 00:56:30,160
Andrew, Alex, thank you. 
Thanks.

