1
00:00:00,040 --> 00:00:03,400
Something we were talking about 
a little bit earlier around Open

2
00:00:03,400 --> 00:00:06,880
the Eyes, release of guaranteed 
capacity. 

3
00:00:06,880 --> 00:00:10,120
That just happened, right? 
I find it really interesting 

4
00:00:10,120 --> 00:00:12,520
because it's not advertised as a
discount mechanism. 

5
00:00:12,520 --> 00:00:16,520
It's, it's advertised as a, you 
know, guaranteed capacity, you 

6
00:00:16,560 --> 00:00:19,320
know, and, and we're just. 
Saying like, if you need this, 

7
00:00:19,320 --> 00:00:22,200
you're always going to have it, 
so you're never going to be 

8
00:00:22,200 --> 00:00:24,560
without resources. 
We're not going to go down on 

9
00:00:24,560 --> 00:00:25,560
you. 
Exactly. 

10
00:00:26,360 --> 00:00:28,840
But from our standpoint, we 
haven't really seen capacity 

11
00:00:28,840 --> 00:00:30,400
issues. 
So it's almost like they're 

12
00:00:30,400 --> 00:00:32,720
they're solving a problem that 
doesn't exist yet. 

13
00:00:32,880 --> 00:00:36,680
That has to be very hard for you
to look at and it probably makes

14
00:00:36,680 --> 00:00:39,320
your job like you have a fixed 
price you can't go over. 

15
00:00:39,320 --> 00:00:41,400
I mean, there's lots of ways to 
address that. 

16
00:00:41,480 --> 00:00:44,240
At the same time you, you see 
the subscription model going 

17
00:00:44,240 --> 00:00:47,440
away and you know, usage base or
a combination of that, those 

18
00:00:47,440 --> 00:00:49,360
things happening. 
And that's the reason why 

19
00:00:49,360 --> 00:00:52,520
because it's fixed, prices just 
don't work anymore. 

20
00:01:04,519 --> 00:01:06,720
All right. 
I'm here with Josh working at 

21
00:01:06,720 --> 00:01:10,000
Superhuman leading the fin OPS. 
I want to hear about your 

22
00:01:10,000 --> 00:01:14,280
journey into tokenomics. 
It's been about a three-year 

23
00:01:14,280 --> 00:01:19,880
journey overall, but but yeah, 
prior to Superhuman, and I'm at 

24
00:01:19,880 --> 00:01:21,960
the time when I joined them, 
they were Grammarly. 

25
00:01:22,720 --> 00:01:26,520
I was at AWS and, you know, had 
no exposure to AI when I was 

26
00:01:26,520 --> 00:01:28,600
really there outside of just 
managing some of the billing 

27
00:01:28,600 --> 00:01:30,840
aspects of it. 
Come to Grammar Lane, you know, 

28
00:01:30,840 --> 00:01:33,240
we have some hosted models that 
I was helping, you know, 

29
00:01:33,640 --> 00:01:35,640
distribute the cost and, and 
report on. 

30
00:01:36,000 --> 00:01:39,240
And then on early 2023, we 
started using this new thing 

31
00:01:39,240 --> 00:01:43,320
called Azure Open AI. 
So I started kind of poking 

32
00:01:43,320 --> 00:01:44,640
around. 
I'm like, I think this is going 

33
00:01:44,640 --> 00:01:47,360
to be kind of big. 
So I want to understand like if 

34
00:01:47,360 --> 00:01:51,960
there's a area where I can help.
So I started kind of poking 

35
00:01:51,960 --> 00:01:54,680
around, asking questions and 
found all this is like a fixed 

36
00:01:54,680 --> 00:01:57,040
capacity that we're, you know, 
we want to make sure we're 

37
00:01:57,040 --> 00:02:00,440
utilizing enough of it. 
And, and, you know, also 

38
00:02:00,440 --> 00:02:04,040
understanding, oh, if we don't 
have enough is this will create 

39
00:02:04,040 --> 00:02:05,680
an outage for some of our 
products. 

40
00:02:05,680 --> 00:02:08,440
So we got it. 
So introduced some new concepts 

41
00:02:08,440 --> 00:02:11,680
like I've never been involved 
with making decisions that could

42
00:02:11,680 --> 00:02:14,120
actually cause an outage on one 
of our products, right. 

43
00:02:14,120 --> 00:02:16,040
That wasn't like a thin up 
thing, but it kind of fell into 

44
00:02:16,040 --> 00:02:18,040
my scope. 
So, but as I started doing that,

45
00:02:18,040 --> 00:02:22,800
I realized after going to some 
Fin OPS Foundation events that I

46
00:02:22,800 --> 00:02:26,000
was one of the only ones doing 
that type of work back then. 

47
00:02:26,920 --> 00:02:31,200
So I certainly encourage anyone 
to, you know, even definitely in

48
00:02:31,200 --> 00:02:34,280
this day and time with AI, just 
most practitioners, if you poke 

49
00:02:34,280 --> 00:02:37,880
around hard enough, you'll find 
a area where you can drop a lot 

50
00:02:37,880 --> 00:02:41,240
of impact for sure and help. 
But it's funny how you started 

51
00:02:41,240 --> 00:02:43,080
poking around and then you 
realized, wait a minute, 

52
00:02:43,080 --> 00:02:47,640
resource management is kind of 
seems like my job in this whole 

53
00:02:47,640 --> 00:02:49,400
new field. 
Yeah, resource management, a 

54
00:02:49,400 --> 00:02:51,320
very, very expensive resource, 
yeah. 

55
00:02:51,760 --> 00:02:53,840
Especially in those days, yeah, 
maybe talk to me about the 

56
00:02:53,840 --> 00:02:57,520
trends and how you saw the 
trends dropping and then now 

57
00:02:57,520 --> 00:03:00,280
reversing of the trend of the 
cost and. 

58
00:03:00,320 --> 00:03:03,400
You look at me a year or two 
ago, I'm like, oh, costs are 

59
00:03:03,400 --> 00:03:06,080
just falling off a Cliff and now
it's just going to be the the 

60
00:03:06,080 --> 00:03:08,920
vendors fighting for market 
share and it's going to be raced

61
00:03:08,920 --> 00:03:12,480
to the bottom. 
So yeah, in the first probably 

62
00:03:12,480 --> 00:03:16,160
year or two sold over 80% drop 
in cost per token. 

63
00:03:17,680 --> 00:03:23,280
And I think some of that was the
vendors getting more mature and 

64
00:03:23,280 --> 00:03:27,800
the infrastructure behind the 
scenes and also them trying to 

65
00:03:27,800 --> 00:03:31,720
gain grab market share. 
And then over the, I'll say 

66
00:03:31,720 --> 00:03:35,040
starting in, you know, really 
last year, but over that 

67
00:03:35,040 --> 00:03:38,560
following year, from 2024 to 
2025, the trend started to 

68
00:03:38,560 --> 00:03:42,480
change and, and it whether it 
started where it was not, you 

69
00:03:42,480 --> 00:03:46,080
know, maybe they were about the 
same to now with, you know, for 

70
00:03:46,080 --> 00:03:50,000
example, GBT 54IS like double 
the cost of five one. 

71
00:03:51,360 --> 00:03:54,400
I think Opus just came out today
and it's like double the cost. 

72
00:03:54,640 --> 00:03:58,640
On top of that, you've got other
skews like web search and and 

73
00:03:59,160 --> 00:04:02,240
reasoning and of these other 
aspects that are just 

74
00:04:02,240 --> 00:04:05,840
compounding the cost. 
So now we're the trends are is 

75
00:04:06,120 --> 00:04:10,360
not only is volume like a hockey
stick right now, but costs are 

76
00:04:10,360 --> 00:04:12,880
rising too. 
No matter like what is pizza 

77
00:04:12,880 --> 00:04:15,840
elsewhere, rates are going up. 
For sure. 

78
00:04:15,960 --> 00:04:20,760
And how did the way that you 
were doing your job back when 

79
00:04:20,760 --> 00:04:24,600
you thought in the next 6 
months, is it just trending to 0

80
00:04:25,000 --> 00:04:28,080
versus now where you're kind of 
looking at it and you're saying 

81
00:04:28,720 --> 00:04:32,640
it may go up, but it definitely 
is not going down? 

82
00:04:33,440 --> 00:04:37,040
It's going to just, it's, it's 
going to exponentially go up. 

83
00:04:38,280 --> 00:04:40,080
It's that's your prediction 
you're calling in here. 

84
00:04:41,120 --> 00:04:43,360
It's only going up. 
It's only going volume is not 

85
00:04:43,360 --> 00:04:46,560
going to go down. 
If it's not going up then as a 

86
00:04:46,560 --> 00:04:49,480
business, then you know that's a
bad for us, right. 

87
00:04:49,480 --> 00:04:53,120
So you know, and ideally we want
more and more users using the 

88
00:04:53,120 --> 00:04:55,440
products, but we just want it to
scale efficiently. 

89
00:04:55,800 --> 00:04:59,480
But the, you know, I would say 
with external LLM that's, that's

90
00:04:59,480 --> 00:05:03,040
it's becoming in my opinion of 
larger and larger financial risk

91
00:05:03,040 --> 00:05:06,480
because of the the rate trends 
that we're seeing here. 

92
00:05:06,480 --> 00:05:08,120
So volume and we know is going 
up. 

93
00:05:08,400 --> 00:05:11,200
We can't have pricing also 
doubling just about every time 

94
00:05:11,200 --> 00:05:12,400
too. 
And then that's where I'm 

95
00:05:12,400 --> 00:05:14,640
looking at exponential 
increases. 

96
00:05:15,000 --> 00:05:17,640
So fascinating. 
So that's making you nervous 

97
00:05:17,640 --> 00:05:20,880
looking at how it's been 
performing and recognizing that 

98
00:05:21,360 --> 00:05:25,160
the volumes going to continue to
grow even if it stays at this 

99
00:05:25,160 --> 00:05:27,880
price that we're at right now. 
That's scary. 

100
00:05:28,000 --> 00:05:33,320
But there's a non 0 probability 
that we're looking at the prices

101
00:05:33,320 --> 00:05:38,120
also increasing. 
And so that forces you now to 

102
00:05:38,840 --> 00:05:42,600
look at different options. 
How does their outlook change? 

103
00:05:42,600 --> 00:05:47,760
The bulk of of Grammarly's 
suggestions are based on open 

104
00:05:47,760 --> 00:05:50,880
source models. 
And that is where I really want 

105
00:05:50,880 --> 00:05:54,840
us to go just from a cost 
aspect, you know, but we're also

106
00:05:54,840 --> 00:05:57,960
an AI company, so we have a 
special set of teams and skills 

107
00:05:57,960 --> 00:06:01,080
that can can make that happen. 
Most companies don't have that. 

108
00:06:01,080 --> 00:06:04,440
But I would say right now, I 
mean, we, we have two sides of 

109
00:06:04,440 --> 00:06:06,600
this AI coin. 
We have our production cost 

110
00:06:06,600 --> 00:06:09,520
which far outweighs our AI 
coding internal cost. 

111
00:06:10,000 --> 00:06:12,280
I would say right now most 
companies are dealing with the 

112
00:06:12,280 --> 00:06:14,240
AI coding side of it and that's 
probably what they should be 

113
00:06:14,240 --> 00:06:16,320
dealing with. 
But for us, it's like both 

114
00:06:16,320 --> 00:06:19,880
sides, we've got them really 
managed to make sure and that 

115
00:06:19,880 --> 00:06:23,840
that both are efficient and 
measuring the value in some way.

116
00:06:23,920 --> 00:06:29,840
Yeah, because the product is 
powered by LMS, yes. 

117
00:06:30,160 --> 00:06:35,160
Again, mostly open source, but 
more and more as you want to 

118
00:06:35,240 --> 00:06:38,240
have very high velocity of 
releasing product and such. 

119
00:06:38,240 --> 00:06:40,960
I mean those Frontier models are
great for that. 

120
00:06:40,960 --> 00:06:44,360
So that's where the growth is 
happening on the external alum 

121
00:06:44,360 --> 00:06:46,400
side. 
And you also mentioned to me 

122
00:06:46,400 --> 00:06:51,080
before we hit record on the 
hidden costs around the price 

123
00:06:51,080 --> 00:06:53,840
increases just with the models. 
Can you talk a bit more about? 

124
00:06:53,840 --> 00:06:56,840
That Oh yeah. 
So when you are looking at say, 

125
00:06:56,840 --> 00:07:01,400
you know, maybe using GBT 51 for
whatever reason and, and five 

126
00:07:01,400 --> 00:07:04,800
four comes out and let's say you
have data residency, U.S. data 

127
00:07:04,800 --> 00:07:07,600
residency requirements. 
So now you, you see the rates 

128
00:07:07,600 --> 00:07:09,600
and you start using it and 
you're like, you know what, we 

129
00:07:09,600 --> 00:07:11,880
do want to use this and we see 
that as two times the cost, but 

130
00:07:11,880 --> 00:07:15,080
we see the value here. 
But then you don't notice the 

131
00:07:15,080 --> 00:07:19,160
fact that there's a 10% data 
residency charge on top of that 

132
00:07:19,160 --> 00:07:21,000
that didn't exist with the five 
one model. 

133
00:07:21,080 --> 00:07:25,360
And that's happening with the 
four Opus 4, I think 46 or 47 

134
00:07:25,360 --> 00:07:29,560
and later. 
And basically Anthropic and Open

135
00:07:29,560 --> 00:07:33,040
AI are both adding this premium 
to US based models. 

136
00:07:33,320 --> 00:07:35,120
You're already seeing this 
before this with the 

137
00:07:35,120 --> 00:07:38,600
hyperscalers, AWS and GCP and 
such. 

138
00:07:38,600 --> 00:07:41,600
But now it's that's another 
premium you're paying a lot of, 

139
00:07:41,600 --> 00:07:43,840
I believe a lot of companies are
kind of overlooking. 

140
00:07:44,080 --> 00:07:47,040
Yeah, you don't realize that 
until you later need it. 

141
00:07:47,040 --> 00:07:50,000
And then you go, oh, we're 
getting charged for that. 

142
00:07:50,120 --> 00:07:51,400
Yeah. 
So whenever you're doing your 

143
00:07:51,400 --> 00:07:53,880
estimations and comparing these 
models, you just make sure you 

144
00:07:53,880 --> 00:07:56,640
look at the fine print about the
this data residency piece 

145
00:07:56,640 --> 00:07:59,280
because at 10% adds up. 
Yeah. 

146
00:07:59,280 --> 00:08:02,480
And Speaking of estimations, I 
know that you've built a cool 

147
00:08:02,480 --> 00:08:04,680
cost calculator. 
You vibed it. 

148
00:08:05,200 --> 00:08:06,560
I vibed it. 
I mean it. 

149
00:08:06,960 --> 00:08:10,240
So I'm not an engineer. 
And I was telling him before 

150
00:08:10,480 --> 00:08:14,280
that, you know, before I started
a grammar, I had really no 

151
00:08:14,280 --> 00:08:18,160
experience with AI at all. 
And which is my kind of poking 

152
00:08:18,160 --> 00:08:20,920
around in curiosity that kind of
put me where I am today. 

153
00:08:20,960 --> 00:08:26,400
So I was took the initiative to 
say, OK, I'm going to space out 

154
00:08:26,400 --> 00:08:30,040
some time in my days to spend 
time building some agents to 

155
00:08:30,280 --> 00:08:32,720
help when I was practice. 
And I really, I knew I wanted to

156
00:08:32,720 --> 00:08:36,120
build some type of calculator 
and didn't really know what it 

157
00:08:36,120 --> 00:08:39,200
was going to look like. 
And I built like a HTML version 

158
00:08:39,200 --> 00:08:43,400
and then a CLI version and it's 
using, We have the LLM proxy 

159
00:08:43,400 --> 00:08:46,000
that already has pricing in the 
back end of it. 

160
00:08:46,000 --> 00:08:49,360
So it's just grabbing that 
information and I did it in 

161
00:08:49,360 --> 00:08:53,080
about 15 minutes. 
And people use it. 

162
00:08:53,880 --> 00:08:57,760
It's just pretty new, but I've 
had our, some people, a part of 

163
00:08:57,760 --> 00:09:00,320
our evaluation team start 
playing around with it and they 

164
00:09:00,320 --> 00:09:03,560
were very happy with it. 
And what this is done, you know,

165
00:09:03,560 --> 00:09:07,080
when you're evaluating and I'm 
just learning this myself too, 

166
00:09:07,080 --> 00:09:11,280
more about this process for the 
workflow at least is that they 

167
00:09:11,280 --> 00:09:15,680
are building these prototypes 
and up till this calculator was 

168
00:09:15,680 --> 00:09:18,080
here, they didn't really have a 
good way to go and estimate 

169
00:09:18,080 --> 00:09:19,360
cost. 
They would build the prototypes,

170
00:09:19,360 --> 00:09:22,240
get a sense of the maybe what 
the request per second will be 

171
00:09:22,240 --> 00:09:24,840
and you know, different steps 
this agent's taking and the 

172
00:09:25,240 --> 00:09:27,320
token types and all of them and 
token volume. 

173
00:09:27,480 --> 00:09:29,880
And they would send it to me in 
a spreadsheet and then I'd run 

174
00:09:29,880 --> 00:09:31,480
the numbers and estimate it for 
them. 

175
00:09:31,920 --> 00:09:34,880
And now they can just hit the 
CLI and it's going to bring in 

176
00:09:34,880 --> 00:09:37,600
OK for this model, this type of 
agent would call system out and 

177
00:09:37,600 --> 00:09:39,720
it's safe to the calculator. 
And then they bring in another 

178
00:09:39,720 --> 00:09:41,600
model, do the same thing, save 
to the calculator. 

179
00:09:41,600 --> 00:09:43,840
And then they now they 
understand before they even 

180
00:09:43,840 --> 00:09:47,440
start doing an experiment that 
what the costs are going to be 

181
00:09:47,440 --> 00:09:50,200
and if it's even scalable. 
And she made the comment to me 

182
00:09:50,200 --> 00:09:53,000
that there have been times where
they did a prototype, went to do

183
00:09:53,000 --> 00:09:55,120
the experiment, was too 
expensive and had to start over 

184
00:09:55,120 --> 00:09:58,240
again. 
So not only is is this helping 

185
00:09:58,240 --> 00:10:02,400
us keep costs down, but also 
speeding up development velocity

186
00:10:02,400 --> 00:10:04,280
because they don't have to go 
back to the board. 

187
00:10:04,280 --> 00:10:06,680
They know what the costs are 
going to be before they get past

188
00:10:06,680 --> 00:10:09,000
the prototype typing stage. 
Yeah. 

189
00:10:09,000 --> 00:10:13,360
And that's just an example of 
how you recognized this is 

190
00:10:13,360 --> 00:10:17,920
something that's happening. 
Can I make it easier on the 

191
00:10:17,920 --> 00:10:20,280
folks who are closer to this 
problem? 

192
00:10:20,360 --> 00:10:23,600
How can we get these tooling? 
I don't want to be in a 

193
00:10:23,600 --> 00:10:24,920
spreadsheet doing estimates 
anymore. 

194
00:10:24,920 --> 00:10:26,440
That's not where my true value 
is. 

195
00:10:26,680 --> 00:10:30,200
My true value is is getting 
maybe building, thinking about 

196
00:10:30,200 --> 00:10:32,080
these tools and building them 
and then putting them in the 

197
00:10:32,080 --> 00:10:35,360
people's hands that are actually
doing the work to make the tool,

198
00:10:35,360 --> 00:10:37,080
the, the agents or what have 
you. 

199
00:10:37,360 --> 00:10:40,720
And then, you know, again, shift
left, get that in early as 

200
00:10:40,720 --> 00:10:42,120
possible. 
I would love for the product 

201
00:10:42,280 --> 00:10:45,160
people to start maybe playing 
around with us if they have the 

202
00:10:45,160 --> 00:10:47,640
data to do it right. 
So I want to be as early in 

203
00:10:47,640 --> 00:10:50,040
development cycle as possible. 
And this really helps with that.

204
00:10:50,240 --> 00:10:53,640
And honestly, this FP and a, my 
counterpart in FP and a 

205
00:10:54,000 --> 00:10:57,920
hopefully she sees this that, 
you know, kind of said, Hey, 

206
00:10:57,920 --> 00:10:59,440
could we book some kind of 
calculator? 

207
00:10:59,440 --> 00:11:01,520
And I'm like, you know, I think 
I can, I don't think it'd be too

208
00:11:01,520 --> 00:11:03,480
bad. 
And and then, you know, about a 

209
00:11:03,480 --> 00:11:07,160
week later I had this thing. 
So, you know, it wasn't even 

210
00:11:07,160 --> 00:11:09,440
totally like me saying I need to
go do this. 

211
00:11:09,440 --> 00:11:11,200
It was like, you know, 
collaboration with the 

212
00:11:11,240 --> 00:11:13,080
financing. 
Oh, I think we should we start 

213
00:11:13,080 --> 00:11:14,880
working toward this. 
And that that piqued my 

214
00:11:14,880 --> 00:11:17,080
curiosity and the luck I could. 
Do it. 

215
00:11:17,200 --> 00:11:21,320
And I like how the evaluation 
team gets to now understand this

216
00:11:21,320 --> 00:11:23,600
and that's probably one of the 
core metrics that they're 

217
00:11:23,600 --> 00:11:26,680
looking at as they're evaluating
if this is feasible. 

218
00:11:26,760 --> 00:11:29,360
Exactly. 
I know another piece of that is 

219
00:11:30,560 --> 00:11:33,120
support request to go to our 
channel, they can send in a 

220
00:11:33,120 --> 00:11:35,760
spreadsheet or what have you if 
they somebody wants. 

221
00:11:36,040 --> 00:11:38,680
And then my support bot is going
to use my calculator I built to 

222
00:11:38,680 --> 00:11:40,800
calculate that. 
And so I that's one thing I 

223
00:11:40,800 --> 00:11:43,000
don't have to go and do anymore.
OK. 

224
00:11:43,000 --> 00:11:48,200
So tell me more about where 
you're thinking of your ability 

225
00:11:48,200 --> 00:11:53,760
to add products or empower the 
engineering team, the rest of 

226
00:11:53,760 --> 00:11:57,280
the organization with this 
knowledge because they're closer

227
00:11:57,280 --> 00:12:01,560
to those problems. 
So I know where the data is, I 

228
00:12:01,560 --> 00:12:05,320
know how to kind of structure 
like what things should properly

229
00:12:05,320 --> 00:12:10,280
cost and and so that's kind of 
the magic I kind of have for 

230
00:12:10,280 --> 00:12:12,040
myself and I just need to take 
that. 

231
00:12:12,080 --> 00:12:14,520
And I think for all 5 thinnest 
practitioners probably have the 

232
00:12:14,520 --> 00:12:17,440
same knowledge, but they need to
take that put it into some kind 

233
00:12:17,440 --> 00:12:19,480
of tooling. 
So it's closer to the 

234
00:12:19,480 --> 00:12:21,160
engineering in the development 
process. 

235
00:12:21,160 --> 00:12:24,440
And then shift their focus up in
the organizational and be 

236
00:12:24,440 --> 00:12:30,960
focused more on strategy and and
and actually just enabling the 

237
00:12:30,960 --> 00:12:34,520
teams to make better decisions 
moving forward on the products, 

238
00:12:34,520 --> 00:12:36,520
no matter, you know, even if it 
is going to be something 

239
00:12:36,520 --> 00:12:38,560
expensive that they have 
understanding all. 

240
00:12:38,560 --> 00:12:40,280
Yeah. 
Well, this is we have a 

241
00:12:40,280 --> 00:12:43,200
reasoning behind this because 
you know, product market fit or 

242
00:12:43,200 --> 00:12:46,320
what have you. 
So if everybody can move forward

243
00:12:46,320 --> 00:12:49,320
from engineering up to in 
product to obviously the C-Suite

244
00:12:49,320 --> 00:12:51,720
with making decisions on knowing
what the costs are ahead of 

245
00:12:51,720 --> 00:12:53,880
time, we're going to be in a 
much better state as a business.

246
00:12:54,440 --> 00:12:57,760
And if we looked back at again, 
going through that journey that 

247
00:12:57,760 --> 00:13:04,360
you've had of when you first 
were encountering AI on Azure 

248
00:13:04,360 --> 00:13:09,560
with Open AI and then versus now
and the maturity that you feel 

249
00:13:09,560 --> 00:13:12,400
like you've been able to come 
into. 

250
00:13:12,720 --> 00:13:15,960
Can you explain some of the 
things that now you're like we 

251
00:13:15,960 --> 00:13:22,080
could not live without XYZ? 
Or say our telemetry streams. 

252
00:13:22,560 --> 00:13:26,680
Telemetry streams. 
From from my perspective, I 

253
00:13:26,680 --> 00:13:29,720
built some telemetry streams. 
So we have LLM proxy as home 

254
00:13:29,720 --> 00:13:33,640
build that that gives me the 
token volume by calling service.

255
00:13:33,640 --> 00:13:36,600
And then a feature is another, 
you know, step below a calling 

256
00:13:36,600 --> 00:13:38,600
service. 
So like, for example, we have 

257
00:13:38,600 --> 00:13:41,960
our main AI chat and there's a 
bunch of connectors for that. 

258
00:13:41,960 --> 00:13:44,160
And that's a bunch of features, 
thousands of features, all those

259
00:13:44,160 --> 00:13:46,320
connectors. 
So I take that information, bump

260
00:13:46,320 --> 00:13:50,120
it into our, our, our cost 
reporting platform, which is a 

261
00:13:50,120 --> 00:13:53,600
third party and, and then 
allocate cost based on the open 

262
00:13:53,600 --> 00:13:57,600
AI bill. 
And I can do that for Bedrock 

263
00:13:57,600 --> 00:14:00,280
and, and any other vendor that 
we're, we're bringing into that,

264
00:14:00,280 --> 00:14:02,920
that tool. 
And all it's doing is looking at

265
00:14:02,920 --> 00:14:04,960
the token type. 
And then, you know, looking at 

266
00:14:05,240 --> 00:14:08,440
that cause versus the token 
volume of each service and all 

267
00:14:08,440 --> 00:14:10,880
these other metadata that's 
using it and then proportionally

268
00:14:10,880 --> 00:14:13,680
allocating it by token type 
service and all those things. 

269
00:14:14,520 --> 00:14:17,800
So without that we would have, 
we would not have visibility to 

270
00:14:17,800 --> 00:14:19,120
actually what these agents. 
Cost. 

271
00:14:19,480 --> 00:14:23,480
And then are you looking at 
spikes and trying to figure out 

272
00:14:23,480 --> 00:14:25,800
what's going on? 
Why did we just spend a lot of 

273
00:14:25,800 --> 00:14:29,440
money here? 
Yes, I've got agents that are 

274
00:14:29,440 --> 00:14:31,080
starting to do that a little bit
for me. 

275
00:14:31,080 --> 00:14:37,960
But I would say our we, we don't
have a lot of spikes scales with

276
00:14:37,960 --> 00:14:39,640
business hours. 
Just about all our production 

277
00:14:39,640 --> 00:14:41,080
stuff just scales with business 
hours. 

278
00:14:41,080 --> 00:14:43,800
Now we'll see a trend all of a 
sudden, you know, cash tokens 

279
00:14:43,800 --> 00:14:45,280
are going down and costs are 
going up. 

280
00:14:45,280 --> 00:14:46,880
Like we'll see things like that 
we need to pick up. 

281
00:14:46,880 --> 00:14:50,120
But it's very, very rare that we
have like a big spike outside 

282
00:14:50,120 --> 00:14:52,720
of, you know, research and use 
cases or experiments that you 

283
00:14:52,720 --> 00:14:54,640
would expect that. 
But usually if it is a giant 

284
00:14:54,640 --> 00:14:57,840
spike, it's usually from my 
experiment that maybe, you know,

285
00:14:57,840 --> 00:15:00,440
was a little higher than 
expected or, you know, research 

286
00:15:00,440 --> 00:15:03,320
that I knew about ahead of time.
Speaking of research, you had 

287
00:15:03,320 --> 00:15:06,520
mentioned before we hit record 
about the research team and how 

288
00:15:06,520 --> 00:15:09,720
they're being empowered also by 
some of the stuff that you're 

289
00:15:09,720 --> 00:15:11,040
doing. 
Can you explain that? 

290
00:15:11,600 --> 00:15:13,680
Well, yeah, I'm, I'm trying to 
actually just today sent 

291
00:15:13,680 --> 00:15:16,000
something to them about sent 
them the calculator and said, 

292
00:15:16,000 --> 00:15:17,240
hey, this is going to be of use 
to you. 

293
00:15:17,240 --> 00:15:19,960
But no, they, I think what we're
talking about is around 

294
00:15:19,960 --> 00:15:22,720
optimization where and it was 
touched on in the keynote today 

295
00:15:22,720 --> 00:15:23,760
too. 
And I thought it was a great, 

296
00:15:23,880 --> 00:15:26,200
great presentation and saying, 
you know, optimization in the 

297
00:15:26,200 --> 00:15:29,280
cloud was just, you know, 
optimizing resources and, you 

298
00:15:29,280 --> 00:15:31,960
know, buying savings plans in 
our eyes and things like that, 

299
00:15:31,960 --> 00:15:34,040
that fin OPS going to handle. 
Or it's just more 

300
00:15:34,040 --> 00:15:36,440
straightforward and it doesn't 
really impact the product at the

301
00:15:36,440 --> 00:15:39,240
end of the day. 
When with LLMS, whenever you're 

302
00:15:39,240 --> 00:15:44,040
changing, OK, you know, cashing 
rates or I'm trying to reduce 

303
00:15:44,040 --> 00:15:47,360
output tokens and input tokens 
and all these things and trying 

304
00:15:47,360 --> 00:15:49,760
to use cheaper models that 
actually impacts the output and 

305
00:15:49,760 --> 00:15:51,280
the user experience. 
Drastically. 

306
00:15:51,440 --> 00:15:54,640
So it is so much more nuanced. 
So I have found much more 

307
00:15:54,640 --> 00:15:58,760
success with optimization when 
this research LED optimization 

308
00:15:58,760 --> 00:16:01,040
where our researchers are 
looking into saying, OK, I think

309
00:16:01,040 --> 00:16:03,600
we could probably use maybe a 
mini model combined with a 

310
00:16:03,600 --> 00:16:07,640
standard model and not impact 
our latency, which latency is a 

311
00:16:07,800 --> 00:16:10,360
very important aspect there. 
You know, all our stuff needs to

312
00:16:10,360 --> 00:16:12,240
be pretty much sub one second 
latency. 

313
00:16:13,280 --> 00:16:16,400
So you know, it's not going to 
impact latency and we're going 

314
00:16:16,400 --> 00:16:19,120
to still getting really good 
outputs and saving money. 

315
00:16:19,120 --> 00:16:22,040
So they, they, you know, maybe 
they'll come out with the 

316
00:16:22,040 --> 00:16:27,000
initial kind of launch of, of an
agent and then from there they 

317
00:16:27,000 --> 00:16:30,040
look for ways to optimize it. 
We do the same thing with our 

318
00:16:30,040 --> 00:16:33,960
open source LLMS and we launched
speculative decoding last year 

319
00:16:33,960 --> 00:16:37,560
and it saved us a bunch of money
and on the on the open source 

320
00:16:37,560 --> 00:16:41,480
models and we're just trying to 
try to try to learn from those, 

321
00:16:41,600 --> 00:16:43,640
those teams and optimizations 
they did there. 

322
00:16:43,640 --> 00:16:46,080
And how can we apply this to 
these external LLMS? 

323
00:16:46,080 --> 00:16:50,360
And and again, it's all usually 
the big heavy optimizations are,

324
00:16:50,560 --> 00:16:52,800
you know, really started from 
the research side. 

325
00:16:53,080 --> 00:16:56,480
And the product research is what
fascinates me the most on that 

326
00:16:56,480 --> 00:16:59,480
because it's not like AI 
researchers. 

327
00:16:59,480 --> 00:17:03,000
It's not. 
It's, as you mentioned, the 

328
00:17:03,000 --> 00:17:09,400
product is so affected by any 
little changes in the AI side of

329
00:17:09,400 --> 00:17:14,920
the house because you're in a 
very unique situation where your

330
00:17:14,920 --> 00:17:19,079
product is driven by AI at the 
end of the day. 

331
00:17:19,079 --> 00:17:22,640
And so you have to make sure 
that that's the highest quality 

332
00:17:22,680 --> 00:17:26,200
product you can put out. 
But also, your main job is 

333
00:17:26,200 --> 00:17:30,600
making sure that it you're just 
not using and bringing a bazooka

334
00:17:30,800 --> 00:17:33,920
every time you need to do the 
smallest little comma update or 

335
00:17:33,920 --> 00:17:35,600
whatever. 
Exactly. 

336
00:17:35,600 --> 00:17:38,080
And we have data behind the 
decisions we've made and we've 

337
00:17:38,440 --> 00:17:40,240
you know and I think most 
companies like this, but we just

338
00:17:40,240 --> 00:17:42,320
like the data. 
You know, right now we have the 

339
00:17:42,320 --> 00:17:45,320
cost data and now we need to 
better tie that to revenue and 

340
00:17:45,320 --> 00:17:48,440
then other and experiments and 
you actually get experiment IDs 

341
00:17:48,440 --> 00:17:51,920
into the telemetry streams. 
I was talking about then seeing 

342
00:17:51,920 --> 00:17:54,840
what the cost of this experiment
was and was it worth it or not. 

343
00:17:54,840 --> 00:17:58,000
So like it's, you know, while 
we've been in this and I've been

344
00:17:58,000 --> 00:18:01,680
doing this work for three years 
now, we're still like getting to

345
00:18:01,680 --> 00:18:05,040
that value phase is hard, 
especially as you acquire 

346
00:18:05,040 --> 00:18:07,600
companies and different products
and different measurements, but.

347
00:18:07,640 --> 00:18:10,160
Bring them into the fold. 
When you say experiments you 

348
00:18:10,160 --> 00:18:13,520
mean doing small little rollouts
on? 

349
00:18:13,520 --> 00:18:16,080
Or do you mean like behind 
close? 

350
00:18:16,200 --> 00:18:18,640
We have a very mature 
experimentation platform where 

351
00:18:18,640 --> 00:18:21,240
we're actually, we'll be running
live experiments on, you know, 

352
00:18:21,480 --> 00:18:26,840
they usually AB test to see, you
know, how maybe a new new change

353
00:18:26,840 --> 00:18:31,400
or new feature impacts. 
It could be in product 

354
00:18:31,400 --> 00:18:35,400
engagement type of rates. 
And and that's how we kind of 

355
00:18:35,400 --> 00:18:39,040
decide on or should this thing 
go full scale or do we need to 

356
00:18:39,040 --> 00:18:42,920
go back to the drawing board? 
And again, it's interesting to 

357
00:18:42,920 --> 00:18:48,360
me because it is product type of
features, but then inside of the

358
00:18:48,360 --> 00:18:52,720
core product, there's probably a
lot of different experiments 

359
00:18:52,720 --> 00:18:55,920
that you're running with all the
models or with the ways that 

360
00:18:55,920 --> 00:18:59,280
you're making those models 
useful. 

361
00:18:59,640 --> 00:19:01,960
It's the same product for me as 
the end user. 

362
00:19:02,080 --> 00:19:04,600
I'm still getting if it's we're 
talking about Grammarly. 

363
00:19:04,600 --> 00:19:08,080
I'm still getting suggestions on
my grammar, but you behind the 

364
00:19:08,080 --> 00:19:11,080
scenes are running these 
experiments that I may or may 

365
00:19:11,080 --> 00:19:12,800
not know about. 
You may not know about and all 

366
00:19:12,800 --> 00:19:17,080
of a sudden you're engaging with
it in a, you know, slightly more

367
00:19:17,080 --> 00:19:20,360
volume and you may not even 
realize it, but then, you know, 

368
00:19:20,560 --> 00:19:24,000
our, our people are picking up 
on, OK, we got more, more, more 

369
00:19:24,000 --> 00:19:26,480
accepts here for this 
suggestion. 

370
00:19:26,480 --> 00:19:29,560
Or, or they actually, you know, 
the clicks on this specific 

371
00:19:29,560 --> 00:19:33,920
agent actually increased over 
time or the amount of time turns

372
00:19:33,920 --> 00:19:35,600
they had with this agent 
increased. 

373
00:19:35,600 --> 00:19:37,920
Like little things like that 
that indicate, hey, this is 

374
00:19:37,920 --> 00:19:39,840
actually making a meaningful 
difference for our customers. 

375
00:19:40,280 --> 00:19:44,320
And you want to be able to, just
to close the circle on this, as 

376
00:19:44,320 --> 00:19:47,280
you mentioned before, when you 
have those experiments, 

377
00:19:47,360 --> 00:19:49,440
recognize the cost of those 
experiments. 

378
00:19:49,560 --> 00:19:51,440
Exactly. 
Tie that to the cost 

379
00:19:51,440 --> 00:19:54,560
fluctuations that we're that I'm
seeing on my side because right 

380
00:19:54,560 --> 00:19:56,840
now it I'll see these spikes and
I do see spikes. 

381
00:19:56,840 --> 00:19:59,200
It's usually around experiments 
and I have to do a lot of 

382
00:19:59,200 --> 00:20:03,280
research or have you know, I'll 
let them go do it for me to 

383
00:20:03,280 --> 00:20:05,240
figure out what exactly is 
causing this. 

384
00:20:05,240 --> 00:20:07,160
And even then, the correlation 
is light. 

385
00:20:08,080 --> 00:20:10,320
It's really hard to to find a 
correlation. 

386
00:20:10,320 --> 00:20:13,720
So you know, we've got, we got a
couple teams collaborating to 

387
00:20:13,720 --> 00:20:17,760
bring that data into our our 
cost reporting so we we can 

388
00:20:17,760 --> 00:20:19,760
actually measure the cost of 
those experiments. 

389
00:20:19,760 --> 00:20:22,840
Because I imagine the teams 
aren't saying like, hey, I'm 

390
00:20:22,840 --> 00:20:25,520
going to run this experiment 
now, like just in case. 

391
00:20:25,720 --> 00:20:28,200
Sometimes, sometimes they'll. 
Give you a heads up. 

392
00:20:28,360 --> 00:20:31,320
I've been very blessed with with
a good relationship with a lot 

393
00:20:31,320 --> 00:20:33,640
of the the teams that do these 
experiments. 

394
00:20:33,640 --> 00:20:35,960
So the larger ones, they usually
come to me and either to 

395
00:20:35,960 --> 00:20:39,080
estimate what this is going to 
cost or just kind of let me 

396
00:20:39,080 --> 00:20:40,360
know, hey, we're going to run 
this thing. 

397
00:20:40,360 --> 00:20:43,320
So I know about the big ones, 
but then they're still small 

398
00:20:43,320 --> 00:20:45,400
ones that are supposed to be 
small, but then they end up 

399
00:20:45,400 --> 00:20:47,000
being a little bigger than 
expected, right? 

400
00:20:47,000 --> 00:20:49,920
But overall, you know, there's 
really good communication 

401
00:20:49,920 --> 00:20:52,680
between us. 
So on those times where you see 

402
00:20:52,680 --> 00:20:55,760
the spikes or when you're 
getting a heads up, is that 

403
00:20:55,760 --> 00:20:59,640
something that now you're 
expecting folks to have a better

404
00:20:59,640 --> 00:21:02,360
idea of because of the cost 
calculator that you created? 

405
00:21:02,400 --> 00:21:04,000
Yeah. 
I mean, ideally they would, they

406
00:21:04,000 --> 00:21:07,480
wouldn't need me to do those 
estimates beforehand or at least

407
00:21:07,480 --> 00:21:10,160
they'll have an idea of what the
costs are going to be before 

408
00:21:10,160 --> 00:21:13,480
they run the experiment. 
And I haven't built this into it

409
00:21:13,480 --> 00:21:16,800
yet, but I'd like to actually 
have that going into a doc, a 

410
00:21:16,800 --> 00:21:18,280
coded doc, you know, one of our 
company. 

411
00:21:18,280 --> 00:21:21,680
But I'd like for whatever 
submissions that they have in 

412
00:21:21,680 --> 00:21:24,400
this, this tool. 
And they can, I'm assuming they,

413
00:21:24,440 --> 00:21:26,320
I'll probably have something 
where they can actually choose, 

414
00:21:26,320 --> 00:21:29,120
OK, this is something to submit.
And then it goes to a dock where

415
00:21:29,120 --> 00:21:33,040
we see all the estimations in 
aggregate and then can pick up 

416
00:21:33,040 --> 00:21:36,160
on any red flags or at least do 
better planning. 

417
00:21:36,160 --> 00:21:39,760
OK, we know we're running this 
thing during this date or within

418
00:21:39,760 --> 00:21:41,280
the next few months. 
We usually don't have dates 

419
00:21:41,280 --> 00:21:44,640
until, you know, pretty late 
into the process and we can 

420
00:21:44,640 --> 00:21:47,800
better plan on, you know, what 
our finances or costs are going 

421
00:21:47,800 --> 00:21:49,760
to be down the line. 
And also compare it with the 

422
00:21:49,760 --> 00:21:55,200
real cost to make sure that hey 
is our calculations, are we 

423
00:21:55,200 --> 00:21:59,160
projecting, is my calculator. 
Broken or is, you know, request 

424
00:21:59,160 --> 00:22:01,320
rates are a little, you know, 
off. 

425
00:22:01,960 --> 00:22:04,520
So yeah, that's true. 
And what are the different 

426
00:22:04,520 --> 00:22:08,200
inputs that you ask folks to 
give you on this cost 

427
00:22:08,200 --> 00:22:11,440
calculator? 
So average request per second 

428
00:22:12,040 --> 00:22:14,920
more minute and then average 
request size. 

429
00:22:14,920 --> 00:22:18,120
So when I say request size, 
that's input tokens, output 

430
00:22:18,120 --> 00:22:21,280
tokens, cashed input. 
And these are all things that I 

431
00:22:21,280 --> 00:22:23,720
said they're kind of building 
the prototype and, and testing 

432
00:22:23,720 --> 00:22:25,280
out the, the outputs that's 
coming. 

433
00:22:25,280 --> 00:22:28,080
You know that they they get 
those token counts from as a 

434
00:22:28,080 --> 00:22:32,440
response body from the the open 
AI or Anthropic, whatever you're

435
00:22:32,440 --> 00:22:34,680
using. 
And I wasn't clear when you were

436
00:22:34,680 --> 00:22:37,560
talking about how you have the 
research, the product research 

437
00:22:37,560 --> 00:22:43,880
team, then they have this 
knowledge of we can do XYZ with 

438
00:22:43,880 --> 00:22:46,960
these models or we can cache 
this, or we can set up a 

439
00:22:46,960 --> 00:22:51,000
different architecture that will
be a better product experience. 

440
00:22:51,240 --> 00:22:55,880
Does that then get propagated 
through the different teams so 

441
00:22:55,880 --> 00:23:03,320
that whatever next team knows oh
I now have the ability to make 

442
00:23:03,320 --> 00:23:07,040
my experiment much cheaper? 
We could be better about that, 

443
00:23:07,040 --> 00:23:10,200
we could be better about 
advertising how teams are doing 

444
00:23:10,200 --> 00:23:12,280
this. 
Well, one other issue there 

445
00:23:12,280 --> 00:23:14,840
though is you know, when you 
look at our mail product, the 

446
00:23:14,840 --> 00:23:19,040
way they use LLMS is way 
different than than the way say 

447
00:23:19,040 --> 00:23:22,240
Grammarly uses LLMS or even our 
new Go product. 

448
00:23:22,320 --> 00:23:25,800
So that is one kind of issue. 
Now I think there should, could,

449
00:23:25,800 --> 00:23:29,080
should be more sharing across 
the, the product or the business

450
00:23:29,080 --> 00:23:32,600
units and, and those products 
because I'm sure they're still 

451
00:23:32,600 --> 00:23:34,240
learnings there. 
But again, you know, these 

452
00:23:34,240 --> 00:23:36,880
acquisitions happen not that 
long ago. 

453
00:23:36,880 --> 00:23:39,640
So it's, it's going to take time
for those things to be sure. 

454
00:23:39,640 --> 00:23:42,200
And I think it's probably part 
of my job to make sure that 

455
00:23:42,200 --> 00:23:48,880
those are broad and broad kind 
of broadly told to the company 

456
00:23:48,880 --> 00:23:51,640
so that other people can't take 
advantage of. 

457
00:23:51,640 --> 00:23:53,160
But that's something I think we 
can improve. 

458
00:23:53,440 --> 00:23:59,160
Well, it does feel like what 
you're saying too around a lot 

459
00:23:59,160 --> 00:24:02,280
of times folks will try and 
let's use the easy example of 

460
00:24:02,280 --> 00:24:07,280
just caching certain results or 
certain tokens so that you can 

461
00:24:07,480 --> 00:24:10,560
spend less. 
But then it ends up having a 

462
00:24:10,560 --> 00:24:12,880
worse product experience some of
the time. 

463
00:24:13,360 --> 00:24:18,960
And so you want to make sure 
that the teams aren't using that

464
00:24:18,960 --> 00:24:24,880
type of architecture at the cost
of like having less token cost 

465
00:24:25,000 --> 00:24:26,120
I. 
Guess Oh, exactly. 

466
00:24:26,120 --> 00:24:30,160
And I mean, I first and foremost
is always the best product for 

467
00:24:30,160 --> 00:24:32,600
our customers. 
Like we, we are not going to 

468
00:24:32,600 --> 00:24:35,800
sacrifice that, especially in 
this very competitive area that 

469
00:24:35,800 --> 00:24:37,120
we're in. 
We're not going to sacrifice 

470
00:24:37,120 --> 00:24:39,680
that at all. 
But it's more about OK, be 

471
00:24:39,680 --> 00:24:41,680
dealing the best product. 
But what are some of the things 

472
00:24:41,680 --> 00:24:44,520
that that are we can still 
deliver the best product, but 

473
00:24:44,520 --> 00:24:48,960
also save money on the back end.
So the team is just because 

474
00:24:48,960 --> 00:24:52,440
these researchers have much 
deeper understanding of these 

475
00:24:52,440 --> 00:24:55,840
LLMS, they're able to go in and 
tweak those certain things or 

476
00:24:55,840 --> 00:24:59,200
know the direction to go. 
And then they are in in the 

477
00:24:59,200 --> 00:25:02,200
position to go in and evaluate 
the outputs and kind of go 

478
00:25:02,200 --> 00:25:03,800
through that process to make 
sure it's all good. 

479
00:25:04,000 --> 00:25:08,960
There's no way a fin OPS person 
could could do that even if they

480
00:25:08,960 --> 00:25:11,440
knew how to adjust the LLMS and 
make it cheaper. 

481
00:25:11,440 --> 00:25:15,560
Evaluating the outputs in a way 
a data scientist or a researcher

482
00:25:15,560 --> 00:25:17,800
or a company with computational 
linguist as well. 

483
00:25:17,800 --> 00:25:21,480
They that's a whole another 
skill set that, you know, fin 

484
00:25:21,480 --> 00:25:24,480
OPS people aren't going to have.
So that's what makes this a lot 

485
00:25:24,480 --> 00:25:26,360
harder to do. 
I mean, I've also tried to use 

486
00:25:26,360 --> 00:25:31,640
claw to optimize claw and our, 
our like clawed API products 

487
00:25:31,640 --> 00:25:35,760
and, and it gives good advice, 
but it's usually still just not 

488
00:25:35,760 --> 00:25:39,920
quite, quite there. 
We've already, you know, tried 

489
00:25:39,920 --> 00:25:43,640
it or the impact wasn't, you 
know, that large, but that could

490
00:25:43,640 --> 00:25:46,200
be another Ave. is just use AI 
to optimize AI. 

491
00:25:46,920 --> 00:25:53,120
So tell me about how this 
reserving capacity kind of cost 

492
00:25:53,120 --> 00:25:55,840
benefit analysis goes in your 
head. 

493
00:25:56,240 --> 00:26:00,680
Do you get cheaper tokens if you
reserve more capacity? 

494
00:26:00,680 --> 00:26:04,680
And is that something you even 
look at, or is it something that

495
00:26:04,680 --> 00:26:08,760
has kind of gone out of fashion?
So early, early on, that was 

496
00:26:08,760 --> 00:26:13,040
really the only reserving 
capacity as you're up in the eye

497
00:26:13,040 --> 00:26:16,280
of the time, PT use was the only
way to get low consistent 

498
00:26:16,280 --> 00:26:18,680
latency. 
And as I mentioned with with 

499
00:26:18,680 --> 00:26:22,920
Grammarly, you know, sub one 
second latency is just hugely 

500
00:26:22,920 --> 00:26:25,200
important for our product. 
So yeah, we need a low 

501
00:26:25,200 --> 00:26:27,160
consistent latency. 
So yeah, in the beginning I was 

502
00:26:27,160 --> 00:26:29,960
managing capacity, having to buy
capacity for the peak. 

503
00:26:30,560 --> 00:26:33,840
And if we didn't, you know, we 
would scale beyond our our peak 

504
00:26:33,840 --> 00:26:37,840
and and have an outage. 
But then over time, we did see 

505
00:26:37,840 --> 00:26:41,760
token rates come down for that 
capacity, but it was a really 

506
00:26:41,760 --> 00:26:46,120
difficult thing to manage. 
One, the actual utilization of 

507
00:26:46,120 --> 00:26:50,080
the capacity you could get, you 
can get metrics to know what 

508
00:26:50,080 --> 00:26:52,840
your utilization was, but 
understanding it beforehand was 

509
00:26:52,840 --> 00:26:56,720
really hard because you know, 
everything's measured in these 

510
00:26:57,160 --> 00:27:00,480
units, abstracted units that 
Azure measures them in. 

511
00:27:00,480 --> 00:27:05,200
And depending on the request 
size or shape that the each 

512
00:27:05,200 --> 00:27:07,520
agent is sending, that 
utilization is going to be a 

513
00:27:07,520 --> 00:27:09,760
little different. 
So knowing that ahead of time 

514
00:27:09,760 --> 00:27:11,920
was tough. 
RL and proxy team actually built

515
00:27:12,360 --> 00:27:16,440
a way they call it shadow 
traffic where it would mimic the

516
00:27:16,440 --> 00:27:20,320
traffic of this thing that was 
going to be coming out to see 

517
00:27:20,520 --> 00:27:24,400
what would in load test the 
actual capacity and see what 

518
00:27:24,400 --> 00:27:27,760
the, you know, peak tokens per 
minute were. 

519
00:27:27,920 --> 00:27:30,840
And then that's how we would 
estimate, OK, we need to add 

520
00:27:30,840 --> 00:27:33,520
this much more capacity for this
new thing that's about to launch

521
00:27:33,520 --> 00:27:35,880
or you know, we can actually 
reduce capacity because things 

522
00:27:35,880 --> 00:27:39,680
have come down. 
So that was back in that day, 

523
00:27:39,680 --> 00:27:42,360
but it was just still very 
difficult to measure ahead of 

524
00:27:42,360 --> 00:27:46,000
time and it was a lot of this 
mathematical lean limbo and 

525
00:27:46,000 --> 00:27:49,120
exhausting. 
Then I want to say it was like 

526
00:27:49,120 --> 00:27:53,440
late 2024, early 2025, Open the 
Eye came out with priority 

527
00:27:53,440 --> 00:27:56,960
processing. 
So this has a late latency 

528
00:27:56,960 --> 00:28:00,760
SLA's, availability SLA's. 
It is double the cost as far as 

529
00:28:00,760 --> 00:28:04,160
from a rate perspective as the 
standard RAID. 

530
00:28:04,440 --> 00:28:06,360
But again, we needed a low 
latency. 

531
00:28:06,640 --> 00:28:09,320
What attracted me to it was I 
didn't have to manage capacity 

532
00:28:09,320 --> 00:28:12,240
anymore because this is 
completely on demand And so it 

533
00:28:12,240 --> 00:28:14,480
scales with our usage. 
Our usage scales with with 

534
00:28:14,480 --> 00:28:17,080
business hours. 
So then I did the numbers and 

535
00:28:17,080 --> 00:28:22,440
looked at, OK, if we had the 
even looked at if we had an 

536
00:28:22,440 --> 00:28:26,920
ideal amount of reserve PT use 
and then also with our proxy 

537
00:28:26,920 --> 00:28:29,280
sent some traffic to priority 
processing. 

538
00:28:29,280 --> 00:28:32,360
Like what you know, could that 
save us money in general, just 

539
00:28:32,360 --> 00:28:35,880
comparing PT US, you know, fully
using PT us versus priority 

540
00:28:35,880 --> 00:28:38,600
processing. 
And really it came down to you 

541
00:28:38,600 --> 00:28:41,440
could see some savings with the 
minis and nanos of the world, 

542
00:28:41,440 --> 00:28:46,600
but you couldn't, it wasn't 
enough to, to need to manage 

543
00:28:46,600 --> 00:28:49,600
that capacity with the standard 
size models. 

544
00:28:49,720 --> 00:28:51,640
There was no savings even with 
using PT us. 

545
00:28:51,640 --> 00:28:54,640
And this is a, this is a year or
so ago, it may have changed, but

546
00:28:54,640 --> 00:28:57,240
I doubt it. 
So priority processing was an 

547
00:28:57,240 --> 00:28:59,800
absolute no brainer. 
Not on top of just making my 

548
00:28:59,800 --> 00:29:03,560
life easier. 
It also allowed it makes much 

549
00:29:03,560 --> 00:29:06,760
easier to allocate cost because 
again, it's just very difficult 

550
00:29:06,760 --> 00:29:10,480
to understand how much of a 
single service he's using of, of

551
00:29:10,480 --> 00:29:13,440
this PTU. 
Because yeah, you could look at 

552
00:29:13,440 --> 00:29:18,320
token volume of raw, but tokens 
are not equal across agents. 

553
00:29:18,320 --> 00:29:21,880
And I took so it was just very 
complex of this and I improved 

554
00:29:21,880 --> 00:29:25,720
our visibility, saved us some 
money at that time at least and 

555
00:29:25,840 --> 00:29:27,640
and and it made my life a lot 
easier. 

556
00:29:27,880 --> 00:29:30,960
I can only imagine you doing 
this complex math like what's 

557
00:29:30,960 --> 00:29:34,960
the right blend of PTU or just 
straight token. 

558
00:29:35,120 --> 00:29:38,040
Pretty large people sheets out 
floating around. 

559
00:29:38,760 --> 00:29:41,720
And then you're like, wait for 
quality of life. 

560
00:29:41,800 --> 00:29:44,600
Let's just do it on this there. 
That should be one of the 

561
00:29:44,600 --> 00:29:47,960
factors too. 
Like you don't have to do this 

562
00:29:47,960 --> 00:29:50,560
crazy math anymore. 
You know what you're going to 

563
00:29:50,560 --> 00:29:53,760
get. 
And it also if it's the same 

564
00:29:53,760 --> 00:29:56,840
more or less the same price then
it's a no brainer. 

565
00:29:57,160 --> 00:29:59,160
Exactly, exactly. 
And then something we were 

566
00:29:59,160 --> 00:30:02,560
talking about a little bit 
earlier around Open the Eyes, 

567
00:30:02,840 --> 00:30:06,480
release of guaranteed capacity. 
That just happened, right? 

568
00:30:07,120 --> 00:30:10,000
I find it really interesting 
because it's not advertised as a

569
00:30:10,000 --> 00:30:13,480
discount mechanism. 
It's, it's advertised as a, you 

570
00:30:13,480 --> 00:30:16,040
know, guaranteed capacity, you 
know, and, and I'm just. 

571
00:30:16,040 --> 00:30:19,040
Saying like, if you need this, 
you're always going to have it, 

572
00:30:19,040 --> 00:30:22,040
so you're never going to be 
without resources. 

573
00:30:22,040 --> 00:30:23,160
We're not going to go down on 
you. 

574
00:30:23,360 --> 00:30:25,920
Exactly. 
But from our standpoint, we 

575
00:30:25,920 --> 00:30:27,720
haven't really seen capacity 
issues. 

576
00:30:27,720 --> 00:30:29,840
So it's almost like they're 
they're solving a problem that 

577
00:30:29,840 --> 00:30:33,280
doesn't exist yet. 
Or they're foreshadowing what's 

578
00:30:33,280 --> 00:30:35,920
coming. 
So I'm very interested to see 

579
00:30:35,920 --> 00:30:38,160
how this play out. 
It could also be first a first 

580
00:30:38,160 --> 00:30:40,520
step toward a savings planner RI
kind of construct. 

581
00:30:40,520 --> 00:30:43,920
But but yeah, nonetheless, I, I 
see a lot of financial risk in 

582
00:30:43,920 --> 00:30:46,360
these external loan providers 
and, and prices are going to 

583
00:30:46,360 --> 00:30:50,280
continue to rise and or they're 
going to want multi year 

584
00:30:50,280 --> 00:30:53,720
commitments. 
Because with the guaranteed 

585
00:30:53,720 --> 00:30:55,920
capacity, you can't do that for 
a month, right? 

586
00:30:55,920 --> 00:30:58,160
Oh. 
No, no, they're multi year is is

587
00:30:58,160 --> 00:31:01,040
what I believe that they're 
looking for there. 

588
00:31:01,640 --> 00:31:03,560
Yeah, yeah. 
Which I don't know 100% for 

589
00:31:03,560 --> 00:31:05,480
certain, but yeah, that's, 
that's my impression. 

590
00:31:05,720 --> 00:31:08,520
Yeah. 
That is something that now 

591
00:31:08,520 --> 00:31:11,840
again, you're thinking about it 
and you're starting to realize 

592
00:31:11,840 --> 00:31:16,520
these cost benefit analysis and 
saying, huh, if we're good 

593
00:31:16,520 --> 00:31:19,400
enough now we're never going to 
change. 

594
00:31:19,800 --> 00:31:24,080
But then later down the line, if
it starts to get a little shaky,

595
00:31:24,440 --> 00:31:28,000
now you have to recognize, is it
worth it for us to get this 

596
00:31:28,080 --> 00:31:31,200
guaranteed capacity? 
Or as you mentioned earlier, 

597
00:31:31,200 --> 00:31:36,320
too, like they're slipping in a 
lot of new costs like the data 

598
00:31:36,320 --> 00:31:40,000
residency, but now there's this 
new cost on this. 

599
00:31:40,000 --> 00:31:43,520
And so they're almost upselling 
you in every way, shape and 

600
00:31:43,520 --> 00:31:45,000
form. 
Exactly. 

601
00:31:45,000 --> 00:31:48,360
And I think it's really hard 
from my perspective at least to 

602
00:31:48,360 --> 00:31:51,280
do any kind of multi year 
commitment against, you know, in

603
00:31:51,280 --> 00:31:55,880
this AI world that we live in 
that is that's really hard to 

604
00:31:55,880 --> 00:31:58,800
look at all the changes in the 
market and say, OK, we're going 

605
00:31:58,800 --> 00:32:01,760
to be using, you know, said 
provider for three years. 

606
00:32:02,680 --> 00:32:04,320
So it's just it's too 
fast-paced. 

607
00:32:04,880 --> 00:32:07,720
So yeah. 
But if there are people that are

608
00:32:07,960 --> 00:32:10,560
probably in position to, you 
know, be able to do that, our 

609
00:32:10,560 --> 00:32:14,360
proxy also gives us the ability 
to kind of use whomever we want 

610
00:32:14,360 --> 00:32:18,760
to use, which is it's nice 
having a Bender agnostic proxy 

611
00:32:18,760 --> 00:32:19,280
for that.
