1
00:00:00,040 --> 00:00:03,880
Over the last few weeks, the New
York Times, Fortune, TechCrunch,

2
00:00:03,880 --> 00:00:07,360
and many more have all covered 
the same new trend. 

3
00:00:08,119 --> 00:00:11,920
It's called token maxing, a new 
trend that's spreading literally

4
00:00:11,920 --> 00:00:15,040
everywhere. 
Yet this month a consultancy 

5
00:00:15,040 --> 00:00:19,160
called Jellyfish just released a
study of 12,000 developers 

6
00:00:19,160 --> 00:00:22,400
across 200 companies. 
The engineers at the bottom of 

7
00:00:22,400 --> 00:00:27,040
this graph that they created 
ship code changes for about 28 P

8
00:00:27,040 --> 00:00:29,520
each. 
The engineers at the top of that

9
00:00:29,520 --> 00:00:34,320
graph ship them for $89.00 each.
The heavy users produce twice as

10
00:00:34,320 --> 00:00:37,600
much output as 319 times the 
cost. 

11
00:00:38,120 --> 00:00:42,040
This is what token maxing is. 
Today I'm going to discuss many 

12
00:00:42,040 --> 00:00:46,040
of the silly stories behind this
new trend with engineers gaming 

13
00:00:46,040 --> 00:00:49,920
leaderboards to encourage them 
to spend lots of money in tokens

14
00:00:50,840 --> 00:00:54,200
to explain what's really 
happening, what token maxing is 

15
00:00:54,200 --> 00:00:57,040
and why the same dynamic is 
coming to every industry and 

16
00:00:57,040 --> 00:01:00,560
job, not just software. 
And importantly, why it's such a

17
00:01:00,560 --> 00:01:03,960
big problem that you must catch 
at work and what you should be 

18
00:01:03,960 --> 00:01:07,240
doing instead. 
This is a loop of Jack Horton. 

19
00:01:07,240 --> 00:01:27,390
I hope you enjoy the show. 
So let's start at the top and 

20
00:01:27,390 --> 00:01:30,750
discuss what the new trends of 
token maxing actually is and and

21
00:01:30,750 --> 00:01:32,230
to be honest, why it's a bit of 
a big problem. 

22
00:01:32,950 --> 00:01:37,550
So for context, in in this 
month, say April 2026, the CTO 

23
00:01:37,550 --> 00:01:41,280
of Uber just got on stage at a 
big industry event and admitted 

24
00:01:41,280 --> 00:01:44,800
that his entire 2026 AI budget 
was completely spent. 

25
00:01:45,960 --> 00:01:48,760
Now four months earlier, he'd 
rolled out Claude code to about 

26
00:01:48,760 --> 00:01:53,720
5000 engineers and and by March,
95% of them were using AI every 

27
00:01:53,720 --> 00:01:57,680
single day, with 70% of the code
they were submitting being AI 

28
00:01:57,680 --> 00:02:00,920
written. 
But only 11% of that was 

29
00:02:00,920 --> 00:02:03,200
actually running in Uber 
systems. 

30
00:02:03,760 --> 00:02:07,880
So 95% of them were using AI. 
70% of code that was being 

31
00:02:07,880 --> 00:02:11,120
submitted were created by AI, 
but only 11% made it to 

32
00:02:11,120 --> 00:02:13,480
production. 
These are very important numbers

33
00:02:13,480 --> 00:02:16,880
and they tell a really good 
story about what's happening 

34
00:02:16,880 --> 00:02:19,000
right now. 
And this is going to be 

35
00:02:19,360 --> 00:02:22,720
something that will be mirrored 
across many different jobs and 

36
00:02:22,720 --> 00:02:25,200
industries. 
As I've said a few times, 

37
00:02:25,800 --> 00:02:29,680
programming is just ahead of the
game and knowledge workers 

38
00:02:29,680 --> 00:02:32,400
generally are going to be going 
through many of the same 

39
00:02:33,040 --> 00:02:36,640
challenges, ideas, new ways of 
working that developers have 

40
00:02:36,640 --> 00:02:38,360
gone through over the last 18 
months. 

41
00:02:38,840 --> 00:02:42,280
So for context, a token is 
roughly 3/4 of a word. 

42
00:02:42,280 --> 00:02:45,280
So when an LLM gives you 3/4 of 
a word, you span the token. 

43
00:02:45,800 --> 00:02:50,040
And to give you a sense of 
scale, the entire text of War 

44
00:02:50,040 --> 00:02:52,520
and Peace, the book is about 
1,000,000 tokens. 

45
00:02:53,280 --> 00:02:56,360
And very funnily, there's 
there's VCs and people blogging 

46
00:02:56,360 --> 00:03:00,800
online and releasing articles 
all the way through 2026 about 

47
00:03:00,800 --> 00:03:05,200
burning about 250 million tokens
in a single day by running a 

48
00:03:05,200 --> 00:03:07,800
swarm of agents all operating in
parallel. 

49
00:03:08,360 --> 00:03:13,360
So 250 million tokens by the 
way, is is about roughly 250 

50
00:03:13,360 --> 00:03:16,920
copies of War and Peace consumed
in a single day by a single 

51
00:03:16,920 --> 00:03:21,280
person. 
I think Jensen Huang of NVIDIA 

52
00:03:21,280 --> 00:03:24,480
said on the All In podcast 
recently that if a half 

53
00:03:24,480 --> 00:03:28,000
$1,000,000 engineer, so they pay
him half $1,000,000 a year, did 

54
00:03:28,000 --> 00:03:32,760
not, and I quote, consume at 
least $250,000 worth of tokens a

55
00:03:32,760 --> 00:03:35,240
year. 
He said he's going to be deeply 

56
00:03:35,240 --> 00:03:37,960
alarmed. 
And by March, the New York Times

57
00:03:37,960 --> 00:03:41,360
had called this a basically a 
new status game all over Silicon

58
00:03:41,360 --> 00:03:43,360
Valley. 
And the measure is quite simple.

59
00:03:43,360 --> 00:03:46,960
Right now, it's just consume, 
consume more tokens, spend more 

60
00:03:46,960 --> 00:03:48,800
and prove that you're doing lots
of stuff. 

61
00:03:49,440 --> 00:03:51,840
Now over the last few weeks, 
this story has taken an 

62
00:03:51,840 --> 00:03:54,760
interesting term because 
engineers at Meta and Microsoft 

63
00:03:54,760 --> 00:03:57,760
and Sales Force and a bunch of 
other big companies have been 

64
00:03:57,760 --> 00:04:00,040
found to be deliberately burning
tokens. 

65
00:04:00,360 --> 00:04:04,440
So running lots of agents trying
to have massive contacts. 

66
00:04:04,440 --> 00:04:08,200
Windows, which costs lots of 
money, going over and over on on

67
00:04:08,320 --> 00:04:11,640
prompts, just looping prompts 
because they worked out that 

68
00:04:11,640 --> 00:04:14,520
their managers were using AI 
consumption as some form of 

69
00:04:14,520 --> 00:04:17,640
performance metric. 
And of course they gained it. 

70
00:04:18,440 --> 00:04:21,399
And those engineers that were at
the top percentage in terms of 

71
00:04:21,399 --> 00:04:24,680
usage and token spend were 
producing about twice the 

72
00:04:24,680 --> 00:04:27,040
output, but at 10 times the 
cost. 

73
00:04:27,560 --> 00:04:31,320
And Meta's internal tool called 
Chordonomics is essentially 

74
00:04:31,320 --> 00:04:32,720
something that went public 
recently. 

75
00:04:33,040 --> 00:04:37,560
So they built this leaderboard 
ranking system and it ranked 

76
00:04:37,560 --> 00:04:39,880
people by the number of tokens 
they'd consumed. 

77
00:04:40,080 --> 00:04:44,000
So about 60 trillion tokens were
used or spent in a 30 day 

78
00:04:44,000 --> 00:04:47,640
window. 
To put that 60 trillion tokens 

79
00:04:47,640 --> 00:04:52,560
into context, if every employee 
at Meta read a full novel every 

80
00:04:52,560 --> 00:04:56,040
single day for 30 years, they 
would collectively consume about

81
00:04:56,040 --> 00:05:01,160
0.1% of that total number. 
And Meta by the way, has 78,000 

82
00:05:01,160 --> 00:05:03,520
employees. 
And that leaderboard handed out 

83
00:05:03,520 --> 00:05:06,680
badges like Token Legend, 
Cachet, Wizard, session 

84
00:05:06,680 --> 00:05:10,000
Immortal. 
And Meta CTO was, you know, 

85
00:05:10,000 --> 00:05:12,760
shouting on the rooftops, 
basically saying that, you know,

86
00:05:12,760 --> 00:05:15,080
they're, they're top ranked 
engineers on these leader boards

87
00:05:15,080 --> 00:05:18,720
have become 5X to 10X more 
productive, saying that this is 

88
00:05:18,720 --> 00:05:21,200
easy, this is like, you know, 
printing money basically. 

89
00:05:22,160 --> 00:05:24,280
And when it went public, of 
course, they then pulled the 

90
00:05:24,280 --> 00:05:26,640
leaderboard because it's quite 
embarrassing. 

91
00:05:26,640 --> 00:05:29,520
And I'll go into why now this is
very developer focused, but 

92
00:05:29,520 --> 00:05:31,800
we're going to see the same in 
every single function in some 

93
00:05:31,800 --> 00:05:34,720
form. 
People would so desperately want

94
00:05:34,720 --> 00:05:37,920
their teams to be using AI that 
they'll just want them to 

95
00:05:37,920 --> 00:05:40,520
consume tokens. 
You know, if token consumption 

96
00:05:40,520 --> 00:05:42,160
is going up, surely outputs 
going up. 

97
00:05:42,840 --> 00:05:45,120
Now the question for me, and I 
think you can probably tell my 

98
00:05:45,120 --> 00:05:48,920
opinion here, is this actually 
making people more productive? 

99
00:06:08,320 --> 00:06:11,640
So if we discuss that point, you
know why I think tokens don't 

100
00:06:11,640 --> 00:06:14,640
mean more output. 
So Jellyfish that that 

101
00:06:14,760 --> 00:06:17,200
consultancy that released that 
study that I mentioned at the 

102
00:06:17,200 --> 00:06:20,400
start of this discussion. 
So they measured pull requests 

103
00:06:20,400 --> 00:06:24,080
shipped per quarter across 
12,000 developers over 200 big 

104
00:06:24,080 --> 00:06:26,520
companies. 
Now a pull request for for 

105
00:06:26,520 --> 00:06:30,440
anyone outside of engineering is
a standard unit of developer 

106
00:06:30,440 --> 00:06:33,320
work, a chunk of code that's 
been submitted for review, it's 

107
00:06:33,320 --> 00:06:35,120
been approved and then added to 
the product. 

108
00:06:35,760 --> 00:06:38,120
So shipping poor request has 
often been, other than merge 

109
00:06:38,120 --> 00:06:42,360
requests, I guess the closest 
thing to let's say maybe a 

110
00:06:42,560 --> 00:06:45,520
universal output. 
Now the engineers that we use in

111
00:06:45,520 --> 00:06:51,320
the least AI burned around $3 of
tokens over 1/4, so over three 

112
00:06:51,320 --> 00:06:53,240
months and shipped 11 poor 
requests each. 

113
00:06:53,760 --> 00:06:58,000
Now the engineers that we're 
using the most burn $1820.00 of 

114
00:06:58,000 --> 00:07:02,800
tokens 1/4 and ship 2313 more 
PRS pull requests. 

115
00:07:02,800 --> 00:07:06,160
So that's twice the output but 
about 600 times the cost. 

116
00:07:06,800 --> 00:07:10,640
And the word output I guess is a
is a loaded word because a lot 

117
00:07:10,640 --> 00:07:13,280
of what the heavy users produce 
doesn't necessarily ship 

118
00:07:13,280 --> 00:07:16,600
everywhere. 
Remember Uber's numbers, 70% of 

119
00:07:16,600 --> 00:07:20,720
that committed code was written 
by AI, but only 11% of that 

120
00:07:20,720 --> 00:07:23,080
actually was running in 
production. 

121
00:07:23,560 --> 00:07:24,760
Now, it's quite easy to 
diagnose. 

122
00:07:24,760 --> 00:07:26,960
You know, just because you're 
producing lots doesn't mean it's

123
00:07:26,960 --> 00:07:29,160
good. 
This is a new area. 

124
00:07:29,320 --> 00:07:31,520
It's an area that everybody's 
trying to innovate from a 

125
00:07:31,520 --> 00:07:35,280
process perspective, a systems 
perspective, trainings and skill

126
00:07:35,280 --> 00:07:37,680
perspective. 
But also there's just natural 

127
00:07:37,680 --> 00:07:40,800
bottlenecks in a process. 
If you think about the theorem 

128
00:07:40,800 --> 00:07:44,080
of constraints, process will 
always move as fast as the 

129
00:07:44,080 --> 00:07:47,560
slowest moving part. 
And right now there's certain 

130
00:07:47,560 --> 00:07:49,360
areas that are just bottlenecks.
You know, review is a 

131
00:07:49,360 --> 00:07:52,320
bottleneck. 
A senior engineer might be able 

132
00:07:52,320 --> 00:07:55,840
to review, say, 2 to 400 lines 
of code in a day, let's say, 

133
00:07:56,240 --> 00:07:58,360
yeah, AI can generate that in 90
seconds. 

134
00:07:59,120 --> 00:08:01,200
Now, people are doing 
automations on top of this and 

135
00:08:01,200 --> 00:08:04,800
automated testing. 
But still, that doesn't mean 

136
00:08:04,800 --> 00:08:08,400
that all of this is quality 
enough to pass QA. 

137
00:08:08,400 --> 00:08:11,840
So that is quality assurance and
testing to actually get inside 

138
00:08:11,840 --> 00:08:13,880
the product. 
And again, just because you're 

139
00:08:13,880 --> 00:08:16,400
burning lots of tokens, are you 
doing it effectively? 

140
00:08:16,880 --> 00:08:20,240
Are you being efficient with 
your token spend, you know, 

141
00:08:20,240 --> 00:08:24,040
putting 1000 words PDF into a 
context window and then asking 

142
00:08:24,040 --> 00:08:25,840
to do, you know, lots of 
questions on that is going to 

143
00:08:25,840 --> 00:08:29,000
use loads of tokens? 
Are you actually achieving 

144
00:08:29,000 --> 00:08:31,200
anything there? 
Now none of this is new as a 

145
00:08:31,200 --> 00:08:33,720
problem. 
There's a principle called good 

146
00:08:33,720 --> 00:08:35,520
hearts law. 
So this is when a measure 

147
00:08:35,520 --> 00:08:38,520
becomes a target, it ceases to 
be a good measure. 

148
00:08:39,240 --> 00:08:43,840
Now a good example is a paper by
someone called Khrushchev, and 

149
00:08:43,840 --> 00:08:45,920
it's about a Soviet candle 
factory. 

150
00:08:46,520 --> 00:08:49,200
And the way they evaluated 
progress was by weight. 

151
00:08:49,360 --> 00:08:51,800
Now, if the chandeliers grew 
heavier and heavier, so all the 

152
00:08:51,800 --> 00:08:56,080
workers were producing heavier 
and heavier chandeliers until 

153
00:08:56,080 --> 00:08:58,640
they started pulling ceilings 
down, they fulfilled the plan. 

154
00:08:59,120 --> 00:09:01,680
Now the question is, is does 
anyone really need that plan? 

155
00:09:02,160 --> 00:09:03,880
Do they actually solve any 
problem? 

156
00:09:03,880 --> 00:09:05,880
No, it doesn't. 
You're just creating heavy 

157
00:09:05,880 --> 00:09:07,800
chandeliers and measuring value 
based on that. 

158
00:09:08,800 --> 00:09:10,000
Bill Gates once said the same 
thing. 

159
00:09:10,000 --> 00:09:13,920
It's like measuring programmers 
by amount of code written is 

160
00:09:13,920 --> 00:09:17,640
like measuring the progress of 
us building a, let's say a plane

161
00:09:17,960 --> 00:09:20,360
by weight has literally no 
value. 

162
00:09:20,840 --> 00:09:24,120
Wells Fargo famously measured 
stuff by cross sells to selling 

163
00:09:24,120 --> 00:09:26,840
more products to the same 
customers and ended up with 

164
00:09:26,840 --> 00:09:28,960
about 3 1/2 million fake 
customers. 

165
00:09:29,480 --> 00:09:32,520
All of this started as a 
reasonable proxy and KPI to try 

166
00:09:32,520 --> 00:09:35,640
and get people to do the right 
things that the organization 

167
00:09:35,640 --> 00:09:39,320
wanted, but obviously stopped 
working the moment everybody 

168
00:09:39,320 --> 00:09:41,160
realised that that's how they 
were being assessed. 

169
00:09:41,160 --> 00:09:43,680
And it doesn't actually deliver 
any value. 

170
00:09:43,680 --> 00:09:45,840
If that measure does not deliver
value and everyone gains that 

171
00:09:45,840 --> 00:09:47,800
system, you're getting more of a
bad thing. 

172
00:09:48,560 --> 00:09:51,080
And token maxing is just the 
newest member of that family. 

173
00:09:51,080 --> 00:09:53,800
It's the new trend. 
Now there are some people that 

174
00:09:53,800 --> 00:09:55,800
push back on this. 
You know, you could argue that 

175
00:09:55,800 --> 00:09:59,600
if your engineers spend 1000 a 
month on say LLMS and ship 10% 

176
00:09:59,600 --> 00:10:03,080
more, that's arguably cheap. 
It's not too expensive, you 

177
00:10:03,080 --> 00:10:06,600
know, Shopify's VP of 
engineering has been making the 

178
00:10:06,600 --> 00:10:10,120
argument very publicly. 
Shopify CEO makes a very similar

179
00:10:10,120 --> 00:10:13,640
point that only through massive 
usage of AI will you get the 

180
00:10:13,640 --> 00:10:16,360
right skills to eventually 
become productive. 

181
00:10:16,360 --> 00:10:18,840
And and the reality is both are 
right. 

182
00:10:19,000 --> 00:10:23,160
If the output is real, then 
token usage is also great 

183
00:10:23,600 --> 00:10:26,600
because you're getting good 
output and clearly lots of it. 

184
00:10:27,360 --> 00:10:30,040
But the jellyfish numbers that 
study that Jellyfish released 

185
00:10:30,360 --> 00:10:34,440
and the example by Uber suggests
that often this just still isn't

186
00:10:34,440 --> 00:10:36,480
the case. 
And I'll come on to some of the 

187
00:10:36,480 --> 00:10:38,080
reasons I think this is the 
case, by the way. 

188
00:10:39,160 --> 00:10:41,320
And one of them, which is one of
the reasons I got so interested 

189
00:10:41,320 --> 00:10:44,400
in this episode, is what mindset
is actually solving for 

190
00:10:44,400 --> 00:10:46,760
internally and as a result has 
been solving for for other 

191
00:10:46,760 --> 00:10:49,400
customers and partners. 
But let's talk about, I guess, 

192
00:10:49,400 --> 00:10:52,040
the reason that things are 
getting so expensive and things 

193
00:10:52,040 --> 00:10:55,760
aren't necessarily improving in 
terms of productive output. 

194
00:10:55,960 --> 00:10:58,720
And as a result, what I think 
people should be focusing on 

195
00:10:58,720 --> 00:11:01,000
instead. 
So as I said, I, I think the 

196
00:11:01,000 --> 00:11:02,880
case for heavy consumption is 
partly right. 

197
00:11:03,640 --> 00:11:04,960
It's going to help people gain 
skills. 

198
00:11:04,960 --> 00:11:08,160
It's going to force people to 
actually adopt this technology 

199
00:11:08,160 --> 00:11:11,000
and get over themselves. 
But this still doesn't explain 

200
00:11:11,000 --> 00:11:13,960
why things are so expensive and 
things still aren't very good in

201
00:11:13,960 --> 00:11:18,000
terms of quality output. 
And I think there's about 3 or 4

202
00:11:18,000 --> 00:11:19,600
reasons. 
And, and the first reason I 

203
00:11:19,600 --> 00:11:22,080
think is, is pricing. 
So that previous example of 

204
00:11:22,080 --> 00:11:25,600
measuring people by a proxy, 
like lines of code written or 

205
00:11:25,600 --> 00:11:29,040
the weight of a chandelier, 
well, when those came about, 

206
00:11:29,040 --> 00:11:32,600
that didn't cost as much. 
So when a programmer in say 1982

207
00:11:32,600 --> 00:11:37,200
row, let's say 500 lines of 
code, those lines sat on a disk 

208
00:11:37,200 --> 00:11:39,040
for nothing, didn't cost people 
money. 

209
00:11:39,560 --> 00:11:41,920
Whereas AI coding tools aren't 
sold by the seat, they're sold 

210
00:11:41,920 --> 00:11:45,480
by the token. 
So Anthropic, for example, move 

211
00:11:45,480 --> 00:11:49,160
their big enterprise contracts 
towards the late end of 2025 

212
00:11:49,160 --> 00:11:53,120
from a flat roughly $200 per 
user model with tokens included 

213
00:11:53,440 --> 00:11:57,120
to a more metered compute model.
So you get $20.00 for a base 

214
00:11:57,120 --> 00:11:58,920
seat and then you get usage 
charge. 

215
00:11:59,640 --> 00:12:02,120
So when legacy contracts all 
over the world have started 

216
00:12:02,320 --> 00:12:06,800
expiring around, you know, 
March, April, May 2026, and all 

217
00:12:06,800 --> 00:12:09,040
these companies have really been
drinking from the kool-aid of 

218
00:12:09,040 --> 00:12:13,200
token maxing, many customers 
bills will be tripling overnight

219
00:12:13,200 --> 00:12:16,080
for Anthropic. 
Let's say a typical big company 

220
00:12:16,640 --> 00:12:22,120
runs maybe 150 to $250 a month 
per person, and then heavy users

221
00:12:22,120 --> 00:12:25,080
run 5 to 10 times that. 
You can see how the numbers 

222
00:12:25,080 --> 00:12:27,640
start to get really expensive 
when you've got 78000 

223
00:12:27,640 --> 00:12:30,200
developers. 
And then when you have some 

224
00:12:30,200 --> 00:12:32,880
agents that are just running 
continuously, doing huge tasks 

225
00:12:33,480 --> 00:12:36,040
and using that entire token 
limit in just a matter of days. 

226
00:12:36,040 --> 00:12:39,160
Well, imagine if you have 7000 
people with those agents. 

227
00:12:40,000 --> 00:12:42,800
Again, if you remember the 
episode I discussed recently, 

228
00:12:42,800 --> 00:12:46,000
the Open Core episode and why 
it's such an important software 

229
00:12:46,160 --> 00:12:49,960
invention and why Jensen Fuang 
is in love with it, is because 

230
00:12:50,120 --> 00:12:53,280
these things are going to 
consume so many tokens, which 

231
00:12:53,280 --> 00:12:57,360
Dr. LLMS business models which 
ultimately benefits NVIDIA 

232
00:12:58,600 --> 00:13:00,480
anyway. 
I think the second reason is 

233
00:13:00,480 --> 00:13:03,720
culture. 
Right now we've got this weird 

234
00:13:03,720 --> 00:13:06,880
culture like Matter and many of 
the big companies where you've 

235
00:13:06,880 --> 00:13:10,160
got this leaderboard of 
consuming tokens, you're Jensen 

236
00:13:10,160 --> 00:13:13,120
Wang and wanting to people to 
spend, you know, quarter of 

237
00:13:13,120 --> 00:13:17,640
$1,000,000 in tokens per year. 
So really this token consumption

238
00:13:17,640 --> 00:13:22,120
has become quite performative in
a way that you know, let's say a

239
00:13:22,120 --> 00:13:26,600
company selling $20.00 per month
per seat per user just didn't 

240
00:13:26,600 --> 00:13:28,320
see and didn't experience in the
same way. 

241
00:13:28,320 --> 00:13:31,600
And when that metric is really 
publicly visible and the boss 

242
00:13:31,600 --> 00:13:34,840
says no limit and you get 
rewarded a pat on your head, 

243
00:13:35,120 --> 00:13:37,400
obviously people will start to 
play it and gain it. 

244
00:13:37,400 --> 00:13:41,160
I think the third reason here is
a hidden quality cost. 

245
00:13:41,760 --> 00:13:46,360
You know, 43% of AI generated 
code requires lots of manual 

246
00:13:46,360 --> 00:13:48,440
debugging after it's gone 
through QA. 

247
00:13:48,440 --> 00:13:51,680
So testing, and there's been 
some studies to say that some 

248
00:13:51,680 --> 00:13:55,520
developers are spending about 38
to 40% of their week on 

249
00:13:55,520 --> 00:13:59,120
debugging and verifying whether 
what the agent has actually 

250
00:13:59,120 --> 00:14:01,680
built is the thing they meant to
build. 

251
00:14:02,120 --> 00:14:04,320
So that heavy token usage by an 
agent. 

252
00:14:04,320 --> 00:14:07,560
So generating lots of words, 
although it seems really 

253
00:14:07,560 --> 00:14:11,280
attractive, it's only real if a 
human clears it up or you have 

254
00:14:11,280 --> 00:14:15,080
some form of really trustworthy 
testing software and suite that 

255
00:14:15,240 --> 00:14:18,560
automates that. 
And these things aren't always 

256
00:14:18,560 --> 00:14:21,000
easy to automate reliably. 
It's not easy. 

257
00:14:21,480 --> 00:14:23,640
And that clean up time isn't 
showing up on the token 

258
00:14:23,640 --> 00:14:25,480
dashboards. 
It shows up when you try to 

259
00:14:25,480 --> 00:14:27,800
deploy. 
And this is going to be the same

260
00:14:27,800 --> 00:14:31,280
across all spaces and industries
for those in every single, let's

261
00:14:31,280 --> 00:14:34,400
say a law firm or a marketing 
team, a sales team, wherever you

262
00:14:34,400 --> 00:14:36,200
are. 
The role of managers is also 

263
00:14:36,200 --> 00:14:37,840
quality assurance and quality 
checking. 

264
00:14:38,040 --> 00:14:40,680
You need to make sure what we're
producing as an organization is 

265
00:14:40,680 --> 00:14:44,040
up to scratch. 
And again, maybe you can get AI 

266
00:14:44,040 --> 00:14:47,280
and go, yeah, AI, go and check 
this, but maybe you also need 

267
00:14:47,280 --> 00:14:50,360
people. 
So you have explosions of output

268
00:14:50,360 --> 00:14:53,040
from some people and then lots 
of people having to verify that.

269
00:15:09,740 --> 00:15:12,620
I think this comes to the other 
problem, which is contacts and 

270
00:15:12,620 --> 00:15:15,660
orientation of agents and then 
being able to reliably produce 

271
00:15:15,660 --> 00:15:17,660
things that you want it to 
produce. 

272
00:15:18,620 --> 00:15:22,020
Now some people are solving this
with skills, with MD files to 

273
00:15:22,020 --> 00:15:24,340
give agents context and rules 
and processes. 

274
00:15:25,280 --> 00:15:27,120
And I've done episodes on that 
recently that have been really, 

275
00:15:27,120 --> 00:15:29,120
really popular. 
But I think still that it's just

276
00:15:29,120 --> 00:15:31,480
a trend of the day. 
It will go away over the next 18

277
00:15:31,480 --> 00:15:35,240
months because ultimately this 
system still relies on people 

278
00:15:35,240 --> 00:15:38,760
updating those documents, 
updating those skills, still 

279
00:15:39,280 --> 00:15:41,680
keeping them so fresh that they 
don't go stale. 

280
00:15:42,240 --> 00:15:45,160
And an agent reliably using that
at scale across large 

281
00:15:45,160 --> 00:15:47,920
organisations, which although 
it's infinitely better where we 

282
00:15:47,920 --> 00:15:50,880
are now than 18 months ago, I 
still don't think it's perfect. 

283
00:15:50,880 --> 00:15:53,320
I think we're a way off from 
what will come as the next 

284
00:15:53,320 --> 00:15:55,080
evolution. 
And mindset is seeing this 

285
00:15:55,080 --> 00:15:57,960
because we're quite far ahead of
most people from a process 

286
00:15:57,960 --> 00:16:01,160
perspective internally. 
And we've already really faced 

287
00:16:01,160 --> 00:16:02,560
that problem. 
We actually built our own tool 

288
00:16:02,560 --> 00:16:05,560
to solve that as an internal 
tool, which the team are 

289
00:16:05,560 --> 00:16:07,720
absolutely living. 
And I'll give you a good 

290
00:16:07,720 --> 00:16:09,320
example. 
So when you ask an agent, say, 

291
00:16:09,320 --> 00:16:12,760
to make a change in a code base,
it might not have all the 

292
00:16:12,760 --> 00:16:15,960
context of your code base itself
because it can't just, you know,

293
00:16:15,960 --> 00:16:18,920
scan all of your code base 
because the context window blows

294
00:16:18,920 --> 00:16:21,560
up. 
So a lot of people create 

295
00:16:21,560 --> 00:16:25,480
markdown files, so essentially 
documents which specify specific

296
00:16:25,480 --> 00:16:29,200
ways things work, how my APIs 
work, how this works, how that 

297
00:16:29,200 --> 00:16:31,560
works. 
They also have processes they 

298
00:16:31,560 --> 00:16:34,080
enshrined. 
So how to deploy, how to push 

299
00:16:34,080 --> 00:16:36,960
tickets to Jira or Gitlabs, your
product project management 

300
00:16:36,960 --> 00:16:39,040
tools. 
You then get people to pass 

301
00:16:39,040 --> 00:16:40,840
information. 
So a product team would go and 

302
00:16:40,840 --> 00:16:43,960
write up what you're trying to 
build in the spec, you pass it 

303
00:16:43,960 --> 00:16:46,160
to engineers and that might go 
into an MD file. 

304
00:16:46,160 --> 00:16:50,920
And then MD file gets given to 
agents and people create like 

305
00:16:50,920 --> 00:16:53,800
we've been doing quite complex 
setups, like really powerful 

306
00:16:53,800 --> 00:16:57,560
setups, clever setups of skills 
and MD files that essentially 

307
00:16:58,480 --> 00:17:01,640
enable an agent to operate and 
do good coding. 

308
00:17:01,880 --> 00:17:03,820
And this is the same that's 
happening now with called Co 

309
00:17:03,820 --> 00:17:06,760
work and all organisations as 
I've covered in other episodes. 

310
00:17:07,480 --> 00:17:08,880
So you create this fantastic 
system. 

311
00:17:09,440 --> 00:17:13,480
However, that doesn't mean the 
agent actually producing great 

312
00:17:13,480 --> 00:17:14,720
work. 
A lot of the time it produces 

313
00:17:14,720 --> 00:17:17,359
random work. 
That doesn't mean that all of 

314
00:17:17,359 --> 00:17:20,200
your engineers are able to 
suddenly use these things really

315
00:17:20,200 --> 00:17:22,160
well. 
It doesn't mean that you can 

316
00:17:22,160 --> 00:17:25,000
easily track all of the key 
important decisions that are 

317
00:17:25,000 --> 00:17:27,040
made. 
Because in a project management 

318
00:17:27,040 --> 00:17:30,160
tool, you have tickets or you 
have tasks, but an agent doesn't

319
00:17:30,160 --> 00:17:34,720
think in that way, doesn't think
in a small, like tiny task, add 

320
00:17:34,720 --> 00:17:37,400
a headline, write subcopy, 
review. 

321
00:17:37,920 --> 00:17:41,880
It thinks in larger chunks of 
work, but those larger chunks of

322
00:17:41,880 --> 00:17:46,160
work, therefore, when they're 
executed, have to rely on the 

323
00:17:46,160 --> 00:17:49,760
agent accessing that MD file 
structure, the skills that 

324
00:17:49,760 --> 00:17:52,640
you've created to actually do 
the work. 

325
00:17:52,960 --> 00:17:54,720
And then you have to test 
whether that's been done. 

326
00:17:55,200 --> 00:17:57,960
Because if it writes 10,000 
words of or lines of code, so 

327
00:17:57,960 --> 00:18:01,440
let's say doesn't mean that all 
of that is actually what was 

328
00:18:01,440 --> 00:18:03,560
intended to be built. 
You often find things that 

329
00:18:03,560 --> 00:18:05,760
accidentally are snook in, you 
know, why did you do that? 

330
00:18:06,360 --> 00:18:07,640
Oh, I thought this was a good 
idea. 

331
00:18:08,560 --> 00:18:09,800
And again, I'm going to keep 
coming to this. 

332
00:18:09,800 --> 00:18:12,440
This is the same thing that's 
going to be happening at scale 

333
00:18:12,640 --> 00:18:15,680
with knowledge workers. 
And you know, three or four 

334
00:18:15,680 --> 00:18:18,760
months in or five months in from
having on D files and skills, 

335
00:18:19,200 --> 00:18:23,560
you will start to find that 
suddenly some of these go stale.

336
00:18:23,560 --> 00:18:25,920
People stop updating them. 
They don't use them properly 

337
00:18:25,920 --> 00:18:28,680
because naturally people just 
drift, especially when it's 

338
00:18:28,680 --> 00:18:30,520
nebulous. 
You know, it's, it's documents 

339
00:18:31,160 --> 00:18:33,680
like these things of skills that
I, you know, I, I've been 

340
00:18:33,720 --> 00:18:35,320
promoting them and they're 
really great. 

341
00:18:35,520 --> 00:18:38,840
But they are just a document 
specifying how to do something. 

342
00:18:39,280 --> 00:18:42,440
And that information must be 
continually updated 24/7 for 

343
00:18:42,440 --> 00:18:44,720
every decision. 
And the more decisions, 

344
00:18:44,720 --> 00:18:48,480
strategies, projects you do, the
more skills and MD files you 

345
00:18:48,480 --> 00:18:50,240
need. 
And it just becomes very bloated

346
00:18:50,240 --> 00:18:52,360
and chaotic. 
Trust me, very quickly. 

347
00:18:53,280 --> 00:18:56,760
So this problem around token 
consumption relies on everyone 

348
00:18:56,760 --> 00:18:58,800
updating. 
This essentially document sets 

349
00:18:59,320 --> 00:19:01,800
that live somewhere. 
But then you need to track who's

350
00:19:01,800 --> 00:19:05,280
allowed to update that document.
Has this person used that skill?

351
00:19:05,280 --> 00:19:06,800
Can I track if they've used that
skill? 

352
00:19:07,200 --> 00:19:09,760
Wait, someone over here updated 
that document which changed our 

353
00:19:09,760 --> 00:19:12,440
process? 
OK, so we need a review process 

354
00:19:12,440 --> 00:19:16,200
to allow only some people to 
update an asset like an MD file 

355
00:19:16,200 --> 00:19:20,000
or a skill, but not others. 
But then again, things go stale.

356
00:19:20,280 --> 00:19:22,600
Work isn't actually a high 
quality anymore. 

357
00:19:23,360 --> 00:19:25,600
So this is kind of what happens 
in a big system when you start 

358
00:19:25,600 --> 00:19:28,560
to really scale it and scale AI 
first processes. 

359
00:19:29,080 --> 00:19:30,720
So yeah, this is it's it's a 
real problem. 

360
00:19:30,720 --> 00:19:33,480
So you get drift from agents 
delivering the wrong things, you

361
00:19:33,480 --> 00:19:36,880
get documents that go outdated, 
people not keeping everything up

362
00:19:36,880 --> 00:19:40,400
to date, and essentially the 
entire system starts to slowly 

363
00:19:40,400 --> 00:19:44,320
decay and performance goes down,
yet token consumption goes way 

364
00:19:44,320 --> 00:19:46,080
up. 
That's honestly one of the 

365
00:19:46,080 --> 00:19:47,920
reasons we've actually solved 
this and been solving this 

366
00:19:47,920 --> 00:19:49,760
internally for ourselves. 
We call it Memex. 

367
00:19:50,320 --> 00:19:53,400
It's essentially a system 
whereby you specify strategies 

368
00:19:53,400 --> 00:19:57,040
at the top and instead of using 
project management tasks, the 

369
00:19:57,160 --> 00:19:59,880
document which has access to 
your code base, which is 

370
00:19:59,880 --> 00:20:03,880
indexed, has access to specific 
files and other function and all

371
00:20:03,880 --> 00:20:07,240
the skills turns that into tasks
which are managed in a different

372
00:20:07,240 --> 00:20:09,120
way. 
So for anyone that's really 

373
00:20:09,120 --> 00:20:11,600
interested in that, do reach out
because we're actually having 

374
00:20:11,600 --> 00:20:14,640
lots of conversations about this
internally to share how to use 

375
00:20:14,640 --> 00:20:18,040
something like this and how to 
implement better AI first 

376
00:20:18,040 --> 00:20:37,840
processes at the moment. 
Now I think that's a good 

377
00:20:37,840 --> 00:20:40,240
opportunity to now go and what I
think people should be measuring

378
00:20:40,240 --> 00:20:42,360
instead, how they should go 
about this. 

379
00:20:42,840 --> 00:20:46,040
Companies that are thinking 
about this really, really 

380
00:20:46,040 --> 00:20:48,760
effectively over them from a 
process perspective and how 

381
00:20:48,760 --> 00:20:51,560
we're solving that with internal
tools like I described. 

382
00:20:52,440 --> 00:20:54,960
I think also the way you measure
output is important. 

383
00:20:55,400 --> 00:20:58,120
So for example Salesforce have 
been trying to measure something

384
00:20:58,120 --> 00:21:03,040
called agentic work units. 
So AW use as a complete counter 

385
00:21:03,040 --> 00:21:06,240
to token maxing instead of 
counting tokens consumed. 

386
00:21:06,240 --> 00:21:10,240
An AWU is essentially a 
completed unit of business 

387
00:21:10,240 --> 00:21:13,800
relevant work. 
So let's say Singapore Airlines 

388
00:21:14,080 --> 00:21:18,840
use AW use to track customer 
resolution time, another company

389
00:21:18,840 --> 00:21:21,280
use them to track the quality of
product recommendations. 

390
00:21:21,840 --> 00:21:25,840
And about 2.4 billion AWUS were 
generated across the system and 

391
00:21:25,840 --> 00:21:30,400
ecosystem in Q42025 alone. 
Now if you're running a team, 

392
00:21:30,720 --> 00:21:33,320
the job is therefore to define 
what your AWU is. 

393
00:21:33,880 --> 00:21:36,640
The unit has to represent a 
complete piece of work you'd 

394
00:21:36,640 --> 00:21:40,200
otherwise have to pay the human 
to do, and it has to track 

395
00:21:40,200 --> 00:21:41,520
whether that work actually 
stuck. 

396
00:21:41,960 --> 00:21:44,080
So for engineering, it's not 
necessarily pull request 

397
00:21:44,080 --> 00:21:47,560
generated or pull requests 
merged into a Co base. 

398
00:21:47,560 --> 00:21:50,800
It's pull requests that made it 
into production and stayed 

399
00:21:50,800 --> 00:21:55,800
there, or bugs per shipped PR. 
So per shipped piece of work, 

400
00:21:55,880 --> 00:21:59,040
how many bugs were generated? 
How many incidents were traced 

401
00:21:59,040 --> 00:22:02,560
back to AI assisted code? 
Or track the humour review hours

402
00:22:02,560 --> 00:22:05,960
your team is burning on AI code 
for sales. 

403
00:22:05,960 --> 00:22:08,920
It could be qualified 
opportunities that progressed a 

404
00:22:08,920 --> 00:22:13,640
sales stage that was AI assisted
versus not illegal. 

405
00:22:13,640 --> 00:22:17,960
Maybe it's contract sign, not 
drafts produced of content 

406
00:22:17,960 --> 00:22:20,200
actually published that drove 
traffic for marketing, not the 

407
00:22:20,200 --> 00:22:22,040
content generated. 
Again. 

408
00:22:22,040 --> 00:22:24,560
I keep coming back to that 
chandelier factory example. 

409
00:22:25,520 --> 00:22:29,280
What a stupid stupid. 
Proxy and token matching I 

410
00:22:29,280 --> 00:22:31,280
believe is the exact same and it
will end. 

411
00:22:31,280 --> 00:22:33,840
And it is a big mistake to get 
obsessed with that. 

412
00:22:34,560 --> 00:22:37,320
It's maybe a good starting place
to get your team using it, but 

413
00:22:37,400 --> 00:22:41,280
it's not going to produce value.
So yeah, those organisations 

414
00:22:41,280 --> 00:22:44,280
that, yes, try and get people to
adopt these tools but care about

415
00:22:44,480 --> 00:22:47,440
how you organise contextual 
information that goes beyond 

416
00:22:47,440 --> 00:22:50,920
just MD files and skills that 
don't try and get people to gain

417
00:22:50,920 --> 00:22:53,960
the system to spend tokens and 
think about what they're really 

418
00:22:53,960 --> 00:22:57,360
trying to measure, are going to 
succeed here and produce actual 

419
00:22:57,360 --> 00:23:01,680
value, not just trash. 
Anyway, that's it for this week.

420
00:23:02,120 --> 00:23:04,360
Thank you for listening and I'll
see you next time.

