1
00:00:00,120 --> 00:00:04,000
Anthropic just released a major 
labour market report. 

2
00:00:04,600 --> 00:00:07,200
It contains a chart showing the 
gap between what AI could 

3
00:00:07,320 --> 00:00:10,880
theoretically do across every 
occupation and what it's 

4
00:00:10,880 --> 00:00:14,560
actually doing. 
And that gap is, well, pretty 

5
00:00:14,560 --> 00:00:16,720
big. 
I think the chart is a very 

6
00:00:16,760 --> 00:00:19,600
honest picture of where things 
stand with AI today. 

7
00:00:19,800 --> 00:00:24,080
Studies have found that AI could
theoretically handle 94% of 

8
00:00:24,080 --> 00:00:28,520
tasks in computing and math is 
doing 36%. 

9
00:00:29,040 --> 00:00:31,680
In legal work apparently could 
handle 90%. 

10
00:00:32,040 --> 00:00:34,400
He's doing 20. 
In most categories, the gap is 

11
00:00:34,440 --> 00:00:37,240
even wider. 
Today, I'm going to cover what 

12
00:00:37,240 --> 00:00:41,880
is really slowing this down and 
why this ceiling may be just too

13
00:00:41,880 --> 00:00:45,040
high to start with, and what 
technology leaps we must make 

14
00:00:45,400 --> 00:00:52,080
for AI agents to truly impact 93
and 94% of tasks across 

15
00:00:52,080 --> 00:00:54,040
different jobs. 
This is in the Loop with Jack 

16
00:00:54,040 --> 00:00:55,760
Horton. 
I hope you enjoy the show. 

17
00:01:13,800 --> 00:01:16,400
start from the top. 
And there's a couple stories 

18
00:01:16,400 --> 00:01:18,600
from this really interesting 
study, but most importantly the 

19
00:01:18,600 --> 00:01:22,000
chart. 
So Anthropic published a report 

20
00:01:22,000 --> 00:01:26,120
called Labor Market Impacts of 
AI, a new measure and early 

21
00:01:26,120 --> 00:01:27,960
evidence. 
And it introduced something they

22
00:01:27,960 --> 00:01:31,800
called observed exposure, which 
is basically a metric that puts 

23
00:01:31,800 --> 00:01:36,360
theoretical AI capability next 
to real world usage data to see 

24
00:01:36,360 --> 00:01:40,640
which jobs are actually being 
impacted to try and get get away

25
00:01:40,640 --> 00:01:43,280
a little bit from the hype and 
the media cycles. 

26
00:01:43,440 --> 00:01:45,840
And the report is summarized 
well with the radar chart. 

27
00:01:46,160 --> 00:01:49,200
So if you can picture those 
almost a circular chart and it 

28
00:01:49,200 --> 00:01:51,560
looks like there's a spider's 
web in the middle of it. 

29
00:01:51,640 --> 00:01:54,520
They used to tell the user at 
school to rate how good I was at

30
00:01:54,520 --> 00:01:57,120
different skills around cooking 
in cooking class. 

31
00:01:57,240 --> 00:02:00,000
And essentially what you see in 
this image is a large blue 

32
00:02:00,600 --> 00:02:03,600
spider diagram that stretches 
out towards the edges for a 

33
00:02:03,600 --> 00:02:07,160
bunch of different occupations 
of management, business, 

34
00:02:07,400 --> 00:02:11,080
computer, architecture, life 
sciences, social services, 

35
00:02:11,080 --> 00:02:13,360
legal. 
And that's a theoretical impact 

36
00:02:13,400 --> 00:02:16,400
of AI on these occupations. 
And I can tell you it's a much, 

37
00:02:16,400 --> 00:02:20,280
much bigger spider web than the 
red one, which is the actual 

38
00:02:20,280 --> 00:02:23,080
current observed usage. 
And this is measured by 

39
00:02:23,080 --> 00:02:25,000
Anthropic's own data. 
And the numbers are really 

40
00:02:25,000 --> 00:02:28,360
interesting between how much AI 
could theoretically transform 

41
00:02:28,360 --> 00:02:31,800
these occupations versus 
observed usage like computer and

42
00:02:31,800 --> 00:02:36,880
maths. 94% theoretical but only 
33% observed. 

43
00:02:37,320 --> 00:02:42,000
Legal, meaning 90% theoretical 
but barely passed. 20% observed.

44
00:02:42,120 --> 00:02:45,280
Almost across every domain. 
It's the same pattern apart from

45
00:02:45,280 --> 00:02:47,760
a bunch of domains that 
basically hasn't been impacted 

46
00:02:47,760 --> 00:02:49,680
at all. 
You've got, you know, grounds 

47
00:02:49,680 --> 00:02:53,000
maintenance, food serving, 
agriculture, construction, 

48
00:02:53,240 --> 00:02:55,400
installation and repair. 
None of those have been touched 

49
00:02:55,400 --> 00:02:57,520
by AI really. 
Obviously Anthropic's 

50
00:02:57,520 --> 00:03:00,040
interpretation is that this is a
major growth story. 

51
00:03:00,160 --> 00:03:03,520
They said, and I quote, as 
capabilities advance, adoption 

52
00:03:03,520 --> 00:03:07,320
spreads and deployment deepens, 
the red area will grow to cover 

53
00:03:07,320 --> 00:03:10,600
the blue area. 
And in their framing, the red is

54
00:03:10,600 --> 00:03:13,840
the blue's younger self. 
Give it time and convergence is 

55
00:03:13,840 --> 00:03:15,760
obviously inevitable. 
However, you could read it 

56
00:03:15,760 --> 00:03:18,640
differently. 
That gap could also be telling 

57
00:03:18,640 --> 00:03:21,320
us where the boundaries are, not
where it's heading. 

58
00:03:21,640 --> 00:03:24,280
And Anthropic is looking at that
distance and seeing potential. 

59
00:03:24,280 --> 00:03:27,560
Whereas, as we know, some people
will see it as a big chasm that 

60
00:03:27,560 --> 00:03:48,130
may or may not ever be crossed. 
So yeah, let's discuss that. 

61
00:03:48,370 --> 00:03:50,610
So from one perspective, you 
could argue that the ceiling is 

62
00:03:50,610 --> 00:03:53,450
actually lower than it looks. 
So that blue shape, that 

63
00:03:53,450 --> 00:03:58,690
theoretical ceiling comes from a
2023 Open AI paper called GPTS 

64
00:03:58,690 --> 00:04:01,130
Are GPTS. 
It's almost three years old now.

65
00:04:01,250 --> 00:04:04,170
The authors went through every 
task in the ONET database, which

66
00:04:04,170 --> 00:04:07,930
is a primary source or US source
of occupational information, 

67
00:04:08,690 --> 00:04:12,210
covering about 900 jobs. 
And in the study, Open AI scored

68
00:04:12,250 --> 00:04:17,440
each 11 if an LLM alone could 
double the speed, 0.5 if it 

69
00:04:17,440 --> 00:04:21,240
could with additional tools, 0 
if it was out of reach. 

70
00:04:21,360 --> 00:04:24,880
Anthropic adopted this as the 
upper bound of their chart. 

71
00:04:25,040 --> 00:04:27,400
The issue is how those scores 
were generated though. 

72
00:04:27,880 --> 00:04:30,200
In the study, they were asking 
whether a task could be speed up

73
00:04:30,200 --> 00:04:33,720
by a model with, say, a 2000 
word input limit and no access 

74
00:04:33,720 --> 00:04:35,200
to current information on the 
Internet. 

75
00:04:36,000 --> 00:04:39,080
That's not measuring what 
happens when the AI has to be 

76
00:04:39,080 --> 00:04:42,440
integrated into a complex 
procurement workflow at bank. 

77
00:04:42,520 --> 00:04:44,920
It's not measuring what happens 
when the output has to pass a 

78
00:04:44,920 --> 00:04:49,760
legal review, or when the data 
lives in 3-4 different systems 

79
00:04:49,760 --> 00:04:52,400
that just don't talk to each 
other, or when the person using 

80
00:04:52,400 --> 00:04:54,800
it doesn't trust it. 
It's measuring what's possible 

81
00:04:54,800 --> 00:04:58,520
in a very clean, you know, very 
controlled hypothetical 

82
00:04:58,520 --> 00:05:02,000
environment, and then trying to 
project that across the entire 

83
00:05:02,160 --> 00:05:06,040
economy of the occupation. 
Anthropic does flag this to be 

84
00:05:06,040 --> 00:05:09,040
fair to them, but their argument
is that the models just have 

85
00:05:09,040 --> 00:05:12,240
improved, so the real ceiling is
probably much higher than the 

86
00:05:12,240 --> 00:05:14,520
bully suggests. 
But you can run that logic in 

87
00:05:14,520 --> 00:05:16,080
the opposite direction too, you 
know. 

88
00:05:16,080 --> 00:05:19,720
The original scoring was 
probably way, way, way generous 

89
00:05:20,040 --> 00:05:22,960
because it was measuring what a 
model could do under perfect 

90
00:05:22,960 --> 00:05:28,080
conditions with no friction, no 
context, no complexity, and also

91
00:05:28,080 --> 00:05:31,680
during a period when hype around
LLMS was at its absolute peak. 

92
00:05:32,240 --> 00:05:34,520
So if anything, you could argue 
the blue is too big, not too 

93
00:05:34,520 --> 00:05:36,320
small. 
See, I, I, I think studies like 

94
00:05:36,320 --> 00:05:39,440
this are really helpful to 
understand, but not obviously 

95
00:05:39,440 --> 00:05:41,680
where I get very interested. 
We should come on to, I guess in

96
00:05:41,680 --> 00:05:44,000
the secondary part of this 
discussion, because I've talked 

97
00:05:44,000 --> 00:05:48,080
a lot about this. 
Benchmarks these days are, well,

98
00:05:48,920 --> 00:05:51,520
just as much about the 
importance of vibes. 

99
00:05:51,680 --> 00:05:55,400
So the vibes of this new LLM and
how it feels and B, the 

100
00:05:55,400 --> 00:05:57,640
complexity of real world 
deployment that that's what 

101
00:05:57,640 --> 00:06:00,320
these benchmarks rarely capture.
So the question I'm actually 

102
00:06:00,320 --> 00:06:06,600
interested in is, how do we make
this translate into a situation 

103
00:06:06,600 --> 00:06:10,360
where conditions aren't perfect?
Because these benchmarks are all

104
00:06:10,360 --> 00:06:12,560
in perfect conditions. 
The chart suggests that the 

105
00:06:12,560 --> 00:06:16,080
answer is not nearly as simple 
as we'd like to think, because 

106
00:06:16,080 --> 00:06:19,400
that red spider's web is much, 
much smaller than the blue 

107
00:06:19,400 --> 00:06:21,120
theoretical one. 
Let's explore that. 

108
00:06:21,280 --> 00:06:23,640
Why isn't AI impacting more 
tasks? 

109
00:06:23,640 --> 00:06:26,520
I want to say something quite 
clearly before we get into this 

110
00:06:26,520 --> 00:06:29,800
discussion because it could 
sound like I feel somewhat 

111
00:06:29,800 --> 00:06:33,160
negative about the situation 
where it's not actually where I 

112
00:06:33,160 --> 00:06:36,960
land. 
I would describe myself as a 

113
00:06:37,240 --> 00:06:40,920
realist argument about the 
situation because I believe AI 

114
00:06:40,920 --> 00:06:43,840
is going to be transformative. 
I work in the industry, we 

115
00:06:43,840 --> 00:06:48,640
actively build systems and 
products to deploy it into big 

116
00:06:48,640 --> 00:06:50,120
companies. 
So yeah, it's more of a 

117
00:06:50,120 --> 00:06:52,560
realistic argument about how 
long transformation actually 

118
00:06:52,560 --> 00:06:55,560
takes before technology becomes 
truly upscale. 

119
00:06:55,760 --> 00:06:58,760
Anyway, there are a few things 
to explain this gap. 

120
00:06:58,760 --> 00:07:03,280
I think in theoretical 
expectation and reality 1 is 

121
00:07:03,280 --> 00:07:06,440
about the technology, which 
something we as in mindset are 

122
00:07:06,480 --> 00:07:09,200
actively working on. 
So it's really interesting and 

123
00:07:09,360 --> 00:07:11,400
can really give a kind of 
insider's perspective on. 

124
00:07:12,120 --> 00:07:14,560
And the others are human 
problems. 

125
00:07:15,080 --> 00:07:18,040
And I think really the human and
organizational problems are the 

126
00:07:18,040 --> 00:07:37,480
hardest to fix. 
Let's start with the people 

127
00:07:37,480 --> 00:07:39,560
problem. 
The first problem is that the 

128
00:07:39,560 --> 00:07:42,560
adoption curve is more of a 
Cliff rather than a slope right 

129
00:07:42,560 --> 00:07:45,560
now in AI open AI S own numbers 
tell us a really good story. 

130
00:07:46,080 --> 00:07:49,880
ChatGPT has about 900 million 
active users or weekly active 

131
00:07:49,880 --> 00:07:52,200
users less now after the quick 
GBT movement. 

132
00:07:52,960 --> 00:07:55,640
And about 50 million of those 
are actually paying power users.

133
00:07:55,640 --> 00:07:59,360
So the top 5% of those paid 
subscribers use reasoning 

134
00:07:59,360 --> 00:08:03,320
capabilities about 7 times more 
than a median paying customer. 

135
00:08:04,040 --> 00:08:07,920
So if we do the maths on that, 
5% of 50 million is 2.5 million,

136
00:08:08,000 --> 00:08:12,800
that's 0.25% of the total user 
base actually using AI in a way 

137
00:08:12,800 --> 00:08:17,840
that could truly reach that 
theoretical expectation of how 

138
00:08:17,840 --> 00:08:20,680
AI will impact an occupation. 
Most people therefore we can 

139
00:08:20,680 --> 00:08:23,920
assume are using it for very 
basic things or in a way that 

140
00:08:23,960 --> 00:08:26,960
isn't going to transform their 
occupation if this were a normal

141
00:08:26,960 --> 00:08:29,760
adoption curve. 
So in technology, when a new 

142
00:08:29,760 --> 00:08:32,520
technology appears, we have we 
see something called an S curve.

143
00:08:32,720 --> 00:08:35,559
So it's a standard model for how
new technologies get adopted. 

144
00:08:35,559 --> 00:08:39,760
It looks like a flattened S on 
its side, which means a slow 

145
00:08:39,760 --> 00:08:42,280
uptake at the start. 
Early adopters are small 

146
00:08:42,280 --> 00:08:44,920
numbers, then a steep 
acceleration in the middle as 

147
00:08:44,920 --> 00:08:47,680
mainstream users pile in, which 
is called the tipping point. 

148
00:08:48,080 --> 00:08:50,960
And then it kind of flattens out
at the top as you hit saturation

149
00:08:51,360 --> 00:08:54,320
and only really the laggards are
left to start using it. 

150
00:08:55,120 --> 00:08:57,760
Well, if we're experiencing that
the middle should be filling by 

151
00:08:57,760 --> 00:09:00,040
now. 
And you could argue there is and

152
00:09:00,040 --> 00:09:03,760
that when you're in the middle 
of it you are very impatient. 

153
00:09:03,920 --> 00:09:06,120
But 3 1/2 years after ChatGPT 
launch and it isn't. 

154
00:09:06,280 --> 00:09:08,440
What we're looking at is close 
to a Cliff edge right now. 

155
00:09:08,440 --> 00:09:11,560
A tiny spike of heavy users 
right at the top and then the 

156
00:09:11,560 --> 00:09:13,520
near vertical drop to everyone 
else. 

157
00:09:13,880 --> 00:09:16,360
So a handful of people have 
obviously figured it out and a 

158
00:09:16,360 --> 00:09:19,080
vast majority of people are just
bounced off or haven't found the

159
00:09:19,080 --> 00:09:21,800
use case or skill set to make it
stick. 

160
00:09:21,840 --> 00:09:24,520
And I'll say this because the 
data seemingly points to it. 

161
00:09:24,560 --> 00:09:26,120
I think probably have a skill 
issue. 

162
00:09:26,240 --> 00:09:28,880
Now, it could be me being 
impatient here. 

163
00:09:29,120 --> 00:09:32,440
And therefore that Cliff will 
become that S curve over time 

164
00:09:32,440 --> 00:09:34,080
and it's just going to take a 
little bit longer. 

165
00:09:34,600 --> 00:09:36,480
We'll have to see. 
But I don't think most people 

166
00:09:36,480 --> 00:09:38,920
have developed the ability to 
use these tools in ways that 

167
00:09:39,120 --> 00:09:40,920
produce seriously meaningful 
results. 

168
00:09:41,640 --> 00:09:45,040
And nobody is teaching them how 
to do this at scale. 

169
00:09:45,560 --> 00:09:46,960
I bet it's not being taught at 
school. 

170
00:09:47,080 --> 00:09:50,320
The second problem is that the 
economy moves far slower than I 

171
00:09:50,320 --> 00:09:51,880
guess the AI industry wants to 
admit. 

172
00:09:52,320 --> 00:09:55,360
So large companies don't RIP out
enterprise systems they're 

173
00:09:55,360 --> 00:09:57,040
running on because something 
better came along. 

174
00:09:57,040 --> 00:10:01,560
They're locked into 1218, two 
year contracts with platforms 

175
00:10:01,560 --> 00:10:03,440
that, you know, meet their needs
quite well. 

176
00:10:03,560 --> 00:10:06,920
Even if those platforms aren't 
AI native and aren't that 

177
00:10:06,920 --> 00:10:10,120
impressive, people will be 
forced into waiting. 

178
00:10:10,280 --> 00:10:12,840
Now that that will last for 
long, you know, by the end of 

179
00:10:12,840 --> 00:10:16,160
that contract, they'll be at 
risk of losing or they'll be 

180
00:10:16,160 --> 00:10:18,960
forced to start integrating 
other companies into that 

181
00:10:18,960 --> 00:10:23,200
solution, which will slowly eat 
and eat and eat away at, I guess

182
00:10:23,240 --> 00:10:25,640
the perceived value of the big 
company. 

183
00:10:25,760 --> 00:10:28,400
So the cycle time for these 
enterprise software deployments 

184
00:10:28,400 --> 00:10:31,760
are measured in years. 
You know, some companies take 

185
00:10:31,760 --> 00:10:34,440
1218 months to deploy your 
technology. 

186
00:10:34,480 --> 00:10:38,760
So therefore there will be a 
slow adoption enterprise level, 

187
00:10:39,200 --> 00:10:42,360
even if they've all adopted 
ChatGPT, there's still 10s of 

188
00:10:42,360 --> 00:10:45,880
thousands of other SAS of 
technology platforms who are 

189
00:10:45,880 --> 00:10:48,160
solving problems and who are not
AI native. 

190
00:10:48,160 --> 00:10:52,640
And you know, only 3% of 
Microsoft 36 fives, 450 million 

191
00:10:52,640 --> 00:10:56,680
users have even adopted or opted
into Copilot yet. 

192
00:10:56,840 --> 00:11:00,200
So again, it will take time. 
The third is that the technology

193
00:11:00,200 --> 00:11:03,120
stuck is actually harder to 
integrate into apps people use 

194
00:11:03,560 --> 00:11:06,160
daily. 
And so I see it as also a supply

195
00:11:06,160 --> 00:11:08,800
problem. 
So what I mean here is that you 

196
00:11:08,800 --> 00:11:10,360
see the supply side every single
day. 

197
00:11:10,720 --> 00:11:14,640
So people might use AI if it was
in front of them in a technology

198
00:11:14,640 --> 00:11:16,560
solution they already use 
everyday, but it's not. 

199
00:11:16,920 --> 00:11:20,120
So they therefore would have to 
learn to be an expert on other 

200
00:11:20,120 --> 00:11:22,880
AI native tools. 
So the claws, the chat, GBTS, 

201
00:11:22,880 --> 00:11:25,280
the many, many others, because 
we work with technology 

202
00:11:25,280 --> 00:11:28,360
companies, SAS companies, 
typically large B to B SAS 

203
00:11:28,360 --> 00:11:30,560
companies deploying agents into 
their products. 

204
00:11:31,160 --> 00:11:33,520
And the complexity is really, 
you know, quite staggering when 

205
00:11:33,520 --> 00:11:37,640
you go from, say, a prototype or
a quick demo, You need an entire

206
00:11:37,640 --> 00:11:41,040
stack to make all of this work. 
You got huge amounts of 

207
00:11:41,040 --> 00:11:43,680
complexity to embed and 
integrate these into solutions. 

208
00:11:44,080 --> 00:11:46,160
You need to identify how it's 
all going to work and even the 

209
00:11:46,560 --> 00:11:48,040
use cases that you want to 
deploy for. 

210
00:11:48,040 --> 00:11:49,960
First, you know, you need an 
agent framework. 

211
00:11:49,960 --> 00:11:52,840
You need reasoning and logic. 
You need to wire this into all 

212
00:11:52,840 --> 00:11:55,360
your systems. 
You need to understand how 

213
00:11:55,360 --> 00:12:00,040
agents understand data and APIs.
You need to build APIs for that.

214
00:12:00,280 --> 00:12:02,280
You need to create MCPS. 
You need to call it, create a 

215
00:12:02,280 --> 00:12:04,760
tool calling layer. 
You need to then have a testing 

216
00:12:04,760 --> 00:12:06,560
layer. 
You need to be able to surface 

217
00:12:06,560 --> 00:12:08,320
the right information and test 
that that's the right 

218
00:12:08,320 --> 00:12:11,280
information because big 
enterprise companies really care

219
00:12:11,280 --> 00:12:13,320
about compliance. 
You need a purpose built 

220
00:12:13,320 --> 00:12:15,520
interface for every type of 
interaction, like widgets 

221
00:12:15,520 --> 00:12:19,400
appearing in charts and charts. 
All this has to then work at 

222
00:12:19,400 --> 00:12:22,440
production scale, not just for, 
you know, 2000 people or 100 

223
00:12:22,440 --> 00:12:24,520
people. 
And I'd say that about 75% of 

224
00:12:24,520 --> 00:12:28,040
the companies we talked to are 
either just getting started or a

225
00:12:28,080 --> 00:12:31,440
little bit into a build. 
And these platforms aren't small

226
00:12:31,440 --> 00:12:35,480
though, you know, companies with
10/15/20 thousand enterprise 

227
00:12:35,480 --> 00:12:38,920
customers worldwide. 
And so without that supply, IE 

228
00:12:38,920 --> 00:12:41,520
the big companies that already 
have millions of users 

229
00:12:41,960 --> 00:12:44,680
integrating this into their 
products in a really effective 

230
00:12:44,680 --> 00:12:47,040
way, it's really hard for most 
people to use AI. 

231
00:12:47,240 --> 00:12:49,080
So the real question therefore 
is will it grow? 

232
00:12:49,160 --> 00:12:51,200
And I'm confident that it will. 
And I think there's a really 

233
00:12:51,200 --> 00:12:54,800
core component to what will make
it grow and that's making the 

234
00:12:54,800 --> 00:12:57,960
technology itself reliable. 
So this comes to the fourth 

235
00:12:57,960 --> 00:13:01,040
reason I think that there's 
still a gap between theoretical 

236
00:13:01,040 --> 00:13:03,880
and real world usage. 
And this is most interesting to 

237
00:13:03,880 --> 00:13:05,880
me because it's what we're 
obsessed, we're trying to solve 

238
00:13:05,880 --> 00:13:07,720
right now. 
It's, you know, something that 

239
00:13:07,720 --> 00:13:09,760
we actively are always in at 
mindset. 

240
00:13:09,880 --> 00:13:11,720
So essentially there was a 
really good paper that 

241
00:13:11,720 --> 00:13:13,680
highlights this point and 
discussion here. 

242
00:13:14,440 --> 00:13:18,480
And it's called Towards a 
Science of AI Agent Reliability.

243
00:13:18,640 --> 00:13:20,640
And it essentially puts rigor 
behind something. 

244
00:13:20,840 --> 00:13:22,640
You know, I've been saying on 
this podcast for a long time 

245
00:13:22,800 --> 00:13:25,040
that the industry is often 
measuring the wrong thing. 

246
00:13:25,640 --> 00:13:29,200
You know, you draw a line 
between capability, So can a 

247
00:13:29,200 --> 00:13:31,600
model do this task and 
reliability? 

248
00:13:31,960 --> 00:13:35,480
Will it do it consistently, 
safely and predictably in 

249
00:13:35,480 --> 00:13:37,320
conditions that are rarely 
perfect? 

250
00:13:37,480 --> 00:13:39,960
And we don't just mean that they
do the right thing most of the 

251
00:13:39,960 --> 00:13:42,280
time. 
We mean something much more. 

252
00:13:42,280 --> 00:13:43,960
And this is what the paper 
breaks down really 

253
00:13:44,040 --> 00:13:46,360
interestingly. 
A colleague and therefore an AI 

254
00:13:46,360 --> 00:13:49,920
agent must be able to #1 be 
consistent. 

255
00:13:50,720 --> 00:13:53,760
This means they get it right 
consistently, and that means not

256
00:13:53,760 --> 00:13:57,760
right today and therefore wrong 
tomorrow on the same task #2 

257
00:13:57,760 --> 00:14:00,400
Robustness. 
They can't just fall apart soon 

258
00:14:00,400 --> 00:14:04,040
as conditions aren't perfect for
them #3 Calibration. 

259
00:14:04,720 --> 00:14:07,400
They're able to tell you when 
the unsure rather than 

260
00:14:07,640 --> 00:14:10,040
confidently guessing and #4 is 
safety. 

261
00:14:10,120 --> 00:14:13,800
So when they do make a mistake, 
their mistakes are fixable, and 

262
00:14:13,800 --> 00:14:15,640
they're not just catastrophic 
failures. 

263
00:14:15,800 --> 00:14:17,720
Again, what you'd expect from a 
colleague. 

264
00:14:17,920 --> 00:14:20,960
And to assess agent reliability 
right now, the authors broke 

265
00:14:20,960 --> 00:14:26,080
down all of this into 12 metrics
across 4 dimensions and tested 

266
00:14:26,080 --> 00:14:31,240
14 different models from Open 
AI, Google, Anthropic, spanning 

267
00:14:31,240 --> 00:14:33,200
a year and a half of product 
releases. 

268
00:14:33,280 --> 00:14:36,120
They run each task five times 
with different sets of 

269
00:14:36,120 --> 00:14:39,480
paraphrase instructions. 
And they also deliberately broke

270
00:14:39,480 --> 00:14:41,760
things. 
So they tried to cause tool 

271
00:14:41,760 --> 00:14:43,640
failures and API calling 
failures. 

272
00:14:44,080 --> 00:14:47,800
They tried to simulate 
environmental faults to see how 

273
00:14:47,800 --> 00:14:50,320
agents handled imperfect 
conditions. 

274
00:14:51,080 --> 00:14:53,560
They asked the models to report 
their own confidence to test 

275
00:14:53,560 --> 00:14:55,160
whether they could tell when 
they're wrong. 

276
00:14:55,760 --> 00:15:00,280
Overall they had 500 total 
benchmark runs and the headline 

277
00:15:00,280 --> 00:15:04,920
finding was that 18 months of 
LLM capability gains produced 

278
00:15:04,920 --> 00:15:07,000
only modest improvements in 
reliability. 

279
00:15:07,160 --> 00:15:10,200
All three major providers 
custard quite closely together, 

280
00:15:10,440 --> 00:15:12,960
suggesting that this is an 
industry wide challenge, not a 

281
00:15:12,960 --> 00:15:15,760
company specific one. 
And for me, I don't think this 

282
00:15:15,760 --> 00:15:18,280
is solved at the model layer, 
it's sold at the application 

283
00:15:18,280 --> 00:15:20,000
layer. 
So for example, on consistency, 

284
00:15:20,000 --> 00:15:22,840
whether the model gave the same 
answer when asked the same thing

285
00:15:22,840 --> 00:15:26,120
twice. 
Scores range from 30 to 75% on 

286
00:15:26,120 --> 00:15:28,960
robustness. 
So whether it can deal with 

287
00:15:28,960 --> 00:15:32,280
imperfect conditions. 
Models handled infrastructure 

288
00:15:32,280 --> 00:15:35,680
failures like big server crashes
quite well, but apart from that 

289
00:15:35,760 --> 00:15:38,840
didn't really do that well on 
predictability, whether the 

290
00:15:38,840 --> 00:15:41,440
model knew it's wrong. 
Performance was also very mixed 

291
00:15:41,600 --> 00:15:44,600
and this speaks to a challenge 
within implementing AI right 

292
00:15:44,600 --> 00:15:46,080
now. 
I think that, you know, mindset 

293
00:15:46,080 --> 00:15:48,800
sees all the time it's in that 
we're actively working through. 

294
00:15:48,960 --> 00:15:51,480
And as I said, it doesn't get 
solved at the model layer. 

295
00:15:51,480 --> 00:15:52,880
It gets solved at the 
application layer. 

296
00:15:53,200 --> 00:15:56,160
And again, models are trying to 
bring in more application 

297
00:15:56,200 --> 00:15:57,960
technologies. 
So they're trying to turn their 

298
00:15:57,960 --> 00:16:00,400
model into an agent itself in 
the way it works. 

299
00:16:00,560 --> 00:16:03,240
And this also is why SAS 
companies have such a big 

300
00:16:03,240 --> 00:16:05,760
opportunity really to create 
perfect agents for their 

301
00:16:05,760 --> 00:16:10,240
specific niche because a general
purpose, Claude, although is 

302
00:16:10,280 --> 00:16:14,480
absolutely incredible and can do
many things, if you make a hyper

303
00:16:14,480 --> 00:16:18,280
specific agent for a specific 
area that you're an expert in 

304
00:16:18,280 --> 00:16:21,000
can be really, really powerful 
and more consistent, more 

305
00:16:21,000 --> 00:16:22,480
robust. 
And the reason that we have to 

306
00:16:22,480 --> 00:16:24,680
deal with this is essentially 
what you're asking is do we want

307
00:16:24,680 --> 00:16:29,040
an agent to have deterministic 
controls or have high inference?

308
00:16:29,040 --> 00:16:31,720
So essentially, can an agent 
decide and reason what to do 

309
00:16:31,720 --> 00:16:37,160
high inference, or do we, as say
the agent creator, give it 

310
00:16:37,160 --> 00:16:41,320
specific instructions that it's 
not really allowed to stray 

311
00:16:41,320 --> 00:16:43,160
from? 
So if you give that agent a step

312
00:16:43,160 --> 00:16:45,680
by step plan to follow, that's 
all it can ever do. 

313
00:16:45,680 --> 00:16:47,640
So if a person comes in and 
wants to go in any other 

314
00:16:47,640 --> 00:16:49,600
direction, it's not going to be 
able to do it. 

315
00:16:50,400 --> 00:16:53,640
Or if you let it reason itself 
and decide, it might go off in 

316
00:16:53,640 --> 00:16:55,720
different direction to solve the
exact same problem. 

317
00:16:55,840 --> 00:16:59,800
Because ultimately an agent, all
it's doing is observing, 

318
00:17:00,240 --> 00:17:02,680
planning, acting, checking, 
adjusting. 

319
00:17:03,040 --> 00:17:06,400
So it makes a plan, it sees how 
that plan goes, sees what data 

320
00:17:06,400 --> 00:17:08,880
and information it gets from 
that plan, checks whether that's

321
00:17:08,880 --> 00:17:11,040
the right thing, and if it 
isn't, adjust. 

322
00:17:11,240 --> 00:17:13,960
I'll check in with the humor so 
that that's what enables agents 

323
00:17:13,960 --> 00:17:16,599
to do really incredible things 
that blow your mind. 

324
00:17:16,720 --> 00:17:19,440
And it will deal with imperfect 
conditions really well. 

325
00:17:19,599 --> 00:17:21,520
It can tell you when it needs 
help. 

326
00:17:22,400 --> 00:17:24,920
So it can act more like a 
person, but it doesn't mean it's

327
00:17:24,920 --> 00:17:26,440
reliable. 
There's loads of different 

328
00:17:26,440 --> 00:17:28,520
things that can go wrong when 
it's reasoning. 

329
00:17:28,520 --> 00:17:31,520
So this is where that trade off 
between consistency and 

330
00:17:31,520 --> 00:17:34,920
flexibility becomes really 
important to rigid a set of 

331
00:17:34,920 --> 00:17:37,840
instructions. 
It can't help people in any ways

332
00:17:37,880 --> 00:17:39,240
other than that step by step 
plan. 

333
00:17:39,960 --> 00:17:42,800
Too much inference means that 
it's likely to be inconsistent. 

334
00:17:42,920 --> 00:17:44,960
And so there's lots of ways to 
try and solve this mindset, 

335
00:17:44,960 --> 00:17:46,360
solving this in really 
interesting ways. 

336
00:17:46,360 --> 00:17:48,560
We're we're creating something 
called blueprints, which is more

337
00:17:48,560 --> 00:17:51,560
like giving an agent a set of 
DNA, which is a set of 

338
00:17:51,560 --> 00:17:54,280
instructions that can last over 
many months. 

339
00:17:54,400 --> 00:17:57,480
So rather than just, you know, a
quick process, you can give it 

340
00:17:57,480 --> 00:17:59,800
the ability to solve, say 50 
problems with what we call 

341
00:17:59,800 --> 00:18:02,280
tools. 
Now a tool might be, you know, a

342
00:18:02,280 --> 00:18:05,000
integration to a separate system
to help a user do something. 

343
00:18:05,200 --> 00:18:07,360
It could be a small set of 
instructions that the agent can 

344
00:18:07,360 --> 00:18:11,880
follow, but a, the blueprint is 
essentially its DNA over many, 

345
00:18:11,880 --> 00:18:13,840
many, many months, like a long 
term objective. 

346
00:18:13,840 --> 00:18:17,520
It might be the HR recruitment, 
an entire employee life cycle of

347
00:18:17,520 --> 00:18:22,360
how it expects to help people, 
or a product discovery journey 

348
00:18:22,360 --> 00:18:25,520
that might take six months. 
So it describes a set of things 

349
00:18:25,520 --> 00:18:28,040
an agent must take people 
through over a long period of 

350
00:18:28,040 --> 00:18:29,440
time. 
So yeah, this is a really big 

351
00:18:29,440 --> 00:18:32,000
area of investment by all these 
big companies. 

352
00:18:32,720 --> 00:18:35,760
It's really, really important to
solve and all of it does take 

353
00:18:35,760 --> 00:18:37,560
time and therefore it takes time
for adoption. 

354
00:18:37,840 --> 00:18:42,600
And we've seen this before. 
In 1987, there was a writer that

355
00:18:42,600 --> 00:18:45,440
said, see the computer age 
everywhere but in the 

356
00:18:45,440 --> 00:18:48,440
productivity statistics. 
And in February, Apollo's chief 

357
00:18:48,440 --> 00:18:50,840
economist said the exact same 
thing about AI. 

358
00:18:50,960 --> 00:18:53,960
So yeah, 15 years into the 
computer revolution, 

359
00:18:54,160 --> 00:18:57,800
productivity actually slowed, 
especially during the IT build 

360
00:18:57,800 --> 00:18:59,680
out. 
It dropped from 2.9% annual 

361
00:18:59,680 --> 00:19:05,400
growth in productivity to 1.1%. 
But then between 1995 and 2005, 

362
00:19:05,400 --> 00:19:07,280
it surged. 
The technology hadn't 

363
00:19:07,280 --> 00:19:11,840
necessarily improve that much. 
The change was the organizations

364
00:19:11,840 --> 00:19:13,600
finally rebuilt how they worked 
with it. 

365
00:19:13,840 --> 00:19:16,440
The technology was super 
consistent, it was accessible, 

366
00:19:17,040 --> 00:19:18,680
and it took a decade and a half 
to do. 

367
00:19:36,990 --> 00:19:38,910
Let's conclude. 
The dominant narrative I think 

368
00:19:38,910 --> 00:19:41,630
says that there's a massive 
market disruption that's 

369
00:19:41,630 --> 00:19:42,870
imminent. 
And I think we're starting to 

370
00:19:42,870 --> 00:19:45,110
see that. 
You know, Dario Amadei, the CEO 

371
00:19:45,110 --> 00:19:47,230
Anthropic, has said this over 
and over again. 

372
00:19:47,310 --> 00:19:50,070
But clearly the story, I think 
this data support is a little 

373
00:19:50,070 --> 00:19:52,720
bit more complicated. 
AI's economic impact is 

374
00:19:52,720 --> 00:19:56,240
definitely real, but held back 
by reliability limits that the 

375
00:19:56,240 --> 00:20:00,080
industry hasn't yet fully solved
and organizational human 

376
00:20:00,080 --> 00:20:03,560
barriers that are just not easy 
to solve because theoretical 

377
00:20:03,560 --> 00:20:06,320
baselines are always more 
aspirational than measurement. 

378
00:20:07,000 --> 00:20:11,560
So yeah, Will AI transform work?
Yes, Computers did reshape the 

379
00:20:11,560 --> 00:20:14,400
economy, but on their own 
schedule, not on what the 

380
00:20:14,400 --> 00:20:17,040
industry wanted it to be. 
And because everything 

381
00:20:17,040 --> 00:20:19,200
surrounded, the technology had 
to mature with it. 

382
00:20:19,800 --> 00:20:21,720
And I think the same thing is 
happening right now. 

383
00:20:22,440 --> 00:20:26,200
And this type of reliability 
research is going to force us 

384
00:20:26,240 --> 00:20:28,960
all in the industry to measure 
what matters. 

385
00:20:29,080 --> 00:20:31,160
So anyway, I hope you enjoyed 
the episode. 

386
00:20:31,160 --> 00:20:34,200
I hope you learned something. 
Thank you for listening and I'll

387
00:20:34,200 --> 00:20:35,360
see you next week.
