1
00:00:00,080 --> 00:00:03,920
So on Wednesday, the US Treasury
Secretary and the chair of the 

2
00:00:03,920 --> 00:00:07,120
Federal Reserve called an 
emergency meeting with CE OS of 

3
00:00:07,120 --> 00:00:10,400
America's biggest banks. 
The subject wasn't interest 

4
00:00:10,400 --> 00:00:13,520
rates or inflation or, you know,
pending World War Three. 

5
00:00:13,520 --> 00:00:16,720
It was an AI model. 
Anthropic has now built 

6
00:00:16,720 --> 00:00:20,120
something that could find and 
exploit security flaws in 

7
00:00:20,120 --> 00:00:23,080
virtually any software, it was 
pointed out, and the financial 

8
00:00:23,080 --> 00:00:26,200
system wanted to know how 
exposed it really was. 

9
00:00:26,800 --> 00:00:29,000
Anthropic have decided not to 
sell this. 

10
00:00:29,400 --> 00:00:32,080
Instead they've called all the 
biggest companies from Apple, 

11
00:00:32,080 --> 00:00:35,880
Amazon, Google, banks, 
governments to form a coalition 

12
00:00:35,880 --> 00:00:38,680
called Project Glass Way. 
And they've kept the model 

13
00:00:38,680 --> 00:00:41,840
behind closed doors, only giving
access to those companies. 

14
00:00:42,640 --> 00:00:45,720
Today I want to walk through 
what this model actually does 

15
00:00:45,720 --> 00:00:49,120
and what it means for us in 
practice, Why the sceptics have 

16
00:00:49,120 --> 00:00:53,360
some really interesting points 
here, and why the way this was 

17
00:00:53,360 --> 00:00:58,320
all handled tells us more about 
but where AI governance is truly

18
00:00:58,320 --> 00:01:00,440
at and what risks we really 
face. 

19
00:01:01,640 --> 00:01:03,160
This is In the Loop with Jack 
Horton. 

20
00:01:03,680 --> 00:01:18,930
Hope you enjoy the show. 
So let's start with what Mythos 

21
00:01:18,970 --> 00:01:21,730
actually did. 
So Mythos is Anthropic's newest 

22
00:01:21,770 --> 00:01:24,210
AI model. 
And what makes it quite unusual 

23
00:01:24,210 --> 00:01:27,480
is that it turned out to be 
extremely good at, you know, one

24
00:01:27,480 --> 00:01:30,480
thing, you know, instead of 
being the best in the world at 

25
00:01:30,480 --> 00:01:32,520
writing, and maybe it was really
good at writing. 

26
00:01:32,880 --> 00:01:36,200
It turned out it was really good
at finding security flaws in 

27
00:01:36,200 --> 00:01:40,680
software, not small business, 
not small theoretical problems 

28
00:01:40,680 --> 00:01:43,800
that real flaws in real 
applications that are out there 

29
00:01:43,800 --> 00:01:46,080
right now and used by basically 
everyone in the world. 

30
00:01:46,560 --> 00:01:51,240
And these bugs have been missed 
by human experts for years. 

31
00:01:51,640 --> 00:01:54,600
A little bit of context here. 
Every bit of software we use, 

32
00:01:54,600 --> 00:01:57,800
you know, our browsers, our 
phones, the servers that handle 

33
00:01:57,800 --> 00:02:00,160
online banking, there's just 
loads of bugs in it. 

34
00:02:00,160 --> 00:02:04,040
And most of these are absolutely
harmless, but some are what 

35
00:02:04,160 --> 00:02:07,320
security researcher would 
describe as zero days. 

36
00:02:07,800 --> 00:02:12,240
So flaws that nobody knew about 
but could let an attacker take 

37
00:02:12,240 --> 00:02:14,440
control of an entire local 
machine. 

38
00:02:14,440 --> 00:02:16,960
So you take control of your 
laptop just because you clicked 

39
00:02:16,960 --> 00:02:19,560
on a website. 
If you heard about a company or 

40
00:02:19,560 --> 00:02:22,240
a hospital or a government being
hacked, it's usually because 

41
00:02:22,720 --> 00:02:26,320
someone found and exploited one 
of these zero day problems. 

42
00:02:26,840 --> 00:02:29,840
And you know, only a small group
of elite researchers typically 

43
00:02:29,840 --> 00:02:33,440
take on the job of finding 
these, whereas Mythos has just 

44
00:02:33,440 --> 00:02:36,160
found them in minutes, you know,
across many different systems 

45
00:02:36,160 --> 00:02:39,320
all at once, and then without 
any human guidance or prompting,

46
00:02:39,600 --> 00:02:41,880
worked out exactly how to 
exploit them. 

47
00:02:41,880 --> 00:02:45,680
And there are a few examples to 
talk through because, you know, 

48
00:02:45,680 --> 00:02:48,840
it really tells us what this 
would mean in practice at scale.

49
00:02:49,240 --> 00:02:52,160
You know, it found a bug in Free
BSD. 

50
00:02:52,160 --> 00:02:55,040
So it's an operating system that
basically runs a lot of Internet

51
00:02:55,040 --> 00:02:58,280
infrastructure and servers, 
networking equipment, basically 

52
00:02:58,280 --> 00:03:01,560
the really boring backbone stuff
that most people just never see.

53
00:03:02,040 --> 00:03:04,320
And that bug has been in the 
code since 2009. 

54
00:03:04,640 --> 00:03:08,080
Ultimate Testing Tools had run 
through it over 5,000,000 times 

55
00:03:08,200 --> 00:03:11,200
and it never caught this bug, 
whereas Mythos found it and 

56
00:03:11,200 --> 00:03:14,800
actually built a multi step 
attack across different network 

57
00:03:14,800 --> 00:03:18,280
requests that were given an 
attacker full control of 

58
00:03:18,280 --> 00:03:20,360
everything. 
So in practical terms, that 

59
00:03:20,360 --> 00:03:22,800
means anyone that could reach 
their server could read every 

60
00:03:22,800 --> 00:03:26,080
file that exists on it, install 
whatever they wanted, and use it

61
00:03:26,080 --> 00:03:29,280
as basically a launchpad to 
attack any of the machines on 

62
00:03:29,280 --> 00:03:32,720
that same network. 
He also found a 27 year old 

63
00:03:32,720 --> 00:03:36,200
networking bugging Open BSD, 
which is another operating 

64
00:03:36,200 --> 00:03:39,040
system known for being one of 
the most security countries in 

65
00:03:39,040 --> 00:03:43,040
the world, and the the bug had 
been there since 1998. 

66
00:03:43,440 --> 00:03:46,520
It found vulnerabilities in 
every major web browser and it 

67
00:03:46,520 --> 00:03:49,880
found a chain of four separate 
bugs and flaws that if you 

68
00:03:49,880 --> 00:03:52,960
string it together and attack, 
meant that if you visited a 

69
00:03:52,960 --> 00:03:56,200
particular website and then 
Tucker could literally take 

70
00:03:56,200 --> 00:03:59,920
control of the browser, the 
security, it can break out of 

71
00:03:59,920 --> 00:04:02,440
the operating security and take 
control of your laptop. 

72
00:04:02,440 --> 00:04:05,000
You wouldn't need to download 
anything or click anything. 

73
00:04:05,280 --> 00:04:07,600
It would literally take control 
of your PC. 

74
00:04:07,920 --> 00:04:11,120
Mythos has found thousands of 
these zero day vulnerabilities 

75
00:04:11,120 --> 00:04:13,640
and few of them, 1% have 
actually been patched so far. 

76
00:04:14,320 --> 00:04:18,160
And Anthropic have published 
cryptographic proof that these 

77
00:04:18,160 --> 00:04:21,519
discoveries exist by basically 
tamper proof time stamps showing

78
00:04:21,519 --> 00:04:23,720
that these things really do 
exist. 

79
00:04:23,720 --> 00:04:26,280
And they found these flaws 
before anyone else did. 

80
00:04:26,880 --> 00:04:28,200
And you know, none of this was 
the goal. 

81
00:04:28,200 --> 00:04:30,640
You know, Anthropic actually set
out to build the best coding 

82
00:04:30,640 --> 00:04:33,640
model in the world. 
And it turns out that also makes

83
00:04:33,640 --> 00:04:35,480
it really good at attacking 
software. 

84
00:04:51,920 --> 00:04:54,120
Now it's worth having a 
conversation about the skeptics 

85
00:04:54,120 --> 00:04:57,360
because there are a lot of 
skeptics out there and write 

86
00:04:57,360 --> 00:04:59,880
this out. 
For example, the headline number

87
00:04:59,880 --> 00:05:04,200
that Anthropic have given out is
thousands of 0 day flaws and 

88
00:05:04,200 --> 00:05:09,000
books and Anthropic as Human 
reviewers check 190 of these 

89
00:05:09,000 --> 00:05:12,760
reports and reviewers agreed 
with the models assessment about

90
00:05:12,760 --> 00:05:15,840
89% of the time, but that means 
there's still thousands that are

91
00:05:15,840 --> 00:05:16,960
unchecked. 
So there's a lot of 

92
00:05:16,960 --> 00:05:21,040
extrapolation there to say that 
there are thousands of 0 day 

93
00:05:21,720 --> 00:05:24,360
bugs and flaws. 
Another example is the security 

94
00:05:24,360 --> 00:05:26,440
research. 
Different labs for example, have

95
00:05:26,440 --> 00:05:29,760
tested some of the 
vulnerabilities on cheaper 

96
00:05:29,760 --> 00:05:33,680
smaller models and they ran 8 
experiments and found eight of 

97
00:05:33,680 --> 00:05:38,720
the problems. 
So it's 100% matching Mythos. 

98
00:05:39,160 --> 00:05:41,960
Now there's there's, again, lots
of nuance and extrapolation that

99
00:05:41,960 --> 00:05:45,120
that was only 8 problems. 
And it was also focused on the 

100
00:05:45,120 --> 00:05:47,000
things that Mythos had already 
found. 

101
00:05:47,000 --> 00:05:50,440
And on top of that, what those 
models definitely didn't do is 

102
00:05:50,440 --> 00:05:53,920
then build an entire attack plan
by itself. 

103
00:05:53,920 --> 00:05:57,120
And another thing that I think 
is most compelling for me is 

104
00:05:57,120 --> 00:05:58,920
that there's just this timing 
aspect. 

105
00:05:58,920 --> 00:06:02,120
So Anthropic is about to go 
towards a stock market listing 

106
00:06:02,120 --> 00:06:05,520
and it'll be valued at about 
$380 billion, apparently, you 

107
00:06:05,520 --> 00:06:08,280
know, in a model that's too 
dangerous to release is 

108
00:06:08,280 --> 00:06:11,720
brilliant PR in the world for 
everyone getting excited and 

109
00:06:11,720 --> 00:06:14,400
putting lots of money in them. 
And there are many arguments out

110
00:06:14,400 --> 00:06:17,800
there saying that Anthropic are 
essentially withholding the best

111
00:06:17,800 --> 00:06:22,920
models so they can charge big 
enterprise huge amounts of money

112
00:06:23,320 --> 00:06:26,200
and essentially keep everyone 
else a generation behind. 

113
00:06:26,200 --> 00:06:29,200
And, and there's good reasoning 
for that, you know, those 

114
00:06:29,200 --> 00:06:33,640
listening, most likely me, our 
team, we're power users and 

115
00:06:33,640 --> 00:06:35,760
power users cost model providers
money. 

116
00:06:36,320 --> 00:06:38,360
You know, we might be paying. 
You know, I've got a couple of 

117
00:06:38,360 --> 00:06:43,320
accounts at $200 per account. 
We definitely use more than $200

118
00:06:43,320 --> 00:06:47,120
worth of credits every month. 
Again, intelligence isn't cheap,

119
00:06:47,680 --> 00:06:50,560
but you know, the long tail of 
you, let's say hundreds of 

120
00:06:50,560 --> 00:06:54,400
millions of users out there that
use it, fraction of the amount 

121
00:06:54,400 --> 00:06:57,640
that maybe I do, That's where 
they make their money. 

122
00:06:57,640 --> 00:07:01,000
So then suddenly finding a 
really good model, charging 

123
00:07:01,000 --> 00:07:03,520
absolute fortunes to big 
enterprise, who wants a 

124
00:07:03,520 --> 00:07:06,760
competitive advantage like that 
and will it be able to pay those

125
00:07:06,960 --> 00:07:09,560
huge sums of money is a very 
clever strategy. 

126
00:07:10,160 --> 00:07:12,560
And there's also an aspect which
is that we have been here 

127
00:07:12,560 --> 00:07:14,960
before. 
You know, in 2019, Open AI 

128
00:07:14,960 --> 00:07:18,080
announced a model called GPT 2, 
and they refused to release it. 

129
00:07:18,440 --> 00:07:19,880
You know, they said it was too 
dangerous. 

130
00:07:20,440 --> 00:07:22,600
It could write so convincingly 
that it's going to fuel 

131
00:07:22,600 --> 00:07:26,480
disinformation and propaganda. 
And so they staged it out over 

132
00:07:26,480 --> 00:07:29,160
nine months, releasing 
progressively larger versions. 

133
00:07:29,160 --> 00:07:31,840
And, you know, by the time the 
model had come out, they had big

134
00:07:31,840 --> 00:07:35,000
test groups and nobody found any
risk in the end. 

135
00:07:35,920 --> 00:07:39,680
But obviously, it was brilliant 
PR for them to raise more money 

136
00:07:39,680 --> 00:07:42,960
and attract more exciting people
to come work for them, come 

137
00:07:42,960 --> 00:07:45,640
invest in them. 
Though what I would add on to 

138
00:07:45,640 --> 00:07:47,400
that is a little bit of a 
nuance, which is, you know, a 

139
00:07:47,560 --> 00:07:49,560
mythos is a bit different. 
You know, Anthropic isn't 

140
00:07:49,560 --> 00:07:52,840
withholding anything. 
They're in fact literally 

141
00:07:52,840 --> 00:07:57,000
publishing all the evidence, all
the reports, and allowing 

142
00:07:57,000 --> 00:08:00,360
anybody to test them themselves.
And arguably you can't really 

143
00:08:00,360 --> 00:08:03,960
have a staged release process 
when the model is able to just 

144
00:08:03,960 --> 00:08:07,320
attack software at will and take
down the entire Internet. 

145
00:08:07,320 --> 00:08:09,400
It's a bit too dangerous to go 
and risk it. 

146
00:08:25,840 --> 00:08:27,320
So, yeah, is this real or is 
this marketing? 

147
00:08:27,320 --> 00:08:31,400
And it's a good question. 
I think it's both and I think 

148
00:08:31,400 --> 00:08:33,760
it's OK to have both. 
You know, two things can be true

149
00:08:33,760 --> 00:08:37,039
at once. 
It probably is quite dangerous 

150
00:08:37,039 --> 00:08:39,880
or they're worried that it's 
dangerous and it's a wonderful 

151
00:08:39,880 --> 00:08:42,159
PR opportunity. 
But personally, I think the 

152
00:08:42,159 --> 00:08:45,800
bigger and more important 
question to ask of instead of is

153
00:08:45,800 --> 00:08:48,600
this PR marketing? 
Because, you know, of all the 

154
00:08:48,600 --> 00:08:51,040
companies, I think Anthropic is 
probably the most ethical. 

155
00:08:51,040 --> 00:08:53,760
You know, we covered the episode
where they fought back against 

156
00:08:53,760 --> 00:08:55,640
the US government and the 
Department of War. 

157
00:08:56,240 --> 00:08:59,240
But the bigger question really 
is who decides that something is

158
00:08:59,240 --> 00:09:02,240
or is not safe? 
At what point does this become 

159
00:09:02,240 --> 00:09:05,280
safe enough to release now? 
So like Anthropic is a good 

160
00:09:05,280 --> 00:09:07,840
example, looked at what they 
built, concludes that gosh, this

161
00:09:07,840 --> 00:09:11,320
was too dangerous for the public
and decided which specific 

162
00:09:11,320 --> 00:09:15,560
companies could get to use this.
And they called it Project Glass

163
00:09:15,560 --> 00:09:18,360
Wing. 
You know, Amazon, Apple, Google,

164
00:09:18,360 --> 00:09:20,960
Microsoft, 40 other 
organizations get access. 

165
00:09:21,640 --> 00:09:25,480
And yes, Anthropica committed 
$100 million in computer credits

166
00:09:25,480 --> 00:09:27,160
to fine and patch any 
vulnerabilities. 

167
00:09:27,480 --> 00:09:30,400
But essentially they are still a
private company making a private

168
00:09:30,400 --> 00:09:33,760
assessment from the private 
coalition, arranged a private 

169
00:09:33,760 --> 00:09:37,040
solution with no, say, treaty 
inspectors, independent 

170
00:09:37,680 --> 00:09:41,520
oversight, no democratic input. 
Yet when it comes to nuclear 

171
00:09:41,640 --> 00:09:44,800
weapons or industries, you have 
international oversight, 

172
00:09:45,360 --> 00:09:48,440
independent governance bodies 
who are deciding who might get 

173
00:09:48,440 --> 00:09:51,680
access, whether they're actually
breaking the rules, whether 

174
00:09:51,680 --> 00:09:53,920
they're being safe. 
Now, I'm not trying to be all 

175
00:09:53,920 --> 00:09:58,280
too European on this in terms of
trying to create huge regulation

176
00:09:58,280 --> 00:10:00,320
that makes it really difficult 
for these organizations to 

177
00:10:00,320 --> 00:10:02,960
operate because that's 
definitely not me. 

178
00:10:03,360 --> 00:10:06,720
However, I am often quite 
balanced when it comes to 

179
00:10:07,080 --> 00:10:12,520
wanting to pair risk and reward.
So maybe they are the right 

180
00:10:12,520 --> 00:10:15,880
organization to take the right 
steps and make the right 

181
00:10:15,880 --> 00:10:19,280
choices, but equally at all 
organizations, the right 

182
00:10:19,280 --> 00:10:21,520
organization. 
I don't trust open AI as far as 

183
00:10:21,520 --> 00:10:23,720
I could throw them. 
See, I think this tells us where

184
00:10:23,720 --> 00:10:28,360
we are in the maturity of AI and
governance generally, that we 

185
00:10:28,360 --> 00:10:31,800
still rely on private companies 
and we still haven't got the 

186
00:10:31,800 --> 00:10:35,840
systems in place which Anthropic
of openly acknowledged to keep 

187
00:10:35,840 --> 00:10:38,640
society safe as these things are
rolled out. 

188
00:10:54,460 --> 00:10:56,620
Anyway, what I think to 
conclude, what sticks for me 

189
00:10:56,620 --> 00:10:59,180
isn't where the mythos is, you 
know, super powerful and, and I 

190
00:10:59,180 --> 00:11:01,060
actually really hope it is 
because it's quite exciting. 

191
00:11:01,060 --> 00:11:03,100
It's going to help me at work, 
it's going to help me at life. 

192
00:11:04,020 --> 00:11:06,500
And to be honest, it probably 
is, or at least close to it. 

193
00:11:06,940 --> 00:11:09,620
But what sticks with me is that 
this is the first time the AI 

194
00:11:09,620 --> 00:11:13,820
industry as a whole has agreed, 
really, that this capabilities 

195
00:11:13,820 --> 00:11:17,720
are too dangerous to release 
based on real results rather 

196
00:11:17,720 --> 00:11:19,160
than, you know, theoretical 
risk. 

197
00:11:19,160 --> 00:11:22,040
And the best way we could manage
this risk was obviously a group 

198
00:11:22,040 --> 00:11:25,400
of private companies deciding 
amongst themselves what to do 

199
00:11:25,400 --> 00:11:27,800
about it. 
So yeah, I think the question on

200
00:11:27,800 --> 00:11:31,680
my mind right now is who decides
what happens next or when the 

201
00:11:31,680 --> 00:11:36,120
next one appears, or on whose 
authority and accountability do 

202
00:11:36,120 --> 00:11:40,080
these things get released? 
Anyway, thanks for listening and

203
00:11:40,360 --> 00:11:41,960
I'll see you next week.
