1
00:00:00,040 --> 00:00:05,160
Imagine you run a massive 
Fortune 500 company and you just

2
00:00:05,160 --> 00:00:08,320
fired your entire software 
development team, not to save 

3
00:00:08,320 --> 00:00:13,040
money, but because starting 
today, every single employee in 

4
00:00:13,040 --> 00:00:16,280
your company is generating their
own bespoke software 

5
00:00:16,280 --> 00:00:18,640
applications just on the fly. 
Yeah. 

6
00:00:18,640 --> 00:00:20,720
Which sounds like absolute 
chaos. 

7
00:00:20,800 --> 00:00:22,720
Exactly. 
A total management nightmare. 

8
00:00:22,960 --> 00:00:25,320
But I mean, looking at the 
developments from just this past

9
00:00:25,320 --> 00:00:28,240
week, the infrastructure to make
this happen isn't science 

10
00:00:28,240 --> 00:00:30,200
fiction anymore. 
It actually exists. 

11
00:00:30,200 --> 00:00:32,360
It does. 
It is a profound structural 

12
00:00:32,360 --> 00:00:36,160
change in computing. 
Today is Friday, April 17th, 

13
00:00:36,200 --> 00:00:39,920
2026, and we are tracking a 
transition that fundamentally 

14
00:00:39,920 --> 00:00:41,680
redefines your relationship with
software. 

15
00:00:42,160 --> 00:00:45,880
The data we're analyzing today 
from the new Stack Latent Space 

16
00:00:45,880 --> 00:00:48,920
Import AI, along with some 
really fresh survey data from 

17
00:00:48,920 --> 00:00:51,760
Stack Overflow, it all points to
the exact same conclusion. 

18
00:00:51,760 --> 00:00:54,040
It's just that we're moving away
from the era of the AI 

19
00:00:54,040 --> 00:00:56,320
assistant, right? 
Precisely, we are entering the 

20
00:00:56,320 --> 00:00:59,960
era of the autonomous AI staff. 
For you, the listener, the 

21
00:00:59,960 --> 00:01:02,880
fundamental shift here is moving
your role from being a human in 

22
00:01:02,880 --> 00:01:04,760
the loop to becoming a human on 
the loop. 

23
00:01:05,080 --> 00:01:06,440
OK, let's unpack this for a 
second. 

24
00:01:06,800 --> 00:01:10,640
Because the human in the loop, 
that's when you ask an AI to 

25
00:01:10,640 --> 00:01:13,680
say, write a specific recipe, 
you check it's work, and then 

26
00:01:13,680 --> 00:01:16,240
you go cook the meal yourself. 
Right, you are micromanaging the

27
00:01:16,240 --> 00:01:17,160
process. 
Exactly. 

28
00:01:17,560 --> 00:01:21,320
But being on the loop, it feels 
like we're moving to having an 

29
00:01:21,480 --> 00:01:24,280
AI staff completely manage your 
kitchen. 

30
00:01:24,360 --> 00:01:26,400
I mean, they are monitoring the 
inventory, ordering the 

31
00:01:26,400 --> 00:01:29,360
groceries, prepping the 
stations, doing the dishes, and 

32
00:01:29,360 --> 00:01:32,720
you are just overseeing the menu
and tasting the final product. 

33
00:01:33,120 --> 00:01:36,120
But to actually pull that off, 
to let software run your kitchen

34
00:01:36,120 --> 00:01:38,920
autonomously, you need insanely 
capable models. 

35
00:01:39,640 --> 00:01:41,560
Which brings us to yesterday's 
massive launch. 

36
00:01:41,560 --> 00:01:45,080
Oh yes, anthropic dropping 
clawed Opus 4.7. 

37
00:01:45,080 --> 00:01:48,200
Yeah, yesterday, April 16th. 
And the timing of that release 

38
00:01:48,200 --> 00:01:50,320
is really the catalyst for this 
entire discussion. 

39
00:01:50,400 --> 00:01:52,960
Late in Space published a 
breakdown almost immediately, I 

40
00:01:52,960 --> 00:01:55,080
think the headline called Opus 
4.7. 

41
00:01:55,080 --> 00:01:57,320
Literally one step better in 
every dimension. 

42
00:01:57,320 --> 00:01:58,960
Wow. 
Yeah, they highlighted the 

43
00:01:58,960 --> 00:02:02,400
significantly upgraded vision 
processing, the expanded state, 

44
00:02:02,400 --> 00:02:05,880
full memory and just a drastic 
improvement in complex 

45
00:02:05,880 --> 00:02:08,440
instruction following over 
version 4.6. 

46
00:02:08,720 --> 00:02:11,200
And developers, I mean they 
didn't even wait for the dust to

47
00:02:11,200 --> 00:02:13,880
settle. 
Simon Willison pushed an update 

48
00:02:13,880 --> 00:02:18,520
to his philanthropic plug in 
version Viro .25 on the exact 

49
00:02:18,520 --> 00:02:20,760
same day. 
He moves incredibly fast. 

50
00:02:20,800 --> 00:02:22,920
He does. 
And he didn't just, you know, 

51
00:02:23,080 --> 00:02:26,400
swap the API endpoint. 
He unlocked the new reasoning 

52
00:02:26,400 --> 00:02:28,960
parameters, specifically that 
thinking effort setting. 

53
00:02:29,320 --> 00:02:31,680
He allowed users to Max it out 
at XI. 

54
00:02:32,960 --> 00:02:35,520
Yeah, plus he added these 
boolean toggles for thinking 

55
00:02:35,520 --> 00:02:38,880
Display and thinking adaptive. 
He even increased the default 

56
00:02:38,880 --> 00:02:40,760
Max tokens to the absolute 
maximum. 

57
00:02:40,760 --> 00:02:43,000
That thinking adaptive 
parameter, though, That is 

58
00:02:43,000 --> 00:02:45,680
particularly revealing about 
where the technology is heading 

59
00:02:45,680 --> 00:02:46,880
right now. 
How so? 

60
00:02:46,960 --> 00:02:50,560
Well, we are no longer forcing a
model to use a static amount of 

61
00:02:50,560 --> 00:02:53,360
compute for every query. 
When you toggle Thinking 

62
00:02:53,360 --> 00:02:56,640
Adaptive, you're allowing the 
model to evaluate the perplexity

63
00:02:56,640 --> 00:02:58,640
of your prompt in real time. 
Oh, I see. 

64
00:02:58,640 --> 00:03:00,160
So it judges how hard the 
question is. 

65
00:03:00,200 --> 00:03:03,000
Exactly. 
If it hits a comp complex high 

66
00:03:03,000 --> 00:03:07,280
entropy logic puzzle, it 
dynamically allocates a massive 

67
00:03:07,280 --> 00:03:10,440
budget of hidden reasoning 
tokens to literally think 

68
00:03:10,440 --> 00:03:13,560
through the problem before 
generating a single word of 

69
00:03:13,560 --> 00:03:15,760
output. 
But if the prompt is simple, it 

70
00:03:15,760 --> 00:03:18,920
just bypasses that layer 
entirely to save you latency and

71
00:03:18,920 --> 00:03:20,960
cost. 
But here is where it gets really

72
00:03:20,960 --> 00:03:23,880
interesting though. 
So you have Opus 4.7, this 

73
00:03:23,880 --> 00:03:27,960
massive, highly sophisticated 
frontier cloud model that can 

74
00:03:27,960 --> 00:03:29,720
dynamically budget its own 
reasoning. 

75
00:03:30,080 --> 00:03:33,840
And yet Simon Willison runs what
calls the Pelican benchmark. 

76
00:03:35,040 --> 00:03:37,520
Yes, the Pelican bench. 
It's just so funny. 

77
00:03:37,520 --> 00:03:41,440
It's his quirky informal image 
generation test where he asks a 

78
00:03:41,440 --> 00:03:44,080
model to generate an image of a 
Pelican riding a bicycle. 

79
00:03:44,080 --> 00:03:47,520
Which tests spatial reasoning, 
anatomical alignment, and prompt

80
00:03:47,520 --> 00:03:50,280
adherence for just totally 
absurd concepts. 

81
00:03:50,280 --> 00:03:52,040
Exactly. 
So he runs this benchmark on the

82
00:03:52,040 --> 00:03:56,040
powerhouse Opus 4.7 and then he 
runs it on a local Alibaba model

83
00:03:56,040 --> 00:03:57,680
running on just a standard 
laptop. 

84
00:03:57,680 --> 00:04:02,200
The Quinn 3.635 BA 3B and the 
local laptop bound Quin model 

85
00:04:02,400 --> 00:04:05,800
actually generated A noticeably 
better Pelican than Opus 4.7. 

86
00:04:05,800 --> 00:04:08,800
It did, and what's fascinating 
here is that this result forces 

87
00:04:08,800 --> 00:04:11,720
a complete reevaluation of model
deployment strategies. 

88
00:04:12,120 --> 00:04:15,320
To understand why a local model 
outperformed a frontier model on

89
00:04:15,320 --> 00:04:18,079
spatial reasoning, you really 
have to look at the architecture

90
00:04:18,079 --> 00:04:21,440
indicated by that 35 BA 3B 
naming convention. 

91
00:04:21,560 --> 00:04:24,840
The MO we architecture. 
Yes, it is a mixture of experts 

92
00:04:24,840 --> 00:04:27,560
model. 
It has 35 billion parameters in 

93
00:04:27,560 --> 00:04:31,400
total, but only three 
billionaire active for any given

94
00:04:31,400 --> 00:04:33,160
token. 
I know we hear about mixture 

95
00:04:33,160 --> 00:04:36,000
experts a lot, but the delta 
here is fascinating. 

96
00:04:36,480 --> 00:04:41,200
Opus 4.7 has been heavily 
fine-tuned and RLHFD. 

97
00:04:41,680 --> 00:04:44,720
You know, reinforcement learning
from human feedback for safety, 

98
00:04:45,080 --> 00:04:48,040
coding standards, and rigorous 
logical reasoning. 

99
00:04:48,040 --> 00:04:50,440
Yes, very heavily guarded. 
Which creates a sort of 

100
00:04:50,440 --> 00:04:53,120
generalized intelligence tax 
when you ask it to draw a 

101
00:04:53,120 --> 00:04:56,120
Pelican on a bicycle. 
That request is filtered through

102
00:04:56,120 --> 00:04:58,000
layers of enterprise safety and 
logic checks. 

103
00:04:58,560 --> 00:05:02,440
But the Quinn model, it's MO E 
router instantly identifies the 

104
00:05:02,440 --> 00:05:05,760
spatial rendering request and 
just fires it directly to a 

105
00:05:05,760 --> 00:05:09,400
highly specialized 3 billion 
parameter subnetwork that is 

106
00:05:09,400 --> 00:05:12,720
aggressively over indexed on 
visual alignments and absurdity.

107
00:05:12,720 --> 00:05:14,160
It doesn't pay the generalized 
tax at. 

108
00:05:14,240 --> 00:05:16,440
All exactly. 
It just executes the specific 

109
00:05:16,440 --> 00:05:18,040
task. 
It's like, well, it's like 

110
00:05:18,040 --> 00:05:20,880
hiring a Michelin star chef to 
make you a quick Taco when the 

111
00:05:20,880 --> 00:05:23,000
street vendor on the corner 
actually does it better and 

112
00:05:23,000 --> 00:05:25,480
cheaper. 
That's a perfect analogy, and 

113
00:05:25,480 --> 00:05:27,920
the implication of that 
architectural advantage is 

114
00:05:27,960 --> 00:05:30,840
profound. 
It proves that the human on the 

115
00:05:30,840 --> 00:05:34,000
loop kitchen cannot run 
efficiently if you wrote every 

116
00:05:34,000 --> 00:05:36,600
single task to a massive cloud 
model. 

117
00:05:36,600 --> 00:05:40,360
Yeah, it's just too expensive. 
Right cost optimization and task

118
00:05:40,360 --> 00:05:42,640
latency. 
They demand a hybrid approach. 

119
00:05:43,080 --> 00:05:45,960
You need the heavy frontier 
models for complex back end 

120
00:05:45,960 --> 00:05:49,480
reasoning, but you need these 
highly specialized, incredibly 

121
00:05:49,480 --> 00:05:52,520
efficient local models for 
narrow execution. 

122
00:05:52,960 --> 00:05:55,840
But wait, how do we decide which
model gets which task? 

123
00:05:56,280 --> 00:05:59,160
Because there was a fatal flaw 
in relying on an army of 

124
00:05:59,160 --> 00:06:02,360
different models, especially 
scrappy local ones, for 

125
00:06:02,360 --> 00:06:04,960
enterprise tasks. 
And that flaw is trust. 

126
00:06:05,240 --> 00:06:07,920
Yes, trust. 
If we are dynamically routing 

127
00:06:07,920 --> 00:06:11,560
complex tasks down to local 3 
billion parameter specialists 

128
00:06:11,560 --> 00:06:14,400
instead of Claude, we suddenly 
have a zero trust environment. 

129
00:06:15,120 --> 00:06:17,800
I mean, if a local Quinn model 
decides to rewrite a core 

130
00:06:17,800 --> 00:06:21,240
database, or an autonomous agent
initiates a financial 

131
00:06:21,240 --> 00:06:23,920
transaction, how do we know it 
hasn't been hijacked? 

132
00:06:24,160 --> 00:06:26,560
We can't just cross our fingers 
and hope the local model isn't 

133
00:06:26,560 --> 00:06:29,320
compromised. 
We can, and the industry 

134
00:06:29,320 --> 00:06:32,800
recognizes that vulnerability, 
which is why the second major 

135
00:06:32,800 --> 00:06:35,960
development in our sources 
tackles identity verification 

136
00:06:35,960 --> 00:06:38,200
head on. 
The new Stack reports that 

137
00:06:38,200 --> 00:06:41,560
Anthropic is rolling out a 
dedicated identity verification 

138
00:06:41,560 --> 00:06:44,840
layer for Claude, and they are 
being explicitly clear about 

139
00:06:44,840 --> 00:06:46,640
this. 
This is not for casual consumer 

140
00:06:46,640 --> 00:06:47,880
use. 
No, definitely not. 

141
00:06:47,880 --> 00:06:50,720
This is engineered specifically 
for high stakes enterprise 

142
00:06:50,720 --> 00:06:54,120
environments to establish 
verifiable audit trails. 

143
00:06:54,120 --> 00:06:56,640
Let's drill into the how of 
that, because audit trails can 

144
00:06:56,640 --> 00:06:59,600
sound like, I don't know, middle
management buzzwords, right? 

145
00:06:59,840 --> 00:07:02,760
We aren't talking about creating
AI bouncers that check a 

146
00:07:02,760 --> 00:07:04,400
password at a door. 
No, not at. 

147
00:07:04,600 --> 00:07:07,360
All we are talking about 
cryptographic software identity.

148
00:07:07,880 --> 00:07:11,640
When an autonomous clawed agent 
touches A repository or triggers

149
00:07:11,640 --> 00:07:16,000
an API, it isn't just leaving a 
text log, it is attaching A 

150
00:07:16,000 --> 00:07:19,360
cryptographically signed 
execution token to the payload. 

151
00:07:19,360 --> 00:07:21,520
Yes. 
So if an enterprise security 

152
00:07:21,520 --> 00:07:24,640
team needs to know exactly which
model altered a piece of code, 

153
00:07:24,960 --> 00:07:27,000
they can verify the signature 
mathematically. 

154
00:07:27,160 --> 00:07:30,400
And that level of mathematical 
certainty is really the only way

155
00:07:30,400 --> 00:07:35,200
to calm enterprise security 
teams, who, rightfully so, view 

156
00:07:35,200 --> 00:07:38,000
agentic AI as a terrifying black
box. 

157
00:07:38,000 --> 00:07:39,880
Oh, totally. 
But identity verification is 

158
00:07:39,880 --> 00:07:42,520
just the defensive baseline. 
We are also seeing the 

159
00:07:42,520 --> 00:07:45,160
deployment of models 
specifically designed for 

160
00:07:45,160 --> 00:07:47,560
offensive security, right? 
Claude Mythos. 

161
00:07:47,600 --> 00:07:51,200
Yes, the notes highlight Claude 
Mythos, which is a cybersecurity

162
00:07:51,200 --> 00:07:54,160
focused model. 
The UKAI Safety Institute 

163
00:07:54,200 --> 00:07:56,480
recently published an 
independent evaluation of 

164
00:07:56,480 --> 00:07:59,800
Mythos, concluding it is quote 
exceptionally effective at 

165
00:07:59,800 --> 00:08:02,520
identifying 0 day 
vulnerabilities and logic. 

166
00:08:02,520 --> 00:08:04,840
Flaws and Open AI is pushing 
into that exhibit. 

167
00:08:04,920 --> 00:08:09,680
Exact same space with GPT 5.4 
Cyber, which they explicitly 

168
00:08:09,680 --> 00:08:12,680
market as being fine-tuned for 
defensive cybersecurity. 

169
00:08:12,680 --> 00:08:14,800
The arms race is definitely on. 
It is. 

170
00:08:14,960 --> 00:08:17,400
We also have partnerships 
forming like the Cloudflare 

171
00:08:17,400 --> 00:08:20,240
Agent Cloud integrating with 
Open AI, and that is 

172
00:08:20,240 --> 00:08:23,560
specifically to offer AI agents 
that come with hard, verifiable 

173
00:08:23,560 --> 00:08:25,360
security guarantees right 
out-of-the-box. 

174
00:08:25,360 --> 00:08:28,320
And the underlying philosophy 
here is captured perfectly by a 

175
00:08:28,320 --> 00:08:30,280
quote from Simon Willison in our
sources. 

176
00:08:30,280 --> 00:08:32,559
Oh I love this quote. 
He notes cybersecurity looks 

177
00:08:32,559 --> 00:08:35,200
like proof of work now. 
Yeah, what he means is that the 

178
00:08:35,200 --> 00:08:38,480
traditional heuristic models of 
cybersecurity are essentially 

179
00:08:38,480 --> 00:08:41,919
dead. 
In an agentic ecosystem, trust 

180
00:08:41,919 --> 00:08:44,920
is established by a model's 
ability to systematically break 

181
00:08:44,920 --> 00:08:47,640
things. 
The models that can autonomously

182
00:08:47,640 --> 00:08:51,000
discover vulnerabilities and 
execute complex adversarial 

183
00:08:51,000 --> 00:08:54,520
attacks are really the only ones
we trust to write the defense 

184
00:08:54,520 --> 00:08:58,920
protocols. 
Import AI Issue 453 dedicated a 

185
00:08:58,920 --> 00:09:01,760
massive deep dive to this exact 
concept. 

186
00:09:02,200 --> 00:09:05,440
You have to subject these 
agentic systems to relentless 

187
00:09:05,440 --> 00:09:08,520
adversarial testing just to 
understand their attack 

188
00:09:08,520 --> 00:09:10,360
surfaces. 
I have to push back on this a 

189
00:09:10,360 --> 00:09:11,680
little bit though. 
OK, go ahead. 

190
00:09:11,920 --> 00:09:15,840
When we talk about cryptographic
execution tokens, adversarial 

191
00:09:15,840 --> 00:09:18,920
proof of work, and strict 
enterprise identity layers, 

192
00:09:19,200 --> 00:09:22,680
doesn't this entirely kill the 
Wild West frictionless fun of 

193
00:09:22,680 --> 00:09:24,800
generative AI? 
It does change the vibe, 

194
00:09:24,800 --> 00:09:26,520
certainly. 
I mean, the whole appeal of this

195
00:09:26,520 --> 00:09:28,520
technology for the last few 
years was that it was 

196
00:09:28,520 --> 00:09:31,280
unconstrained. 
You just type a prompt and watch

197
00:09:31,280 --> 00:09:33,880
it run. 
Are we just creating AI bouncers

198
00:09:33,880 --> 00:09:36,080
to check the ID's of other AI 
models? 

199
00:09:36,200 --> 00:09:38,880
Now we're burying it in 
bureaucratic security protocols.

200
00:09:39,040 --> 00:09:42,120
I understand the nostalgia for 
that frictionless environment, I

201
00:09:42,120 --> 00:09:44,920
really do. 
But consumer experimentation and

202
00:09:44,920 --> 00:09:47,480
enterprise infrastructure 
operate under completely 

203
00:09:47,480 --> 00:09:50,040
different physics. 
Yeah, unconstrained generation 

204
00:09:50,040 --> 00:09:52,720
is fantastic if you are 
brainstorming marketing copy. 

205
00:09:53,360 --> 00:09:55,840
But if you are a healthcare 
provider using an autonomous 

206
00:09:55,840 --> 00:09:59,600
agent to route patient data or a
financial institution updating 

207
00:09:59,600 --> 00:10:03,040
transaction ledgers, 
unconstrained is basically 

208
00:10:03,040 --> 00:10:05,880
synonymous with uninsurable. 
Wow, uninsurable. 

209
00:10:06,240 --> 00:10:09,560
The why behind these rigid 
protocols is sheer survival. 

210
00:10:10,160 --> 00:10:13,960
When models execute tasks 
autonomously, a single 

211
00:10:13,960 --> 00:10:18,160
hallucination or hijacked prompt
can corrupt an entire database. 

212
00:10:18,720 --> 00:10:21,560
If we connect this to the bigger
picture, you just can't have a 

213
00:10:21,560 --> 00:10:23,280
genetic workflows without 
identity. 

214
00:10:23,480 --> 00:10:26,240
You cannot scale them without 
replacing blind trust with 

215
00:10:26,240 --> 00:10:29,880
cryptographic verification. 
Fairpoint, if the AI kitchen 

216
00:10:29,880 --> 00:10:32,480
staff is going to operate the 
deep fryer without me watching, 

217
00:10:32,480 --> 00:10:35,240
I want mathematical proof that 
they know this AC protocol. 

218
00:10:35,480 --> 00:10:38,880
So we solve the routing problem 
by mixing cloud and local 

219
00:10:38,880 --> 00:10:41,080
models. 
We solve the trust problem with 

220
00:10:41,080 --> 00:10:43,280
cryptographic identity and 
adversarial testing. 

221
00:10:43,760 --> 00:10:46,680
But that brings us to the 
communication bottleneck. 

222
00:10:46,680 --> 00:10:48,960
A big one. 
Yeah, if I have a secure Opus 

223
00:10:48,960 --> 00:10:52,680
model, a secure Quinn model, 
secure quad Methos model, all 

224
00:10:52,680 --> 00:10:55,720
working on my behalf, how do 
they actually share state? 

225
00:10:55,720 --> 00:10:58,640
How do they pass data back and 
forth without me basically 

226
00:10:58,640 --> 00:11:00,680
acting as the middleman? 
And that bottleneck was 

227
00:11:00,680 --> 00:11:04,560
addressed head on at the MCP 
Summit in New York City on April

228
00:11:04,560 --> 00:11:08,160
16th, right? 
The Model Context Protocol, or 

229
00:11:08,160 --> 00:11:11,440
MCP, has been simmering in the 
background for a while, but the 

230
00:11:11,440 --> 00:11:14,680
new Stacks coverage of the 
Summit signals a massive turning

231
00:11:14,680 --> 00:11:17,240
point. 
Clara Liguori, a senior 

232
00:11:17,240 --> 00:11:20,680
principal software engineer at 
AWS, delivered a presentation 

233
00:11:20,680 --> 00:11:23,760
that functionally committed 
Amazon to MCP as the universal 

234
00:11:23,760 --> 00:11:26,120
standard. 
This is such a critical shift 

235
00:11:26,240 --> 00:11:28,680
because initially MCP was 
designed just for tool calling, 

236
00:11:28,720 --> 00:11:29,200
right? 
Exactly. 

237
00:11:29,200 --> 00:11:32,560
It was the protocol on LLM used 
to securely query a local 

238
00:11:32,560 --> 00:11:34,600
database. 
Or, you know, ping a weather 

239
00:11:34,600 --> 00:11:38,080
API. 
But the summit positioned MCP as

240
00:11:38,080 --> 00:11:39,960
an agent to agent communication 
standard. 

241
00:11:40,240 --> 00:11:42,680
And to understand why that 
matters we have to look at how 

242
00:11:42,680 --> 00:11:46,200
these models used to interact. 
Historically, if one model 

243
00:11:46,200 --> 00:11:48,960
needed to pass work to another, 
you basically had to have model 

244
00:11:49,040 --> 00:11:52,920
A spit out raw text and model B 
had to scrape and interpret that

245
00:11:52,920 --> 00:11:55,960
text. 
Which was incredibly fragile and

246
00:11:55,960 --> 00:11:59,920
prone to massive context loss. 
MCP changes the underlying 

247
00:11:59,920 --> 00:12:03,080
mechanics completely. 
It provides a universal Jason 

248
00:12:03,080 --> 00:12:07,160
schema for context injection and
tool calling it standardizes the

249
00:12:07,160 --> 00:12:08,800
payload. 
So it's not guessing anymore. 

250
00:12:09,080 --> 00:12:12,240
Right, so a local kin model 
working on a visual layout 

251
00:12:12,240 --> 00:12:16,760
doesn't just pass text to Opus 
4.7, it passes a structured 

252
00:12:16,760 --> 00:12:20,800
machine readable semantic state 
directly into Opus context 

253
00:12:20,800 --> 00:12:22,720
window. 
They share a unified 

254
00:12:22,720 --> 00:12:25,000
understanding of the tasks 
current environment. 

255
00:12:25,240 --> 00:12:27,640
So what does this all mean? 
If these models are all 

256
00:12:27,640 --> 00:12:31,320
different workers in a factory, 
MCP is like the universal UN 

257
00:12:31,320 --> 00:12:33,960
translation headset that lets 
them seamlessly talk to each 

258
00:12:33,960 --> 00:12:36,880
other to finish your project. 
That is exactly what it is and 

259
00:12:36,880 --> 00:12:38,320
we are already seeing this in 
the wild. 

260
00:12:38,440 --> 00:12:41,480
Cloud Code natively supports 
this multi agent MCP 

261
00:12:41,480 --> 00:12:44,480
coordination right now and the 
developer community is 

262
00:12:44,480 --> 00:12:46,360
aggressively upscaling to match 
it. 

263
00:12:46,720 --> 00:12:50,080
GitHub just dropped a series of 
secure code game challenges 

264
00:12:50,080 --> 00:12:54,000
focused entirely on a AI 
security vulnerabilities. 

265
00:12:54,000 --> 00:12:56,920
Basically teaching how these 
models interact over these 

266
00:12:56,920 --> 00:13:00,400
protocols without exposing 
vectors for injection attacks. 

267
00:13:00,400 --> 00:13:03,000
Right, exactly. 
And over 10,000 developers have 

268
00:13:03,000 --> 00:13:04,320
already cleared those 
challenges. 

269
00:13:04,320 --> 00:13:07,320
That's a lot of buy in. 
It is the rapid adoption by 

270
00:13:07,320 --> 00:13:10,320
developers highlights that 
standardized communication was 

271
00:13:10,320 --> 00:13:13,760
really the final missing piece 
required to unlock the ultimate 

272
00:13:13,760 --> 00:13:16,000
benefit for the listener. 
Which is customized personal 

273
00:13:16,000 --> 00:13:18,600
tools, yes. 
We have highly capable models 

274
00:13:18,600 --> 00:13:21,680
for reasoning and execution. 
We have the cryptographic 

275
00:13:21,680 --> 00:13:25,880
identity layers to secure them, 
and now with MCP we have the 

276
00:13:25,880 --> 00:13:28,840
universal schema allowing them 
to coordinate autonomously. 

277
00:13:29,120 --> 00:13:31,560
When you stack those 3 pillars 
together, the resulting 

278
00:13:31,560 --> 00:13:34,480
structure changes the very 
definition of software itself. 

279
00:13:34,680 --> 00:13:37,160
Which brings us to the core 
thesis of our deep dive today, 

280
00:13:37,440 --> 00:13:41,240
the rise of Personal software. 
On April 15th, The New Stack 

281
00:13:41,240 --> 00:13:44,400
published a piece arguing that 
the ultimate end game of AI 

282
00:13:44,560 --> 00:13:48,160
isn't about helping startups 
build the next massive 1,000,000

283
00:13:48,160 --> 00:13:50,040
user saws application. 
No, not at. 

284
00:13:50,200 --> 00:13:52,960
All the end game is eliminating 
the overhead of scale, 

285
00:13:52,960 --> 00:13:56,520
deployment and maintenance so 
entirely that you, the 

286
00:13:56,520 --> 00:13:59,480
individual listener, can build 
complex software applications 

287
00:13:59,840 --> 00:14:02,400
just for yourself. 
And the economic implications of

288
00:14:02,400 --> 00:14:04,320
personal software are just 
staggering. 

289
00:14:04,480 --> 00:14:07,640
Traditional software development
is defined by compromises. 

290
00:14:07,760 --> 00:14:09,640
Always. 
You have to build features that 

291
00:14:09,640 --> 00:14:12,840
appeal to a mass market to 
justify the enormous cost of 

292
00:14:12,840 --> 00:14:15,680
engineering, deploying and 
maintaining the code base. 

293
00:14:16,240 --> 00:14:19,640
But if a coordinated team of AI 
agents can write, secure and 

294
00:14:19,640 --> 00:14:23,600
deploy an application at near 0 
marginal cost, the need for mass

295
00:14:23,600 --> 00:14:26,440
market appeal simply vanishes. 
Yeah, the sources site, this 

296
00:14:26,640 --> 00:14:29,680
brilliant example of a GitHub 
engineer who used the GitHub 

297
00:14:29,680 --> 00:14:33,560
Copilot CLI to build a hyper 
specific organizational command 

298
00:14:33,560 --> 00:14:36,360
center, right. 
This wasn't a product designed 

299
00:14:36,360 --> 00:14:38,880
for onboarding new users and no 
marketing site. 

300
00:14:39,240 --> 00:14:43,360
It was a tool built to perfectly
map to the exact neurological 

301
00:14:43,360 --> 00:14:45,880
quirks of that single engineers 
workflow. 

302
00:14:46,240 --> 00:14:48,360
It is software with an audience 
of 1. 

303
00:14:48,640 --> 00:14:51,520
And we are seeing major 
enterprise platforms tear their 

304
00:14:51,520 --> 00:14:55,040
own products down to the studs 
just to prepare for this shift. 

305
00:14:55,040 --> 00:14:58,000
Oh, like Notion? 
Yes, if you listen to the Latent

306
00:14:58,000 --> 00:15:00,540
Space podcast from earlier this 
week featuring Notions Co 

307
00:15:00,540 --> 00:15:03,920
founder Simon Last and their 
head of AI, Sarah Saks, they 

308
00:15:03,920 --> 00:15:07,080
detail what they call the 
software factory Future. 

309
00:15:07,120 --> 00:15:09,720
Yeah. 
To prepare for a world where AI 

310
00:15:09,720 --> 00:15:13,320
agents build personal software, 
Notion went through five 

311
00:15:13,320 --> 00:15:17,160
complete architecture rebuilds 
and utilized over 100 internal 

312
00:15:17,160 --> 00:15:18,600
tools. 
I've rebuilds. 

313
00:15:18,600 --> 00:15:21,040
That is wild. 
It's a massive undertaking. 

314
00:15:21,360 --> 00:15:24,720
I mean, you can't just slap an 
AI chatbot on top of a legacy 

315
00:15:24,720 --> 00:15:26,560
database and call it a software 
factory. 

316
00:15:26,800 --> 00:15:30,200
If I want AI Agent to instantly 
spin up a custom dashboard for 

317
00:15:30,200 --> 00:15:33,360
my project, the underlying 
architecture has to support that

318
00:15:33,360 --> 00:15:36,600
dynamic manipulation natively. 
That is exactly what Notion 

319
00:15:36,600 --> 00:15:39,200
realized. 
They had to transition away from

320
00:15:39,200 --> 00:15:44,040
rigid static data schemas and 
build a hyper flexible block 

321
00:15:44,040 --> 00:15:46,880
based graph architecture. 
They had to expose their 

322
00:15:46,880 --> 00:15:49,800
primitive API's so that 
autonomous agents could 

323
00:15:49,800 --> 00:15:53,440
programmatically assemble, 
manipulate and tear down user 

324
00:15:53,440 --> 00:15:57,360
interfaces on the fly without 
corrupting the underlying data 

325
00:15:57,360 --> 00:15:59,000
layer. 
They basically rebuilt their 

326
00:15:59,000 --> 00:16:02,640
entire foundation so that the AI
AI could be the developer. 

327
00:16:02,800 --> 00:16:05,800
Because when the AI is the 
developer, the signal to noise 

328
00:16:05,800 --> 00:16:07,960
ratio hits absolute zero. 
Yes. 

329
00:16:08,240 --> 00:16:09,960
Think about how software usually
works. 

330
00:16:10,120 --> 00:16:12,720
You have a product manager 
guessing what a user might need,

331
00:16:12,720 --> 00:16:16,440
a developer trying to interpret 
the product managers spec, and 

332
00:16:16,440 --> 00:16:19,280
then the actual user who just 
has to adapt their workflow to 

333
00:16:19,280 --> 00:16:21,000
whatever the developer finally 
ships. 

334
00:16:21,000 --> 00:16:22,960
Which is so much friction. 
Exactly. 

335
00:16:23,800 --> 00:16:26,400
With personal software, the 
operator, the user, and the 

336
00:16:26,400 --> 00:16:28,680
domain expert are all the exact 
same person. 

337
00:16:28,720 --> 00:16:31,200
It's you. 
You become the CEO of your own 

338
00:16:31,200 --> 00:16:34,640
software factory, and this 
transition is actively being 

339
00:16:34,640 --> 00:16:36,840
measured right now. 
By the Stack Overflow survey. 

340
00:16:37,080 --> 00:16:40,600
Yes, the survey that dropped on 
the 15th specifically asked 

341
00:16:40,600 --> 00:16:43,440
developers to classify their 
current daily workflows. 

342
00:16:43,760 --> 00:16:46,080
Are they a human in the loop or 
a human on the loop? 

343
00:16:46,200 --> 00:16:49,160
And the results. 
The metrics show a massive 

344
00:16:49,160 --> 00:16:52,760
accelerating migration from 
developers identifying as a 

345
00:16:52,760 --> 00:16:56,000
human in the loop writing the 
code to a human on the loop 

346
00:16:56,040 --> 00:16:57,960
orchestrating the agents that 
write the code. 

347
00:16:58,040 --> 00:17:00,200
And if you are stepping into 
that human on the loop 

348
00:17:00,200 --> 00:17:03,520
management role this week, we 
actually have an incredibly 

349
00:17:03,520 --> 00:17:06,440
actionable pro tip pulled 
directly from the developer 

350
00:17:06,440 --> 00:17:09,119
dispatches in our sources. 
This is a really good one. 

351
00:17:09,160 --> 00:17:11,680
Yeah, so if you are running a 
genetic clawed code session 

352
00:17:11,680 --> 00:17:14,119
specifically where you have 
turned off the safety 

353
00:17:14,119 --> 00:17:18,040
confirmations to let the AI act 
completely autonomously, you 

354
00:17:18,040 --> 00:17:21,560
need to implement a strict 
oversight rule in your claw EE 

355
00:17:21,560 --> 00:17:24,920
dot MD file. 
The rule is simple, but it is 

356
00:17:24,920 --> 00:17:27,520
critical for maintaining 
architectural integrity. 

357
00:17:27,760 --> 00:17:30,920
You instruct the model that 
before executes any action that 

358
00:17:30,920 --> 00:17:34,320
touches more than two files 
simultaneously, it must write a 

359
00:17:34,320 --> 00:17:36,720
one line intense summary as a 
code comment. 

360
00:17:36,960 --> 00:17:40,320
Let's explain why that specific 
mechanism is so powerful. 

361
00:17:41,400 --> 00:17:44,880
By forcing the model to 
articulate its intent before it 

362
00:17:44,880 --> 00:17:48,040
executes a multi file change, 
you are essentially functioning 

363
00:17:48,160 --> 00:17:50,080
as your own enterprise security 
team. 

364
00:17:50,200 --> 00:17:52,920
Exactly. 
First off, you force the LLM to 

365
00:17:52,920 --> 00:17:55,360
commit its reasoning to the 
context window, which 

366
00:17:55,360 --> 00:17:58,120
significantly reduces the chance
of hallucinations during 

367
00:17:58,120 --> 00:18:00,560
execution. 
It has to think before it acts. 

368
00:18:01,240 --> 00:18:05,040
Second, you're generating a 
machine readable audit trail. 

369
00:18:05,440 --> 00:18:08,560
Because the intent summaries are
uniformly formatted, they are 

370
00:18:08,560 --> 00:18:10,640
highly graphical for post 
session review. 

371
00:18:10,720 --> 00:18:12,800
It makes oversight incredibly 
easy. 

372
00:18:12,840 --> 00:18:14,960
Totally. 
When you step back into your 

373
00:18:14,960 --> 00:18:17,880
oversight role at the end of the
day, you can run a simple search

374
00:18:17,880 --> 00:18:21,920
query and instantly review every
major architectural decision the

375
00:18:21,920 --> 00:18:23,720
autonomous agents made while you
were away. 

376
00:18:24,120 --> 00:18:26,760
You maintain total visibility 
over the factory floor. 

377
00:18:26,920 --> 00:18:30,120
It is a perfect example of how 
the human on the loop interacts 

378
00:18:30,120 --> 00:18:32,640
with the system. 
You aren't writing the code, but

379
00:18:32,640 --> 00:18:35,240
you are strictly defining the 
operational parameters. 

380
00:18:35,320 --> 00:18:38,920
So if I can instantly spin up a 
perfectly customized app just 

381
00:18:38,920 --> 00:18:42,440
for my own daily workflow, does 
this mean the death of mass 

382
00:18:42,440 --> 00:18:45,000
market sauce apps? 
Well, this raises an important 

383
00:18:45,000 --> 00:18:46,760
question about the future of the
industry. 

384
00:18:46,760 --> 00:18:49,680
The traditional mass market 
sauce model will face an 

385
00:18:49,680 --> 00:18:51,680
existential threat at the 
interface level. 

386
00:18:51,680 --> 00:18:55,320
Absolutely right. 
A generic off the shelf project 

387
00:18:55,320 --> 00:18:58,880
management tool simply cannot 
compete with a bespoke tool 

388
00:18:58,880 --> 00:19:02,320
tailored to perfectly match your
specific cognitive habits. 

389
00:19:02,480 --> 00:19:05,480
It's no contest. 
However, mass market saws won't 

390
00:19:05,480 --> 00:19:09,360
disappear completely, it will be
forced to pivot and retreat down

391
00:19:09,360 --> 00:19:11,280
the stack. 
What do you mean Companies will 

392
00:19:11,280 --> 00:19:15,480
pivot away from selling user 
interfaces and focus entirely on

393
00:19:15,480 --> 00:19:18,640
selling highly secure, heavily 
compliant back end 

394
00:19:18,640 --> 00:19:21,480
infrastructure. 
The API's, the data lakes, the 

395
00:19:21,480 --> 00:19:23,640
raw materials. 
The back end will remains 

396
00:19:23,640 --> 00:19:26,840
standardized but the front end 
interface and logic will be 

397
00:19:26,840 --> 00:19:30,000
generated on the family as hyper
personalized software. 

398
00:19:30,000 --> 00:19:31,760
Wow. 
OK, to quickly bring this all 

399
00:19:31,760 --> 00:19:33,760
together for you, we were 
watching a fundamental 

400
00:19:33,760 --> 00:19:36,520
reorganization of computing. 
We started the week with 

401
00:19:36,520 --> 00:19:40,160
Anthropic launching Opus 4.7, 
only to learn that specialized 

402
00:19:40,160 --> 00:19:44,280
local models like Quinn 3.6 can 
wrap specific spatial tasks 

403
00:19:44,440 --> 00:19:47,240
faster and better. 
We saw the deployment of AI 

404
00:19:47,240 --> 00:19:50,840
bouncers, checking ID's, 
cryptographic identity layers 

405
00:19:51,120 --> 00:19:54,360
and adversarial proof of work to
ensure we can actually trust 

406
00:19:54,360 --> 00:19:55,920
these agents to operate 
securely. 

407
00:19:55,920 --> 00:19:58,840
Crucial infrastructure. 
And we watched AWS cement the 

408
00:19:58,840 --> 00:20:02,280
MCP translation headset 
standard, giving these disparate

409
00:20:02,280 --> 00:20:05,200
models a unified schema to 
communicate seamlessly. 

410
00:20:06,000 --> 00:20:08,120
And the result of all that 
infrastructure is that you are 

411
00:20:08,120 --> 00:20:11,240
no longer just a user of 
software, you are the domain 

412
00:20:11,240 --> 00:20:14,960
expert overseeing a vast, 
capable digital workforce 

413
00:20:14,960 --> 00:20:16,800
building tools specifically for 
you. 

414
00:20:17,160 --> 00:20:20,800
It is an incredibly empowering 
position to be in, but as we 

415
00:20:20,800 --> 00:20:24,040
transition into this era of 
hyper personalized software 

416
00:20:24,040 --> 00:20:27,480
tailored exactly to our own 
brains and workflows, I want to 

417
00:20:27,480 --> 00:20:29,040
leave you with a final thought 
to Mull over. 

418
00:20:29,640 --> 00:20:32,360
If we all build custom 
dashboards, unique workflows, 

419
00:20:32,360 --> 00:20:35,520
and bespoke interfaces, do we 
lose the shared language? 

420
00:20:35,600 --> 00:20:37,400
Language of work? 
Oh wow, When you and your 

421
00:20:37,400 --> 00:20:40,600
colleagues no longer use the 
same standardized tools, how 

422
00:20:40,600 --> 00:20:42,080
will we collaborate in the 
future? 

423
00:20:42,200 --> 00:20:44,600
That is, yeah, that is a 
fascinating problem. 

424
00:20:45,000 --> 00:20:49,000
If everyone is the head chef of 
their own completely customized 

425
00:20:49,120 --> 00:20:52,080
AI run kitchen, how do we ever 
coordinate a cook a meal 

426
00:20:52,080 --> 00:20:54,920
together, somebody to chew on as
you start spinning up your own 

427
00:20:54,920 --> 00:20:56,480
autonomous software factory this
weekend. 

428
00:20:56,760 --> 00:20:59,000
Thank you so much for joining us
for this deep dive. 

429
00:20:59,000 --> 00:21:00,000
We will see you next time.
