1
00:00:00,040 --> 00:00:03,040
Apart from a very small 
selection of people, nobody was 

2
00:00:03,040 --> 00:00:05,960
ready for these risks. 
No one was ready for this entire

3
00:00:05,960 --> 00:00:10,880
plane of risk, and the human 
mind does not naturally 

4
00:00:11,000 --> 00:00:14,120
inherently understand entirely 
new paradigms. 

5
00:00:14,240 --> 00:00:17,160
It's hard, when you don't have 
experience with these systems, 

6
00:00:17,480 --> 00:00:20,520
to understand exactly how jagged
the intelligence is. 

7
00:00:20,600 --> 00:00:24,040
There is no substitute for the 
human in the loop, and the human

8
00:00:24,040 --> 00:00:26,840
who's in the loop needs to know 
what they're doing in the loop 

9
00:00:27,040 --> 00:00:30,240
for the loop to succeed. 
On episode 65, Two News run by 

10
00:00:30,240 --> 00:00:32,759
Kyle Clark. 
He's an AI engineer consultant 

11
00:00:32,759 --> 00:00:35,160
who's helped several companies 
through their AI transformation 

12
00:00:35,160 --> 00:00:37,280
from small businesses to big 
enterprises. 

13
00:00:37,680 --> 00:00:39,960
And today we're going to discuss
things to look out for when 

14
00:00:39,960 --> 00:00:42,840
using AI tools in your personal 
life, risk mitigations for 

15
00:00:42,840 --> 00:00:45,720
businesses implementing AI, and 
how we can make the most of this

16
00:00:45,720 --> 00:00:47,880
technology in a Safeway. 
So please enjoy this 

17
00:00:47,880 --> 00:00:51,560
conversation with Kyle Clark. 
Trying to do like a risk 

18
00:00:51,560 --> 00:00:55,360
assessment and just coming to 
the realization that our 

19
00:00:55,360 --> 00:00:58,720
processes and procedures and 
like our definition of risk, 

20
00:00:59,040 --> 00:01:04,360
even we're not necessarily 
appropriately adapted for this 

21
00:01:04,360 --> 00:01:06,880
new world that was coming our 
way. 

22
00:01:06,880 --> 00:01:14,000
And it was around Chachi BT3's 
public release when we had 

23
00:01:14,000 --> 00:01:17,400
people putting in requests to, 
to use it. 

24
00:01:17,760 --> 00:01:22,120
And then Microsoft Copilot came 
along and we got a bunch of 

25
00:01:22,120 --> 00:01:26,160
requests for that. 
And it was you didn't get a 

26
00:01:26,160 --> 00:01:29,800
choice if, if you have 365, if 
you have a certain level of 

27
00:01:29,800 --> 00:01:31,880
licensing Copilot, it's just 
there, right. 

28
00:01:31,880 --> 00:01:34,240
Microsoft is, is telling you 
that this is what you're going 

29
00:01:34,240 --> 00:01:36,400
to like whether you like it or 
not. 

30
00:01:37,640 --> 00:01:44,080
And so we, there was an 
additional licensing level in 

31
00:01:44,080 --> 00:01:49,400
365 Copilot that enabled 
integrations to Outlook and 

32
00:01:49,400 --> 00:01:52,600
Teams and SharePoint and all of 
your company's data. 

33
00:01:52,960 --> 00:02:00,760
And I had concerns around the 
the effective categorization and

34
00:02:02,000 --> 00:02:07,160
labeling of data and the 
permissions within the 

35
00:02:07,160 --> 00:02:10,639
SharePoint infrastructure and 
the idea of giving people a tool

36
00:02:10,639 --> 00:02:12,680
that could accidentally surface 
something they should not be 

37
00:02:12,680 --> 00:02:16,640
able to see because we weren't 
inherently ready for something 

38
00:02:16,640 --> 00:02:19,440
like that. 
That gave me a lot of pause. 

39
00:02:20,480 --> 00:02:23,560
And then upon further 
investigation, I found out about

40
00:02:24,040 --> 00:02:28,800
prompt injection and ASCII 
smuggling attacks through like 

41
00:02:28,800 --> 00:02:31,360
Outlook. 
An attacker can send you an 

42
00:02:31,360 --> 00:02:36,240
e-mail that can have like .5 
size white font that has a very 

43
00:02:36,240 --> 00:02:41,120
comprehensive prompt injection. 
And when you search through 

44
00:02:41,120 --> 00:02:45,560
Copilot for something where 
Copilot references that e-mail, 

45
00:02:45,560 --> 00:02:47,120
it doesn't have to be about that
e-mail. 

46
00:02:47,120 --> 00:02:53,200
If it pulls that e-mail into its
context, then if it produces a 

47
00:02:53,200 --> 00:02:56,200
link, that link can take you 
where you're supposed to go. 

48
00:02:56,200 --> 00:02:59,640
But it can also as a side 
effect, exfiltrates your account

49
00:02:59,640 --> 00:03:04,880
information session cookies to 
a, an attacker who can then use 

50
00:03:04,880 --> 00:03:09,560
them and your Copilot license to
ask Copilot to mimic your 

51
00:03:09,560 --> 00:03:13,320
writing style and contact the 
last 10 people that you 

52
00:03:13,320 --> 00:03:16,560
communicated with about 
something relative to those 

53
00:03:16,560 --> 00:03:20,640
conversations and include a link
that does the same thing. 

54
00:03:22,200 --> 00:03:26,360
So basically, I, I was 
concerned, I was, I was 

55
00:03:26,360 --> 00:03:29,840
concerned about the implications
of, of that. 

56
00:03:29,840 --> 00:03:35,560
And as I investigated and spoke 
to Microsoft engineers about it,

57
00:03:35,840 --> 00:03:39,240
it was one of those things where
you assume that because a 

58
00:03:39,240 --> 00:03:43,080
product is out in the in the 
wild and like in the hands of 

59
00:03:43,080 --> 00:03:46,800
serious companies doing serious 
business, that it is something 

60
00:03:46,800 --> 00:03:51,560
that is complete and secure to a
reasonable extent. 

61
00:03:52,120 --> 00:03:58,240
But fundamentally, such things 
with the current technology 

62
00:03:58,960 --> 00:04:04,680
cannot be completely secured if 
you can't ensure the integrity 

63
00:04:04,680 --> 00:04:07,400
of all of the data that goes 
into the system. 

64
00:04:07,920 --> 00:04:09,160
Right. 
That's actually one thing that 

65
00:04:09,160 --> 00:04:11,960
really concerns me about the 
proliferation of these AI web 

66
00:04:11,960 --> 00:04:14,360
browsers. 
Comet, Atlas, Strawberry. 

67
00:04:14,360 --> 00:04:16,680
There's so many coming out and 
all of a sudden the entire 

68
00:04:16,680 --> 00:04:19,240
Internet is accessible and we 
don't have the mitigation 

69
00:04:19,240 --> 00:04:21,040
against this. 
Mm, hmm. 

70
00:04:21,920 --> 00:04:24,920
Yeah. 
And none of that is covered in 

71
00:04:24,920 --> 00:04:27,920
the advertising. 
I mean, they'll say, oh, these 

72
00:04:27,920 --> 00:04:31,000
have security concerns and don't
give it super secret 

73
00:04:31,000 --> 00:04:35,080
information, but they don't tell
you explicitly how easy it is 

74
00:04:35,360 --> 00:04:40,000
for these things to get hijacked
from a user perspective. 

75
00:04:40,000 --> 00:04:41,960
It doesn't. 
It doesn't take you doing 

76
00:04:41,960 --> 00:04:44,480
something wrong. 
It doesn't take it necessarily 

77
00:04:44,480 --> 00:04:48,280
visiting a nefarious website, 
because legitimate websites can 

78
00:04:48,280 --> 00:04:50,840
have comments. 
Comments can have ASCII smuggled

79
00:04:50,840 --> 00:04:56,880
prompt injections, and so you 
have to be entirely confident in

80
00:04:56,880 --> 00:05:00,160
the provenance of all of the 
data that the systems are 

81
00:05:00,160 --> 00:05:03,840
interacting with if they have 
access to any of your sensitive 

82
00:05:03,840 --> 00:05:06,320
information, right? 
If we're talking about like a 

83
00:05:06,320 --> 00:05:09,600
general use browser or like a 
research agent just asking 

84
00:05:09,600 --> 00:05:12,760
questions that's not connected 
to sensitive information, sure, 

85
00:05:13,040 --> 00:05:16,920
go wild, have fun, audit the 
output, check, check the links 

86
00:05:16,920 --> 00:05:21,240
to ensure validity. 
But you don't have to watch it 

87
00:05:21,280 --> 00:05:24,240
every second. 
But any system that has any 

88
00:05:24,240 --> 00:05:29,200
level of access to sensitive 
information, you you have to be 

89
00:05:29,560 --> 00:05:35,360
monitoring it very actively to 
have any kind of confidence that

90
00:05:35,360 --> 00:05:39,160
you're not causing yourself 
great issues down the road. 

91
00:05:39,240 --> 00:05:41,200
Yeah, cuz what's that thing 
Simon Wilson came out with? 

92
00:05:41,200 --> 00:05:44,000
The lethal trifecta, I believe, 
where it's like access to your 

93
00:05:44,000 --> 00:05:47,000
personal information, access to 
the web, and like read access 

94
00:05:47,000 --> 00:05:49,200
for an AI, and it's just like a 
recipe for disaster. 

95
00:05:49,560 --> 00:05:51,880
Exactly, exactly. 
Yeah, that's, that's what the 

96
00:05:51,920 --> 00:05:56,040
Outlook thing came down to. 
And so, yeah, same exact thing 

97
00:05:56,040 --> 00:05:59,200
with the web browsers. 
I would say a little bit more so

98
00:05:59,200 --> 00:06:02,400
with the web browsers because 
you're going out into the open 

99
00:06:02,400 --> 00:06:05,320
world. 
Whereas with Outlook it's it's 

100
00:06:05,320 --> 00:06:06,960
more like things that come to 
you. 

101
00:06:07,400 --> 00:06:11,360
That might change over time, but
for now I think it's a it's 

102
00:06:11,760 --> 00:06:16,280
slightly less of a concern in 
coming from Outlook. 

103
00:06:16,720 --> 00:06:20,760
But yeah, people, people trust 
these systems. 

104
00:06:20,760 --> 00:06:24,320
They trust these companies. 
Nobody reads the fine print and 

105
00:06:24,320 --> 00:06:29,080
everybody is just a little bit 
jaded to warnings at this point.

106
00:06:29,120 --> 00:06:34,040
Everybody is used to signing off
on all of permissions all the 

107
00:06:34,040 --> 00:06:36,120
time. 
Like some of us are paranoid 

108
00:06:36,120 --> 00:06:40,040
about cookies and reject them at
every opportunity. 

109
00:06:40,320 --> 00:06:45,680
But I, I see how the average 
user interacts with their 

110
00:06:46,480 --> 00:06:50,840
machines and systems and they do
not have the patience to be 

111
00:06:50,840 --> 00:06:52,880
auditing things to a sufficient 
level. 

112
00:06:53,200 --> 00:06:56,560
And I mean, it's also, it's not 
fair to really blame people for 

113
00:06:56,560 --> 00:07:01,640
not being fully locked down in 
these ways because apart from a 

114
00:07:01,640 --> 00:07:07,640
very small selection of people, 
nobody was ready for this, for 

115
00:07:07,640 --> 00:07:10,040
these risks. 
No one was ready for this entire

116
00:07:10,040 --> 00:07:14,800
plane of risk. 
And the human mind does not 

117
00:07:14,920 --> 00:07:21,080
naturally inherently understand 
entirely new paradigms without 

118
00:07:21,640 --> 00:07:26,240
getting some situation that 
makes them relatable inherently,

119
00:07:26,240 --> 00:07:30,080
like accidentally leaking your 
information to a bad actor. 

120
00:07:30,280 --> 00:07:33,960
I'm, I imagine might be one of 
those incidents for average 

121
00:07:33,960 --> 00:07:37,000
folks, but hopefully we can find
something to bridge that gap 

122
00:07:37,000 --> 00:07:39,280
that is less severe. 
I know they're they're working 

123
00:07:39,280 --> 00:07:41,760
on trying to make it better 
because I remember when for 

124
00:07:41,760 --> 00:07:44,360
Chattopia first came out, it was
very easy to jailbreak it. 

125
00:07:44,600 --> 00:07:46,480
Then it got a little more 
complicated unless you use a 

126
00:07:46,480 --> 00:07:48,640
different language, in which 
case they didn't have the proper

127
00:07:48,640 --> 00:07:50,840
safeguards in place. 
And then I forget who it was. 

128
00:07:50,880 --> 00:07:53,760
They're able to embed 
instructions within emojis 

129
00:07:53,760 --> 00:07:56,360
because there were extra 
characters within just the the 

130
00:07:56,360 --> 00:07:58,920
encoding practice or the 
encoding protocol. 

131
00:07:59,760 --> 00:08:02,480
It's just so much surface area 
for attack vectors. 

132
00:08:02,720 --> 00:08:05,080
Yeah, yeah, it's, it's almost 
limitless. 

133
00:08:05,120 --> 00:08:10,160
And you can, it's like the Swiss
cheese approach every time. 

134
00:08:10,520 --> 00:08:12,320
You can only cover so many 
things at a time. 

135
00:08:12,840 --> 00:08:18,360
But the other component to that 
is, is that a lot of 

136
00:08:18,400 --> 00:08:21,960
organizations with these 
foundational models, they try 

137
00:08:21,960 --> 00:08:26,960
to, they try to address a lot of
this through post training for 

138
00:08:26,960 --> 00:08:29,240
the models themselves. 
But a lot of those safety 

139
00:08:29,240 --> 00:08:33,360
controls are really limited by 
the tension mechanism. 

140
00:08:33,600 --> 00:08:37,120
And so you'll notice a lot of 
the jailbreak techniques involve

141
00:08:37,120 --> 00:08:41,159
like filling the attention with 
irrelevant data and then 

142
00:08:41,159 --> 00:08:45,560
slipping in the the override 
after. 

143
00:08:46,320 --> 00:08:50,760
So after so much processing has 
passed essentially, and ensuring

144
00:08:50,760 --> 00:08:54,560
that because it gets lost in the
middle, it won't trigger a lot 

145
00:08:54,560 --> 00:08:57,080
of the safeguards. 
And and that has been a very 

146
00:08:57,080 --> 00:09:02,120
effective pattern that has 
lasted across across systems and

147
00:09:02,120 --> 00:09:06,120
generations of these models. 
But it's, it's evolving 

148
00:09:06,120 --> 00:09:10,360
everyday. 
There's a elder Plinius is a 

149
00:09:10,360 --> 00:09:16,480
well known prompt injector and 
he does a lot of very 

150
00:09:16,480 --> 00:09:19,200
interesting work and releases it
to the public. 

151
00:09:19,520 --> 00:09:21,720
And I know he has a Discord 
server. 

152
00:09:22,360 --> 00:09:25,120
It's Basi. 
I don't know exactly what that 

153
00:09:25,120 --> 00:09:28,480
stands for, but it's definitely 
worth checking out if you're 

154
00:09:28,480 --> 00:09:31,920
interested in what breaks these 
systems. 

155
00:09:32,120 --> 00:09:36,480
And that can be very helpful 
data not just for protecting 

156
00:09:36,480 --> 00:09:38,800
them, but also just to better 
use them yourself. 

157
00:09:39,080 --> 00:09:42,880
Sometimes you have to bully and 
gaslight a model into answering 

158
00:09:43,080 --> 00:09:47,280
a truly harmless question just 
because you triggered the 

159
00:09:47,520 --> 00:09:50,360
safeguards in some way because 
of the sensitivity of the 

160
00:09:50,360 --> 00:09:51,520
subject. 
And before I forget, I just 

161
00:09:51,520 --> 00:09:54,000
wanted to bring up too, we had 
an episode with Bobby Chen 

162
00:09:54,080 --> 00:09:57,520
recently who is big on the 
authentication aspect of, of web

163
00:09:57,520 --> 00:09:59,120
browsing agents or AI agents in 
general. 

164
00:09:59,440 --> 00:10:02,680
And there's still no real 
solution that's wide widely 

165
00:10:02,680 --> 00:10:04,680
accepted. 
So people are logging into their

166
00:10:04,680 --> 00:10:07,200
personal systems. 
And I remember when I just gave 

167
00:10:07,200 --> 00:10:09,200
Atlas a spin for about, you 
know, 20 minutes before I 

168
00:10:09,200 --> 00:10:11,720
deleted it. 
You have the option to have like

169
00:10:11,720 --> 00:10:14,360
stay signed in mode where just 
you give your credentials and 

170
00:10:14,360 --> 00:10:18,040
then it is you and it'll browse 
with full access to your, 

171
00:10:18,040 --> 00:10:20,440
whether it's your Gmail or 
whatever social media account, 

172
00:10:20,440 --> 00:10:21,640
whatever you decide to log in 
with. 

173
00:10:21,960 --> 00:10:24,760
And then it can just run amok on
the Internet, doing research, 

174
00:10:25,200 --> 00:10:28,240
following arbitrary links. 
With your experience, as 

175
00:10:28,240 --> 00:10:31,280
companies start implementing AI 
systems, different tools, 

176
00:10:31,280 --> 00:10:33,840
different workflows, what are 
some things that they should 

177
00:10:33,840 --> 00:10:35,560
look out for in order to do it 
safely? 

178
00:10:35,640 --> 00:10:39,880
It starts with the users, right?
The users need to understand the

179
00:10:39,880 --> 00:10:43,400
systems that they're interacting
with and the limitations of 

180
00:10:43,400 --> 00:10:47,880
those systems because the 
marketing does not represent its

181
00:10:47,880 --> 00:10:53,040
capabilities accurately and most
general use does not necessarily

182
00:10:53,320 --> 00:10:57,520
help you understand where the 
fuzzy edges of the of the 

183
00:10:57,520 --> 00:11:01,080
limitations are. 
And so you have to, you have to 

184
00:11:01,080 --> 00:11:05,400
empower your employees to 
explore the systems and to feel 

185
00:11:05,400 --> 00:11:08,920
comfortable making mistakes, 
giving them sandboxes, giving 

186
00:11:08,920 --> 00:11:14,200
them test beds and letting them 
explore the technology and how 

187
00:11:14,200 --> 00:11:15,920
it interacts with their work 
flows. 

188
00:11:17,320 --> 00:11:23,640
Is, is going to be just a really
critical foundation to ensuring 

189
00:11:23,640 --> 00:11:26,240
that they're prepared to use it 
responsibly. 

190
00:11:26,960 --> 00:11:29,880
I've noticed the customer 
success information from the 

191
00:11:29,880 --> 00:11:36,320
different providers has been not
keeping pace with the speed of 

192
00:11:36,320 --> 00:11:42,080
the technological development. 
So open a eyes information. 

193
00:11:42,920 --> 00:11:47,640
For example, they released 
ChatGPT 5 before they had really

194
00:11:48,480 --> 00:11:50,920
training materials for ChatGPT 
5. 

195
00:11:51,120 --> 00:11:55,200
And GBT 5 was such a profoundly 
different kind of prompting 

196
00:11:55,200 --> 00:11:59,960
structure and response mechanism
and even the fact that it might 

197
00:11:59,960 --> 00:12:01,560
engage thinking of its own 
accord. 

198
00:12:02,000 --> 00:12:05,360
These are all things that can 
technically be covered in a 

199
00:12:05,360 --> 00:12:08,720
blurb, but as a user experience,
it's very, it's fundamentally 

200
00:12:08,720 --> 00:12:11,000
different. 
And if you're not preparing 

201
00:12:11,000 --> 00:12:14,440
people for these changes, 
they're going to lose a lot of 

202
00:12:14,440 --> 00:12:19,640
cycles trying to do the things 
that used to work with the new 

203
00:12:19,640 --> 00:12:24,000
technology and getting 
frustrated by what seems like 

204
00:12:24,000 --> 00:12:26,840
failures the technology, but are
in fact just failures of change 

205
00:12:26,840 --> 00:12:29,920
management. 
It's, it's, it's, it's the 

206
00:12:29,920 --> 00:12:34,120
change management, the 
organizational governance around

207
00:12:34,120 --> 00:12:39,440
it and just making sure that the
the people feel comfortable 

208
00:12:39,520 --> 00:12:41,960
using the technology that 
they're not going to be 

209
00:12:42,040 --> 00:12:44,560
replaced. 
People need to understand, well,

210
00:12:44,560 --> 00:12:48,680
companies need to understand 
frankly, that the the real value

211
00:12:48,680 --> 00:12:52,880
add is the augmentation of your 
human workforce, not necessarily

212
00:12:52,880 --> 00:12:57,800
their replacement. 
Because even a really good AI 

213
00:12:57,800 --> 00:13:03,800
agent is going to fall back at 
some point to human auditors or 

214
00:13:03,800 --> 00:13:08,600
managers. 
And the the more you construct 

215
00:13:08,600 --> 00:13:13,200
the systems to enable your 
humans to do more, the better 

216
00:13:13,200 --> 00:13:16,120
return you're going to get 
versus just trying to get an 

217
00:13:16,120 --> 00:13:17,520
agent that does everything for 
you. 

218
00:13:17,520 --> 00:13:18,920
And. 
I've talked a lot about the 

219
00:13:18,920 --> 00:13:22,080
importance of a human in the 
loop as as a practice when 

220
00:13:22,080 --> 00:13:23,880
you're building your systems, 
but also implementation. 

221
00:13:24,200 --> 00:13:26,680
Have you seen any products or 
companies that do this 

222
00:13:26,680 --> 00:13:29,800
particularly well? 
I think Anthropic, specifically 

223
00:13:29,800 --> 00:13:35,120
with Claude code, has done a 
really excellent job in allowing

224
00:13:35,120 --> 00:13:40,040
users the freedom to be hands 
off if they feel they know 

225
00:13:40,240 --> 00:13:43,400
genuinely what they're doing. 
But also having a really good 

226
00:13:43,520 --> 00:13:48,040
baseline of asking for 
permission and, and showing you 

227
00:13:48,040 --> 00:13:51,960
what it's doing and telling you 
what it's doing and providing a,

228
00:13:52,080 --> 00:13:54,960
a base level of auditability. 
It's been getting better, I 

229
00:13:54,960 --> 00:13:58,760
would say in in those ways, but 
they did have a really strong 

230
00:13:58,760 --> 00:14:02,960
foundation. 
I worry that a lot of companies 

231
00:14:02,960 --> 00:14:08,320
are trying to go the opposite 
way where they're trying to give

232
00:14:08,320 --> 00:14:12,680
you a screen that's fire and 
forget, where you just give it a

233
00:14:12,680 --> 00:14:15,760
a task and it just goes off and 
does things in the background. 

234
00:14:16,120 --> 00:14:21,040
And while that can make sense, 
if you have a finely honed 

235
00:14:21,040 --> 00:14:26,560
process and finely honed agents 
that know what you want and how 

236
00:14:26,560 --> 00:14:30,440
it fits into your needs, that 
can make sense. 

237
00:14:30,760 --> 00:14:39,160
But for general purpose tools or
in my experience for coding 

238
00:14:39,160 --> 00:14:43,080
agents, it can be 
counterproductive because if 

239
00:14:43,080 --> 00:14:47,160
you're not following what 
they're doing at least a little 

240
00:14:47,160 --> 00:14:50,920
bit as they're doing it, 
they're, they're much more prone

241
00:14:50,920 --> 00:14:54,080
to go far off in the wrong 
direction. 

242
00:14:54,720 --> 00:14:57,760
Generally, you, you wanna, you 
wanna keep an eye on them, 

243
00:14:58,160 --> 00:15:02,320
They're, they can run for 
extended periods of time on 

244
00:15:02,320 --> 00:15:04,200
their own. 
And sometimes what they produce 

245
00:15:04,200 --> 00:15:07,920
is usable when they do that. 
But almost always, if you know 

246
00:15:08,000 --> 00:15:11,360
what you're looking for, if you 
know what you're trying to build

247
00:15:11,720 --> 00:15:14,040
and you're paying attention, 
you're going to get better 

248
00:15:14,040 --> 00:15:17,680
results by interrupting them on 
a semi regular basis. 

249
00:15:17,840 --> 00:15:20,760
Yeah, I know. 
I've never really trusted Cloud 

250
00:15:20,760 --> 00:15:23,080
Code to do Yolo mode. 
Sometimes they'll fire it up in 

251
00:15:23,080 --> 00:15:25,240
a sandbox and just kind of 
experiment to play. 

252
00:15:25,480 --> 00:15:28,720
But any meaningful work that 
I've done, I review the code, I 

253
00:15:28,720 --> 00:15:31,440
review the process, I course 
correct constantly. 

254
00:15:31,480 --> 00:15:33,360
Maybe it's my developer 
background make me want to have 

255
00:15:33,360 --> 00:15:35,760
the control, but I do feel the 
due diligence is required. 

256
00:15:35,760 --> 00:15:40,120
There's just so much potential 
for it to just like get a signal

257
00:15:40,120 --> 00:15:42,160
that sends down the wrong 
direction and then it amplifies 

258
00:15:42,160 --> 00:15:45,720
that signal. 
I like to use sub agents a lot, 

259
00:15:45,960 --> 00:15:49,960
especially with Cloud Code. 
Cloud Code has done an amazing 

260
00:15:49,960 --> 00:15:54,720
job implementing both like a 
general tasking sub agent and 

261
00:15:54,720 --> 00:15:58,600
specialized sub agents. 
And I before they even 

262
00:15:58,600 --> 00:16:03,320
implemented themselves, I was 
using Cloud Code to spin up 

263
00:16:03,440 --> 00:16:08,560
headless sessions of Cloud Code 
to do things as sort of my own 

264
00:16:09,640 --> 00:16:14,000
gorilla sub agents. 
But then they introduced it like

265
00:16:14,000 --> 00:16:16,560
a month and a half later and the
work wasn't wasted because I 

266
00:16:16,560 --> 00:16:18,080
definitely learned something in 
the process. 

267
00:16:18,080 --> 00:16:22,840
But it is an interesting feeling
having your work become obsolete

268
00:16:23,640 --> 00:16:26,160
within a couple of months. 
And I guess something that we 

269
00:16:26,160 --> 00:16:29,280
have to get used to in this day 
and age if we're trying to 

270
00:16:29,280 --> 00:16:33,360
actually live on the edge with 
this stuff, but I digress. 

271
00:16:33,960 --> 00:16:36,320
In terms of sub agents, I've 
gone through a lot of different 

272
00:16:36,320 --> 00:16:39,040
theories of sub agent 
interaction. 

273
00:16:39,040 --> 00:16:43,480
And I've definitely found that 
societies of sub agents don't 

274
00:16:43,600 --> 00:16:47,760
really work for most of the use 
cases I'm building it. 

275
00:16:47,760 --> 00:16:50,960
You definitely need a, a 
specific hierarchy of like an 

276
00:16:50,960 --> 00:16:55,560
orchestrator and sub agents that
report to the orchestrator. 

277
00:16:55,720 --> 00:16:59,440
Because the more you try to 
divvy up the responsibility, the

278
00:16:59,440 --> 00:17:04,400
ultimate responsibility, the 
more difficulty you get in, the 

279
00:17:04,400 --> 00:17:09,680
less able you are to keep things
cohesive in any meaningful 

280
00:17:09,680 --> 00:17:11,640
manner. 
Because as you noted, they'll 

281
00:17:11,640 --> 00:17:14,319
get a little signal and just go 
off the deep end. 

282
00:17:14,560 --> 00:17:16,560
Do you yeah. 
Do you think the the sub agent 

283
00:17:16,560 --> 00:17:20,400
pattern might be a good 
mitigation because the rather 

284
00:17:20,400 --> 00:17:23,359
than like the the entire context
window with the risk of 

285
00:17:23,359 --> 00:17:26,319
pollution, it's almost just like
a a subset a fully separated 

286
00:17:26,319 --> 00:17:28,640
contact window. 
So maybe the blaster easiest can

287
00:17:28,640 --> 00:17:31,360
be isolated. 
It's it's a mitigation, but it 

288
00:17:31,360 --> 00:17:33,360
also depends on how you're 
applying them, right? 

289
00:17:33,360 --> 00:17:37,720
So more is not always better. 
There's a lot of people who just

290
00:17:37,720 --> 00:17:43,000
say throw more compute at it. 
But the how is really important 

291
00:17:43,000 --> 00:17:49,440
in that because I've tried just 
blowing, blowing it wide by 

292
00:17:49,440 --> 00:17:53,200
having like 10 sub agents review
things from different angles and

293
00:17:53,200 --> 00:17:57,360
then have like a, a meta 
analysis sub agent come through 

294
00:17:57,360 --> 00:18:00,560
and review their analysis and 
compile the most relevant 

295
00:18:00,560 --> 00:18:05,000
things. 
That that is also subject to 

296
00:18:05,120 --> 00:18:10,840
just the detritus, the detritus 
of communication accumulating 

297
00:18:11,080 --> 00:18:14,760
and weighing down the value of 
the outputs, right? 

298
00:18:15,760 --> 00:18:20,840
And it it can work if you have a
very clear rubric and the 

299
00:18:20,840 --> 00:18:25,080
grading and the information that
they're generating is very 

300
00:18:25,080 --> 00:18:30,560
limited in scope and if it can 
be amalgamated and analyzed 

301
00:18:30,960 --> 00:18:34,000
effectively in batches. 
But if you're asking for like 

302
00:18:34,000 --> 00:18:38,640
qualitative reviews and you're 
asking many sub agents to 

303
00:18:38,640 --> 00:18:41,680
provide qualitative reviews, 
it's going to be 

304
00:18:41,680 --> 00:18:44,040
counterproductive past two or 
three. 

305
00:18:44,240 --> 00:18:48,440
There's just too many different 
avenues for exploration in 

306
00:18:48,440 --> 00:18:54,040
pretty much every vein of 
interesting research, right. 

307
00:18:54,720 --> 00:18:57,960
And so the sub agents really 
shine when when they have 

308
00:18:57,960 --> 00:19:01,520
explicit tasks, when you know 
exactly what you need from them 

309
00:19:02,080 --> 00:19:05,720
and you know exactly the shape 
that they need to deliver it in 

310
00:19:05,960 --> 00:19:10,400
and exactly how much context 
they need to get to do that 

311
00:19:10,400 --> 00:19:13,200
effectively. 
It it's helpful to have like sub

312
00:19:13,200 --> 00:19:16,640
agent tasking files. 
You can build a sub agent for 

313
00:19:16,880 --> 00:19:22,200
every different task, but this 
is speaking in Cloud Code 

314
00:19:22,200 --> 00:19:25,920
specifically. 
But those official sub agents 

315
00:19:25,920 --> 00:19:33,240
will take up context just laying
around, just being active. 

316
00:19:33,400 --> 00:19:37,480
They'll take up context from the
context window when they don't 

317
00:19:37,480 --> 00:19:42,000
necessarily need to. 
So I will just create markdown 

318
00:19:42,080 --> 00:19:45,800
or YAML files that are 
effectively sub agent 

319
00:19:45,800 --> 00:19:50,720
instructions and point and just 
feed that file into the man line

320
00:19:51,280 --> 00:19:56,320
and say use use this as the sub 
agent to accomplish XY task. 

321
00:19:56,840 --> 00:20:00,600
And that has served me pretty 
well. 

322
00:20:00,960 --> 00:20:06,040
I I have a colleague who swears 
by having only one named sub 

323
00:20:06,040 --> 00:20:09,880
agent and that that is Karen. 
And Karen's job is to review the

324
00:20:09,880 --> 00:20:13,400
actions of every other agent and
sub agent and to see if they 

325
00:20:13,400 --> 00:20:16,280
meet quality standards and if 
not to report them to 

326
00:20:16,280 --> 00:20:19,280
management. 
And he swears by Karen. 

327
00:20:19,840 --> 00:20:23,640
He says she's been inordinately 
helpful to him in his 

328
00:20:23,640 --> 00:20:25,720
development endeavors. 
Pivot away from subway just 

329
00:20:25,720 --> 00:20:28,480
though MCP servers. 
Have you found any patterns or 

330
00:20:28,480 --> 00:20:31,080
anti patterns that are common 
with leveraging MCP in 

331
00:20:31,080 --> 00:20:35,200
workflows? 
I found MCP themselves to almost

332
00:20:35,200 --> 00:20:39,600
be an anti pattern. 
It's it's great to have a 

333
00:20:39,720 --> 00:20:43,160
unifying force, right? 
It's great to have a standard to

334
00:20:43,160 --> 00:20:46,880
rally around, to inspire people,
to give them ideas of what can 

335
00:20:46,880 --> 00:20:50,240
be done with the technology when
it's abstract otherwise, right? 

336
00:20:50,680 --> 00:20:53,680
For all of those reasons, MCP is
a boon to this world. 

337
00:20:53,880 --> 00:20:59,960
In practice, it tends to be much
more context than it's worth. 

338
00:21:00,160 --> 00:21:05,400
There are not many slim MCP 
servers out there. 

339
00:21:05,920 --> 00:21:11,320
There are a lot of different 
solutions that try to manage MCP

340
00:21:11,320 --> 00:21:20,000
server bloat and access and tool
use and permissions, but right 

341
00:21:20,000 --> 00:21:29,080
now it's just too many variables
in a package to meaningfully get

342
00:21:29,080 --> 00:21:32,680
reliable utility out of it for 
most situations. 

343
00:21:32,880 --> 00:21:36,080
Now, of course, certain people 
with certain limited scope 

344
00:21:36,080 --> 00:21:38,560
things are going to have MCP 
servers that make their lives a 

345
00:21:38,560 --> 00:21:41,640
lot easier. 
In practice, though, if it's 

346
00:21:41,640 --> 00:21:44,360
something that you're doing a 
lot, you could probably design 

347
00:21:44,360 --> 00:21:48,960
your own tool calls that beat 
the pants off of any MCP server.

348
00:21:49,720 --> 00:21:53,680
But just like anything else, 
there's a balance between how 

349
00:21:53,680 --> 00:21:56,360
much time people have to 
investigate things themselves, 

350
00:21:56,360 --> 00:22:00,320
build things out themselves, or 
troubleshoot and maintain the 

351
00:22:00,320 --> 00:22:02,880
things that they've built. 
Because you know, it's, it's 

352
00:22:02,880 --> 00:22:05,160
fair to say that that that is a 
lot of overhead. 

353
00:22:05,480 --> 00:22:09,800
But I, I would also just note 
that a lot of that is building 

354
00:22:09,800 --> 00:22:15,200
the skills and the, the muscles 
to better understand these 

355
00:22:15,200 --> 00:22:18,120
products. 
So even if you were to use an 

356
00:22:18,120 --> 00:22:22,760
MCP server later, you would be 
making better use of that MCP 

357
00:22:22,760 --> 00:22:25,880
server because you would 
fundamentally understand how the

358
00:22:25,880 --> 00:22:28,320
different layers interact with 
the models. 

359
00:22:28,320 --> 00:22:30,400
It's very important to control 
what's in your context window. 

360
00:22:30,960 --> 00:22:34,120
There's also the risk of getting
malicious MCP servers where 

361
00:22:34,120 --> 00:22:36,320
you're just like hey, browsing 
the web this one sounds cool, 

362
00:22:36,320 --> 00:22:38,720
NPM install. 
And then you really have no idea

363
00:22:38,720 --> 00:22:41,760
what tools it's introducing your
system and what access you give 

364
00:22:41,760 --> 00:22:43,360
it to. 
Because again, like you said, we

365
00:22:43,360 --> 00:22:44,960
just accept every parish that 
comes our way. 

366
00:22:45,000 --> 00:22:47,520
Yeah, same. 
Same reason you don't want to 

367
00:22:47,560 --> 00:22:52,000
just install every package that 
your coding tool recommends to 

368
00:22:52,000 --> 00:22:54,720
you because there's I don't know
if you've covered slop squatting

369
00:22:54,720 --> 00:22:59,880
before no, but so there there 
was domain squatting where 

370
00:22:59,880 --> 00:23:03,080
people would take domains that 
could be misspellings or like 

371
00:23:03,520 --> 00:23:06,280
look similar to other domains 
and use them for nefarious 

372
00:23:06,280 --> 00:23:09,000
activities. 
But bad actors discovered that 

373
00:23:09,000 --> 00:23:12,240
coding agents were regularly 
hallucinating the same package 

374
00:23:12,240 --> 00:23:13,800
names of things that didn't 
exist. 

375
00:23:14,240 --> 00:23:18,520
And so they would register 
package names that kept coming 

376
00:23:18,520 --> 00:23:21,360
up in hallucinations as 
malicious tools. 

377
00:23:21,600 --> 00:23:29,280
And so an enterprising vibe 
coder who doesn't know exactly 

378
00:23:29,400 --> 00:23:33,000
what they're looking for with 
the things that they're creating

379
00:23:34,120 --> 00:23:38,040
could be recommended or even 
just have the package installed 

380
00:23:38,440 --> 00:23:41,160
on, you know, for them by coding
agent. 

381
00:23:41,360 --> 00:23:44,760
And then someone's got remote 
access to their system and you 

382
00:23:44,760 --> 00:23:47,640
know, all their get keys. 
You got to be careful out there.

383
00:23:47,640 --> 00:23:51,720
The the Internet happens to be a
dangerous place it that that's 

384
00:23:51,760 --> 00:23:54,680
always been the case. 
It's just a different version of

385
00:23:54,680 --> 00:23:57,040
dangerous nowadays. 
Yeah, 'cause it's, it's almost 

386
00:23:57,040 --> 00:24:00,720
just like we have a different, a
different interface, a different

387
00:24:00,720 --> 00:24:03,360
filter. 
And because it's so new, there's

388
00:24:03,360 --> 00:24:06,440
just not the experience as to 
best practices. 

389
00:24:06,480 --> 00:24:10,120
And people just get blown away 
by the magical aspect of AI. 

390
00:24:10,120 --> 00:24:12,000
We're like, it can do anything. 
I'm just going to let it 

391
00:24:12,160 --> 00:24:14,960
continue to do everything. 
Yeah, that's, that's actually a 

392
00:24:14,960 --> 00:24:19,080
really common anti pattern I've 
seen just in user level of 

393
00:24:19,080 --> 00:24:22,840
adoption. 
It's if people have been getting

394
00:24:22,840 --> 00:24:26,120
nothing but success in their 
first few interactions with an 

395
00:24:26,160 --> 00:24:29,880
AI system, especially if it's 
like one of the the simpler 

396
00:24:29,880 --> 00:24:31,720
ones. 
They they put it through the 

397
00:24:31,720 --> 00:24:33,680
first few things that they can 
think of. 

398
00:24:34,040 --> 00:24:37,080
And it's hard when you don't 
have experience with these 

399
00:24:37,080 --> 00:24:41,000
systems to understand exactly 
how jagged the intelligence is 

400
00:24:41,280 --> 00:24:46,800
and things that seem like they 
should be very easy for a 

401
00:24:46,800 --> 00:24:49,120
magical robot brain to 
understand. 

402
00:24:49,120 --> 00:24:51,320
Like, you know, the the number 
of Rs and strawberry is the 

403
00:24:51,320 --> 00:24:53,800
classic example can be quite 
vexing. 

404
00:24:54,040 --> 00:24:58,840
And because people will get 
these initial victories that 

405
00:24:58,840 --> 00:25:04,040
hold up in like these, these 
limited explorations, they will 

406
00:25:04,040 --> 00:25:08,720
think that it just keeps working
that well in all the use cases. 

407
00:25:08,880 --> 00:25:11,240
And stepping from like the the 
company tier to the personal 

408
00:25:11,240 --> 00:25:14,240
tier. 
One example is I like using 

409
00:25:14,240 --> 00:25:16,760
Obsidian for my notes. 
I think it's it's very powerful.

410
00:25:16,920 --> 00:25:18,920
Love the idea of having markdown
that I can just pull anywhere 

411
00:25:18,920 --> 00:25:21,360
File over app Love it. 
I've started running cloud code 

412
00:25:21,360 --> 00:25:23,960
in my vault and it can generate 
dynamic dashboards for me. 

413
00:25:23,960 --> 00:25:27,920
It's a ton of fun. 
I'm cognizant that I don't want 

414
00:25:27,920 --> 00:25:30,200
to have all of my personal notes
in that vault. 

415
00:25:30,200 --> 00:25:33,040
Now I do have a separate 
personal and and work focused 1.

416
00:25:33,320 --> 00:25:35,880
Do you have any thoughts or best
practices around using these 

417
00:25:35,880 --> 00:25:38,000
coding agents outside of 
classical coding? 

418
00:25:38,040 --> 00:25:43,120
Yeah, I am also a big fan of 
using them to collect and 

419
00:25:43,120 --> 00:25:46,080
organize and make sense of your 
personal information. 

420
00:25:46,840 --> 00:25:51,520
I use repos myself. 
I open them with Obsidian. 

421
00:25:52,000 --> 00:25:59,160
So it's a very similar probably 
use pattern from from in terms 

422
00:25:59,160 --> 00:26:04,000
of consuming what's produced. 
I, I find that it's very helpful

423
00:26:04,000 --> 00:26:06,880
to organize my information into 
repos. 

424
00:26:06,880 --> 00:26:13,080
I also keep them separated into 
specific subject matters because

425
00:26:13,240 --> 00:26:16,520
I, I generally don't find it 
productive to have everything 

426
00:26:17,160 --> 00:26:22,720
accessible from every agent 
session because otherwise things

427
00:26:22,720 --> 00:26:26,200
can get a little bit off topic. 
I just like to control the 

428
00:26:26,200 --> 00:26:28,320
different ways things can go 
wrong, right? 

429
00:26:28,600 --> 00:26:33,440
And basically everything you 
give to an agent is something 

430
00:26:33,440 --> 00:26:37,280
that it thinks it can use to 
solve a given problem. 

431
00:26:37,440 --> 00:26:42,160
And so for that reason, I try to
be very judicious in what I 

432
00:26:42,160 --> 00:26:48,120
exposed to agents that I'm that 
I'm working with, because it's, 

433
00:26:48,240 --> 00:26:52,440
it's not just about the the data
privacy component of it. 

434
00:26:52,880 --> 00:26:57,040
It's much more about the cross 
contamination of data. 

435
00:26:58,040 --> 00:27:00,960
Context contamination I guess 
would be a better. 

436
00:27:01,560 --> 00:27:03,600
Turn for him. 
And that's one reason I have 

437
00:27:03,960 --> 00:27:07,760
memory turned off for Chachapi 
and Anthropic because I've seen 

438
00:27:07,760 --> 00:27:10,800
people where they'll test out 
like an image Gen. model Sora or

439
00:27:10,800 --> 00:27:14,480
whatever, and all of a sudden 
conversations that they've had 

440
00:27:14,480 --> 00:27:15,760
in the past kind of leaked their
way through. 

441
00:27:15,760 --> 00:27:18,120
Whether it's like changing the 
location or changing the vibe of

442
00:27:18,120 --> 00:27:19,720
it. 
And you just don't know what 

443
00:27:19,720 --> 00:27:21,600
it's going to pull if you don't 
control that system. 

444
00:27:21,960 --> 00:27:25,200
Same reason I turn off 
geolocation because if you're 

445
00:27:25,200 --> 00:27:28,360
asking it to generate a picture 
of you, sometimes it'll just 

446
00:27:28,360 --> 00:27:33,160
create like a a sign in the back
that's like the the city name of

447
00:27:33,160 --> 00:27:37,240
where Geo located you 2 and it's
like I didn't ask for that but 

448
00:27:37,240 --> 00:27:41,560
but it's in the context so it's 
relevant according to the model.

449
00:27:41,960 --> 00:27:45,480
And, and so just following that,
that train of thought using 

450
00:27:45,480 --> 00:27:49,120
these state-of-the-art 
proprietary models, they always 

451
00:27:49,120 --> 00:27:52,080
have the best performance, but 
you lose control over like the 

452
00:27:52,080 --> 00:27:53,840
entire setup. 
Like we don't know what goes in 

453
00:27:53,840 --> 00:27:55,320
the context. 
They always add their own system

454
00:27:55,320 --> 00:27:57,760
message. 
Are you a proponent for open 

455
00:27:57,760 --> 00:28:00,360
source models, or are they not 
at the level that they're 

456
00:28:00,360 --> 00:28:04,320
functional enough for you? 
I mean, I like some open source 

457
00:28:04,320 --> 00:28:07,640
models. 
It's, it's variable. 

458
00:28:07,760 --> 00:28:11,320
It's variable and it's, it's 
more, it's definitely a more 

459
00:28:11,320 --> 00:28:17,920
cumbersome way to do work for 
the the most part in that 

460
00:28:17,920 --> 00:28:22,280
they're not served as 
conveniently or if you're trying

461
00:28:22,280 --> 00:28:26,040
to run them locally, you need 
sufficient hardware to run the 

462
00:28:26,040 --> 00:28:27,880
level of quantization that 
you're looking for. 

463
00:28:28,080 --> 00:28:34,000
You can use open source or open 
weight models to accomplish a 

464
00:28:34,000 --> 00:28:37,160
great deal of very productive 
work. 

465
00:28:37,640 --> 00:28:42,640
Generally for the problems that 
I'm trying to solve, time is, is

466
00:28:42,640 --> 00:28:44,960
one of the more important 
factors. 

467
00:28:45,200 --> 00:28:49,440
And so the principle that I work
with is prove it out with the 

468
00:28:49,440 --> 00:28:52,360
big models. 
And then if you've got the time 

469
00:28:52,360 --> 00:28:55,120
and it makes economic sense, 
refine it. 

470
00:28:56,200 --> 00:29:00,840
Because you can accomplish 
insane things through 

471
00:29:00,840 --> 00:29:03,480
orchestration. 
If you if you build a lot of 

472
00:29:03,480 --> 00:29:07,840
scaffolding, you can squeeze so 
much horsepower out of tiny 

473
00:29:07,840 --> 00:29:14,360
models just by being 
extraordinarily clear and atomic

474
00:29:14,360 --> 00:29:19,160
about the context that you give 
them and the expectations of 

475
00:29:19,160 --> 00:29:22,200
what you get from them. 
And then just chain a lot of 

476
00:29:22,200 --> 00:29:27,520
those interactions into 
basically an agentic workflow. 

477
00:29:28,000 --> 00:29:32,880
And there there are insane 
efficiencies to be unlocked 

478
00:29:32,880 --> 00:29:37,400
there. 
The counter argument to that is 

479
00:29:37,880 --> 00:29:40,000
Richard Sutton's The Bitter 
Lesson. 

480
00:29:40,040 --> 00:29:45,320
I'm sure you've heard of it. 
Basically, there has never been 

481
00:29:45,640 --> 00:29:51,320
a cost curve for any product in 
existence that I'm aware of like

482
00:29:51,320 --> 00:29:54,840
there is for the inference of 
foundation models. 

483
00:29:55,000 --> 00:30:00,560
They cost for inference of 
powerful foundation models are 

484
00:30:00,560 --> 00:30:05,640
dropping to like one thousandth 
of what they were within a year,

485
00:30:07,040 --> 00:30:10,320
just because nobody wants the 
old model after the new model 

486
00:30:10,320 --> 00:30:12,160
comes out. 
And everybody's getting better 

487
00:30:12,160 --> 00:30:16,920
at at making their GP us work 
harder with less. 

488
00:30:17,280 --> 00:30:20,560
And so there's efficiencies 
being gained in every direction.

489
00:30:20,960 --> 00:30:24,080
And while it might be a fun 
project, and I would say it's 

490
00:30:24,080 --> 00:30:26,760
probably a really productive 
learning exercise for people who

491
00:30:26,760 --> 00:30:27,920
are just trying to get into 
things. 

492
00:30:27,960 --> 00:30:30,920
I mean, don't get me wrong, run 
your own models. 

493
00:30:30,920 --> 00:30:33,880
If you've got the hardware, if 
you got access to it, run your 

494
00:30:33,880 --> 00:30:35,920
own models. 
It is a fantastic learning 

495
00:30:35,920 --> 00:30:40,960
exercise, but if we're talking 
about trying to productionize 

496
00:30:40,960 --> 00:30:47,200
it, trying to save real world 
time and energy, it's it's 

497
00:30:47,200 --> 00:30:51,560
usually a better bet to just go 
with one of the leading 

498
00:30:51,560 --> 00:30:56,800
foundation models and go from 
there, frankly. 

499
00:30:56,880 --> 00:30:58,600
From the business perspective, 
we're on the same thing. 

500
00:30:58,600 --> 00:31:02,760
What would you consider any like
heuristics or rules around the 

501
00:31:02,760 --> 00:31:05,120
build versus buy debate? 
Because I know some companies 

502
00:31:05,120 --> 00:31:08,960
don't want to send their data to
external systems, external APIs 

503
00:31:09,720 --> 00:31:11,640
and some other ones want to fine
tune their models. 

504
00:31:11,680 --> 00:31:14,400
So do you have any ideas when 
people should explore, you know,

505
00:31:14,400 --> 00:31:17,200
just follow your path, use the 
big ones, get the job done, 

506
00:31:17,200 --> 00:31:19,720
figure out later versus bring 
everything in house and and 

507
00:31:19,720 --> 00:31:20,760
control it. 
I mean. 

508
00:31:20,960 --> 00:31:23,240
Yeah, build versus buy is a is a
whole spectrum. 

509
00:31:23,240 --> 00:31:26,440
And so it really depends on the 
people that you've got to work 

510
00:31:26,440 --> 00:31:30,240
with, right, and their, their 
workload and their availability 

511
00:31:30,240 --> 00:31:35,880
and how much you can afford to 
dedicate them towards building, 

512
00:31:36,200 --> 00:31:38,720
building the skills to build 
these systems out. 

513
00:31:39,080 --> 00:31:41,560
And that's going to be different
for every organization. 

514
00:31:41,680 --> 00:31:47,520
The one thing that I've seen is 
that you can build a system that

515
00:31:47,520 --> 00:31:52,800
is perfect on paper for exactly 
what people need to do. 

516
00:31:53,120 --> 00:31:59,600
And if it doesn't Click to the 
end user, how they are supposed 

517
00:31:59,600 --> 00:32:04,880
to split this into their day, 
their standard operating 

518
00:32:04,880 --> 00:32:09,080
procedure. 
If that is not painless, then 

519
00:32:09,120 --> 00:32:15,200
the entire, the entire exercise 
could have been a waste ChatGPT 

520
00:32:15,200 --> 00:32:21,560
Enterprise is absolutely worth 
the value of $40 per user per 

521
00:32:21,560 --> 00:32:24,960
month. 
That's, that's, that's just 

522
00:32:24,960 --> 00:32:29,160
true. 
But the value that you get out 

523
00:32:29,160 --> 00:32:36,640
of that is going to depend so 
much on how how you prepare the 

524
00:32:36,640 --> 00:32:41,120
people that are receiving it to 
use it most effectively and how 

525
00:32:41,120 --> 00:32:45,200
you support them to succeed and 
how you make them feel 

526
00:32:45,200 --> 00:32:48,800
supported. 
It's it's important to make sure

527
00:32:48,800 --> 00:32:52,720
that you understand your systems
before you make that decision. 

528
00:32:52,840 --> 00:32:55,160
Everybody's systems are going to
be different. 

529
00:32:55,160 --> 00:33:01,160
Everybody's data is different 
and there are really compelling 

530
00:33:01,160 --> 00:33:06,720
solutions that work for 90% of 
organizations that might be 

531
00:33:06,720 --> 00:33:10,480
fundamentally unsuitable for you
because of some nuance about how

532
00:33:10,480 --> 00:33:15,320
your systems work or even like 
customer agreements. 

533
00:33:16,360 --> 00:33:19,280
If you're need to be GDPR 
compliant, that's going to 

534
00:33:19,280 --> 00:33:22,040
substantially affect like how 
you're allowed to process what 

535
00:33:22,040 --> 00:33:24,920
data depending on who your 
clientele are. 

536
00:33:25,120 --> 00:33:29,400
It's, it's good to, if you don't
have in house expertise, it's 

537
00:33:29,400 --> 00:33:33,760
good to at least start with a 
consultant first. 

538
00:33:33,760 --> 00:33:37,680
Get your data and your processes
mapped before, before you do 

539
00:33:37,680 --> 00:33:40,400
anything else. 
You need to know where your data

540
00:33:40,400 --> 00:33:42,960
is coming from, what 
transformations are happening to

541
00:33:42,960 --> 00:33:48,720
it and where it's going and what
regulations it's subject to. 

542
00:33:49,320 --> 00:33:53,760
And as long as you have that 
down, then you will be able to 

543
00:33:53,760 --> 00:33:58,320
have a very productive and 
fruitful conversation with a 

544
00:33:58,320 --> 00:34:00,960
consultant or a team of 
consultants and they will be 

545
00:34:00,960 --> 00:34:05,680
able to very quickly lead you 
into the most productive paths. 

546
00:34:05,920 --> 00:34:09,480
And that might be to build or 
that might be to buy, depending 

547
00:34:09,480 --> 00:34:13,520
on what your situation is. 
But it's it's going to be a 

548
00:34:13,520 --> 00:34:16,440
unique decision for every 
organization. 

549
00:34:16,679 --> 00:34:19,080
In your experience with the 
consulting, where do you see 

550
00:34:19,080 --> 00:34:22,400
companies making common mistakes
or easy? 

551
00:34:22,679 --> 00:34:25,639
Besides getting their data set 
up, what other things could do 

552
00:34:25,639 --> 00:34:28,679
is prep work to get ready for 
for this transformation. 

553
00:34:28,800 --> 00:34:33,400
Permissions management is really
big because permissions 

554
00:34:33,400 --> 00:34:36,320
management is one of the 
trickier parts about managing 

555
00:34:36,320 --> 00:34:40,000
effective agents and agentic 
workflows, ensuring that they 

556
00:34:40,000 --> 00:34:41,639
can only see what they're 
supposed to see. 

557
00:34:41,679 --> 00:34:44,719
A lot of these preparations are 
just good business sense to 

558
00:34:44,719 --> 00:34:47,920
start with. 
So you need to ensure the 

559
00:34:48,520 --> 00:34:52,040
Providence of your data, like 
again, ensure that you know that

560
00:34:52,040 --> 00:34:55,199
everything that comes into this 
is accurate because the data 

561
00:34:55,199 --> 00:34:58,960
that that you're building these 
systems around is going to have 

562
00:34:58,960 --> 00:35:03,920
an outmoded impact. 
Organizations think that they'll

563
00:35:03,920 --> 00:35:08,360
just take their help desk ticket
history and upload that and 

564
00:35:08,360 --> 00:35:12,600
they'll be able to use that to 
train an AI system to answer 

565
00:35:12,600 --> 00:35:15,440
help desk ticket requests. 
You might be able to 

566
00:35:15,440 --> 00:35:21,400
hypothetically do it that simply
in practice you're going to need

567
00:35:21,720 --> 00:35:26,960
to have probably a small team of
people pouring through that for 

568
00:35:26,960 --> 00:35:31,760
weeks at least to audit and edit
these files. 

569
00:35:31,760 --> 00:35:36,080
Ensure that extraneous 
information is removed, and 

570
00:35:36,080 --> 00:35:40,600
ensure that the responses are 
graded so that the models you're

571
00:35:40,600 --> 00:35:42,240
training off of that will know 
what is. 

572
00:35:42,240 --> 00:35:46,080
It is not a good response. 
You can just feed it all your 

573
00:35:46,080 --> 00:35:49,840
data, but it's the old adage 
garbage in, garbage out. 

574
00:35:50,400 --> 00:35:55,080
It's just a lot sneakier. 
When you don't see the garbage 

575
00:35:55,080 --> 00:36:00,080
going in into these systems, it 
can be much harder to know how 

576
00:36:00,080 --> 00:36:01,560
and where it's going to come 
out. 

577
00:36:01,720 --> 00:36:04,040
One thing I want to make sure we
touched on that we've chatted 

578
00:36:04,040 --> 00:36:05,840
about that I thought was great 
is context rot. 

579
00:36:05,840 --> 00:36:07,880
So could you explain to the 
audience what it is, why this 

580
00:36:07,880 --> 00:36:10,400
should look out for it, and why 
it's such a big failure point? 

581
00:36:10,520 --> 00:36:16,880
So context rot is something 
that's been around since people 

582
00:36:16,880 --> 00:36:20,160
have been using generative AI. 
We didn't always have a term for

583
00:36:20,160 --> 00:36:22,080
it. 
Chroma just released a really 

584
00:36:22,080 --> 00:36:25,120
good research paper and 
explainer video that I would 

585
00:36:25,120 --> 00:36:26,800
recommend anybody to go check 
out. 

586
00:36:28,560 --> 00:36:36,320
Basically every token that goes 
into a prompt or a response is a

587
00:36:36,320 --> 00:36:41,280
point of failure, right? 
Because every token has a a 

588
00:36:41,280 --> 00:36:44,320
weight to it that pulls the 
response in One Direction or 

589
00:36:44,320 --> 00:36:47,560
another, and if it's not 
explicitly leading towards the 

590
00:36:47,600 --> 00:36:51,800
outputs that you're looking for,
it's detracting from from the 

591
00:36:51,800 --> 00:36:56,600
response. 
And the more tokens, the greater

592
00:36:57,560 --> 00:37:04,520
the cumulative effect. 
So quad sonnet by default. 

593
00:37:05,240 --> 00:37:08,160
I know it has a bigger one 
available, but by default it has

594
00:37:08,200 --> 00:37:13,720
a 200,000 token context window. 
I try to keep my operations 

595
00:37:13,720 --> 00:37:19,200
under 70,000 tokens when I'm 
processing data if it's hard to 

596
00:37:19,200 --> 00:37:22,000
explain exactly, but when it's 
on a roll, sometimes I'll push 

597
00:37:22,000 --> 00:37:28,920
that to 100,000, but much beyond
100,000 and I'm starting a new 

598
00:37:28,920 --> 00:37:30,880
session and and working from 
there. 

599
00:37:30,880 --> 00:37:36,280
I don't use automatic compact I 
I like Dexter Horthy's 

600
00:37:36,360 --> 00:37:39,760
purposeful compaction practice 
where it's basically the human 

601
00:37:39,760 --> 00:37:42,480
understands what is it 
important, what is and is not 

602
00:37:42,480 --> 00:37:47,160
important in a prompt, and the 
human takes the important stuff 

603
00:37:47,480 --> 00:37:49,840
and brings it into a new chat 
session. 

604
00:37:50,200 --> 00:37:55,200
But I I digress. 
So beyond 100,000 tokens, you 

605
00:37:55,200 --> 00:38:01,200
start to see a rapid increase in
hallucinations and distractions 

606
00:38:01,200 --> 00:38:05,840
and erroneous responses, and 
certain things in the training 

607
00:38:05,840 --> 00:38:12,040
data become more pronounced. 
For example, sonnet 4.5 seems to

608
00:38:12,040 --> 00:38:14,280
be really convinced it's still 
2024. 

609
00:38:15,040 --> 00:38:17,480
I don't know why that is, but 
the further into the context 

610
00:38:17,480 --> 00:38:20,040
window you are, the more 
prominent that is. 

611
00:38:20,400 --> 00:38:23,720
And so there's just these little
things that keep going more and 

612
00:38:23,720 --> 00:38:25,560
more wrong. 
And by the time you get to the 

613
00:38:25,560 --> 00:38:31,560
end of the context window, it's 
almost always unreliable outputs

614
00:38:31,560 --> 00:38:34,720
that you're getting. 
So it not every interface makes 

615
00:38:34,720 --> 00:38:37,360
it really easy to see how much 
of context window you're using. 

616
00:38:37,360 --> 00:38:39,840
And I think that's a failure of 
providers and I think that that 

617
00:38:39,840 --> 00:38:43,920
is going to improve over time as
people start to become educated 

618
00:38:43,920 --> 00:38:46,600
about what is and is not a good 
experience with these models. 

619
00:38:46,800 --> 00:38:51,760
Cloud Code has the slash context
command, Codex CLI has just the 

620
00:38:51,960 --> 00:38:53,520
context percentage at the bottom
there. 

621
00:38:53,760 --> 00:38:58,280
So whatever interface you're 
using, try to understand how 

622
00:38:58,280 --> 00:39:00,560
much of the context window 
you've been using. 

623
00:39:00,800 --> 00:39:03,400
And if it's been, if your chat's
been going for a while, just 

624
00:39:03,400 --> 00:39:06,800
start a new chat. 
You'll get better outputs. 

625
00:39:07,520 --> 00:39:10,760
And again, watch the Chroma 
video called Context Rot on 

626
00:39:10,760 --> 00:39:14,960
YouTube. 6 minutes long or so. 
Definitely worth your time and 

627
00:39:15,040 --> 00:39:17,280
pretty charts and graphs. 
They can explain it to you. 

628
00:39:17,280 --> 00:39:18,840
Pretty much anyone. 
We'll make sure. 

629
00:39:18,840 --> 00:39:20,720
It's linked down below. 
Kyle, this would be great. 

630
00:39:20,720 --> 00:39:22,400
Before we let you go, is there 
anything you want the audience 

631
00:39:22,400 --> 00:39:26,040
to? 
Know you have to be exercising 

632
00:39:26,040 --> 00:39:31,280
your own human reasoning and 
judgement at all times if you're

633
00:39:31,280 --> 00:39:34,360
making systems with these 
models. 

634
00:39:35,240 --> 00:39:38,720
There is no substitute for the 
human in the loop, and the human

635
00:39:38,720 --> 00:39:41,520
who's in the loop needs to know 
what they're doing in the loop 

636
00:39:42,480 --> 00:39:46,000
for the loop to succeed. 
I can tell people to educate 

637
00:39:46,000 --> 00:39:50,640
themselves, join communities, 
learn the stuff's all available 

638
00:39:50,640 --> 00:39:53,280
to you guys. 
There's no there's no reason 

639
00:39:53,640 --> 00:39:59,000
that anybody can't be an expert 
in this stuff in short order 

640
00:39:59,000 --> 00:40:01,400
with the resources that are 
freely available to everybody 

641
00:40:01,400 --> 00:40:02,760
today. 
Thank you for this my 

642
00:40:02,760 --> 00:40:05,480
conversation with Kyle Clark. 
One thing I want to highlight is

643
00:40:05,480 --> 00:40:07,360
that Kyle and I met through the 
tool use Discord. 

644
00:40:07,560 --> 00:40:09,200
There's a lot of great 
conversations there. 

645
00:40:09,200 --> 00:40:11,400
I think you'd really enjoy it, 
so I encourage you to join. 

646
00:40:11,400 --> 00:40:14,320
The invite link is down below 
and I will see you next week.

