1
00:00:00,320 --> 00:00:06,920
All right, here we go. 
Welcome to the Embedded AI 

2
00:00:06,920 --> 00:00:09,440
Podcast. 
I'm one of your hosts, a real 

3
00:00:09,440 --> 00:00:13,480
person, host Ryan Torvick. 
And I'm Luke and Johnny. 

4
00:00:14,440 --> 00:00:20,280
And yeah, we're, we're here to 
talk about AI and embedded space

5
00:00:20,320 --> 00:00:24,240
all the way from fog computing 
through vibe coding and beyond. 

6
00:00:25,120 --> 00:00:28,360
And today we are talking about 
what were you talking about? 

7
00:00:29,200 --> 00:00:30,880
We're talking about context 
management, I think. 

8
00:00:31,600 --> 00:00:33,600
Context management. 
OK, all right. 

9
00:00:34,400 --> 00:00:37,520
It's like I've got a bunch of 
tokens and then I put the tokens

10
00:00:37,520 --> 00:00:40,360
in the machine and then the 
machine does things. 

11
00:00:40,360 --> 00:00:42,760
That's context. 
Management, I think, I think 

12
00:00:42,760 --> 00:00:50,160
that's A1 armed bandit. 
So all right, let's let's let's 

13
00:00:50,760 --> 00:00:52,960
set some context for the 
listeners, right? 

14
00:00:55,280 --> 00:00:56,720
Oh, this is going to be. 
This is going to be good. 

15
00:00:56,720 --> 00:00:59,640
This is going to be good. 
OK, let's set some context up. 

16
00:01:00,240 --> 00:01:02,720
Yeah. 
So what, what are we talking 

17
00:01:02,720 --> 00:01:04,480
about when we're talking about 
context management? 

18
00:01:04,480 --> 00:01:08,280
And you mentioned those, those 
tokens, what, what is a token in

19
00:01:08,480 --> 00:01:13,400
like the context of LLMS? 
And the context of LLMS, Oh my 

20
00:01:13,560 --> 00:01:17,200
gosh. 
So I mean a token like it's like

21
00:01:17,240 --> 00:01:20,200
a unit. 
It's like a unit of, of, of 

22
00:01:20,200 --> 00:01:23,080
processing power that you've 
used and that you're then 

23
00:01:23,080 --> 00:01:26,720
holding on to. 
And if you think about like, you

24
00:01:26,720 --> 00:01:28,760
know, if you're writing, writing
something down on a piece of 

25
00:01:28,760 --> 00:01:32,000
paper, like the more you write 
on the piece of paper, the more 

26
00:01:32,000 --> 00:01:33,960
the paper gets filled up, the 
harder it is to understand it. 

27
00:01:33,960 --> 00:01:36,360
And so if you think about a 
token as I've written this 

28
00:01:36,360 --> 00:01:39,600
number down on the paper, and 
then every time you write 

29
00:01:39,600 --> 00:01:41,920
another token, you're you're 
kind of filling up that paper 

30
00:01:41,920 --> 00:01:44,960
with more stuff. 
Yeah, so, so in other words, I 

31
00:01:44,960 --> 00:01:49,920
guess it's, it's like 1, one 
element that an element LLM 

32
00:01:50,000 --> 00:01:52,400
thinks about, right. 
So you know, if you, if you type

33
00:01:52,400 --> 00:01:55,920
a sentence into the LLM, then it
will break it down into tokens. 

34
00:01:56,240 --> 00:01:59,440
And a token may be a word, but 
it may also be a fraction of a 

35
00:01:59,440 --> 00:02:02,000
word. 
I guess it could even be several

36
00:02:02,000 --> 00:02:04,240
small words together. 
But I guess that's fairly 

37
00:02:04,600 --> 00:02:06,440
unusual. 
So anyway. 

38
00:02:07,160 --> 00:02:10,440
It seems like the that bumps in 
the chunking that we talked 

39
00:02:10,440 --> 00:02:13,240
about with rags. 
It's like how you break up the. 

40
00:02:14,320 --> 00:02:17,800
How you break up the data that 
you're receiving determines what

41
00:02:17,800 --> 00:02:20,560
your tokens are. 
Yeah, true, but but in the 

42
00:02:20,560 --> 00:02:25,760
context of of regular LLMS, it's
going to be words or fractions 

43
00:02:25,760 --> 00:02:27,120
of words. 
OK. 

44
00:02:27,800 --> 00:02:30,880
Is is going to be a token, So 
anything you input into the LLM 

45
00:02:30,880 --> 00:02:35,200
will be split up into tokens and
those will be fed into the 

46
00:02:35,200 --> 00:02:40,160
neural network and then it does 
it's magic and and outcome more 

47
00:02:40,160 --> 00:02:44,280
tokens, right. 
So anything dumps will be will 

48
00:02:44,280 --> 00:02:46,240
be. 
Sometimes you can observe that, 

49
00:02:46,240 --> 00:02:50,040
like if you're watching ChatGPT 
or something generate it will 

50
00:02:50,040 --> 00:02:54,320
sort of generate words piece by 
piece and each piece that you 

51
00:02:54,320 --> 00:02:58,440
see is a single token. 
So sometimes, sometimes it will 

52
00:02:58,440 --> 00:03:02,160
spit out entire words in in in 
one go, but also sometimes you 

53
00:03:02,160 --> 00:03:06,320
can sort of observe it, build a 
word out of 2-3 chunks and 

54
00:03:06,320 --> 00:03:07,040
that's. 
Why? 

55
00:03:07,280 --> 00:03:08,360
Each one of them will have been 
a talk. 

56
00:03:09,000 --> 00:03:11,320
OK. 
And and that's why when you see 

57
00:03:11,320 --> 00:03:13,840
like you see input tokens and 
output tokens. 

58
00:03:14,280 --> 00:03:17,120
And so the input tokens are what
I'm sending up to the LLM, and 

59
00:03:17,120 --> 00:03:19,440
the output tokens are what's 
coming back from the LLM. 

60
00:03:19,880 --> 00:03:20,960
That's exactly right. 
Yeah. 

61
00:03:20,960 --> 00:03:25,120
So you in other words, a token 
is, is the only thing that an 

62
00:03:25,120 --> 00:03:27,400
LLM can work with in the 1st 
place. 

63
00:03:27,400 --> 00:03:32,800
Like you pass in tokens and then
tokens sort of percolate through

64
00:03:32,800 --> 00:03:36,160
the layers and then outcome 
tokens at the other end. 

65
00:03:36,640 --> 00:03:40,040
And in in the LMS that we used 
to, they are large language 

66
00:03:40,040 --> 00:03:44,160
models, so they they tend to 
work on input tokens that that 

67
00:03:44,200 --> 00:03:47,080
represent language. 
And if you're, and if you're 

68
00:03:47,080 --> 00:03:51,320
putting money into this, this 
process, you're generally being 

69
00:03:51,320 --> 00:03:53,680
charged by the number of tokens 
you're processing. 

70
00:03:54,400 --> 00:03:56,160
That's exactly right. 
So you're being charged, but 

71
00:03:56,160 --> 00:04:00,320
also maybe more importantly for 
the context of context 

72
00:04:00,320 --> 00:04:06,320
management is that I promise 
that wasn't even intentional 

73
00:04:06,320 --> 00:04:08,440
this. 
Is such a useful word. 

74
00:04:09,720 --> 00:04:13,320
It is, isn't it? 
The the point is that all LLMS 

75
00:04:13,320 --> 00:04:15,520
have what is called a context 
window. 

76
00:04:15,520 --> 00:04:17,519
It's sort of their working 
memory, right? 

77
00:04:17,519 --> 00:04:22,360
It is what they can keep in the 
in the electronic brains and and

78
00:04:22,360 --> 00:04:25,800
reason about. 
And, and like the way that you 

79
00:04:25,800 --> 00:04:27,920
can think about it also as a 
human being, like as you're 

80
00:04:27,920 --> 00:04:30,560
trying to do work, you can only 
think about so many things at 

81
00:04:30,560 --> 00:04:34,680
the same time. 
And you, I'm sure have plenty of

82
00:04:34,680 --> 00:04:37,640
times going, like, OK, I just 
need to sit down and compress 

83
00:04:37,640 --> 00:04:39,320
this for a second. 
I'm going to go walk away from 

84
00:04:39,320 --> 00:04:42,480
them, from the keyboard and just
let this kind of sink in and 

85
00:04:42,480 --> 00:04:45,640
percolate in and kind of like 
get rid of the stuff that I 

86
00:04:45,640 --> 00:04:47,720
don't need to think about 
anymore and just focus on the 

87
00:04:47,840 --> 00:04:49,800
the important parts of the stuff
that I've just learned over the 

88
00:04:49,800 --> 00:04:52,880
past hour or two. 
Exactly. 

89
00:04:52,880 --> 00:04:56,520
And so the question is why did 
we want to talk about context 

90
00:04:56,520 --> 00:04:59,840
Management Today? 
The reason is that I find, or I 

91
00:04:59,840 --> 00:05:06,360
guess we find that we think more
about managing context even than

92
00:05:06,360 --> 00:05:10,960
than sort of crafting prompts. 
Or I mean, you of course prompts

93
00:05:10,960 --> 00:05:13,680
are exactly, you know, the input
tokens. 

94
00:05:14,280 --> 00:05:17,800
But the point is you know, the, 
the explicit instructions that 

95
00:05:17,800 --> 00:05:22,920
it that you give to an LLM will 
often only be a very small part 

96
00:05:22,960 --> 00:05:25,960
of the overall instructions that
you pass to it. 

97
00:05:25,960 --> 00:05:29,480
And a much larger chunk of that 
will be context. 

98
00:05:29,480 --> 00:05:32,640
You know that the files that it 
is supposed to operate on, the 

99
00:05:32,680 --> 00:05:35,520
requirements that it's supposed 
to keep in mind, etcetera, 

100
00:05:35,520 --> 00:05:37,880
etcetera. 
And so, Long story short, I find

101
00:05:37,880 --> 00:05:41,960
myself worrying about context 
quite a lot. 

102
00:05:42,480 --> 00:05:44,840
Yeah. 
Well, and I think, I think the, 

103
00:05:45,000 --> 00:05:48,320
the, the problems that the 
biggest problem that you have 

104
00:05:48,440 --> 00:05:52,640
with not, if you're not managing
your context effectively is the,

105
00:05:52,920 --> 00:05:56,600
the hallucinations is that it 
starts spitting out stuff that 

106
00:05:56,600 --> 00:05:59,000
just doesn't make sense. 
It doesn't necessarily work. 

107
00:05:59,000 --> 00:06:01,400
It's not going to go in the 
direction you want to go. 

108
00:06:01,560 --> 00:06:05,400
And there's kind of a sweet spot
where you know, if you are early

109
00:06:05,400 --> 00:06:07,240
on, you don't have any context 
in there yet. 

110
00:06:07,240 --> 00:06:09,480
You're not really getting good 
results yet until you have 

111
00:06:09,480 --> 00:06:12,880
filled up enough context for it 
to, to be able to be effective 

112
00:06:12,960 --> 00:06:16,160
at actually giving you back the,
the data that you're looking 

113
00:06:16,160 --> 00:06:18,080
for. 
But then if you go over that 

114
00:06:18,080 --> 00:06:22,840
other upper boundary, then it 
starts like getting confused and

115
00:06:22,840 --> 00:06:24,800
starts it's like, oh, what what 
is happening? 

116
00:06:24,800 --> 00:06:27,320
Like this is not helpful. 
Like stop, stop being helpful. 

117
00:06:27,320 --> 00:06:30,000
And you want to shake it, but 
it's inside a computer. 

118
00:06:30,000 --> 00:06:31,880
I guess you can shake your 
computer, but that isn't really,

119
00:06:31,880 --> 00:06:35,240
I don't know, doesn't have that 
visceral like grabbing somebody 

120
00:06:35,240 --> 00:06:37,640
by the shirt and shaking them, 
you know, shaking his computer. 

121
00:06:37,640 --> 00:06:41,600
Stuff there, there, there was 
there was a statistic probably 

122
00:06:41,600 --> 00:06:45,920
one or two decades ago now about
how many people have have hit 

123
00:06:45,920 --> 00:06:48,480
the computer in frustration yet.
And it was a surprisingly large 

124
00:06:48,480 --> 00:06:53,880
number, probably like half of 
our computer users had admitted 

125
00:06:53,880 --> 00:06:58,040
to hitting their computer. 
And the study rather cheekily 

126
00:06:58,080 --> 00:07:00,360
added that most of them had hit 
the monitor, even though the 

127
00:07:00,360 --> 00:07:04,080
monitor really, you know, 
couldn't be blamed at all. 

128
00:07:05,600 --> 00:07:08,160
I will say my my wife does not 
like to be in the house while 

129
00:07:08,160 --> 00:07:10,960
I'm typing because she says I 
sound angry on the keyboard. 

130
00:07:14,360 --> 00:07:17,600
Right. 
So anyway that that's the thing,

131
00:07:17,600 --> 00:07:20,520
right? 
We, we need to give the LLS 

132
00:07:20,520 --> 00:07:23,600
stuff to work with. 
And you know, in the context of 

133
00:07:23,600 --> 00:07:26,480
embedded systems development, 
very often it will be source 

134
00:07:26,480 --> 00:07:30,000
code, for instance, or it will 
be requirements documentation or

135
00:07:30,240 --> 00:07:35,640
stuff along those lines. 
And then we want the LLM to work

136
00:07:35,640 --> 00:07:39,360
with that, essentially transform
it in some way, you know, write 

137
00:07:39,360 --> 00:07:43,000
code according to the 
specification or rewrite the 

138
00:07:43,000 --> 00:07:46,760
code according to certain goals 
or that sort of thing, extend 

139
00:07:46,760 --> 00:07:50,760
the code and it kind of, you 
know, the only thing it has to 

140
00:07:50,760 --> 00:07:55,120
work with at all is its context.
It's not like, like you and me 

141
00:07:55,120 --> 00:07:58,560
who might be aware of the 
entirety of like the, the source

142
00:07:58,560 --> 00:08:02,360
code repository. 
The LLM does not have a concept 

143
00:08:02,360 --> 00:08:04,720
of that. 
It will only understand what is 

144
00:08:04,800 --> 00:08:07,400
in its context. 
This is why context management 

145
00:08:07,400 --> 00:08:10,080
is is really crucial and and 
sometimes also really hard 

146
00:08:10,080 --> 00:08:13,720
because like if you give the LLM
the wrong things to work with. 

147
00:08:14,360 --> 00:08:16,840
Yep. 
You know, what are you going to 

148
00:08:16,840 --> 00:08:19,320
expect? 
Of course it's going to, you 

149
00:08:19,320 --> 00:08:22,760
know. 
Right, well, and even, you know,

150
00:08:22,760 --> 00:08:25,600
I've been working on like, OK, 
it's, and I'm watching it write 

151
00:08:25,600 --> 00:08:28,800
tests and we talked about the 
mocking stuff, you know, last 

152
00:08:28,800 --> 00:08:32,720
week and you know, I, I've 
watched it write tests and it's 

153
00:08:32,720 --> 00:08:34,679
like, and I'm looking at it 
like, why didn't you look, 

154
00:08:34,679 --> 00:08:37,559
there's a test suite directory 
right there. 

155
00:08:37,640 --> 00:08:41,200
Why didn't you put the tests in 
the test suite directory instead

156
00:08:41,200 --> 00:08:43,799
of just putting it on the in the
root of the project? 

157
00:08:43,799 --> 00:08:45,480
Why did you put in the root of 
the project not in the test 

158
00:08:45,480 --> 00:08:47,520
suite directory? 
Oh, it's because it didn't have 

159
00:08:47,520 --> 00:08:49,960
that in its context. 
I didn't specifically tell it to

160
00:08:49,960 --> 00:08:52,640
have that in its context. 
But then as you say, if you give

161
00:08:52,640 --> 00:08:54,560
it too much stuff, it starts 
dropping things out of its 

162
00:08:54,560 --> 00:08:55,720
hands. 
It's like it only fits so much 

163
00:08:55,720 --> 00:08:58,280
stuff in the bag. 
You can only carry so many logs 

164
00:08:58,280 --> 00:08:59,960
of wood before you start 
dropping them on the ground. 

165
00:09:00,480 --> 00:09:04,760
It's even funnier than that 
because you know, and LLM of 

166
00:09:04,760 --> 00:09:11,280
course, is also just a human. 
So, and maybe you know this 

167
00:09:11,280 --> 00:09:16,280
about about your fellow humans 
that we tend to remember 

168
00:09:16,280 --> 00:09:19,360
particularly well the things 
that was said sort of at the 

169
00:09:19,360 --> 00:09:23,040
start of conversation or at the 
end of a conversation. 

170
00:09:23,240 --> 00:09:27,040
But the stuff in the middle kind
of like, you know, isn't quite 

171
00:09:27,040 --> 00:09:28,880
as relevant. 
And interestingly enough, the 

172
00:09:28,880 --> 00:09:32,360
same is also true for LLM. 
So maybe now we're getting into 

173
00:09:32,400 --> 00:09:33,720
what context management is all 
about. 

174
00:09:33,720 --> 00:09:37,320
Like it's not even about what do
you put in there, but also. 

175
00:09:37,880 --> 00:09:40,480
When? 
Where do you put it and, and how

176
00:09:40,520 --> 00:09:42,720
do you relate the different 
parts together? 

177
00:09:43,160 --> 00:09:44,760
And so this is one of the magic 
tricks. 

178
00:09:45,000 --> 00:09:49,320
If you want the LLM to be to, to
pay particular attention to 

179
00:09:49,920 --> 00:09:53,320
something, you should have put 
it either at the beginning or at

180
00:09:53,320 --> 00:09:56,840
the end of your prompt. 
So it will be sort of very fresh

181
00:09:56,880 --> 00:10:00,680
in the LL Ms. mind, if you like.
Oh my gosh, I've never even 

182
00:10:00,680 --> 00:10:03,160
thought because we talked about 
that with like performances. 

183
00:10:03,160 --> 00:10:05,520
If you're, if you're playing a 
piece of music, make sure the 

184
00:10:05,520 --> 00:10:08,880
beginning is solid, make sure 
the end is solid. 

185
00:10:09,240 --> 00:10:10,880
And then when you get to the 
middle, that's fine. 

186
00:10:10,880 --> 00:10:13,600
Just just do as good as you can 
in the middle, but make sure the

187
00:10:13,680 --> 00:10:16,360
outside the edges are solid 
because that's what people are 

188
00:10:16,360 --> 00:10:18,400
going to remember. 
And to think about it that way 

189
00:10:18,400 --> 00:10:20,800
for the LLM too is like OK here 
and here. 

190
00:10:21,280 --> 00:10:24,480
OK, and, and you, and, and this 
is, you know, you can observe 

191
00:10:24,480 --> 00:10:28,240
this in, in many cases, like for
instance, I've got very strict 

192
00:10:28,840 --> 00:10:33,280
rules for my, for my LLM. 
Like I, I exclusively work in a 

193
00:10:33,280 --> 00:10:36,640
test driven style. 
And it's going to be very good 

194
00:10:36,640 --> 00:10:39,800
at following this directive the 
first couple of turns around. 

195
00:10:40,520 --> 00:10:43,480
And then at some point it will 
just start to forget because it,

196
00:10:43,480 --> 00:10:47,520
it really gets drowned out by 
the evolving context. 

197
00:10:47,520 --> 00:10:51,120
But you know, more and more 
stuff being in its tiny little 

198
00:10:51,120 --> 00:10:54,000
electronic brain. 
And then I need to remind it and

199
00:10:54,000 --> 00:10:57,440
say, well, stop. 
Here's the prompt, remember? 

200
00:10:57,560 --> 00:10:59,280
Oh, yes, yes, thank you. 
I you're right. 

201
00:10:59,280 --> 00:11:01,080
I should have done it. 
You're absolutely right. 

202
00:11:01,080 --> 00:11:03,200
That's you're completely right. 
Yeah. 

203
00:11:03,680 --> 00:11:07,240
So that is, I feel something 
that that trips people up for a 

204
00:11:07,240 --> 00:11:10,480
while until they recognize that 
this is something that happens. 

205
00:11:10,840 --> 00:11:14,160
You know, you think you've given
the LLM explicit instructions 

206
00:11:14,160 --> 00:11:19,800
and you're you expect computers 
to be. 

207
00:11:20,240 --> 00:11:22,240
Infinite. 
Yeah. 

208
00:11:22,240 --> 00:11:25,440
And you know, if you give a 
computer an instruction, then 

209
00:11:25,440 --> 00:11:28,080
you expect this instruction to 
be sort of absolute and be 

210
00:11:29,760 --> 00:11:32,880
absolutely followed and always 
being followed the same way. 

211
00:11:32,880 --> 00:11:35,360
But this is not true for LLMS. 
Because you. 

212
00:11:36,080 --> 00:11:38,760
Know they are stochastic 
machines in the 1st place and 

213
00:11:39,840 --> 00:11:43,800
and the context is a moving 
target and so certain things 

214
00:11:43,800 --> 00:11:48,680
start to move out of the context
get drowned out by other stuff. 

215
00:11:49,640 --> 00:11:51,440
I. 
Wonder, Yeah. 

216
00:11:51,960 --> 00:11:55,880
Yeah, yeah, yeah. 
I, I wonder because we are used 

217
00:11:55,880 --> 00:11:58,480
to doing deterministic 
development. 

218
00:11:58,480 --> 00:12:01,320
Like that's what I, I give it 
you this instruction, this if 

219
00:12:01,320 --> 00:12:04,080
statement happens and this for 
loop happens and it does it so 

220
00:12:04,080 --> 00:12:06,200
many times because that's 
exactly what I told you to do. 

221
00:12:06,640 --> 00:12:09,040
So you're expecting the computer
to behave deterministically 

222
00:12:09,040 --> 00:12:11,200
because that's how we're used to
determine interacting with the 

223
00:12:11,200 --> 00:12:13,320
computer, especially when 
writing software. 

224
00:12:13,560 --> 00:12:18,840
But then these LLMS, they don't.
While it's the same foundational

225
00:12:18,840 --> 00:12:21,360
stuff, there's a lot of 
randomness tossed into it that 

226
00:12:21,360 --> 00:12:23,360
makes it so that it is a less 
predictable thing. 

227
00:12:23,360 --> 00:12:26,160
So you have to remember that 
you're working with a non 

228
00:12:26,160 --> 00:12:28,840
deterministic solution now. 
Super cool. 

229
00:12:29,440 --> 00:12:32,160
Exactly. 
So this is, this is maybe 11 

230
00:12:32,160 --> 00:12:36,920
interesting insight into context
management expect that that your

231
00:12:36,920 --> 00:12:39,920
context is going to sort of 
deteriorate over over time and 

232
00:12:39,920 --> 00:12:42,880
the things are of particular 
importance to you, you need to 

233
00:12:42,880 --> 00:12:46,880
pull them back in periodically. 
Or you need to if you. 

234
00:12:47,000 --> 00:12:49,120
If you don't need it until 
later, don't give it until 

235
00:12:49,120 --> 00:12:53,000
later. 
Oh yes, very much don't that 

236
00:12:53,000 --> 00:12:55,560
that's a good point. 
That's that's also if I feel 

237
00:12:55,600 --> 00:12:59,160
sort of a rookie mistake 
throwing in all of the stuff the

238
00:12:59,160 --> 00:13:03,080
LLM might possibly need. 
And then it will, you know, it 

239
00:13:03,080 --> 00:13:05,920
will lose the plot, it will 
start missing the forest for the

240
00:13:05,920 --> 00:13:08,640
trees. 
So that's another important 

241
00:13:08,640 --> 00:13:12,880
aspect of context management, 
sort of restraining yourself in 

242
00:13:12,920 --> 00:13:15,680
what you give the LLM. 
Of course, if you don't give it 

243
00:13:15,880 --> 00:13:18,280
certain information, then it 
will obviously not be able to 

244
00:13:18,280 --> 00:13:21,480
work with that and will not give
you the results that you were 

245
00:13:21,480 --> 00:13:24,920
hoping for maybe. 
But if you give it that 

246
00:13:24,920 --> 00:13:27,160
information plus ten other 
things, then it will just be 

247
00:13:27,760 --> 00:13:30,760
overwhelmed. 
Yeah, yeah. 

248
00:13:31,560 --> 00:13:34,480
OK. 
So, so you know, what are some 

249
00:13:34,480 --> 00:13:38,920
of the ways that because, and 
it's, it's evolved over the last

250
00:13:38,920 --> 00:13:41,800
several months, I would say. 
So I think there was initially 

251
00:13:41,800 --> 00:13:44,760
the first time I was using COD, 
it had a button that you could 

252
00:13:44,760 --> 00:13:48,040
click and you could, it would 
say, OK, compress it, compress 

253
00:13:48,040 --> 00:13:50,120
all the context. 
And now I've noticed when I'm 

254
00:13:50,120 --> 00:13:52,400
using COD code, it's doing that 
automatically when it gets to a 

255
00:13:52,400 --> 00:13:54,680
certain threshold. 
So let's talk about compression.

256
00:13:54,680 --> 00:13:57,560
What is that? 
And I mean, it's not perfect. 

257
00:13:58,840 --> 00:14:02,760
No, but OK, so maybe, maybe we 
need to back up a little bit on 

258
00:14:03,560 --> 00:14:07,200
think about and think about what
happens as you're working and 

259
00:14:07,200 --> 00:14:09,960
iterating together with, let's 
say called code. 

260
00:14:10,240 --> 00:14:13,480
You know you're prompting it. 
It will generate a reply. 

261
00:14:14,120 --> 00:14:18,080
And now this old prompt and the 
old reply will be part of the 

262
00:14:18,080 --> 00:14:20,280
context. 
And then when you give a follow 

263
00:14:20,280 --> 00:14:23,640
up instruction, then you know, 
then it will have the old 

264
00:14:23,640 --> 00:14:26,040
instruction, the old reply and 
the new instruction. 

265
00:14:26,200 --> 00:14:28,400
And then it will generate the 
new reply and so on and so 

266
00:14:28,400 --> 00:14:31,640
forth. 
So it will start to fill up it's

267
00:14:31,640 --> 00:14:35,080
context window with the history 
of this conversation. 

268
00:14:35,560 --> 00:14:39,120
And so at some point, quite 
unsurprisingly, the context 

269
00:14:39,120 --> 00:14:42,720
window will will fill up. 
Right It is. 

270
00:14:43,440 --> 00:14:45,600
Modern LL NS are much better 
than than the other ones like 

271
00:14:45,600 --> 00:14:48,320
Chachi. 
BT3 I think had like 16 K 

272
00:14:48,320 --> 00:14:54,000
context window which is not bad 
anyway, but but now modern ones 

273
00:14:54,000 --> 00:15:00,160
have like 128 K, 200 Ki. 
Think what's it called? 

274
00:15:00,240 --> 00:15:05,200
And Gemini has 1,000,001 million
tokens of context window. 

275
00:15:07,000 --> 00:15:09,960
But even this big context window
will at some point of course 

276
00:15:09,960 --> 00:15:12,600
fill up. 
And then we get to what you 

277
00:15:12,600 --> 00:15:15,280
mentioned, compression or 
compaction, which is essentially

278
00:15:15,280 --> 00:15:16,840
just telling the LLM, you know 
what? 

279
00:15:17,000 --> 00:15:21,400
Summarize the entire field of 
the conversation, flush your 

280
00:15:21,400 --> 00:15:26,360
context window and stick that 
summary at the top so that we 

281
00:15:26,560 --> 00:15:29,840
don't have a heartbreak from the
from the previous conversation. 

282
00:15:30,240 --> 00:15:34,200
But we freed up our context 
window by by dropping stuff that

283
00:15:34,320 --> 00:15:38,880
we hopefully don't need anymore.
Well, and that's, that's we 

284
00:15:38,880 --> 00:15:41,280
hopefully don't need any more as
the interesting part because 

285
00:15:41,920 --> 00:15:44,800
generally the the LLM is the one
doing the compression. 

286
00:15:44,800 --> 00:15:47,440
So the LLM is the one deciding 
what it needs going forward and 

287
00:15:47,440 --> 00:15:48,640
what it doesn't need going 
forward. 

288
00:15:49,160 --> 00:15:52,400
And I've definitely seen 
equations where it's like, OK, 

289
00:15:52,400 --> 00:15:56,240
well thanks for summarizing. 
As a summary, you are taking 

290
00:15:56,240 --> 00:15:58,160
things out. 
That's just how summaries work. 

291
00:15:58,520 --> 00:16:01,600
You're you're leaving some of 
the context behind. 

292
00:16:02,560 --> 00:16:04,200
Exactly. 
And hopefully you're dropping 

293
00:16:04,200 --> 00:16:07,640
the right stuff. 
And of course, I this is part of

294
00:16:07,640 --> 00:16:09,480
it. 
Like you can for example, with 

295
00:16:09,480 --> 00:16:12,080
Cloud code, you can give it 
instructions when when 

296
00:16:12,160 --> 00:16:14,920
compacting the context, you can 
tell you know this is the 

297
00:16:14,920 --> 00:16:17,720
direction I want to go in the 
next couple of turns. 

298
00:16:18,000 --> 00:16:21,360
Be sure to leave ample 
information about that in there 

299
00:16:22,000 --> 00:16:24,840
because the LLM of course can't 
know what you intend until until

300
00:16:24,840 --> 00:16:29,440
no, unless you tell it unless 
you tell it Well, at least we 

301
00:16:29,440 --> 00:16:34,120
have compaction like because in 
Claude web you don't have a 

302
00:16:34,120 --> 00:16:37,920
compaction function. 
It drives me nuts because. 

303
00:16:39,360 --> 00:16:41,360
No. 
You know, I've, I've, I've been 

304
00:16:41,360 --> 00:16:46,040
working on something fairly 
sophisticated with it and you 

305
00:16:46,040 --> 00:16:48,680
know, that takes a while and 
it's just a lot of stuff and it 

306
00:16:48,680 --> 00:16:50,720
fills up the context window and 
then at some point it will say 

307
00:16:51,240 --> 00:16:53,000
context exhausted. 
Done. 

308
00:16:53,840 --> 00:16:54,520
Done. 
Yeah. 

309
00:16:56,320 --> 00:16:59,720
It is so frustrating. 
And then what I do, what do I 

310
00:16:59,720 --> 00:17:01,520
do? 
Like I'm stuck now. 

311
00:17:02,040 --> 00:17:05,440
Yeah, because you need, you 
need, you need the summary of 

312
00:17:05,440 --> 00:17:06,839
that context window and it's 
like. 

313
00:17:06,920 --> 00:17:11,760
Yeah, won't give it to you at 
least at least in some code like

314
00:17:11,760 --> 00:17:15,560
the the programming interface it
it has this explicit mechanism 

315
00:17:15,560 --> 00:17:18,240
for compaction. 
And in fact, like as you said, 

316
00:17:18,839 --> 00:17:21,599
it will auto compact when it 
runs out of context. 

317
00:17:22,359 --> 00:17:26,680
But I think it's wise to 
deliberately compact when you 

318
00:17:26,680 --> 00:17:30,720
know, now is a good, you know, a
good place in the conversation. 

319
00:17:31,080 --> 00:17:33,160
OK, fine, this particular task 
is done. 

320
00:17:33,400 --> 00:17:35,880
Let's compact now and start 
fresh. 

321
00:17:36,440 --> 00:17:38,800
Or in fact, very often I don't 
even compact. 

322
00:17:38,800 --> 00:17:43,400
I just say slash new as it goes 
in, in cloud and and I just I 

323
00:17:43,400 --> 00:17:46,760
flush the entire context and I 
say fine, we start with a clean 

324
00:17:46,760 --> 00:17:49,440
slate. 
I'm going to pull in the data 

325
00:17:49,440 --> 00:17:51,640
that I want for the next couple 
of turns. 

326
00:17:52,120 --> 00:17:55,920
Yeah. 
So Long story short, there is, 

327
00:17:55,920 --> 00:17:59,040
there is. 
You keep fiddling with the 

328
00:17:59,040 --> 00:18:04,920
context because our, you know, 
our ability to, to, to work with

329
00:18:04,920 --> 00:18:07,520
context is just much greater as 
humans. 

330
00:18:08,000 --> 00:18:12,320
And so we need to preselect for 
the LLM what it needs to know 

331
00:18:12,320 --> 00:18:16,400
right now versus what it can 
safely ignore at at this point. 

332
00:18:17,080 --> 00:18:19,160
And that's if you're willing to 
trust the compression. 

333
00:18:19,160 --> 00:18:22,160
And I've seen it as, as I 
mentioned, like it's, you know, 

334
00:18:22,240 --> 00:18:25,320
OK, in general. 
That's what I've done too, was 

335
00:18:25,320 --> 00:18:26,920
what you said was just I'm just 
going to start over. 

336
00:18:26,920 --> 00:18:30,040
I'm just going to go to a new, 
new window altogether and then 

337
00:18:30,040 --> 00:18:32,840
bring over what I think is 
important because I have not had

338
00:18:32,840 --> 00:18:36,000
a lot of success in letting that
compaction go. 

339
00:18:36,000 --> 00:18:38,080
And then going forward, you 
know, we, we were talking about 

340
00:18:38,080 --> 00:18:40,840
movies versus books and like the
books this big, the movies only 

341
00:18:40,840 --> 00:18:43,280
two hours long. 
It just compression in there 

342
00:18:43,520 --> 00:18:45,840
trying to get the story in and 
trying to get what's important 

343
00:18:45,840 --> 00:18:48,480
out of it to fit in there. 
And it's, and you can go back 

344
00:18:48,480 --> 00:18:50,800
and forth, which is better. 
But like really, you had to 

345
00:18:50,800 --> 00:18:53,120
compact it in order to get it 
into this new format. 

346
00:18:53,600 --> 00:18:57,320
So what about agents? 
So we've instead of compressing 

347
00:18:57,320 --> 00:19:00,960
things, what about agents? 
So we've got different like 

348
00:19:00,960 --> 00:19:03,760
robots that you can send out to 
do different tasks, right? 

349
00:19:04,440 --> 00:19:06,000
Yeah, but so let's talk about 
what? 

350
00:19:06,280 --> 00:19:08,040
When we're talking about agents,
what do we mean? 

351
00:19:08,040 --> 00:19:10,840
I feel like everybody throws 
around the word agent. 

352
00:19:11,400 --> 00:19:13,920
Well, yeah, we're doing an 
agentic AI to make agentic 

353
00:19:13,920 --> 00:19:17,400
decisions and for our context, 
so that we have a better context

354
00:19:17,400 --> 00:19:19,160
of what the agentic AI can do 
for us. 

355
00:19:19,720 --> 00:19:23,160
So, so right. 
So what is what is an agent as 

356
00:19:23,160 --> 00:19:26,720
far as you're concerned, you 
know, for this sort of the 

357
00:19:26,720 --> 00:19:28,000
podcast? 
Yeah. 

358
00:19:28,000 --> 00:19:31,960
So an agent is, is is just an 
instance of an lol, of an lol 

359
00:19:31,960 --> 00:19:34,480
that you're interacting with 
that's been given a specific 

360
00:19:34,480 --> 00:19:38,960
prompt ahead of time to say, OK,
this is the direction I want you

361
00:19:38,960 --> 00:19:41,680
to start and you're going to be 
pointed in this direction. 

362
00:19:41,680 --> 00:19:42,960
This is how I want you to 
behave. 

363
00:19:42,960 --> 00:19:45,400
And you give it that context, 
give it that context. 

364
00:19:45,400 --> 00:19:48,080
You give it a prompt at the 
beginning to say, like, set 

365
00:19:48,080 --> 00:19:50,800
yourself up to be so that you 
were intending to be this type 

366
00:19:50,800 --> 00:19:52,560
of person. 
You know, we, you do this with, 

367
00:19:52,600 --> 00:19:54,560
with I'm running, running 
JavaScript. 

368
00:19:54,560 --> 00:19:56,600
You do it a little bit where 
you're like, you are an expert 

369
00:19:56,600 --> 00:20:01,240
JavaScript engineer, you know, 
level 6 at a major JavaScript 

370
00:20:01,240 --> 00:20:04,560
writing company and you have 15 
years of experience doing this. 

371
00:20:04,560 --> 00:20:07,200
You know, you give it that kind 
of initial, I'm going to say 

372
00:20:07,200 --> 00:20:09,680
context again, you give that 
initial direction. 

373
00:20:12,000 --> 00:20:15,320
To kind of like frame what how 
it interprets the next set of of

374
00:20:15,320 --> 00:20:17,160
instructions that you give in? 
Yeah. 

375
00:20:18,120 --> 00:20:21,000
Exactly. 
So as you said, an agent will be

376
00:20:21,400 --> 00:20:24,560
a new instance essentially of of
an LM or you know, of a 

377
00:20:24,560 --> 00:20:30,680
conversation, let's say. 
And the purpose that that serves

378
00:20:30,680 --> 00:20:33,920
really is to to clearly separate
those two contexts. 

379
00:20:33,920 --> 00:20:36,560
Like you might have sort of a 
main thread of your conversation

380
00:20:36,560 --> 00:20:38,840
that you're working in and you 
don't want to. 

381
00:20:39,440 --> 00:20:42,440
OK, let's make this, let's make 
this more concrete. 

382
00:20:42,480 --> 00:20:46,880
For example, I have got an 
agent, I guess it would be 

383
00:20:46,880 --> 00:20:51,040
called, which is a test review 
agent. 

384
00:20:51,480 --> 00:20:56,040
So whenever new tests are 
generated, I run this, this test

385
00:20:56,040 --> 00:20:59,360
review agent and it looks, looks
the tests over according to 

386
00:20:59,920 --> 00:21:02,120
things that are important to me.
You know, I, I want my 

387
00:21:02,120 --> 00:21:03,840
assertions in a, in a particular
style. 

388
00:21:03,840 --> 00:21:08,480
I want a certain level of 
specificity versus elasticity. 

389
00:21:08,800 --> 00:21:11,840
You know, I, you know, I, I 
don't want my assertions to be 

390
00:21:11,840 --> 00:21:14,040
too strict. 
So the tests keep breaking left 

391
00:21:14,040 --> 00:21:16,720
and right, but I also don't want
them so, so loose that they end 

392
00:21:16,720 --> 00:21:20,560
up not being very actionable, 
etcetera, etcetera. 

393
00:21:20,560 --> 00:21:22,680
So there's, there's a lot of 
stuff in there I can't even 

394
00:21:22,680 --> 00:21:26,800
remember, but it's like 2-3 
pages worth of criteria and 

395
00:21:26,800 --> 00:21:30,760
examples that I've collected. 
And the point is I could take 

396
00:21:30,760 --> 00:21:34,120
this prompt and stick it into my
main conversation, but what 

397
00:21:34,120 --> 00:21:38,760
would happen is that I sort of 
dilute the context of this main 

398
00:21:38,760 --> 00:21:40,600
conversation, which is maybe 
concerned with, you know, 

399
00:21:40,880 --> 00:21:44,200
implementing a feature. 
And now I start talking about a 

400
00:21:44,200 --> 00:21:48,440
tangent, essentially, you know, 
talking about testing and what's

401
00:21:48,440 --> 00:21:51,320
important to me with with in 
relation to testing, etcetera, 

402
00:21:51,320 --> 00:21:54,560
etcetera. 
That will throw off the LLM. 

403
00:21:55,080 --> 00:21:58,440
But now it's nicely segregated. 
You know, there's, there's an 

404
00:21:58,440 --> 00:22:02,160
entirely new thread of 
conversation, a new context 

405
00:22:02,160 --> 00:22:05,080
window and that can turn through
its task. 

406
00:22:05,400 --> 00:22:07,920
And then just essentially like a
function call, right? 

407
00:22:08,120 --> 00:22:11,640
Return it's observations, you 
know, this test is not so good 

408
00:22:11,640 --> 00:22:13,880
because this that test is not so
good because that blah, blah, 

409
00:22:13,880 --> 00:22:16,320
blah. 
And return that to the main 

410
00:22:16,320 --> 00:22:19,720
conversation thread where it's 
sort of still on topic. 

411
00:22:20,280 --> 00:22:24,040
You know, we just implemented a 
feature and the the associated 

412
00:22:24,040 --> 00:22:25,800
tests and now we get feedback on
the tests. 

413
00:22:26,280 --> 00:22:28,720
And now we can use that feedback
in the context of the main 

414
00:22:28,720 --> 00:22:32,680
thread without polluting the 
main thread with, you know, this

415
00:22:33,120 --> 00:22:35,240
topic change. 
Yep, Yep. 

416
00:22:35,880 --> 00:22:37,920
Yeah. 
I think polluting is the right 

417
00:22:38,000 --> 00:22:40,400
is the right terminology there. 
Like, that is definitely what it

418
00:22:40,400 --> 00:22:42,040
feels like you're doing. 
You're just throwing so much 

419
00:22:42,040 --> 00:22:45,440
extra stuff in here that the the
main thread, if you were to 

420
00:22:45,440 --> 00:22:47,360
leave that in there, it's just 
going to get confused about 

421
00:22:47,360 --> 00:22:48,560
what, what am I supposed to do 
next? 

422
00:22:48,560 --> 00:22:51,640
What am I supposed to do next? 
When you talk about that recency

423
00:22:51,640 --> 00:22:55,240
bias, like, OK, I guess then I'm
going to look at tests and then 

424
00:22:55,240 --> 00:22:56,440
we're going to focus on that 
now. 

425
00:22:56,440 --> 00:22:59,400
Like, that's not what I want. 
And then you start shaking your 

426
00:22:59,400 --> 00:23:01,440
hands at it again. 
That's not what I wanted you to 

427
00:23:01,440 --> 00:23:04,200
do. 
Do what I meant, not what I told

428
00:23:04,200 --> 00:23:06,720
you. 
Yes, indeed. 

429
00:23:07,440 --> 00:23:11,240
So now we've, we've talked about
compaction, we've talked about 

430
00:23:11,240 --> 00:23:14,280
agents or sort of, you know, 
having side conversations 

431
00:23:14,280 --> 00:23:19,520
actually on the side. 
Maybe we should also talk about 

432
00:23:20,160 --> 00:23:23,440
other sources of of context. 
You know, one that we already 

433
00:23:23,440 --> 00:23:25,920
talked about is files. 
Of course you're going to in 

434
00:23:25,920 --> 00:23:28,160
the, you know, as you're 
programming, you're going to 

435
00:23:28,160 --> 00:23:30,720
pull in files. 
You know, this file is something

436
00:23:30,720 --> 00:23:33,760
that we're working with. 
This file contains, you know the

437
00:23:33,760 --> 00:23:38,480
constants that we are trying to 
use, you know in in this new 

438
00:23:38,480 --> 00:23:40,320
functionality. 
So we need to pull in that that 

439
00:23:40,520 --> 00:23:42,400
header file that defines the 
consents or whatever. 

440
00:23:43,200 --> 00:23:45,800
Well, there's also, so there's, 
there's code files, but then 

441
00:23:45,800 --> 00:23:47,760
there's also like design 
documents. 

442
00:23:47,760 --> 00:23:49,600
If I have a design document that
I'm talking about for the 

443
00:23:49,600 --> 00:23:52,760
overall product that we're 
working on, I can, I can say 

444
00:23:52,840 --> 00:23:56,720
bring in this design document or
bring in this section of the 

445
00:23:56,720 --> 00:23:58,800
design document. 
If I can specifically say I just

446
00:23:58,800 --> 00:24:00,640
want this part of the file, 
don't read the whole file. 

447
00:24:00,640 --> 00:24:02,720
I don't want you getting 
overwhelmed, but read this 

448
00:24:02,720 --> 00:24:05,760
section of this file and then 
bring that into your context 

449
00:24:05,760 --> 00:24:08,120
too. 
That is actually an excellent 

450
00:24:08,120 --> 00:24:10,600
point. 
Like I also do that, you know, 

451
00:24:10,640 --> 00:24:14,800
pull in design documents or 
specifications or whatever, but 

452
00:24:15,960 --> 00:24:18,800
most frequently you'd really 
only want a particular section 

453
00:24:18,800 --> 00:24:21,200
of it. 
So like you say, it makes sense 

454
00:24:21,200 --> 00:24:26,640
to say only pull in this section
and either you have a modern LLM

455
00:24:26,640 --> 00:24:29,920
and it's fairly good about 
specifically reading a chunk of 

456
00:24:29,920 --> 00:24:32,880
a file because remember, as far 
as the LLM concerned, there is 

457
00:24:32,880 --> 00:24:34,640
no, there is no such thing as a 
file. 

458
00:24:34,680 --> 00:24:37,880
It, it, it's all flat text. 
You know, it's, it's all been 

459
00:24:37,880 --> 00:24:41,720
sort of piped into standard 
input as far as the LLM is 

460
00:24:41,720 --> 00:24:44,120
concerned. 
And any notion of files are, are

461
00:24:45,200 --> 00:24:47,360
sort of arbitrary. 
So you can, wasn't it? 

462
00:24:47,520 --> 00:24:49,680
And you should be with. 
Jobs that did that. 

463
00:24:49,680 --> 00:24:51,160
He said that files are a thing 
in the past. 

464
00:24:51,160 --> 00:24:52,360
We should stop talking about 
files. 

465
00:24:55,960 --> 00:24:58,960
It's just input man. 
Just no, no, no no. 

466
00:24:59,000 --> 00:25:01,160
I I I'm a Unix nerd. 
As far as I'm concerned, 

467
00:25:01,160 --> 00:25:06,160
everything is a file go away. 
Oh my gosh. 

468
00:25:06,920 --> 00:25:09,160
Anyway, so. 
This is so. 

469
00:25:09,560 --> 00:25:13,040
This is. 
So I've even done like, so we 

470
00:25:13,040 --> 00:25:15,520
talk about design documents. 
I've even done like, OK, here's 

471
00:25:15,520 --> 00:25:18,920
my development guide for this 
project and like, well, I'll 

472
00:25:18,920 --> 00:25:21,080
have it, I'll have it. 
And I actually, I had to know, 

473
00:25:21,520 --> 00:25:24,720
you know, call it or whatever, 
dump out like, OK, based on the 

474
00:25:24,720 --> 00:25:27,320
prompt that I'm giving you and 
go through all the files that 

475
00:25:27,320 --> 00:25:31,520
I've got here, generate a, a 
design document or not a design 

476
00:25:31,520 --> 00:25:34,560
document, a developer's guide 
for this project. 

477
00:25:34,560 --> 00:25:37,360
So that it's got in there. 
This is where the, the, the 

478
00:25:37,360 --> 00:25:39,760
tests are. 
This is where the source code 

479
00:25:39,760 --> 00:25:41,400
goes. 
This is how you know, this is 

480
00:25:41,400 --> 00:25:43,440
where the documentation goes. 
This is kind of the format that 

481
00:25:43,440 --> 00:25:46,000
we're expecting to do. 
This is, these are the commands 

482
00:25:46,000 --> 00:25:47,200
you should run in order to build
it. 

483
00:25:47,200 --> 00:25:49,880
These are the commands that you 
should run in order to pretty 

484
00:25:49,880 --> 00:25:52,760
add up to make sure that the the
formatting and and like having 

485
00:25:52,760 --> 00:25:56,520
all that in there in a a 
developer's guide and then 

486
00:25:56,520 --> 00:25:58,880
telling it to ingest the 
developer's guide when it goes 

487
00:25:58,880 --> 00:26:00,880
to do something or at least a 
section of it. 

488
00:26:03,040 --> 00:26:06,760
Yeah, I actually, I go even 
further than that sometimes if 

489
00:26:06,760 --> 00:26:10,440
I've got a lot of context and I 
know it's going to be a 

490
00:26:10,440 --> 00:26:15,520
challenge to not overwhelm the 
LLM that I that I specifically 

491
00:26:15,520 --> 00:26:18,360
prepare sort of problem specific
documentation. 

492
00:26:18,360 --> 00:26:22,240
I say, fine, take out the 
relevant chunks from the overall

493
00:26:22,240 --> 00:26:27,160
documentation, stick them in a 
new like markdown document that 

494
00:26:27,280 --> 00:26:30,160
sort of distills the relevant 
information. 

495
00:26:30,720 --> 00:26:33,520
And then I start a new 
conversation, you know, flush 

496
00:26:33,520 --> 00:26:38,040
the prompt, flush the context 
and say, OK, pull in this this 

497
00:26:38,040 --> 00:26:41,760
document with relevant context 
and then let's start going over 

498
00:26:41,760 --> 00:26:42,920
it. 
Yeah. 

499
00:26:42,960 --> 00:26:46,680
So that that it can really go as
far as that. 

500
00:26:47,240 --> 00:26:49,960
The use of temporary files I 
think is a really important 

501
00:26:50,480 --> 00:26:53,120
thing to be comfortable with of 
like creating these temporary 

502
00:26:53,120 --> 00:26:57,080
files that contain small pieces 
of context that you'd like to 

503
00:26:57,080 --> 00:27:01,560
bring in from outside and bring 
it into your into your context. 

504
00:27:02,640 --> 00:27:07,160
So yeah, I mean, the downside is
I'm, I'm actually kind of 

505
00:27:07,160 --> 00:27:11,520
annoyed by, you know, how 
eagerly LLMS create context or 

506
00:27:11,760 --> 00:27:17,600
no, create new text files that 
sort of just clutter my project.

507
00:27:18,280 --> 00:27:21,560
But I've found that it's a good 
habit to just have a place 

508
00:27:21,560 --> 00:27:24,400
designated for them. 
It's like, maybe I'll do a 

509
00:27:24,440 --> 00:27:27,360
certain spike and then I'll 
create a directory for that 

510
00:27:27,360 --> 00:27:32,320
spike and then all of the spike 
specific documentation, 

511
00:27:32,320 --> 00:27:34,080
distillation, etcetera goes in 
there. 

512
00:27:34,440 --> 00:27:37,040
And then once I'm done, I know I
can just throw it away. 

513
00:27:38,000 --> 00:27:40,280
Yeah. 
Well, and the, the thing that I 

514
00:27:40,280 --> 00:27:45,360
always come back to because one 
of the, one of the problems that

515
00:27:46,520 --> 00:27:49,960
we have as a group, group of 
organisms of individuals trying 

516
00:27:49,960 --> 00:27:52,280
to accomplish the task is that 
sharing that institutional 

517
00:27:52,280 --> 00:27:54,400
knowledge about why did I make 
this decision? 

518
00:27:54,400 --> 00:27:56,000
Why did the code turn out this 
way? 

519
00:27:56,280 --> 00:27:59,160
And I see a lot of the, those 
temporary files is like 

520
00:27:59,440 --> 00:28:03,280
capturing that, that 
documentation of why we're, why 

521
00:28:03,280 --> 00:28:06,520
the LLM is going to do this 
specific task in this specific 

522
00:28:06,520 --> 00:28:08,720
direction. 
And I'm like, I'm really, I'm 

523
00:28:08,720 --> 00:28:11,240
like a hoarder. 
Like I want to store that file 

524
00:28:11,240 --> 00:28:13,600
so that we can look at it later 
so we can come back to it. 

525
00:28:13,920 --> 00:28:16,760
But then if I did that for every
temporary file they created so 

526
00:28:16,760 --> 00:28:21,200
it can kind of so it can 
contextualize itself when it's 

527
00:28:21,200 --> 00:28:24,040
trying to do a task like, then I
just be covered in in all these 

528
00:28:24,040 --> 00:28:26,880
extra documentation files that 
are probably out of date by the 

529
00:28:26,880 --> 00:28:31,960
time I read them next time. 
You might be interested to know 

530
00:28:31,960 --> 00:28:35,600
that I have an agent 
specifically dedicated to 

531
00:28:35,600 --> 00:28:39,000
distilling information from 
those temporary files and then 

532
00:28:39,000 --> 00:28:43,520
sticking them into permanent 
information files. 

533
00:28:43,760 --> 00:28:45,640
Like you know, in my in my 
documentation folder or 

534
00:28:45,640 --> 00:28:47,000
something. 
Yeah. 

535
00:28:47,960 --> 00:28:51,360
So this is one of the last steps
I do when I, when I close out a 

536
00:28:51,360 --> 00:28:54,560
particular feature that I 
implemented, I say, OK, now 

537
00:28:54,560 --> 00:28:58,320
let's let's look over all of the
things that we thought about and

538
00:28:58,320 --> 00:29:01,680
that we collected. 
And let's make sure to not lose 

539
00:29:01,680 --> 00:29:05,080
the things that that continue to
do to be relevant and throw out 

540
00:29:05,080 --> 00:29:07,240
the rest. 
And throw out the rest. 

541
00:29:09,200 --> 00:29:13,360
Compact the context into a nice 
file and then stave it off. 

542
00:29:16,240 --> 00:29:20,200
All right, what about so? 
So actually I've, I've done, 

543
00:29:20,400 --> 00:29:24,600
I've done agents with rue code, 
like rue code has a very clear 

544
00:29:24,600 --> 00:29:27,800
like, OK, it's got agents. 
You can configure them in the 

545
00:29:27,800 --> 00:29:30,920
GUI like there's a lot of 
specifically because it's it's 

546
00:29:30,920 --> 00:29:32,320
very agentic when it tries to do
that. 

547
00:29:32,320 --> 00:29:35,360
When I've done stuff with cloud 
code, I don't it's not as 

548
00:29:35,360 --> 00:29:37,400
straightforward. 
How do you actually launch 

549
00:29:37,400 --> 00:29:39,080
another agent to go do 
something? 

550
00:29:39,920 --> 00:29:46,920
So that I I'm I'm sort of 
laughing to myself because I'm 

551
00:29:46,960 --> 00:29:50,200
I'm fairly sick of the term 
agentic because everybody sort 

552
00:29:50,200 --> 00:29:53,440
of bandies is about. 
I'm going to embedded world 

553
00:29:53,440 --> 00:29:54,880
soon, man. 
Like we're going to talk about 

554
00:29:54,880 --> 00:29:58,400
agentic everything next week. 
Oh yeah. 

555
00:29:59,640 --> 00:30:04,200
But the the point is so Cloud 
Code specifically I think likes 

556
00:30:04,200 --> 00:30:08,760
to have its own sort of agentic 
workflow with in itself. 

557
00:30:09,040 --> 00:30:12,960
Tries to create a small To Do 
List and then kind of work 

558
00:30:12,960 --> 00:30:15,400
through that. 
Yeah, yeah, for sure. 

559
00:30:16,000 --> 00:30:18,320
Which is, I think, their 
interpretation of agentic. 

560
00:30:18,320 --> 00:30:20,760
Because, you know, like, let's 
be honest, nobody knows what it 

561
00:30:20,760 --> 00:30:22,360
means, which is why I keep 
asking. 

562
00:30:23,120 --> 00:30:25,000
When? 
When you say agentic, what do we

563
00:30:25,000 --> 00:30:27,880
mean? 
I think this is what code code 

564
00:30:28,000 --> 00:30:33,240
means, that it sort of creates a
plan for itself and then, you 

565
00:30:33,560 --> 00:30:37,200
know, feeds itself that plan 
sort of keeps keeps its own 

566
00:30:37,200 --> 00:30:41,800
little checklist of things handy
in its working memory and then 

567
00:30:41,800 --> 00:30:43,520
and then tries to work through 
it. 

568
00:30:43,720 --> 00:30:49,480
So it doesn't lose context, 
doesn't lose direction and is I 

569
00:30:49,520 --> 00:30:53,080
guess sometimes it works. 
But there is the other thing 

570
00:30:53,080 --> 00:30:57,000
which we already hinted at, 
which is explicit. 

571
00:30:57,520 --> 00:31:00,520
I could I guess you could call 
them agents, you know slash 

572
00:31:00,520 --> 00:31:02,640
commands which you can just 
define yourself. 

573
00:31:03,480 --> 00:31:08,200
They are just prompts that you 
store in, in a pre find 

574
00:31:08,200 --> 00:31:10,720
location. 
And if you call them using 

575
00:31:10,720 --> 00:31:14,120
slash, you know, whatever the 
prompts name is, then they will 

576
00:31:14,360 --> 00:31:20,760
split off a new chunk of context
and work within that context and

577
00:31:20,760 --> 00:31:25,880
then eventually return their 
output back to the main line 

578
00:31:26,400 --> 00:31:28,280
context window. 
So like a function call, it's, 

579
00:31:28,280 --> 00:31:30,480
it's really sort of like a 
function call, which is why I'm 

580
00:31:30,920 --> 00:31:33,120
redacting to call it an agent 
because I, I think of it like a 

581
00:31:33,120 --> 00:31:36,040
function and I've got functions 
like that for making commits or 

582
00:31:36,040 --> 00:31:39,040
for reviewing code or for 
distending documentation, 

583
00:31:39,040 --> 00:31:43,240
etcetera, etcetera, that are 
very handy and very explicitly 

584
00:31:43,240 --> 00:31:48,200
don't where I'm explicitly not 
interested in in the, in the 

585
00:31:48,360 --> 00:31:51,320
context that is generated as 
this agent runs. 

586
00:31:51,400 --> 00:31:54,680
I'm only interested perhaps in 
the output, perhaps not even 

587
00:31:54,680 --> 00:31:56,560
that. 
And I guess could I could do 

588
00:31:56,560 --> 00:31:59,200
that by hand by, you know, 
manually starting a new plot 

589
00:31:59,200 --> 00:32:02,120
code instance and sticking the 
prompt in there. 

590
00:32:02,400 --> 00:32:04,880
But this is just a convenient. 
For one thing it's a convenient 

591
00:32:04,880 --> 00:32:08,520
shorthand, and for another it 
allows you to get return values 

592
00:32:08,520 --> 00:32:10,600
from this quote UN quote 
function call. 

593
00:32:11,080 --> 00:32:12,240
Yep. 
All right. 

594
00:32:12,240 --> 00:32:14,160
So the slash commands are good 
and you can set those up. 

595
00:32:14,720 --> 00:32:17,840
What about we, we talked about a
couple times ago, we talked 

596
00:32:17,840 --> 00:32:20,160
about rag. 
So like, that's another way of 

597
00:32:20,160 --> 00:32:24,480
getting stuff into your context 
window from the outside world. 

598
00:32:25,120 --> 00:32:28,680
Yeah, and that and that can be 
very efficient because you don't

599
00:32:29,120 --> 00:32:32,080
ingest like an entire 
specification file you you 

600
00:32:32,080 --> 00:32:36,640
specifically pull in, you know 
only those two requirements that

601
00:32:36,640 --> 00:32:38,320
were the search hits from the 
rug. 

602
00:32:38,880 --> 00:32:42,640
And so that for one thing, it 
will be more more token 

603
00:32:42,640 --> 00:32:46,400
efficient, which is cheaper, 
faster and maybe more reliable. 

604
00:32:46,920 --> 00:32:50,040
And like you said, you know 
hopefully leads to few 

605
00:32:50,040 --> 00:32:51,480
hallucinations, right? 
Because it is. 

606
00:32:51,840 --> 00:32:57,600
It is less for the LLM to to 
latch onto, but you don't want 

607
00:32:57,600 --> 00:32:58,080
it to. 
Maybe. 

608
00:32:58,600 --> 00:33:01,680
Yeah. 
So there's rags, and then 

609
00:33:01,680 --> 00:33:05,800
there's the model context 
protocol MCP. 

610
00:33:06,840 --> 00:33:11,920
What the earth is that? 
That's a function call, right? 

611
00:33:14,520 --> 00:33:16,440
Yeah, but I guess it's more than
that, isn't it? 

612
00:33:17,240 --> 00:33:18,960
Oh, it gets you out. 
That's the. 

613
00:33:19,080 --> 00:33:23,280
The important thing is that it 
gets you executing whatever it 

614
00:33:23,280 --> 00:33:25,080
is you want. 
So if you can run it in a Docker

615
00:33:25,080 --> 00:33:29,400
container, then you can do it. 
And that's really what the NCP 

616
00:33:29,400 --> 00:33:32,040
does. 
Is it it it, it drops you out of

617
00:33:32,040 --> 00:33:36,240
the LLM into a Docker container 
that you can now run whatever 

618
00:33:36,520 --> 00:33:40,720
commands, whatever you can reach
out to, whatever web services 

619
00:33:40,720 --> 00:33:44,680
like, it just, it lets you out 
of the LLM and then it lets you 

620
00:33:44,680 --> 00:33:47,800
dump whatever you come out with 
back into the LLM. 

621
00:33:48,640 --> 00:33:50,120
Yeah, but I think it's more than
that. 

622
00:33:50,200 --> 00:33:54,560
Like, so firstly, you know, 
Docker container is is either 

623
00:33:54,560 --> 00:33:56,760
here or there, maybe maybe it 
uses a Docker container, maybe 

624
00:33:56,760 --> 00:33:57,920
it doesn't, doesn't really 
matter. 

625
00:33:58,200 --> 00:34:04,000
The point is that it is a 
standard interface that the LLM 

626
00:34:04,000 --> 00:34:08,159
understands and that contains 
its own explanation in ways that

627
00:34:08,159 --> 00:34:10,800
the LLM understands, which in 
many cases is just words, you 

628
00:34:10,800 --> 00:34:13,719
know, English language, this 
being a large language model 

629
00:34:14,360 --> 00:34:18,679
that tells you, OK, I am, you 
know, I am an MCP that can 

630
00:34:18,760 --> 00:34:22,440
interface with gyra forsake of 
argument. 

631
00:34:24,239 --> 00:34:28,600
And you know, I, I, I can, I, I 
can offer you the following 

632
00:34:28,600 --> 00:34:30,400
function. 
So it's, it's sort of like, like

633
00:34:30,400 --> 00:34:34,280
an API, except it isn't. 
It's an interaction interface 

634
00:34:34,280 --> 00:34:39,560
that is specifically designed 
for LLMS and enables them to 

635
00:34:39,560 --> 00:34:44,800
pull in information slash 
context from all kinds of other 

636
00:34:44,800 --> 00:34:49,920
sources without you having to 
like implement the Gyra API or 

637
00:34:49,920 --> 00:34:55,520
craft like curl commands and you
know, interact with, with with 

638
00:34:55,639 --> 00:34:58,920
Gyra that way and then maybe 
ingesting whatever curl returned

639
00:34:58,920 --> 00:35:03,360
and then trying to pause it. 
So it can be very, very 

640
00:35:03,360 --> 00:35:08,200
interesting because it's makes 
certain functionality much more 

641
00:35:08,200 --> 00:35:10,840
approachable to an LLM than it 
would otherwise be. 

642
00:35:11,520 --> 00:35:15,360
Yep. 
So, yeah, so the the MCP and 

643
00:35:15,360 --> 00:35:18,080
then there aren't like there are
a lot of people that have 

644
00:35:18,080 --> 00:35:20,640
published MCPS that you can go 
and just pull off of. 

645
00:35:20,640 --> 00:35:23,720
So, you know, if you want to use
a Jira MCP, you do not need to 

646
00:35:23,720 --> 00:35:25,200
go and write your own. 
There's one out there. 

647
00:35:25,520 --> 00:35:27,960
They have ones that are yeah, 
yeah. 

648
00:35:28,400 --> 00:35:30,280
You know, if you want to 
interact with Git, there's the 

649
00:35:30,280 --> 00:35:33,200
one that I've used a lot of this
brave search. 

650
00:35:33,360 --> 00:35:35,440
That's one thing that we 
haven't, that we haven't touched

651
00:35:35,440 --> 00:35:41,920
on is the fact that the, what 
the LLM knows about the history 

652
00:35:41,920 --> 00:35:45,960
of, you know, humanity or the 
world stopped at the day that it

653
00:35:45,960 --> 00:35:48,360
came out. 
And so if you ask it anything 

654
00:35:48,360 --> 00:35:51,000
that's like I, I've run into 
issues where I'm asking it to 

655
00:35:51,000 --> 00:35:55,480
use a library and it's giving 
me, you know, an old version of 

656
00:35:55,480 --> 00:35:57,240
the library that's two or three 
months old. 

657
00:35:57,560 --> 00:36:00,120
And I'm like, no, I want to use 
the new Golf. 

658
00:36:00,320 --> 00:36:03,600
And so I need to go tell it to 
use an MCP to go find the new 

659
00:36:03,600 --> 00:36:06,360
version of this library and then
the API for that new version of 

660
00:36:06,360 --> 00:36:08,200
the library and use that 
instead. 

661
00:36:08,600 --> 00:36:10,520
Because it doesn't know about 
it, because that's there was a 

662
00:36:10,520 --> 00:36:14,240
cut off date, it's birth date or
anything after its birth date, 

663
00:36:14,240 --> 00:36:19,720
which is backwards. 
But yeah, it doesn't know it. 

664
00:36:19,720 --> 00:36:22,680
So you have to, you have to 
augment it with like, OK, here's

665
00:36:22,680 --> 00:36:27,280
mod, here's more, more newer. 
There's more newer information 

666
00:36:27,280 --> 00:36:30,840
that you can then use to help 
you understand and solve the 

667
00:36:31,000 --> 00:36:32,600
problem. 
That's an interesting point, 

668
00:36:32,640 --> 00:36:34,800
right? 
You know, in in that sense, the 

669
00:36:34,800 --> 00:36:37,760
entire knowledge of the LLM can 
also be considered part of the 

670
00:36:37,760 --> 00:36:40,120
context, right? 
And it ages like you said. 

671
00:36:41,760 --> 00:36:45,760
And I've, I've spotted this in a
couple of places where it was 

672
00:36:46,120 --> 00:36:53,200
not, no, where it actually knew 
the new API for a given thing, 

673
00:36:53,600 --> 00:36:57,240
but that was so new that it 
tended to disregard it and use a

674
00:36:57,280 --> 00:37:01,000
deprecated API. 
And then the code would, and it 

675
00:37:01,000 --> 00:37:03,440
would try to compile and it 
would fail and then it would 

676
00:37:03,440 --> 00:37:06,840
say, oh, that's right, I need to
use the new one and then rewrite

677
00:37:06,840 --> 00:37:09,240
it in the proper way. 
And it kept doing that 5 or 6 

678
00:37:09,240 --> 00:37:13,120
times, trying to use the old 
one, watching it fail, saying, 

679
00:37:13,160 --> 00:37:15,360
oh, I need to use the new one, 
rewrite. 

680
00:37:15,360 --> 00:37:19,600
And then it worked. 
That's quite remarkable how it, 

681
00:37:20,400 --> 00:37:23,480
you know, it, its own memory, 
its own context was sort of 

682
00:37:23,480 --> 00:37:28,040
biased towards the the more 
familiar but now outdated 

683
00:37:28,040 --> 00:37:31,320
version, even though it was 
aware of the newer version and 

684
00:37:31,320 --> 00:37:34,720
presumably also sort of was 
aware that it was the newer 

685
00:37:34,720 --> 00:37:37,440
version. 
But it was, you know, it just it

686
00:37:37,440 --> 00:37:40,240
defaulted to the old version. 
Well, there's, there's 

687
00:37:40,240 --> 00:37:43,280
significantly more of the old 
versions data out there. 

688
00:37:43,280 --> 00:37:45,960
So if you think about, you know,
solutions on Stack Overflow, 

689
00:37:45,960 --> 00:37:48,200
you're going to find, well, 
hopefully not anymore, But like 

690
00:37:48,480 --> 00:37:51,640
there was a long time where you 
would find the solution to a 

691
00:37:51,640 --> 00:37:55,600
Python question in Python 2 much
more readily than you would find

692
00:37:55,600 --> 00:37:56,840
it in Python 3. 
Like. 

693
00:37:57,360 --> 00:38:00,720
And so I would make sense that 
if you are an LLM and you don't 

694
00:38:00,720 --> 00:38:03,320
really know the difference or 
why you should do one or the 

695
00:38:03,320 --> 00:38:06,680
other, that you're constantly 
trying to use Python 2 even 

696
00:38:06,680 --> 00:38:09,000
though you know, you as a human 
being or no, no, no, no, I have 

697
00:38:09,000 --> 00:38:11,120
to use Python 3 because of this,
this, this and this. 

698
00:38:11,600 --> 00:38:14,840
But like the the majority of the
data that was available. 

699
00:38:14,960 --> 00:38:17,320
So if you're just looking at 
quantity of data, it's for 

700
00:38:17,320 --> 00:38:19,360
Python 2. 
Hopefully that's not the case 

701
00:38:19,360 --> 00:38:20,960
anymore. 
Like, come on, get with the 

702
00:38:20,960 --> 00:38:22,440
Times, 2025. 
Let's go. 

703
00:38:22,600 --> 00:38:27,320
Yeah, but, but especially you as
security expert should also be 

704
00:38:27,320 --> 00:38:29,560
very familiar with the 
phenomenon that it that it 

705
00:38:29,560 --> 00:38:35,440
defaults to in many cases to to 
UN to insecure code just because

706
00:38:35,800 --> 00:38:40,520
that was in the in the examples 
that didn't make any attempt to 

707
00:38:40,520 --> 00:38:43,760
be secure. 
They just tried to be very clear

708
00:38:43,760 --> 00:38:46,760
and obvious and not cluttered. 
Yes, there's. 

709
00:38:47,120 --> 00:38:50,680
So, so widespread that the 
actual good way of doing it, the

710
00:38:50,680 --> 00:38:54,160
secure, the professional way of 
doing it is sort of drowned out 

711
00:38:54,160 --> 00:38:58,120
by not not even inferior code in
that sense, but but obviously 

712
00:38:58,120 --> 00:39:01,520
example code that actively 
disregarded all of those 

713
00:39:01,520 --> 00:39:05,880
concerns to gain clarity. 
Yeah, the, the security 

714
00:39:05,880 --> 00:39:09,320
implementations of using an LLM 
to do work at all is still 

715
00:39:10,960 --> 00:39:13,600
there's, there's a lot we can 
talk about it when it, when, 

716
00:39:13,600 --> 00:39:17,880
when, when it's time. 
Like even MCPS even, you know, 

717
00:39:18,360 --> 00:39:20,640
getting data into the LLM from 
the outside. 

718
00:39:20,880 --> 00:39:24,560
These, the, the companies that 
craft these LLMS have, have 

719
00:39:24,560 --> 00:39:27,680
spent a decent amount of time 
trying to make sure that the, 

720
00:39:28,560 --> 00:39:30,960
the data they're using and the 
results they're getting are not 

721
00:39:31,040 --> 00:39:35,120
polluted in a certain bias that 
is inappropriate for human 

722
00:39:35,120 --> 00:39:38,000
consumption. 
But when you were pulling in 

723
00:39:38,000 --> 00:39:40,680
random data from the Internet, 
you threw an MCP. 

724
00:39:40,680 --> 00:39:42,480
There's nothing stopping you 
from getting. 

725
00:39:42,480 --> 00:39:44,840
So there's just, there's just so
many security concerns that are 

726
00:39:44,840 --> 00:39:47,080
with all of the things that 
we're talking about. 

727
00:39:48,160 --> 00:39:51,320
But we're not even think talking
about things like, you know, 

728
00:39:51,560 --> 00:39:55,320
token injection or something. 
We're really talking about the 

729
00:39:55,320 --> 00:39:58,640
thing that that it generates. 
All right, let's talk other 

730
00:39:58,640 --> 00:40:01,960
source of context. 
So we had MCP but but also 

731
00:40:02,360 --> 00:40:05,920
something that can also be very 
useful is ingesting. 

732
00:40:06,400 --> 00:40:10,120
Things like log files or 
standard output of a process, 

733
00:40:11,240 --> 00:40:15,400
that can often be very helpful. 
But also, especially in the case

734
00:40:15,400 --> 00:40:18,680
of log files, it may pollute 
your context. 

735
00:40:18,680 --> 00:40:22,320
So in in many cases, if I, if I 
knew that I was looking for a 

736
00:40:22,320 --> 00:40:25,040
particular situation for a 
particular error message, for 

737
00:40:25,040 --> 00:40:29,240
instance, I would deliberately, 
for example, grab for this. 

738
00:40:29,240 --> 00:40:32,680
I'd say that worked particularly
well for example, with Ada, 

739
00:40:32,680 --> 00:40:36,000
where I could explicitly tell it
which command to run and then 

740
00:40:36,000 --> 00:40:37,560
afterwards ingest the standard 
output. 

741
00:40:38,880 --> 00:40:45,880
They, you know, slash run cat 
log file dot TXT pipe grip, you 

742
00:40:45,880 --> 00:40:47,880
know, error or you know, 
whatever. 

743
00:40:48,200 --> 00:40:53,480
And then I would have a nice and
compact input for my context 

744
00:40:53,480 --> 00:40:58,120
that firstly didn't didn't tax 
my context window size too much 

745
00:40:58,480 --> 00:41:02,960
and secondly didn't contain 
confusing information that was 

746
00:41:02,960 --> 00:41:05,200
not relevant to the to the 
situation at hand. 

747
00:41:05,960 --> 00:41:11,000
Yep, I think the I don't know, 
it was really surreal the first 

748
00:41:11,000 --> 00:41:16,400
time it happened where the LLM 
went to go run the command to 

749
00:41:16,400 --> 00:41:21,240
flash and compiled project to my
board and then went to run it 

750
00:41:21,520 --> 00:41:24,480
and was listening to the serial 
output. 

751
00:41:24,480 --> 00:41:27,040
It was reading the serial output
from like, because it's just 

752
00:41:27,040 --> 00:41:29,920
you're just using, you know, 
shell commands to do that. 

753
00:41:30,040 --> 00:41:33,280
And so it knew, OK, here's the 
command to flash it and then the

754
00:41:33,280 --> 00:41:35,800
flash would fail and then we'll 
do all these other OK, Now I've 

755
00:41:35,800 --> 00:41:38,000
done a hard reset. 
Now it's OK, it's worked now. 

756
00:41:38,240 --> 00:41:40,200
Now I'm going to run it and I'm 
going to listen on the to the 

757
00:41:40,200 --> 00:41:43,080
serial and I'm going to read the
output that's printed back from 

758
00:41:43,080 --> 00:41:45,800
the statement to see how the The
thing is actually functioning. 

759
00:41:46,240 --> 00:41:48,720
The only it couldn't do LE DS, 
I'll give it that. 

760
00:41:48,720 --> 00:41:52,360
But but it could definitely read
off AUR. 

761
00:41:52,360 --> 00:41:54,800
If you were printing off AUR 
like, it could do that and then 

762
00:41:54,800 --> 00:41:57,240
interpret that data to make a 
decision. 

763
00:41:57,240 --> 00:42:01,200
About what you can, you know, 
for example, Claude can also 

764
00:42:01,480 --> 00:42:05,680
look at images. 
So you could take like a webcam,

765
00:42:05,680 --> 00:42:06,840
take a webcam picture. 
Oh. 

766
00:42:07,280 --> 00:42:11,400
I haven't tried this yet. 
Yeah, I haven't tried it in, in 

767
00:42:11,400 --> 00:42:14,240
the context of embedded systems,
but I've tried it in the context

768
00:42:14,280 --> 00:42:19,520
of of website design, which I'm 
very bad at, but the LLMS are 

769
00:42:19,520 --> 00:42:22,320
fairly good at. 
And like, I'll have it generate 

770
00:42:22,320 --> 00:42:25,400
a new increment of the website 
and the website would, you know,

771
00:42:25,520 --> 00:42:27,560
look mostly OK, but there's 
going to be something off. 

772
00:42:27,760 --> 00:42:30,000
And then I just take a 
screenshot and I say, look at 

773
00:42:30,000 --> 00:42:32,720
this like this is misaligned. 
I would say, Oh yes, I can see 

774
00:42:32,720 --> 00:42:34,960
it. 
Let me craft the correct CSS or 

775
00:42:34,960 --> 00:42:37,560
something. 
And it works shockingly well, 

776
00:42:38,360 --> 00:42:39,720
certainly much better than I 
could do it. 

777
00:42:40,280 --> 00:42:44,640
Oh my gosh, yeah. 
So, so there's lots of different

778
00:42:44,640 --> 00:42:48,400
ways of getting data into that 
context window and then managing

779
00:42:48,400 --> 00:42:51,080
it so that you're not polluting 
it too much, You're not going 

780
00:42:51,080 --> 00:42:53,200
too far in this direction, but 
then you're still getting the 

781
00:42:53,200 --> 00:42:55,600
the data that you actually need 
and there. 

782
00:42:56,800 --> 00:42:59,360
Exactly. 
So I guess to sum it up right, 

783
00:42:59,600 --> 00:43:01,080
there is such a thing as a 
context window. 

784
00:43:01,440 --> 00:43:04,440
Yes, it's very big, but also 
it's very much finite and you 

785
00:43:04,440 --> 00:43:06,440
will exhaust it much quicker 
than you want to. 

786
00:43:06,800 --> 00:43:10,920
And even then, not, not all 
places in the context window are

787
00:43:10,920 --> 00:43:13,520
created equal, right? 
What's at the beginning of the 

788
00:43:13,520 --> 00:43:17,440
context window is particularly 
interesting to the LLM. 

789
00:43:17,440 --> 00:43:19,680
What's what's what's at the end 
of the context window is 

790
00:43:19,680 --> 00:43:22,040
particularly interesting to the 
LLM, and the stuff in the middle

791
00:43:22,040 --> 00:43:24,840
gets kind of drowned out. 
Yeah, yeah. 

792
00:43:25,320 --> 00:43:28,960
So it makes sense to very 
deliberately sort of groom your 

793
00:43:28,960 --> 00:43:33,040
context window or your context 
and say, OK, what do I put in? 

794
00:43:33,040 --> 00:43:36,560
What do I pull out? 
When is the right time to 

795
00:43:36,880 --> 00:43:41,760
compact my context, or even to 
like throw it all out and start 

796
00:43:41,760 --> 00:43:45,120
fresh? 
Yep, in as often as you can. 

797
00:43:45,960 --> 00:43:50,160
Yep, and then and then getting 
data into your context without 

798
00:43:50,160 --> 00:43:52,240
having to pollute it too much, 
including files. 

799
00:43:52,240 --> 00:43:56,240
Parts of files using slash 
commands to send agents out to 

800
00:43:56,240 --> 00:43:59,960
go do specific tasks to get the 
output, or just carry out a 

801
00:43:59,960 --> 00:44:03,560
specific thing to do without 
having to pollute your existing 

802
00:44:03,560 --> 00:44:05,960
context window in your main 
thread. 

803
00:44:06,280 --> 00:44:08,480
And there was one thing that we 
hadn't mentioned that there's 

804
00:44:08,480 --> 00:44:10,800
something that I come back to a 
lot, and this is what we've 

805
00:44:10,800 --> 00:44:15,880
talked about a lot too, is like 
if you, you as the user need to 

806
00:44:15,880 --> 00:44:19,520
be a better product manager and 
to think globally about the 

807
00:44:19,520 --> 00:44:22,760
problem you're trying to solve. 
Take into account that the LLM 

808
00:44:22,760 --> 00:44:24,520
is going to be able to do a lot 
of things, but it's going to 

809
00:44:24,520 --> 00:44:28,680
have limitations. 
And don't try, don't, don't give

810
00:44:28,680 --> 00:44:30,600
it a task that's going to blow 
out the context window. 

811
00:44:30,600 --> 00:44:33,120
Make sure that you are breaking 
down the problem into 

812
00:44:33,240 --> 00:44:36,720
sufficiently small enough steps 
that the LOL can take care of 

813
00:44:36,720 --> 00:44:38,920
before you one at a time. 
Exactly. 

814
00:44:39,240 --> 00:44:40,600
That's, that's a very good 
point. 

815
00:44:40,600 --> 00:44:43,880
Like, and I don't know whether 
I've said it on the podcast 

816
00:44:43,880 --> 00:44:47,040
before, but I, I, I'm saying it 
fairly frequently. 

817
00:44:47,040 --> 00:44:48,760
I used to be a requirements 
engineer. 

818
00:44:49,600 --> 00:44:54,160
I find myself going back to my 
requirements engineering mindset

819
00:44:54,160 --> 00:44:57,200
and really thinking about, OK, 
what am I talking about? 

820
00:44:57,560 --> 00:45:01,480
What do I need to define, what 
context do I need to give and so

821
00:45:01,480 --> 00:45:06,000
on and so on. 
You know, because everybody 

822
00:45:06,000 --> 00:45:09,720
needs the proper information to 
do good work, whether they are a

823
00:45:09,720 --> 00:45:12,480
human or an LLM or, I don't 
know, a dog. 

824
00:45:12,480 --> 00:45:15,560
I guess because on the Internet 
nobody knows if you're a dog. 

825
00:45:19,440 --> 00:45:21,640
Absolutely. 
All right, Luca. 

826
00:45:21,720 --> 00:45:23,120
Excellent chatting with you, 
Sir. 

827
00:45:24,080 --> 00:45:27,280
That was fantastic. 
This is Ben, the Embedded AI 

828
00:45:27,280 --> 00:45:29,320
Podcast. 
I'm Ryan Torvik. 

829
00:45:29,920 --> 00:45:32,040
And I'm Luke and Johnny. 
We'll see you next time. 

830
00:45:33,240 --> 00:45:36,440
See you. 
Hi Luca here I have a quick 

831
00:45:36,440 --> 00:45:38,720
announcement. 
I've just launched one more new 

832
00:45:38,720 --> 00:45:42,720
training platform, the Embedded 
AI Academy at Embedded AI dot 

833
00:45:42,720 --> 00:45:44,800
Academy. 
It feels like up I've seen in 

834
00:45:44,800 --> 00:45:48,080
the market training that 
approaches AI specifically from 

835
00:45:48,080 --> 00:45:50,400
an embedded systems angle. 
If you're listening to this 

836
00:45:50,400 --> 00:45:54,160
podcast, it should be right up 
your alley and you get 25% off 

837
00:45:54,160 --> 00:45:57,520
any booking made through the end
of 2025 if you use the code 

838
00:45:57,840 --> 00:46:01,760
embedded AI Podcast 25 all in 
capital letters. 

839
00:46:02,120 --> 00:46:05,760
Again, the website is at 
Embedded AI dot Academy.

