1
00:00:00,720 --> 00:00:03,640
Your model says it can handle a 
million tokens. 

2
00:00:03,840 --> 00:00:08,800
It can't. 
On Gemini family models, basic 

3
00:00:08,800 --> 00:00:14,000
random word retention starts 
breaking between 500 and 750 

4
00:00:14,000 --> 00:00:16,720
words. 
That's not a million tokens. 

5
00:00:17,000 --> 00:00:21,120
That's barely 2 pages of text. 
You feed the thing a long 

6
00:00:21,120 --> 00:00:25,640
transcript, Ask a specific 
question, and what comes back is

7
00:00:25,640 --> 00:00:29,520
a confident hallucination or a 
refusal, or just the wrong 

8
00:00:29,520 --> 00:00:32,000
passage. 
The window your provider's 

9
00:00:32,000 --> 00:00:35,600
selling you and the window you 
can actually rely on are two 

10
00:00:35,600 --> 00:00:39,240
very different numbers, and 
almost nobody on your team is 

11
00:00:39,240 --> 00:00:42,360
measuring the gap between them. 
Here's the claim upfront. 

12
00:00:42,960 --> 00:00:46,560
Context rot is the silent decay 
of a model's reasoning as you 

13
00:00:46,560 --> 00:00:49,640
fill the window long before you 
hit the advertise limit. 

14
00:00:50,200 --> 00:00:52,280
The marketing department sells 
you a vault. 

15
00:00:53,040 --> 00:00:55,240
What you actually get is a leaky
bucket. 

16
00:00:56,680 --> 00:01:00,000
And if you're building 
production systems, agents, long

17
00:01:00,000 --> 00:01:03,000
running work flows, anything 
where prompts accumulate over 

18
00:01:03,000 --> 00:01:06,480
time, you are almost certainly 
operating in the rotted zone 

19
00:01:06,480 --> 00:01:10,080
without knowing it. 
Context length is a vanity 

20
00:01:10,080 --> 00:01:16,360
metric. 100,000 tokens is not as
reliable as 500 tokens, and 

21
00:01:16,360 --> 00:01:19,800
pretending it is will break your
system in ways your dashboards 

22
00:01:19,800 --> 00:01:22,880
won't catch. 
I've been running infrastructure

23
00:01:22,880 --> 00:01:26,680
for over a decade and I can tell
you this is the most common 

24
00:01:26,680 --> 00:01:30,040
unreported failure mode in 
production LLM systems right 

25
00:01:30,040 --> 00:01:33,560
now. 
It hides in plain sight because 

26
00:01:33,560 --> 00:01:36,880
the symptoms look like model 
quality issues, not context 

27
00:01:36,880 --> 00:01:39,320
issues. 
Let me walk through what the 

28
00:01:39,320 --> 00:01:43,920
research actually shows. 
Three sources, all recent, all 

29
00:01:43,920 --> 00:01:48,240
practitioner relevant. 
The No Lima paper dropped in 

30
00:01:48,240 --> 00:01:51,000
July. 
The acronym stands for No 

31
00:01:51,000 --> 00:01:55,960
Literal Matching, which is the 
whole .13 models, all claiming 

32
00:01:55,960 --> 00:02:00,240
at least 128,000 tokens. 
The test removes literal lexical

33
00:02:00,240 --> 00:02:04,120
overlap between the question and
the answer, meaning the model 

34
00:02:04,120 --> 00:02:07,160
can't just grab for the right 
words, it has to actually 

35
00:02:07,160 --> 00:02:12,520
understand what's in context. 
At 32,000 tokens, 11 of 13 

36
00:02:12,520 --> 00:02:16,720
models dropped below 50% of 
their short context baseline. 

37
00:02:17,080 --> 00:02:22,880
Read that again, 11 out of 13. 
This isn't a rare failure, this 

38
00:02:22,880 --> 00:02:27,400
is the common case. 
GPT for All goes from 99.3% 

39
00:02:27,400 --> 00:02:33,000
accuracy on short context down 
to 69.7 at 32,000 tokens. 

40
00:02:33,520 --> 00:02:36,720
That's a 30 point tax on 
intelligence for using 1/4 of 

41
00:02:36,720 --> 00:02:39,960
the advertised window. 
And the standard needle in a 

42
00:02:39,960 --> 00:02:44,720
haystack test misses all of this
because needle in a haystack is 

43
00:02:44,720 --> 00:02:48,000
basically grep. 
You hide A sentence about a 

44
00:02:48,000 --> 00:02:52,200
magic grass sandwich in 100,000 
tokens of fluff, ask the model 

45
00:02:52,200 --> 00:02:56,360
to find it, and it does. 
The labs publish the chart, 

46
00:02:56,680 --> 00:03:01,680
check the box, ship the feature,
but finding an exact string is 

47
00:03:01,680 --> 00:03:04,600
not reasoning. 
The moment you force the model 

48
00:03:04,600 --> 00:03:07,800
to use its internal 
representation instead of string

49
00:03:07,800 --> 00:03:10,080
matching, the whole thing falls 
apart. 

50
00:03:11,240 --> 00:03:20,920
Then there's the Chroma study. 
Their team ran 18 models through

51
00:03:20,920 --> 00:03:24,480
long memeval. 
The setup is sharper than most 

52
00:03:24,480 --> 00:03:27,680
evils. 
They compare focused inputs of 

53
00:03:27,680 --> 00:03:33,600
about 300 tokens against a full 
chat history of about 113,000 

54
00:03:33,600 --> 00:03:36,200
tokens. 
Same task, same difficulty. 

55
00:03:36,680 --> 00:03:38,760
The only variable is input 
length. 

56
00:03:39,680 --> 00:03:42,040
This is the honest version of 
what we actually do in 

57
00:03:42,040 --> 00:03:44,760
production. 
Long rambling context with a 

58
00:03:44,760 --> 00:03:49,440
specific question at the end. 
Every model degrades linear 

59
00:03:49,440 --> 00:03:53,000
decay of reliability across the 
board, Not one exception. 

60
00:03:53,880 --> 00:03:56,840
And the Gemini family 
specifically fails on random 

61
00:03:56,840 --> 00:04:00,680
word repetition. 
Between 500 and 750 words, which

62
00:04:00,680 --> 00:04:06,840
is the number I opened with, GPT
4.1 develops a 2 1/2% refusal 

63
00:04:06,840 --> 00:04:09,720
rate. 
Around 2500 words, the model 

64
00:04:09,720 --> 00:04:12,920
just stops trying. 
Not because the question got 

65
00:04:12,920 --> 00:04:14,920
harder, because the room got 
louder. 

66
00:04:15,520 --> 00:04:17,480
And here's the part that really 
got me. 

67
00:04:17,959 --> 00:04:21,079
Because the Chroma team held 
test difficulty constant and 

68
00:04:21,079 --> 00:04:24,720
only varied input length, they 
can isolate length itself as the

69
00:04:24,720 --> 00:04:27,720
direct 'cause this isn't 
confounded with question 

70
00:04:27,720 --> 00:04:30,720
complexity. 
It's not that longer context 

71
00:04:30,720 --> 00:04:32,480
happened to have harder 
questions in them. 

72
00:04:33,120 --> 00:04:36,440
It's that the same question, 
same difficulty, gets less 

73
00:04:36,440 --> 00:04:38,880
reliable as the context around 
it grows. 

74
00:04:40,320 --> 00:04:43,360
The third source is Dan Cleary 
at Promptub August. 

75
00:04:44,160 --> 00:04:48,160
He looks at distractors 
specifically runs the same task 

76
00:04:48,160 --> 00:04:53,120
with 01 and four distractors. 
Every added distractor compounds

77
00:04:53,120 --> 00:04:57,040
the accuracy loss. 
A distractor isn't random noise,

78
00:04:57,080 --> 00:05:00,720
it's irrelevant, but similar 
content like another company's 

79
00:05:00,720 --> 00:05:02,800
earnings report. 
When you're asking about Apple, 

80
00:05:03,600 --> 00:05:07,320
the model has to load both sets 
of facts, compare them, reject 

81
00:05:07,320 --> 00:05:10,160
the wrong one. 
That rejection task eats 

82
00:05:10,160 --> 00:05:14,760
attention at 4 distractors. 
Claude Force Sonnets reasoning 

83
00:05:14,760 --> 00:05:18,440
capability degrades even with a 
million token window. 

84
00:05:19,840 --> 00:05:23,240
The degradation isn't uniform 
across model families either, 

85
00:05:23,400 --> 00:05:26,000
which matters if you're picking 
a provider for a gentic 

86
00:05:26,000 --> 00:05:30,200
workflows. 
Some models survive 1 distractor

87
00:05:30,280 --> 00:05:33,720
and collapse at four. 
Others are more graceful. 

88
00:05:34,200 --> 00:05:38,520
You won't know which one you 
have until you test, and to make

89
00:05:38,520 --> 00:05:42,360
sure the results weren't biased 
by an unreliable judge, they 

90
00:05:42,360 --> 00:05:47,800
used the GPT 4.1 aligned grader 
that had over 99% agreement with

91
00:05:47,800 --> 00:05:50,640
human graders. 
So these numbers aren't soft. 

92
00:05:51,080 --> 00:05:53,560
The common thread across all 
three is blunt. 

93
00:05:53,680 --> 00:05:57,280
You can fit the data, you can't 
use the data. 

94
00:05:57,800 --> 00:06:02,000
Physical capacity and functional
capacity are not the same thing,

95
00:06:02,360 --> 00:06:04,800
and the industry has been 
pretending they are. 

96
00:06:06,120 --> 00:06:08,960
Now the part that really matters
if you're running agents. 

97
00:06:09,400 --> 00:06:11,800
An agent isn't reading one 
prompt once. 

98
00:06:11,800 --> 00:06:17,440
It's looping accumulating state 
appending tool outputs, dragging

99
00:06:17,440 --> 00:06:20,360
a growing tail of its own 
reasoning steps behind it. 

100
00:06:21,680 --> 00:06:24,040
Every turn eats into the safe 
zone. 

101
00:06:24,760 --> 00:06:29,960
If GPT 4 O is already taking a 
30 point accuracy hit at 32,000 

102
00:06:29,960 --> 00:06:33,640
tokens, what happens at the 10th
iteration of a React loop? 

103
00:06:34,160 --> 00:06:39,440
The 15th, The 20th. 
By then, the agent's own history

104
00:06:39,440 --> 00:06:42,320
is the distractor. 
You're not fighting the external

105
00:06:42,320 --> 00:06:45,280
data anymore, you're fighting 
the weight of the agent's own 

106
00:06:45,280 --> 00:06:48,680
previous thoughts. 
The system is operating in the 

107
00:06:48,680 --> 00:06:52,160
high failure regime for most of 
its lifespan, and nobody's 

108
00:06:52,160 --> 00:06:54,920
measuring it because the only 
thing the logs say is response 

109
00:06:54,920 --> 00:06:57,960
returned. 
Of course the response returned.

110
00:06:58,280 --> 00:07:01,040
Whether the response is any good
is a completely separate 

111
00:07:01,040 --> 00:07:03,480
question. 
I've seen agent systems in 

112
00:07:03,480 --> 00:07:07,680
production where the PD9 prompt 
length was five times the P50 

113
00:07:07,920 --> 00:07:10,280
and nobody on the team could 
tell you what the accuracy 

114
00:07:10,280 --> 00:07:15,720
looked like at PD9 versus P50. 
That gap between it ran and it 

115
00:07:15,720 --> 00:07:17,880
worked is where context route 
lives. 

116
00:07:18,480 --> 00:07:21,760
Let me tell you why this happens
mechanically, because if you 

117
00:07:21,760 --> 00:07:25,520
just take my word for it, you'll
forget it in a week. 3 reasons 

118
00:07:25,760 --> 00:07:28,520
they all stack. 
Attention dilution. 

119
00:07:28,760 --> 00:07:30,920
Every token attends to every 
other token. 

120
00:07:31,520 --> 00:07:35,880
That's a dense quadratic matrix.
The softmax function that 

121
00:07:35,880 --> 00:07:39,400
determines attention weights 
distributes a fixed amount of 

122
00:07:39,400 --> 00:07:41,800
probability mass across the 
whole sequence. 

123
00:07:42,360 --> 00:07:45,720
When you have a million tokens, 
the slice of attention landing 

124
00:07:45,720 --> 00:07:50,520
on the right one is microscopic.
The signal gets buried under the

125
00:07:50,520 --> 00:07:54,360
noise floor of everything else. 
Imagine trying to hear one 

126
00:07:54,360 --> 00:07:58,760
specific voice in a stadium. 
Doesn't matter how loud that 

127
00:07:58,760 --> 00:08:02,040
voice is, at some point the 
crowd drowns it out. 

128
00:08:02,840 --> 00:08:04,760
That's what's happening inside 
the model. 

129
00:08:05,520 --> 00:08:09,080
There are attempts to fix this. 
Sparse attention, local 

130
00:08:09,080 --> 00:08:12,840
attention, sliding windows, 
things like the long former or 

131
00:08:12,840 --> 00:08:16,080
Big Bird style patterns. 
They help at the edges. 

132
00:08:16,600 --> 00:08:18,520
They don't change the underlying
math. 

133
00:08:19,120 --> 00:08:22,880
Dense global attention still 
pays the dilution tax whenever 

134
00:08:22,880 --> 00:08:25,720
the model needs to route 
information between distant 

135
00:08:25,720 --> 00:08:29,600
tokens, and that's exactly the 
case where long context is 

136
00:08:29,600 --> 00:08:33,760
supposed to shine. 
The next thing that hit me when 

137
00:08:33,760 --> 00:08:36,440
I was digging into this is 
positional encoding. 

138
00:08:36,919 --> 00:08:40,120
Most modern models use some 
flavor of Rotary position 

139
00:08:40,120 --> 00:08:43,760
embedding row P. 
These worked by rotating the 

140
00:08:43,760 --> 00:08:45,760
token embeddings based on 
position. 

141
00:08:46,360 --> 00:08:48,880
The model learns those rotations
during training. 

142
00:08:49,720 --> 00:08:53,080
Problem is, training happens at 
specific context lengths. 

143
00:08:53,800 --> 00:08:57,920
When you push past what the 
model saw in training, it has to

144
00:08:57,920 --> 00:09:01,360
extrapolate. 
An extrapolation in positional 

145
00:09:01,360 --> 00:09:05,240
space is lossy. 
The resolution gets fuzzy at the

146
00:09:05,280 --> 00:09:08,120
edges. 
The model loses its ability to 

147
00:09:08,120 --> 00:09:12,760
distinguish between token 5000 
and token 50,000 with the 

148
00:09:12,760 --> 00:09:14,480
precision it had during 
training. 

149
00:09:15,080 --> 00:09:18,440
Labs use interpolation tricks to
extend these embeddings. 

150
00:09:18,440 --> 00:09:23,320
YARN, dynamic NTK, position 
interpolation, there's a whole 

151
00:09:23,320 --> 00:09:25,440
zoo of them. 
They help. 

152
00:09:26,240 --> 00:09:29,560
They don't fully restore the 
clarity, and every time the 

153
00:09:29,560 --> 00:09:33,280
window grows another order of 
magnitude, the high frequency 

154
00:09:33,280 --> 00:09:36,720
components of the positional 
signal start to alias. 

155
00:09:37,080 --> 00:09:39,880
The encoding literally runs out 
of expressiveness. 

156
00:09:40,720 --> 00:09:43,560
At that point, the model is 
guessing about where things are 

157
00:09:43,560 --> 00:09:46,800
in the sequence. 
Another failure mode, the 

158
00:09:46,800 --> 00:09:49,800
dashboards won't catch. 
And then there's training 

159
00:09:49,800 --> 00:09:53,600
distribution mismatch. 
This one's almost too obvious 

160
00:09:53,600 --> 00:09:56,560
once you see it. 
Instruction tuning usually 

161
00:09:56,560 --> 00:09:59,680
happens on sequences of 1 or 
2000 tokens. 

162
00:10:00,360 --> 00:10:02,960
The model learns the answer is 
usually nearby. 

163
00:10:03,720 --> 00:10:07,320
Then you ship it with 100,000 
token window and expected to 

164
00:10:07,320 --> 00:10:09,280
maintain focus across the whole 
thing. 

165
00:10:09,520 --> 00:10:12,640
It was never trained on that. 
The lost in the middle 

166
00:10:12,640 --> 00:10:14,760
phenomenon falls out of this 
directly. 

167
00:10:15,640 --> 00:10:19,360
Most long documents in the pre 
training data have the important

168
00:10:19,360 --> 00:10:21,440
information at the beginning or 
the end. 

169
00:10:21,840 --> 00:10:24,640
News articles, papers, even code
files. 

170
00:10:25,120 --> 00:10:28,800
The middle is connective tissue.
The model learns attention 

171
00:10:28,800 --> 00:10:30,720
patterns that favor the 
extremes. 

172
00:10:31,280 --> 00:10:36,360
So when you bury the answer at 
position 50,000 of 100,000, the 

173
00:10:36,360 --> 00:10:39,560
attention pattern trained for 
short sequences doesn't trigger 

174
00:10:39,560 --> 00:10:43,680
the way you need it to, and the 
fine tuning data is even worse. 

175
00:10:44,480 --> 00:10:47,640
Chat data sets are typically a 
few 100 tokens. 

176
00:10:48,160 --> 00:10:51,520
If 95% of the instruction 
following data the model ever 

177
00:10:51,520 --> 00:10:54,720
saw was under 2000 tokens, 
you're asking a lot. 

178
00:10:54,720 --> 00:10:58,160
When the production input is 50 
times that, those 3 mechanisms 

179
00:10:58,160 --> 00:11:00,040
stack. 
They're not alternatives. 

180
00:11:00,560 --> 00:11:04,640
Attention, dilution, positional 
encoding, breakdown, and 

181
00:11:04,640 --> 00:11:08,400
training distribution mismatch 
all compound, which is why the 

182
00:11:08,400 --> 00:11:11,440
degradation isn't linear and 
isn't predictable. 

183
00:11:11,440 --> 00:11:15,600
Without testing, a model might 
look fine at 10,000 tokens, 

184
00:11:16,040 --> 00:11:20,800
mediocre at 30,000, and 
completely unreliable at 100,000

185
00:11:21,000 --> 00:11:24,440
in a way you can't extrapolate 
from any single data point. 

186
00:11:25,240 --> 00:11:27,280
So what do you actually do 
Monday? 

187
00:11:29,040 --> 00:11:33,320
Here's the playbook. 
Build a per workload context 

188
00:11:33,320 --> 00:11:36,680
length, evolve. 
You can't trust the spec sheet 

189
00:11:36,720 --> 00:11:39,080
and you can't trust generic 
benchmarks. 

190
00:11:39,680 --> 00:11:42,880
You need to run your specific 
task at 10 different input 

191
00:11:42,880 --> 00:11:45,640
lengths. 
Hold difficulty constant and 

192
00:11:45,640 --> 00:11:49,880
plot accuracy versus length. 
The curve will tell you your rot

193
00:11:49,880 --> 00:11:52,040
threshold for that model on that
task. 

194
00:11:53,440 --> 00:11:57,480
Below the threshold, safe zone 
above, you're in the casino. 

195
00:11:58,040 --> 00:12:00,240
Without this eval you are flying
blind. 

196
00:12:00,880 --> 00:12:04,520
I've seen team ship agents where
the token count in production 

197
00:12:04,760 --> 00:12:08,320
was three times what they tested
at and nobody noticed until 

198
00:12:08,320 --> 00:12:10,080
support tickets started piling 
up. 

199
00:12:10,920 --> 00:12:15,320
The fix is simple in principle. 
Run a sweep, pick the longest 

200
00:12:15,320 --> 00:12:18,480
length where accuracy stays 
above your bar, and make that 

201
00:12:18,480 --> 00:12:21,720
your hard ceiling. 
Anything above the ceiling 

202
00:12:21,760 --> 00:12:25,360
either gets retrieved down, 
summarized down, or rejected. 

203
00:12:26,120 --> 00:12:29,680
After that, default to retrieval
first pipelines. 

204
00:12:29,680 --> 00:12:33,320
Even when the data fits, the 
temptation to just dump 

205
00:12:33,320 --> 00:12:34,960
everything in the window is 
real. 

206
00:12:35,280 --> 00:12:37,120
It feels efficient. 
It isn't. 

207
00:12:38,040 --> 00:12:41,520
A rag pipeline that sends the 
five most relevant chunks will 

208
00:12:41,520 --> 00:12:45,040
outperform A brute force long 
context prompt almost every 

209
00:12:45,040 --> 00:12:47,720
time. 
Every token in your prompt that 

210
00:12:47,720 --> 00:12:50,960
isn't actively helping answer 
the question is a potential 

211
00:12:50,960 --> 00:12:53,440
distractor and distractors 
compound. 

212
00:12:53,880 --> 00:12:57,160
Dense context beats long 
context, always. 

213
00:12:57,680 --> 00:13:01,080
And if you're worried about RAG 
being too complex, just remember

214
00:13:01,080 --> 00:13:02,720
what the alternative is costing 
you. 

215
00:13:03,160 --> 00:13:06,440
Every hallucination from a 
rotted context is a debugging 

216
00:13:06,440 --> 00:13:08,680
session. 
Every silent refusal is a 

217
00:13:08,680 --> 00:13:11,920
customer complaint. 
The operational cost of an 

218
00:13:11,920 --> 00:13:16,240
unreliable long context pipeline
is much higher than the 

219
00:13:16,240 --> 00:13:20,840
operational cost of a well tuned
RAG pipeline, even when the RAG 

220
00:13:20,840 --> 00:13:24,680
pipeline has more moving parts. 
Here's the one I think most 

221
00:13:24,680 --> 00:13:28,600
teams skip. 
Get ruthless about chat history.

222
00:13:29,400 --> 00:13:31,840
Stop letting it grow until it 
hits the limit. 

223
00:13:32,080 --> 00:13:35,440
Summarize older terms. 
Use a sliding window. 

224
00:13:35,720 --> 00:13:39,520
Separate long term memory which 
lives in a vector database from 

225
00:13:39,520 --> 00:13:42,320
working context, which is what 
the model is thinking about 

226
00:13:42,320 --> 00:13:45,120
right now. 
If the user said something 3 

227
00:13:45,120 --> 00:13:48,680
weeks ago that isn't relevant to
the current question, it doesn't

228
00:13:48,680 --> 00:13:50,200
go in the prompt. 
Period. 

229
00:13:51,120 --> 00:13:54,120
The cost of leaving it in is 
silent accuracy loss. 

230
00:13:55,040 --> 00:13:58,440
The cost of taking it out is A5 
line summarization step. 

231
00:13:58,880 --> 00:14:03,120
That's not a hard trade, and if 
you want to get fancy, use a 

232
00:14:03,120 --> 00:14:07,760
hierarchical memory. 
Recent turns go in verbatim, mid

233
00:14:07,760 --> 00:14:10,440
range turns get summarized into 
paragraphs. 

234
00:14:10,640 --> 00:14:14,320
Old turns get condensed to facts
in a structured store that you 

235
00:14:14,320 --> 00:14:18,760
retrieve from on demand. 
That's three layers, each with a

236
00:14:18,760 --> 00:14:22,960
different compression ratio, and
each one protects the model from

237
00:14:22,960 --> 00:14:27,040
the worst of the rot, something 
people miss constantly. 

238
00:14:27,480 --> 00:14:30,560
Summarization and reasoning are 
different tasks. 

239
00:14:31,160 --> 00:14:34,840
Long context windows are 
actually fine at summarizing a 

240
00:14:34,840 --> 00:14:39,040
single coherent document because
global attention can see the 

241
00:14:39,040 --> 00:14:41,840
whole structure. 
They are a trap for multi 

242
00:14:41,840 --> 00:14:45,480
document reasoning where you 
need to synthesize facts across 

243
00:14:45,480 --> 00:14:48,080
sources. 
If the 2 facts you need to 

244
00:14:48,080 --> 00:14:52,080
connect are 10,000 tokens apart,
the model is much more likely to

245
00:14:52,080 --> 00:14:55,280
miss the connection than if 
you'd retrieve them and put them

246
00:14:55,280 --> 00:14:58,480
side by side. 
Know which task you're doing. 

247
00:14:58,560 --> 00:15:01,120
If it's summarization, long 
context is fine. 

248
00:15:01,720 --> 00:15:05,320
If it's multi hop reasoning over
separate documents, retrieve and

249
00:15:05,320 --> 00:15:08,480
pair don't make the model work 
harder than it has to. 

250
00:15:08,840 --> 00:15:12,320
And the last piece, 
Observability log input token 

251
00:15:12,320 --> 00:15:16,240
length and every request as a 
first class metadata field, not 

252
00:15:16,240 --> 00:15:18,560
optional. 
When a user reports a 

253
00:15:18,560 --> 00:15:22,320
hallucination, the first thing 
you should be able to check is 

254
00:15:22,320 --> 00:15:28,680
the prompt length, Build length,
stratified metrics instead of 1 

255
00:15:28,680 --> 00:15:34,200
aggregate accuracy number, show 
me accuracy at 1010 thousand, 

256
00:15:34,240 --> 00:15:39,800
50,000 and 100,000 tokens. 
Insert Canary fax at random 

257
00:15:39,800 --> 00:15:43,440
depths in production prompts and
check if the model can recall 

258
00:15:43,440 --> 00:15:47,880
them alongside its main task. 
If the Canary fails, the rot 

259
00:15:47,880 --> 00:15:53,800
started and run eval on replay. 
Take production logs, strip them

260
00:15:53,800 --> 00:15:57,320
down to just the retrieved 
chunks using a smaller context 

261
00:15:57,640 --> 00:16:02,280
and see if the answer changes. 
If smaller context beats bigger,

262
00:16:02,480 --> 00:16:05,400
you found a failure point. 
Track it over time. 

263
00:16:06,000 --> 00:16:09,520
I'd put all of this on one 
dashboard and treat it the way 

264
00:16:09,520 --> 00:16:13,720
you treat P-95 latency. 
It's just another production 

265
00:16:13,720 --> 00:16:16,680
signal that tells you whether 
the system is actually working. 

266
00:16:18,080 --> 00:16:21,200
Let me make the Canary idea 
concrete, because it's the most 

267
00:16:21,200 --> 00:16:25,800
useful and the least adopted. 
Take a short factual statement 

268
00:16:26,200 --> 00:16:28,880
that the model couldn't 
plausibly know without being 

269
00:16:28,880 --> 00:16:33,680
told something like the internal
project code name for the Q3 

270
00:16:33,680 --> 00:16:38,480
migration was Blueberry. 
Insert that string at a random 

271
00:16:38,480 --> 00:16:42,760
position inside your production 
context, a different position on

272
00:16:42,760 --> 00:16:45,560
each request. 
At inference time. 

273
00:16:45,800 --> 00:16:49,920
Run 2 calls in parallel, one 
with the Canary and one without,

274
00:16:50,600 --> 00:16:55,640
and the one with the Canary. 
Also ask the model at the end to

275
00:16:55,640 --> 00:17:00,160
quote back the Canary string. 
If it can't, the context is 

276
00:17:00,160 --> 00:17:04,240
rotted for that request. 
You now have a live signal per 

277
00:17:04,240 --> 00:17:08,240
request of whether the model is 
reliably using the context 

278
00:17:08,240 --> 00:17:10,960
you're giving it. 
You don't need to wait for a 

279
00:17:10,960 --> 00:17:14,640
user to complain, you don't need
a label evil set. 

280
00:17:15,119 --> 00:17:18,480
You have a cheap continuous 
integrity check that scales with

281
00:17:18,480 --> 00:17:22,520
your traffic, and the cost is 
one extra field in your prompt 

282
00:17:22,880 --> 00:17:25,040
and one extra output token to 
check. 

283
00:17:25,640 --> 00:17:29,640
The teams that adopt this end up
with a dashboard that looks like

284
00:17:29,640 --> 00:17:33,680
uptime monitoring, green when 
the Canary recalls, red when it 

285
00:17:33,680 --> 00:17:37,160
doesn't, and a historical view 
of when the rot started creeping

286
00:17:37,160 --> 00:17:39,040
in for different models and 
workloads. 

287
00:17:39,600 --> 00:17:42,960
That's the kind of visibility 
you need, and the kind nobody's 

288
00:17:42,960 --> 00:17:46,240
shipping by default. 
None of this is exotic. 

289
00:17:46,680 --> 00:17:50,440
It's the same observability 
discipline we apply everywhere 

290
00:17:50,440 --> 00:17:55,200
else in infrastructure. 
Latency percentiles, error rates

291
00:17:55,200 --> 00:17:57,800
by input class, synthetic 
monitoring. 

292
00:17:58,240 --> 00:18:01,520
LLM outputs just happened to be 
one more signal that needs the 

293
00:18:01,520 --> 00:18:03,960
same treatment. 
The industry's been weirdly 

294
00:18:03,960 --> 00:18:06,440
allergic to treating model calls
like the infrastructure they 

295
00:18:06,440 --> 00:18:08,640
are. 
Probably because everyone's 

296
00:18:08,640 --> 00:18:12,080
excited about the magic. 
The magic gets less magical when

297
00:18:12,080 --> 00:18:15,640
you measure it, and once you 
measure it, you realize how 

298
00:18:15,640 --> 00:18:19,400
little of the advertised 
capability is actually available

299
00:18:19,400 --> 00:18:22,160
to you in production. 
Here's my take. 

300
00:18:23,080 --> 00:18:26,160
The million token window is a 
benchmark win, not an 

301
00:18:26,160 --> 00:18:29,360
engineering primitive. 
The labs are racing for the 

302
00:18:29,360 --> 00:18:32,320
biggest number because that's 
what looks good on the landing 

303
00:18:32,320 --> 00:18:35,800
page. 
Bigger number moves API credits,

304
00:18:36,160 --> 00:18:38,240
bigger number moves stock 
prices. 

305
00:18:38,520 --> 00:18:43,000
More reliable is harder to 
market, So what gets shipped is 

306
00:18:43,000 --> 00:18:45,040
what can be measured in 
marketing copy. 

307
00:18:45,760 --> 00:18:48,720
And your production agents 
silently rot because the 

308
00:18:48,720 --> 00:18:52,240
incentive structure that 
produced these models wasn't 

309
00:18:52,240 --> 00:18:54,840
optimizing for the thing you 
actually need. 

310
00:18:55,720 --> 00:18:58,880
The most sophisticated move you 
can make in your AI stack right 

311
00:18:58,880 --> 00:19:03,160
now isn't finding a model with a
bigger window, it's figuring out

312
00:19:03,160 --> 00:19:05,640
how to use less of the one you 
already have. 

313
00:19:06,480 --> 00:19:08,840
Efficiency is the new 
reliability. 

314
00:19:09,280 --> 00:19:12,400
Everyone chasing bigger is 
chasing the wrong number. 

315
00:19:12,920 --> 00:19:17,040
One last thing. 
Two years ago every rag vendor 

316
00:19:17,320 --> 00:19:20,680
was getting panic calls from 
their customers asking if 

317
00:19:20,680 --> 00:19:23,120
infinite context was going to 
kill them. 

318
00:19:23,920 --> 00:19:27,960
The pitch was the models are 
getting so big, retrieval is 

319
00:19:27,960 --> 00:19:31,080
obsolete. 
It was wrong then and it's wrong

320
00:19:31,080 --> 00:19:33,920
now. 
If anything, these massive 

321
00:19:33,960 --> 00:19:37,760
unreliable windows make the job 
of an infrastructure engineer 

322
00:19:37,760 --> 00:19:41,360
harder because they're 
introducing non deterministic 

323
00:19:41,360 --> 00:19:43,520
failures that are length 
dependent. 

324
00:19:44,240 --> 00:19:47,640
A model that fails consistently 
at a specific input is 

325
00:19:47,640 --> 00:19:50,960
debuggable. 
A model that works until it 

326
00:19:50,960 --> 00:19:55,720
doesn't based on a token count 
you don't track is a much worse 

327
00:19:55,720 --> 00:19:59,160
kind of broken. 
This is the tax of vanity 

328
00:19:59,160 --> 00:20:01,880
metrics. 
We built the tools to sell the 

329
00:20:01,880 --> 00:20:04,480
spec sheet. 
Now we have to build the tools 

330
00:20:04,480 --> 00:20:10,200
to work around the spec sheet. 
RAG isn't obsolete, it's back. 

331
00:20:10,520 --> 00:20:13,720
It's more important than ever, 
and the teams that figured this 

332
00:20:13,720 --> 00:20:17,160
out early are running circles 
around the one still chasing the

333
00:20:17,160 --> 00:20:21,240
million token dream. 
So context rot. 

334
00:20:22,280 --> 00:20:28,360
GPT 4 O at 69.7% accuracy on 1/4
of its advertised window is the 

335
00:20:28,360 --> 00:20:30,960
only number you need to start 
rethinking your architecture 

336
00:20:31,280 --> 00:20:34,840
retrieval first. 
Dense context length, Stratified

337
00:20:34,840 --> 00:20:37,120
evils? 
Canary prompts aggressive 

338
00:20:37,120 --> 00:20:40,400
history pruning. 
Trust your own numbers over the 

339
00:20:40,400 --> 00:20:44,560
claims on a provider's homepage.
The goal isn't how much you can 

340
00:20:44,560 --> 00:20:47,480
fit, it's how little you need to
get the right answer. 

341
00:20:48,040 --> 00:20:50,920
If you're doing it right, you 
should be using less context as 

342
00:20:50,920 --> 00:20:54,280
your system matures, not more. 
That's the one for this episode.

343
00:20:55,120 --> 00:20:58,320
I'm MO. 
This is the practical AI digest.

344
00:20:58,880 --> 00:21:01,880
If you got something out of 
this, send it to one person who 

345
00:21:01,880 --> 00:21:05,360
would actually use it. 
No subscribe campaign, no rating

346
00:21:05,360 --> 00:21:06,920
farm. 
Just one person. 

347
00:21:07,040 --> 00:21:08,200
See you on the next one.
