1
00:00:00,160 --> 00:00:02,640
Here is something that will 
sound completely 

2
00:00:02,640 --> 00:00:06,360
counterintuitive if you have 
been following the AI hardware 

3
00:00:06,360 --> 00:00:09,120
race. 
The biggest bottleneck in 

4
00:00:09,120 --> 00:00:12,600
running large language models 
right now is not compute. 

5
00:00:13,120 --> 00:00:16,640
It is not the number of GPU's 
you have, it is memory. 

6
00:00:17,160 --> 00:00:21,120
Specifically, it is a data 
structure called the KV cache. 

7
00:00:21,400 --> 00:00:25,040
And if you are running any kind 
of production LLM system, it is 

8
00:00:25,040 --> 00:00:28,200
almost certainly the thing that 
is quietly eating your budget, 

9
00:00:28,520 --> 00:00:31,800
throttling your throughput, and 
limiting how many users you can 

10
00:00:31,800 --> 00:00:35,320
actually serve. 
I am MO and this is the 

11
00:00:35,320 --> 00:00:38,960
practical AI digest. 
Today. 

12
00:00:38,960 --> 00:00:42,760
We are talking about KV cache 
compression, the memory wall 

13
00:00:42,760 --> 00:00:45,200
that nobody in the industry 
wants to have an honest 

14
00:00:45,200 --> 00:00:48,640
conversation about, and two 
recent breakthroughs that might 

15
00:00:48,640 --> 00:00:51,600
actually change the game. 
Let me set the stage with a 

16
00:00:51,600 --> 00:00:53,960
number that should make you 
uncomfortable. 

17
00:00:54,680 --> 00:00:58,360
A single instance of a large 
language model processing 

18
00:00:58,360 --> 00:01:03,040
100,000 tokens of context can 
consume 40 gigabytes of GPU 

19
00:01:03,040 --> 00:01:05,840
memory just for the KV cache 
alone. 

20
00:01:06,600 --> 00:01:10,040
That is roughly half the high 
bandwidth memory on an 

21
00:01:10,040 --> 00:01:12,960
enterprise grade GPU like the 
A-100. 

22
00:01:13,640 --> 00:01:17,640
Not the model weights, not the 
activations, just the cache of 

23
00:01:17,640 --> 00:01:20,520
keys and values that the 
attention mechanism needs to 

24
00:01:20,520 --> 00:01:24,560
look back at previous tokens. 
And that is for one user, one 

25
00:01:24,560 --> 00:01:27,480
request. 
Now multiply that by the number 

26
00:01:27,480 --> 00:01:30,560
of concurrent users you need to 
serve, and you start to 

27
00:01:30,560 --> 00:01:33,440
understand why inference costs 
are where they are. 

28
00:01:34,520 --> 00:01:37,840
The reason this happens goes 
back to how Transformers work at

29
00:01:37,840 --> 00:01:41,160
a fundamental level. 
Every time a model generates a 

30
00:01:41,160 --> 00:01:45,000
new token, it needs to attend to
all the previous tokens in the 

31
00:01:45,000 --> 00:01:47,880
sequence. 
The naive approach would be to 

32
00:01:47,880 --> 00:01:51,560
recompute the attention for 
every single previous token on 

33
00:01:51,560 --> 00:01:54,440
every generation step, which 
would be catastrophically 

34
00:01:54,440 --> 00:01:58,200
expensive. 
So instead we cache the key and 

35
00:01:58,200 --> 00:02:02,120
value projections from each 
layer for each token and reuse 

36
00:02:02,120 --> 00:02:05,080
them. 
That is the KV cache. 

37
00:02:05,880 --> 00:02:09,000
It is a brilliant optimization 
that makes autoregressive 

38
00:02:09,000 --> 00:02:12,800
generation tractable, but it 
comes with a cost that scales 

39
00:02:12,800 --> 00:02:17,080
linearly with sequence length, 
linearly with the number of 

40
00:02:17,080 --> 00:02:21,080
layers, and linearly with the 
hidden dimension. 

41
00:02:21,720 --> 00:02:26,040
For a 70 billion parameter model
with 80 layers, you are looking 

42
00:02:26,040 --> 00:02:28,360
at about a MB per token per 
layer. 

43
00:02:29,160 --> 00:02:33,360
At 100,000 tokens, that is 8 
gigabytes per layer times 80 

44
00:02:33,360 --> 00:02:36,040
layers. 
The math gets ugly fast. 

45
00:02:36,720 --> 00:02:41,200
This is what infrastructure 
people call the memory wall, and

46
00:02:41,200 --> 00:02:46,960
it is not a theoretical concern.
WEK published an analysis in 

47
00:02:46,960 --> 00:02:51,640
April 2026 framing this as the 
defining infrastructure 

48
00:02:51,640 --> 00:02:54,080
challenge of the current AI 
generation. 

49
00:02:54,720 --> 00:02:57,000
Their argument is sharp and I 
think correct. 

50
00:02:57,640 --> 00:03:01,840
Organizations have been throwing
GP US at the inference problem, 

51
00:03:02,200 --> 00:03:05,920
but GP US do not solve a memory 
bandwidth bottleneck. 

52
00:03:06,480 --> 00:03:09,360
You can have all the floating 
point operations per second in 

53
00:03:09,360 --> 00:03:12,960
the world, and if your attention
mechanism is waiting on memory 

54
00:03:13,160 --> 00:03:16,440
reads from a bloated KV cache, 
you are memory bound. 

55
00:03:17,360 --> 00:03:21,520
The GPU cores are sitting idle 
waiting for data, and you are 

56
00:03:21,520 --> 00:03:24,160
paying for that idle time at GPU
prices. 

57
00:03:24,800 --> 00:03:27,400
The practical impact hits in 
three ways. 

58
00:03:27,880 --> 00:03:31,000
Throughput goes down because you
can fit fewer concurrent 

59
00:03:31,000 --> 00:03:34,720
requests in memory. 
Latency goes up because memory 

60
00:03:34,720 --> 00:03:39,160
access patterns get less cache 
friendly as the KV cache grows, 

61
00:03:39,640 --> 00:03:44,560
and cost goes up because you 
need more GPUs not for compute, 

62
00:03:44,640 --> 00:03:48,120
but for memory. 
This is why a company serving a 

63
00:03:48,120 --> 00:03:51,680
chatbot to 1,000,000 users is 
not spending most of its 

64
00:03:51,680 --> 00:03:55,600
inference budget on matrix 
multiplications, it is spending 

65
00:03:55,600 --> 00:03:58,920
it on storing and fetching the 
conversational context for each 

66
00:03:58,920 --> 00:04:01,160
user. 
And there is a compounding 

67
00:04:01,160 --> 00:04:05,960
dynamic that makes this worse. 
Over time, as models get larger 

68
00:04:06,160 --> 00:04:10,160
and context windows get longer, 
the KV cache grows in both 

69
00:04:10,160 --> 00:04:14,400
dimensions simultaneously. 
A model with twice the layers 

70
00:04:14,400 --> 00:04:18,560
and twice the context window has
four times the KV cache. 

71
00:04:19,200 --> 00:04:23,520
Meanwhile, GPU memory has been 
growing linearly, roughly 

72
00:04:23,520 --> 00:04:25,600
doubling every two to three 
years. 

73
00:04:26,200 --> 00:04:29,560
The KV cache is growing 
quadratically while the hardware

74
00:04:29,560 --> 00:04:33,000
is growing linearly. 
That is the definition of a 

75
00:04:33,000 --> 00:04:35,440
wall. 
There is also a subtlety here 

76
00:04:35,440 --> 00:04:37,280
that gets lost in the 
benchmarks. 

77
00:04:37,720 --> 00:04:41,520
The memory wall does not just 
affect peak memory, it affects 

78
00:04:41,520 --> 00:04:46,880
memory bandwidth utilization. 
Modern GPU's like the H-100 have

79
00:04:46,880 --> 00:04:51,520
enormous compute throughput over
a petaflop of FP8, but their 

80
00:04:51,520 --> 00:04:55,120
memory bandwidth tops out around
3 terabytes per second. 

81
00:04:55,920 --> 00:04:59,800
When your attention mechanism 
has to read a 40 GB KV cache for

82
00:04:59,800 --> 00:05:03,680
every single generated token, 
you are doing 40 gigabytes times

83
00:05:03,680 --> 00:05:06,520
the number of attention heads of
memory reads per token. 

84
00:05:07,280 --> 00:05:11,160
At 3 terabytes per second, that 
is over 13 milliseconds per 

85
00:05:11,160 --> 00:05:15,880
token just for memory access. 
For comparison, the actual 

86
00:05:15,880 --> 00:05:18,400
matrix multiplications take a 
fraction of that. 

87
00:05:19,160 --> 00:05:22,600
The model is literally waiting 
for data, not crunching numbers.

88
00:05:23,200 --> 00:05:26,280
This is why you see companies 
like Rock building custom chips 

89
00:05:26,280 --> 00:05:30,040
with massive on chips RAM, why 
Cerebras has their wafer scale 

90
00:05:30,040 --> 00:05:33,520
approach, and why even Nvidia's 
Blackwell architecture 

91
00:05:33,520 --> 00:05:36,480
prioritizes memory bandwidth 
over raw flops. 

92
00:05:37,120 --> 00:05:41,200
The industry has figured out 
that compute is abundant, memory

93
00:05:41,200 --> 00:05:45,680
is scarce, and the KV cache is 
where that scarcity hurts the 

94
00:05:45,680 --> 00:05:48,080
most. 
So what do you actually do about

95
00:05:48,080 --> 00:05:51,960
it? 
The field has been working on 

96
00:05:51,960 --> 00:05:56,520
this hard, and 2025 and early 
2026 have produced some 

97
00:05:56,520 --> 00:06:00,240
genuinely promising approaches. 
I want to walk through two of 

98
00:06:00,240 --> 00:06:03,880
them in detail because they 
represent fundamentally 

99
00:06:03,880 --> 00:06:06,960
different strategies, and 
understanding why they work 

100
00:06:06,960 --> 00:06:09,680
tells you something important 
about where inference 

101
00:06:09,680 --> 00:06:14,400
optimization is heading. 
The first is KVTC, which stands 

102
00:06:14,400 --> 00:06:20,600
for KV Cache Transform Coding. 
This was published at ICLR 2026,

103
00:06:20,840 --> 00:06:23,880
and it takes an approach 
borrowed from signal processing,

104
00:06:24,440 --> 00:06:27,600
specifically from how we 
compress images and video. 

105
00:06:28,680 --> 00:06:32,840
The insight is elegant. 
The raw key and value vectors in

106
00:06:32,840 --> 00:06:36,200
the KV cache have significant 
statistical redundancy. 

107
00:06:36,920 --> 00:06:41,120
Adjacent tokens tend to produce 
similar key vectors because the 

108
00:06:41,120 --> 00:06:43,880
underlying representations 
change gradually. 

109
00:06:44,680 --> 00:06:48,400
Different attention heads within
the same layer often encode 

110
00:06:48,560 --> 00:06:51,520
partially overlapping 
information, and the 

111
00:06:51,520 --> 00:06:55,600
distribution of values across 
the hidden dimension is highly 

112
00:06:55,600 --> 00:06:59,000
non uniform, with some 
dimensions carrying much more 

113
00:06:59,000 --> 00:07:03,600
information than others. 
KVTC exploits all three of these

114
00:07:03,600 --> 00:07:08,720
redundancies in a pipeline. 
First, it applies PCA based 

115
00:07:08,720 --> 00:07:11,760
feature decorrelation to remove 
the cross dimension 

116
00:07:11,760 --> 00:07:15,000
correlations. 
This is essentially rotating the

117
00:07:15,000 --> 00:07:19,040
coordinate system so that each 
dimension is as independent as 

118
00:07:19,040 --> 00:07:23,480
possible, which is exactly what 
transform coding does in JPEG 

119
00:07:23,520 --> 00:07:26,840
and video codecs. 
Then it applies adaptive 

120
00:07:26,840 --> 00:07:31,120
quantization, allocating more 
bits to high variance dimensions

121
00:07:31,400 --> 00:07:33,680
and fewer bits to low variance 
ones. 

122
00:07:34,440 --> 00:07:38,120
And finally it runs entropy 
coding on the quantized values 

123
00:07:38,120 --> 00:07:40,920
to squeeze out the remaining 
statistical redundancy. 

124
00:07:41,640 --> 00:07:44,920
The results are striking. 
On standard reasoning 

125
00:07:44,920 --> 00:07:49,680
benchmarks, KVTC achieves up to 
20 times compression with 

126
00:07:49,680 --> 00:07:54,080
negligible accuracy loss. 
On some specific workloads it 

127
00:07:54,080 --> 00:07:57,880
hits 40 times or higher. 
And critically, it is training 

128
00:07:57,880 --> 00:08:00,440
free. 
You do not need to fine tune the

129
00:08:00,440 --> 00:08:02,640
model. 
You do not need calibration 

130
00:08:02,640 --> 00:08:05,560
data. 
You apply it at inference time 

131
00:08:05,600 --> 00:08:08,040
and it works. 
The compression and 

132
00:08:08,040 --> 00:08:11,920
decompression are fast enough 
that the time spent compressing 

133
00:08:12,200 --> 00:08:15,680
is more than offset by the 
memory bandwidth you save on 

134
00:08:15,680 --> 00:08:21,320
cache reads. 20 times 
compression on a 40 GB KV cache 

135
00:08:21,320 --> 00:08:23,840
means you are down to 2 
gigabytes. 

136
00:08:25,120 --> 00:08:28,240
That changes the economics of 
inference entirely. 

137
00:08:28,800 --> 00:08:32,159
You could serve 10 times as many
concurrent users on the same 

138
00:08:32,159 --> 00:08:35,799
hardware, or you could extend 
your context window by an order 

139
00:08:35,799 --> 00:08:38,039
of magnitude without adding 
GPU's. 

140
00:08:38,799 --> 00:08:41,840
There is a nuance in the KVTC 
approach that I think is 

141
00:08:41,840 --> 00:08:45,720
underappreciated. 
The PCA rotation step is not a 

142
00:08:45,720 --> 00:08:47,800
one-size-fits-all 
transformation. 

143
00:08:48,120 --> 00:08:51,240
The authors compute the 
principal components per layer 

144
00:08:51,240 --> 00:08:54,400
and per attention head because 
different layers encode 

145
00:08:54,400 --> 00:08:56,720
fundamentally different kinds of
information. 

146
00:08:57,320 --> 00:09:00,480
Early layers tend to capture 
syntactic and positional 

147
00:09:00,480 --> 00:09:03,160
patterns where the key vectors 
cluster tightly. 

148
00:09:04,120 --> 00:09:07,440
Middle layers handle semantic 
composition, where the variance 

149
00:09:07,440 --> 00:09:11,880
is distributed more evenly. 
Late layers handle task specific

150
00:09:11,880 --> 00:09:14,080
reasoning, where a few 
dimensions dominate. 

151
00:09:14,800 --> 00:09:19,320
By adapting the transform per 
layer, KVTC avoids the common 

152
00:09:19,320 --> 00:09:22,920
trap of compression methods that
work well on average but blow up

153
00:09:22,920 --> 00:09:25,080
on the layers that matter most 
for accuracy. 

154
00:09:26,000 --> 00:09:28,680
The entropy coding step is also 
worth a closer look. 

155
00:09:29,440 --> 00:09:33,120
After quantization, the 
distribution of quantized values

156
00:09:33,120 --> 00:09:36,520
is highly non uniform. 
Some quantization bins are hit 

157
00:09:36,520 --> 00:09:40,040
far more often than others, and 
that non uniformity is free 

158
00:09:40,040 --> 00:09:45,160
compression if you exploit it. 
KVTC uses arithmetic coding, the

159
00:09:45,160 --> 00:09:48,880
same family of algorithms that 
powers modern video codecs like 

160
00:09:48,880 --> 00:09:52,600
H265, to squeeze out this 
remaining redundancy. 

161
00:09:53,960 --> 00:09:56,800
The overhead of the coding 
itself is negligible because the

162
00:09:56,800 --> 00:10:00,000
decoder runs on the GPU's tensor
cores during the attention 

163
00:10:00,000 --> 00:10:03,040
computation. 
You are effectively getting free

164
00:10:03,040 --> 00:10:06,520
compression from statistical 
structure that was always there,

165
00:10:06,840 --> 00:10:08,720
but nobody was bothering to 
exploit. 

166
00:10:09,280 --> 00:10:11,840
One more thing about KVTC that 
surprised me. 

167
00:10:12,600 --> 00:10:15,520
The authors show that their 
compression actually improves 

168
00:10:15,520 --> 00:10:18,600
cache locality. 
In some cases, the compressed 

169
00:10:18,600 --> 00:10:22,800
cache entries are smaller, which
means more of them fit in the 

170
00:10:22,800 --> 00:10:25,840
GPU's L2 cache during the 
attention pass. 

171
00:10:26,520 --> 00:10:29,640
This creates A virtuous cycle 
where compression reduces memory

172
00:10:29,640 --> 00:10:33,520
footprint, which improves cache 
hit rates, which reduces memory 

173
00:10:33,520 --> 00:10:36,760
bandwidth pressure, which speeds
up the attention computation 

174
00:10:36,760 --> 00:10:39,200
beyond what the raw compression 
ratio would predict. 

175
00:10:39,800 --> 00:10:42,880
The measured speed up is 
sometimes higher than the 

176
00:10:42,880 --> 00:10:46,520
compression ratio alone would 
explain, and this locality 

177
00:10:46,520 --> 00:10:51,720
effect is why the second 
breakthrough is Turboquant from 

178
00:10:51,720 --> 00:10:55,280
Google Research, also at ICLR 
2026. 

179
00:10:55,880 --> 00:10:57,920
Turboquant takes a different 
approach. 

180
00:10:58,440 --> 00:11:02,480
Instead of the full transform 
coding pipeline, it focuses 

181
00:11:02,480 --> 00:11:06,840
specifically on quantization, 
but does it in a mathematically 

182
00:11:06,840 --> 00:11:10,840
novel way that eliminates the 
accuracy penalties that previous

183
00:11:10,840 --> 00:11:12,960
quantization methods suffered 
from. 

184
00:11:14,160 --> 00:11:16,960
The key innovation is a 
technique called polar quant. 

185
00:11:17,520 --> 00:11:21,040
Traditional quantization methods
need to store per block 

186
00:11:21,040 --> 00:11:25,240
normalization constants, scaling
factors that map the quantized 

187
00:11:25,240 --> 00:11:28,120
integers back to the original 
floating point range. 

188
00:11:29,120 --> 00:11:32,080
These constants take up space 
and introduce error at the 

189
00:11:32,080 --> 00:11:36,160
boundaries between blocks. 
Polar quant eliminates them 

190
00:11:36,160 --> 00:11:40,720
entirely by representing values 
in polar coordinates where the 

191
00:11:40,720 --> 00:11:44,280
magnitude and angle can be 
quantized independently with 

192
00:11:44,280 --> 00:11:47,480
tighter bounds. 
It sounds like a small change, 

193
00:11:47,800 --> 00:11:51,160
but it removes an entire class 
of quantization artifacts. 

194
00:11:51,680 --> 00:11:55,080
On top of Polar Quant, 
Turboquant adds something called

195
00:11:55,080 --> 00:11:58,840
Quantize Johnson Lindenstrauss, 
or QJL. 

196
00:11:59,600 --> 00:12:02,920
This is a dimensionality 
reduction technique that 

197
00:12:02,920 --> 00:12:06,760
projects the key vectors into a 
lower dimensional space while 

198
00:12:06,760 --> 00:12:10,080
preserving the relative 
distances between them, which is

199
00:12:10,080 --> 00:12:12,560
what the attention mechanism 
actually cares about. 

200
00:12:13,160 --> 00:12:16,520
The Johnson Lindenstrauss lemma 
guarantees that random 

201
00:12:16,520 --> 00:12:20,040
projections preserve distances 
with high probability, and 

202
00:12:20,040 --> 00:12:23,680
Turboquan shows that you can 
quantize the projected vectors 

203
00:12:23,840 --> 00:12:27,040
to very low bit widths without 
violating this guarantee. 

204
00:12:27,960 --> 00:12:31,600
The combined result is 3 bits 
per element in the KV cache. 

205
00:12:32,320 --> 00:12:38,440
That is a six times memory 
reduction on NVIDIA H-100 GPU's.

206
00:12:38,560 --> 00:12:41,920
The attention computation itself
gets up to 8 times faster 

207
00:12:41,920 --> 00:12:45,280
because the compressed cache 
fits better in the GPU's on chip

208
00:12:45,280 --> 00:12:48,360
memory hierarchy. 
And here's the part that made me

209
00:12:48,360 --> 00:12:53,160
sit up. 0 measurable accuracy 
degradation, not small 

210
00:12:53,200 --> 00:12:58,320
degradation, not acceptable 
degradation. 0 on long bench, on

211
00:12:58,320 --> 00:13:01,960
needle in haystack tests, on 
standard reasoning benchmarks. 

212
00:13:02,240 --> 00:13:05,440
No calibration data required, no
fine tuning. 

213
00:13:06,160 --> 00:13:09,160
You apply it at inference time 
and it just works. 

214
00:13:09,480 --> 00:13:13,800
Now you might be wondering how 
GULO validated the 0 degradation

215
00:13:13,800 --> 00:13:17,320
claim, because that sounds too 
good to be true and honestly 

216
00:13:17,480 --> 00:13:19,760
when I first read it I was 
skeptical too. 

217
00:13:20,120 --> 00:13:22,800
But the key is in how attention 
actually works. 

218
00:13:23,360 --> 00:13:27,120
The attention mechanism computes
dot products between query 

219
00:13:27,120 --> 00:13:31,040
vectors and key vectors, then 
uses those scores to weight the 

220
00:13:31,040 --> 00:13:34,560
value vectors. 
What matters for the output is 

221
00:13:34,560 --> 00:13:38,240
not the exact values of the 
keys, but the relative ordering 

222
00:13:38,240 --> 00:13:41,760
of the dot product scores. 
If your compression preserves 

223
00:13:41,760 --> 00:13:46,240
the ranking, the ARD Max, the 
softmax distribution, then the 

224
00:13:46,240 --> 00:13:48,240
output is functionally 
identical. 

225
00:13:48,760 --> 00:13:53,200
Polar quant and QJL are designed
specifically to preserve these 

226
00:13:53,200 --> 00:13:57,120
relative distances, and the 
empirical results confirm that 

227
00:13:57,120 --> 00:14:00,040
the preserved ordering 
translates to preserved output 

228
00:14:00,040 --> 00:14:02,760
quality across a wide range of 
tasks. 

229
00:14:03,360 --> 00:14:06,040
There is also a practical 
consideration that the paper 

230
00:14:06,040 --> 00:14:10,120
addresses what happens at the 
boundaries when you compress mid

231
00:14:10,120 --> 00:14:12,160
sequence. 
Do you get artifacts at the 

232
00:14:12,160 --> 00:14:15,640
transition between compressed 
and uncompressed cache entries? 

233
00:14:16,440 --> 00:14:19,440
Turboquan handles this by 
compressing the entire cache 

234
00:14:19,440 --> 00:14:21,840
uniformly so there are no 
boundary effects. 

235
00:14:22,440 --> 00:14:25,840
Every entry gets the same 
treatment, which means the 

236
00:14:25,840 --> 00:14:28,040
compression is invisible to the 
model. 

237
00:14:29,000 --> 00:14:32,520
I want to be clear about why 
this matters beyond the obvious 

238
00:14:32,520 --> 00:14:36,040
memory savings. 
The memory wall is not just a 

239
00:14:36,040 --> 00:14:38,880
cost problem, it is a capability
ceiling. 

240
00:14:39,480 --> 00:14:42,760
Today the practical limit on 
context length for most 

241
00:14:42,760 --> 00:14:46,560
production deployments is not 
the models advertised context 

242
00:14:46,560 --> 00:14:50,400
window, it is how much KV cache 
memory you can afford. 

243
00:14:51,200 --> 00:14:55,040
A model that advertises 
1,000,000 token context window 

244
00:14:55,360 --> 00:15:00,000
but requires a TB of KV cache 
memory at that length is not 

245
00:15:00,000 --> 00:15:03,680
really a million token model for
anyone running real workloads. 

246
00:15:04,520 --> 00:15:08,440
KV cache compression directly 
translates into longer usable 

247
00:15:08,440 --> 00:15:10,840
context at the same hardware 
cost. 

248
00:15:11,640 --> 00:15:14,840
It also changes the calculus on 
model architecture. 

249
00:15:16,880 --> 00:15:20,320
One of the reasons mixture of 
experts models are popular is 

250
00:15:20,320 --> 00:15:23,960
that they reduce the per token 
compute cost, but they do not 

251
00:15:23,960 --> 00:15:28,600
reduce the KV cache size because
every token still needs full key

252
00:15:28,760 --> 00:15:31,160
and value projections across all
layers. 

253
00:15:31,840 --> 00:15:35,280
With aggressive KV cache 
compression, dense models become

254
00:15:35,280 --> 00:15:38,640
more competitive again because 
their main disadvantage, higher 

255
00:15:38,640 --> 00:15:41,560
memory usage per token gets 
compressed away. 

256
00:15:42,000 --> 00:15:45,440
I also want to put this in an 
economic frame because the 

257
00:15:45,440 --> 00:15:49,240
dollar amounts clarify the 
urgency in a way that technical 

258
00:15:49,240 --> 00:15:54,560
metrics sometimes do not. 
A single H-100 GPU costs roughly

259
00:15:54,560 --> 00:15:59,600
$30,000, or about $4.00 per hour
on the major cloud providers. 

260
00:16:00,280 --> 00:16:05,440
If your KV cache is consuming 
half the GPU memory, then half 

261
00:16:05,440 --> 00:16:09,680
of that $4.00 per hour is going 
to store in context, not 

262
00:16:09,680 --> 00:16:14,120
generating tokens. 
For a company running 100 GPUs 

263
00:16:14,120 --> 00:16:20,920
for inference, that is $200 per 
hour, roughly $175,000 per month

264
00:16:21,320 --> 00:16:26,040
just on KV cache memory. 
A six times compression from 

265
00:16:26,040 --> 00:16:28,640
Turboquant cuts that to under 
30,000. 

266
00:16:29,400 --> 00:16:33,840
A 20 times compression from KVTC
cuts it to under 9000. 

267
00:16:34,440 --> 00:16:36,440
These are not theoretical 
savings. 

268
00:16:36,880 --> 00:16:39,840
They are the difference between 
an inference deployment that is 

269
00:16:39,880 --> 00:16:43,080
economically viable and one that
is hemorrhaging money. 

270
00:16:43,600 --> 00:16:47,760
And it compounds. 
When you free up GPU memory from

271
00:16:47,760 --> 00:16:52,040
KV cache, you can fit more 
concurrent requests for GPU, 

272
00:16:52,320 --> 00:16:55,840
which means you need fewer GPU's
for the same throughput. 

273
00:16:56,600 --> 00:17:00,720
Fewer GPUs means less data 
center power, less cooling, 

274
00:17:00,920 --> 00:17:04,319
fewer network switches, less 
operational overhead. 

275
00:17:04,920 --> 00:17:07,240
The KV cache is the first 
domino. 

276
00:17:07,560 --> 00:17:10,720
Compress it and everything 
downstream gets cheaper. 

277
00:17:11,160 --> 00:17:13,680
So what does the practitioner 
playbook look like? 

278
00:17:13,680 --> 00:17:17,160
If you are running inference in 
production today, let me give 

279
00:17:17,160 --> 00:17:21,240
you 5 concrete things. 
Benchmark your KV cache memory 

280
00:17:21,240 --> 00:17:24,880
usage per request at your actual
workload lengths. 

281
00:17:25,160 --> 00:17:30,360
Not the models maximum context. 
Your actual 50 and P99 context 

282
00:17:30,360 --> 00:17:33,760
lengths. 
Most teams I have talked to do 

283
00:17:33,760 --> 00:17:37,360
not know this number and it is 
the single most important metric

284
00:17:37,360 --> 00:17:42,280
for capacity planning. 
Evaluate KVTC or Turboquant on 

285
00:17:42,280 --> 00:17:44,120
your specific model and 
workload. 

286
00:17:44,480 --> 00:17:47,920
Both are training free and can 
be tested without modifying your

287
00:17:47,920 --> 00:17:50,920
model. 
KBTC gives higher compression 

288
00:17:50,920 --> 00:17:54,000
ratios but with more 
computational overhead. 

289
00:17:54,720 --> 00:17:58,600
Turboquant gives lower 
compression but with faster 

290
00:17:58,600 --> 00:18:01,400
decompression and 0 accuracy 
risk. 

291
00:18:01,920 --> 00:18:04,880
The right choice depends on 
whether you are memory bound or 

292
00:18:04,880 --> 00:18:07,760
compute bound. 
If you are not ready for a 

293
00:18:07,760 --> 00:18:11,640
compression library, start with 
the low hanging fruit sliding 

294
00:18:11,640 --> 00:18:15,760
window attention on your KV 
cache where you only keep the 

295
00:18:15,760 --> 00:18:19,320
most recent end tokens plus a 
few anchor tokens from the 

296
00:18:19,320 --> 00:18:22,040
beginning. 
This is not as principled as the

297
00:18:22,040 --> 00:18:25,680
compression approaches, but it 
is trivial to implement and gets

298
00:18:25,680 --> 00:18:28,840
you meaningful memory savings 
for conversational workloads. 

299
00:18:29,400 --> 00:18:33,200
Set up memory observability. 
You should be tracking KV cache 

300
00:18:33,200 --> 00:18:37,720
size per request, peak memory 
usage for GPU, and the ratio of 

301
00:18:37,720 --> 00:18:41,560
KV cache memory to total GPU 
memory utilization. 

302
00:18:42,200 --> 00:18:45,760
If that ratio is above 50%, you 
are leaving throughput on the 

303
00:18:45,760 --> 00:18:48,520
table. 
Plan for the convergence, 

304
00:18:49,160 --> 00:18:52,720
Google's official turboquant 
implementation is expected in 

305
00:18:52,720 --> 00:18:56,960
the second quarter of 2026, and 
the major inference frameworks 

306
00:18:57,000 --> 00:19:02,000
VLLM, Tensor, RTLLM, and the 
rest will integrate these 

307
00:19:02,000 --> 00:19:03,960
techniques within a release 
cycle or two. 

308
00:19:04,400 --> 00:19:08,320
The teams that understand the 
memory wall now will be ready to

309
00:19:08,320 --> 00:19:11,680
deploy compression the day it 
ships in their stack. 

310
00:19:12,320 --> 00:19:16,880
The teams that do not will be 6 
months behind paying six times 

311
00:19:16,880 --> 00:19:21,520
more for memory than they need 
to and do not sleep on the 2nd 

312
00:19:21,520 --> 00:19:25,280
order effects. 
Once KV cache compression is 

313
00:19:25,280 --> 00:19:28,920
standard, the entire model 
serving landscape shifts. 

314
00:19:30,160 --> 00:19:33,920
Longer context windows become 
practical on smaller hardware. 

315
00:19:35,000 --> 00:19:37,800
Edge deployment of large models 
becomes feasible. 

316
00:19:38,400 --> 00:19:41,920
The cost per query drops enough 
that use cases that were 

317
00:19:41,920 --> 00:19:45,720
previously uneconomical. 
Real time document analysis 

318
00:19:45,800 --> 00:19:50,440
always on coding assistance with
full repository context. 

319
00:19:50,720 --> 00:19:53,520
Persistent multi turn agents 
that remember everything. 

320
00:19:53,840 --> 00:19:55,920
All of those become commercially
viable. 

321
00:19:56,720 --> 00:20:00,960
The memory wall is not just a 
technical constraint, it is a 

322
00:20:00,960 --> 00:20:04,120
capability gate and we are about
to open it. 

323
00:20:05,200 --> 00:20:08,560
There is one more thing I want 
to address because I know some 

324
00:20:08,560 --> 00:20:12,280
of you are thinking it. 
What about multi query attention

325
00:20:12,280 --> 00:20:16,280
and grouped query attention? 
These are architectural changes 

326
00:20:16,280 --> 00:20:20,760
that share key and value heads 
across multiple query heads and 

327
00:20:20,760 --> 00:20:26,200
they do reduce KV cache size. 
GQA with eight groups instead of

328
00:20:26,200 --> 00:20:30,440
full multi head attention cuts 
the KV cache by the number of 

329
00:20:30,440 --> 00:20:33,520
query heads divided by the 
number of key value groups. 

330
00:20:34,280 --> 00:20:39,040
Lama 3, Gemma 2 and most recent 
models use some form of this. 

331
00:20:39,600 --> 00:20:43,000
But here is the thing. 
GQA is an architectural decision

332
00:20:43,000 --> 00:20:46,240
made at training time. 
You cannot retrofit it onto an 

333
00:20:46,240 --> 00:20:49,880
existing model. 
And even with GQA the KV cache 

334
00:20:49,880 --> 00:20:52,160
still scales linearly with 
sequence length. 

335
00:20:52,840 --> 00:20:55,480
At long enough context you hit 
the wall again. 

336
00:20:57,520 --> 00:21:01,400
The compression techniques I 
have been talking about, KVTC 

337
00:21:01,400 --> 00:21:05,080
and Turboquant, work on top of 
whatever attention architecture 

338
00:21:05,080 --> 00:21:07,960
the model uses. 
They are complementary, not 

339
00:21:07,960 --> 00:21:13,240
competing a model with GQA plus 
turboquant gets the benefits of 

340
00:21:13,240 --> 00:21:16,920
both. 
The memory wall is real, it is 

341
00:21:16,920 --> 00:21:20,720
expensive, and it is solvable. 
The research is ahead of the 

342
00:21:20,720 --> 00:21:24,360
production tooling right now, 
but that gap is closing fast. 

343
00:21:25,120 --> 00:21:28,800
Most important thing you can do 
today is stop thinking about 

344
00:21:28,960 --> 00:21:32,720
inference optimization as a 
compute problem and start 

345
00:21:32,720 --> 00:21:37,720
thinking about it as a memory 
problem, because the GPU is not 

346
00:21:37,720 --> 00:21:40,400
your bottleneck, The KV cache 
is. 

347
00:21:41,240 --> 00:21:44,400
That's the one for this episode 
in MO. 

348
00:21:44,480 --> 00:21:48,760
This is the practical AI digest.
If you got something out of 

349
00:21:48,760 --> 00:21:51,600
this, send it to one person 
who'd actually use it. 

350
00:21:52,360 --> 00:21:55,920
No subscribe campaign, no rating
farm, just one person. 

351
00:21:56,120 --> 00:21:57,160
See you on the next one.
