1
00:00:00,040 --> 00:00:04,440
Welcome back to the Deep Dive. 
So I was looking at a photo of a

2
00:00:04,440 --> 00:00:08,320
modern beta center the other day
and it just, it struck me. 

3
00:00:08,480 --> 00:00:12,200
It didn't look anything like the
server rooms I used to see in, 

4
00:00:12,200 --> 00:00:15,280
you know, corporate offices. 
The clean white rooms with all 

5
00:00:15,280 --> 00:00:18,040
the blinking blue lights. 
Read The Sterile Environment. 

6
00:00:18,040 --> 00:00:21,680
Yeah, this this looks like the 
engine room of a nuclear 

7
00:00:21,680 --> 00:00:23,880
submarine. 
It basically is an engine room. 

8
00:00:23,880 --> 00:00:28,000
It really felt like it just rows
and rows of these black monolith

9
00:00:28,000 --> 00:00:30,560
cabinets in this massive 
plumbing and structure 

10
00:00:30,560 --> 00:00:34,120
everywhere for liquid cooling, 
and you could almost feel this 

11
00:00:34,120 --> 00:00:38,960
low vibrating hum of incredible 
energy just from the picture. 

12
00:00:38,960 --> 00:00:43,160
It felt heavy. 
It is heavy physically and in 

13
00:00:43,160 --> 00:00:45,760
terms of power draw. 
And that's really what triggered

14
00:00:45,760 --> 00:00:48,400
today's deep dive. 
We spend so much time analyzing 

15
00:00:48,400 --> 00:00:52,360
the mind of AI, the algorithms, 
prompt engineering, the ethics. 

16
00:00:52,360 --> 00:00:54,560
Chatbots that write poetry. 
Exactly. 

17
00:00:54,800 --> 00:00:57,880
But we almost never stopped to 
talk about the brain, the actual

18
00:00:57,880 --> 00:01:00,600
physical matter that makes any 
of that thinking possible. 

19
00:01:00,800 --> 00:01:02,160
Which is, you know, kind of 
ironic. 

20
00:01:02,280 --> 00:01:04,640
Yeah, because without the 
specific breakthroughs in 

21
00:01:04,640 --> 00:01:07,800
silicon and physics we've seen 
over the last decade, none of 

22
00:01:07,800 --> 00:01:09,800
the software we talked about 
would even exist. 

23
00:01:09,800 --> 00:01:11,680
Not at all. 
Jet GPT would still just be a 

24
00:01:11,680 --> 00:01:13,680
theoretical paper on a dusty 
shelf somewhere. 

25
00:01:14,000 --> 00:01:17,160
So that is our mission today. 
We are stripping away the hype 

26
00:01:17,160 --> 00:01:19,480
to look at the metal. 
We're digging into the A 

27
00:01:19,600 --> 00:01:25,560
Sixteens AI hardware explained 
guide and also the latest info Q

28
00:01:25,560 --> 00:01:29,160
Trends report to really 
understand the physical engine 

29
00:01:29,160 --> 00:01:30,880
of artificial intelligence, 
right? 

30
00:01:31,200 --> 00:01:33,000
We are not talking about prompts
today. 

31
00:01:33,000 --> 00:01:35,680
We were talking about the 
silicon that makes them think. 

32
00:01:35,760 --> 00:01:38,160
The Foundation. 
And I'll be honest, I'm kind of 

33
00:01:38,160 --> 00:01:40,680
relieved we're doing this 
because I want to finally 

34
00:01:40,680 --> 00:01:43,840
understand why NVIDIA is so 
important beyond just, you know,

35
00:01:43,840 --> 00:01:45,880
the stock price and it. 
Really is the bottleneck of the 

36
00:01:45,880 --> 00:01:47,960
future. 
I mean, software is infinite, 

37
00:01:47,960 --> 00:01:50,320
you can always write more code. 
But physics? 

38
00:01:50,600 --> 00:01:54,440
Physics is finite. 
And the hook here, and this is 

39
00:01:54,440 --> 00:01:56,600
the thing that really surprised 
me when I was reading through 

40
00:01:56,600 --> 00:02:01,000
the sources, is that this entire
revolution is, well, essentially

41
00:02:01,000 --> 00:02:03,000
an accident. 
It's completely accidental. 

42
00:02:03,000 --> 00:02:04,560
That's the amazing part of the 
story. 

43
00:02:04,680 --> 00:02:08,199
That part just blew my mind. 
The hardware that's driving the 

44
00:02:08,199 --> 00:02:10,400
most advanced technology in 
human history. 

45
00:02:11,000 --> 00:02:13,960
It wasn't built for scientists. 
No, it wasn't built for AI 

46
00:02:13,960 --> 00:02:15,960
researchers at all. 
It was built for teenagers. 

47
00:02:15,960 --> 00:02:18,600
Exactly. 
It was built for gamers trying 

48
00:02:18,600 --> 00:02:21,400
to get, you know, ridiculously 
high frame rates on Call of 

49
00:02:21,400 --> 00:02:23,000
Duty. 
And then later for the crypto 

50
00:02:23,000 --> 00:02:24,800
miners? 
And later for the crypto guys 

51
00:02:24,800 --> 00:02:26,280
mining Bitcoin in their 
basements. 

52
00:02:26,280 --> 00:02:30,800
Yeah, the gaming KC and the 
Bitcoin mining rig, those were 

53
00:02:30,800 --> 00:02:33,360
the prototypes for the AI 
supercomputers we see today. 

54
00:02:33,560 --> 00:02:36,600
OK, so help me bridge that gap. 
How do we go from shooting 

55
00:02:36,600 --> 00:02:41,520
zombies and 4K resolution to 
building large language models? 

56
00:02:41,720 --> 00:02:45,000
What is the architectural link 
between a video game and a 

57
00:02:45,000 --> 00:02:48,200
neural network? 
It all comes down to a really 

58
00:02:48,200 --> 00:02:52,440
fundamental difference in how 
computers will not think, but 

59
00:02:52,440 --> 00:02:56,240
how they process instructions. 
To understand AI hardware, you 

60
00:02:56,240 --> 00:02:59,080
have to understand the 
difference between the CPU, the 

61
00:02:59,080 --> 00:03:01,240
central processing unit that's 
probably in your laptop right 

62
00:03:01,240 --> 00:03:04,040
now, and the GPU, the graphics 
processing unit. 

63
00:03:04,280 --> 00:03:07,760
Let's start with the CPU then. 
This is the Intel or AMD chip, 

64
00:03:07,840 --> 00:03:10,880
the brain of personal computing 
for what, 40 years now? 

65
00:03:11,120 --> 00:03:13,200
Right. 
And you can think of the CPU as 

66
00:03:13,200 --> 00:03:17,200
a genius mathematician, or maybe
a really sharp high-powered 

67
00:03:17,200 --> 00:03:19,320
executive. 
OK, I like that analogy. 

68
00:03:19,320 --> 00:03:21,000
It's designed to be incredibly 
versatile. 

69
00:03:21,200 --> 00:03:23,840
You can run your operating 
system, open a huge spreadsheet,

70
00:03:23,840 --> 00:03:26,480
play a song, check your e-mail, 
manage your WI. 

71
00:03:26,480 --> 00:03:28,400
Fi all at the same time, 
basically. 

72
00:03:28,400 --> 00:03:30,720
Yeah, it switches between all 
these different tasks 

73
00:03:30,720 --> 00:03:33,720
constantly. 
It's the ultimate multitasker. 

74
00:03:33,720 --> 00:03:36,920
So it's a generalist. 
It is, but structurally it's a 

75
00:03:36,920 --> 00:03:40,280
sequential thinker. 
A classic CPU is designed to 

76
00:03:40,280 --> 00:03:44,080
take a complex logic problem and
work through it step by step. 

77
00:03:44,080 --> 00:03:46,720
One instruction at a time. 
One instruction at a time. 

78
00:03:46,720 --> 00:03:51,200
Historically now, even modern 
CPUs with multiple cores, say 12

79
00:03:51,200 --> 00:03:54,480
or 16 cores, are still just a 
small team of those genius 

80
00:03:54,480 --> 00:03:57,120
mathematicians. 
They're optimized for latency. 

81
00:03:57,280 --> 00:03:59,440
Latency, meaning speed on one 
task. 

82
00:03:59,440 --> 00:04:01,120
Exactly. 
They want to finish one hard 

83
00:04:01,120 --> 00:04:03,200
task as fast as possible so they
can get to the next one. 

84
00:04:03,360 --> 00:04:06,840
So if you give a CPU a calculus 
problem or a complex logic tree 

85
00:04:06,840 --> 00:04:09,280
like if this then that, but it's
not, then go over here. 

86
00:04:09,280 --> 00:04:11,240
It solves it instantly 
brilliantly. 

87
00:04:11,240 --> 00:04:14,360
Yeah, now let's look at the GPU,
the graphics processing unit. 

88
00:04:14,560 --> 00:04:16,480
The other side of the coin. 
Completely. 

89
00:04:16,959 --> 00:04:21,920
If the CPU is that genius 
mathematician, the GPU is an 

90
00:04:21,920 --> 00:04:25,520
army of 10,000 laborers. 
None of them are particularly 

91
00:04:25,520 --> 00:04:27,440
smart on their own. 
They can't run an operating 

92
00:04:27,440 --> 00:04:29,320
system. 
They're terrible at complex 

93
00:04:29,320 --> 00:04:32,400
logic branches or jumping 
between different tasks. 

94
00:04:32,560 --> 00:04:35,080
But there are just a ton of. 
Them massive amounts. 

95
00:04:35,520 --> 00:04:40,160
While ACPU might do say 10s of 
operations per cycle, a modern 

96
00:04:40,200 --> 00:04:44,440
AI accelerator or GPU can 
execute over 100,000 

97
00:04:44,440 --> 00:04:47,800
instructions per cycle. 100,000 
that's that's a completely 

98
00:04:47,800 --> 00:04:49,560
different scale. 
It's a different universe. 

99
00:04:49,720 --> 00:04:52,600
They are built for throughput. 
They're designed to do 1 simple 

100
00:04:52,600 --> 00:04:55,840
thing, but do it to a massive 
amount of data all at once. 

101
00:04:55,840 --> 00:04:58,360
And this is where the gaming 
connection comes back in. 

102
00:04:58,440 --> 00:05:01,880
Because rendering a video game, 
that's not a logic puzzle, 

103
00:05:01,880 --> 00:05:03,280
right? 
Think about your screen. 

104
00:05:03,680 --> 00:05:08,200
A 4K screen has roughly 8 
million pixels to render a 

105
00:05:08,200 --> 00:05:11,600
single frame of video game, and 
you want sixty of those every 

106
00:05:11,600 --> 00:05:13,480
second. 
You have to calculate the color,

107
00:05:13,600 --> 00:05:16,400
the lighting, the texture for 
every single one of those 8 

108
00:05:16,400 --> 00:05:18,000
million pixels that the the 
exact same time. 

109
00:05:18,000 --> 00:05:19,800
You can't do it 1 by 1. 
You can't. 

110
00:05:19,800 --> 00:05:22,960
You don't want to calculate 
pixel one, then pixel 2, then 

111
00:05:22,960 --> 00:05:25,040
pixel three. 
It would take forever. 

112
00:05:25,280 --> 00:05:26,840
You need to do the whole screen 
at once. 

113
00:05:26,840 --> 00:05:29,680
So it's a parallel problem. 
It is the definition of a 

114
00:05:29,680 --> 00:05:32,520
parallel problem. 
It's a massive, relatively dumb 

115
00:05:32,680 --> 00:05:35,680
math problem where you perform 
the exact same calculation 

116
00:05:35,680 --> 00:05:39,480
millions of times 
simultaneously, and it turns out

117
00:05:40,080 --> 00:05:43,320
that is exactly mathematically 
what a neural network is. 

118
00:05:43,320 --> 00:05:44,640
OK, I want to double click on 
that. 

119
00:05:44,640 --> 00:05:48,160
Why is a neural network like a 
video game screen? 

120
00:05:48,200 --> 00:05:51,200
That feels like a leap. 
It's because deep learning, at 

121
00:05:51,200 --> 00:05:55,320
its very core isn't a complex 
logic tree like old software. 

122
00:05:55,520 --> 00:05:57,840
It's not a bunch of isn't 
statements, right? 

123
00:05:57,840 --> 00:06:01,200
It represents information as 
massive stacks of matrix 

124
00:06:01,200 --> 00:06:03,720
multiplications. 
It's just rows and columns of 

125
00:06:03,720 --> 00:06:05,520
numbers. 
We call them matrices being 

126
00:06:05,520 --> 00:06:08,160
multiplied against each other 
over and over again to find 

127
00:06:08,160 --> 00:06:10,520
patterns. 
So when I asked Strat GPTA 

128
00:06:10,520 --> 00:06:14,040
question, it's not thinking in a
straight line like ACPU 

129
00:06:14,120 --> 00:06:15,600
processing a script. 
Not at all. 

130
00:06:15,800 --> 00:06:18,680
It's activating billions of 
parameters all at once to 

131
00:06:18,680 --> 00:06:20,440
predict the next word. 
Correct. 

132
00:06:20,520 --> 00:06:24,680
It's taking a massive grid of 
numbers that's your input and 

133
00:06:24,680 --> 00:06:28,240
multiplying it through layers of
other massive grids of number. 

134
00:06:28,280 --> 00:06:29,640
Those are the weights of the 
model. 

135
00:06:29,760 --> 00:06:32,280
You don't need the genius 
mathematician for that. 

136
00:06:32,320 --> 00:06:35,160
You need the army of laborers. 
You need the army of laborers 

137
00:06:35,360 --> 00:06:37,960
doing massive simultaneous 
arithmetic. 

138
00:06:38,280 --> 00:06:40,440
You know, this reminds reminds 
me of that famous Marc 

139
00:06:40,440 --> 00:06:44,920
Andreessen quote from what 2011 
Software is Eating the World? 

140
00:06:45,760 --> 00:06:49,400
But looking at these reports and
seeing how the demands of neural

141
00:06:49,400 --> 00:06:53,760
networks are just completely 
reshaping the entire 

142
00:06:53,760 --> 00:06:57,520
semiconductor industry, it feels
like the dynamic has shifted. 

143
00:06:57,520 --> 00:07:00,280
It feels like hardware is now 
eating software. 

144
00:07:00,280 --> 00:07:02,880
I'd say it's a symbiotic 
relationship, but you're right, 

145
00:07:02,880 --> 00:07:05,280
the hardware is now dictating 
the capability. 

146
00:07:05,720 --> 00:07:07,800
And if you're a software 
engineer listening to this, the 

147
00:07:07,800 --> 00:07:10,840
big take away here is that you 
have to stop thinking in loops, 

148
00:07:10,880 --> 00:07:13,120
which is that sequential CPU 
style of. 

149
00:07:13,120 --> 00:07:14,920
Thinking and start thinking in 
tensors. 

150
00:07:14,920 --> 00:07:16,680
And start thinking in tensors 
absolutely. 

151
00:07:16,680 --> 00:07:18,680
OK, let's define that word, 
because we hear it everywhere. 

152
00:07:18,840 --> 00:07:22,760
Tensor flow, tensor cores, TPU. 
It sounds incredibly technical, 

153
00:07:22,760 --> 00:07:25,520
but what actually is a tensor? 
It really sounds like sci-fi 

154
00:07:25,520 --> 00:07:27,960
jargon, doesn't it? 
But it's actually just geometry.

155
00:07:28,800 --> 00:07:32,240
I mean, you know what vector is.
Yeah, a line of numbers like a 

156
00:07:32,240 --> 00:07:36,040
list in an array, right? 
A1D array and a matrix. 

157
00:07:36,280 --> 00:07:39,880
That's a grid of numbers, like a
spreadsheet, rows and columns 

158
00:07:39,880 --> 00:07:42,400
SO2DA. 
Tensor is just the next step up.

159
00:07:42,400 --> 00:07:45,680
It's a multi dimensional array. 
It could be a cube of numbers 

160
00:07:45,680 --> 00:07:46,680
that would be 3D. 
Yeah. 

161
00:07:47,080 --> 00:07:49,120
Or a structure with even more 
dimensions. 

162
00:07:49,600 --> 00:07:52,320
It's just a generic term for a 
container of numbers. 

163
00:07:52,320 --> 00:07:54,160
So it's just a bucket for 
numbers. 

164
00:07:54,320 --> 00:07:56,880
Essentially, yes. 
And the reason we talk about 

165
00:07:56,880 --> 00:08:00,160
them so much is that companies 
literally started building chips

166
00:08:00,680 --> 00:08:03,600
specifically to move these 
containers around. 

167
00:08:04,080 --> 00:08:07,320
You mentioned Google's TPU, the 
Tensor Processing Unit. 

168
00:08:07,320 --> 00:08:09,720
The name is literal. 
The name is completely literal. 

169
00:08:09,960 --> 00:08:12,840
It is a processor physically 
designed to manipulate these 

170
00:08:12,840 --> 00:08:15,960
specific data containers and 
multiply them against each other

171
00:08:15,960 --> 00:08:18,880
as fast as humanly possible. 
So we have the architecture down

172
00:08:19,000 --> 00:08:21,480
the T-21 because it's a parallel
math engine. 

173
00:08:21,840 --> 00:08:24,200
OK, now let's talk about the 
landscape of players. 

174
00:08:24,560 --> 00:08:26,840
We can't have a conversation 
about AI hardware without 

175
00:08:26,840 --> 00:08:30,720
talking about the king of the 
hill, NVIDIA. 

176
00:08:31,160 --> 00:08:34,080
Their data center revenue is 
just astronomical. 

177
00:08:34,559 --> 00:08:36,600
But here's something I've always
found a little puzzling. 

178
00:08:37,280 --> 00:08:41,039
If the hardware is just about 
parallel math, why hasn't Intel 

179
00:08:41,039 --> 00:08:43,840
or AMD or Qualcomm just caught 
up? 

180
00:08:44,039 --> 00:08:47,160
That's a great question. 
I mean why is NVIDIA a trillion 

181
00:08:47,160 --> 00:08:49,680
dollar company and everyone else
seems to be fighting for the 

182
00:08:49,680 --> 00:08:52,320
scraps? 
Is their silicon really that 

183
00:08:52,320 --> 00:08:54,440
much better? 
And this is probably the most 

184
00:08:54,440 --> 00:08:56,840
misunderstood part of the entire
NVIDIA story. 

185
00:08:56,840 --> 00:09:00,080
Everyone looks at the H-100 or 
the new Blackwell chips and they

186
00:09:00,080 --> 00:09:01,640
think the magic is in the 
silicon. 

187
00:09:01,680 --> 00:09:03,680
And it's not. 
Well, the silicon is amazing, 

188
00:09:03,680 --> 00:09:05,680
don't get me wrong. 
It's a marvel of engineering, 

189
00:09:06,400 --> 00:09:08,240
but that's not the Moat. 
So what is? 

190
00:09:08,640 --> 00:09:11,840
The Moat is software. 
You're talking about CUD. 

191
00:09:11,840 --> 00:09:14,240
CUDA, Exactly. 
It stands for Compute Unified 

192
00:09:14,240 --> 00:09:16,720
Device Architecture. 
It's the software layer, the 

193
00:09:16,720 --> 00:09:19,600
programming model that allows 
developers to actually talk to 

194
00:09:19,600 --> 00:09:21,520
the GPU. 
And they've been working on this

195
00:09:21,520 --> 00:09:24,320
for a while. 
NVIDIA released this almost 20 

196
00:09:24,320 --> 00:09:26,960
years ago, way way before the AI
boom. 

197
00:09:27,280 --> 00:09:30,480
They spent decades building out 
libraries, tools, and 

198
00:09:30,480 --> 00:09:33,360
optimization protocols that let 
programmers easily send 

199
00:09:33,360 --> 00:09:35,480
instructions to the graphics 
card for general purpose 

200
00:09:35,480 --> 00:09:37,960
computing. 
So it's the ecosystem, it's not 

201
00:09:37,960 --> 00:09:40,480
just the chip. 
It is 100% the ecosystem. 

202
00:09:40,720 --> 00:09:42,920
If I'm a researcher at a 
university and I download an 

203
00:09:42,920 --> 00:09:47,320
open source model from GitHub 
today, the chances are what, 99%

204
00:09:47,520 --> 00:09:49,480
that the code is written to run 
on CDA. 

205
00:09:49,480 --> 00:09:51,840
Which means it runs 
out-of-the-box on NVIDIA 

206
00:09:51,840 --> 00:09:54,360
hardware. 
Instantly out-of-the-box, no 

207
00:09:54,360 --> 00:09:55,880
hassle. 
And if I try to run that same 

208
00:09:55,880 --> 00:09:59,680
code on say an AMD chip. 
It breaks, or it one's terribly 

209
00:09:59,680 --> 00:10:01,320
slowly. 
You have to use these clunky 

210
00:10:01,320 --> 00:10:05,040
translation layers or you have 
to manually rewrite the code for

211
00:10:05,040 --> 00:10:08,760
AM DS competing platform ROCM. 
Which nobody wants to do that. 

212
00:10:08,760 --> 00:10:11,480
Friction is massive. 
It's vendor locking at the 

213
00:10:11,480 --> 00:10:15,040
deepest, most fundamental level.
The hardware is great, sure, but

214
00:10:15,040 --> 00:10:17,960
the software ecosystem is what 
keeps the competition at Bay. 

215
00:10:18,520 --> 00:10:20,400
Engineers want things to just 
work. 

216
00:10:20,680 --> 00:10:24,800
But the competition is coming. 
The sources list a few key 

217
00:10:24,800 --> 00:10:26,960
players that are trying to chip 
away at this lead. 

218
00:10:27,640 --> 00:10:30,600
We mentioned Google's TPU, which
they mostly use internally. 

219
00:10:30,960 --> 00:10:32,800
Who else is making serious 
moves? 

220
00:10:33,040 --> 00:10:36,960
Well, AWS, Amazon Web Services 
is a fascinating one, right? 

221
00:10:37,000 --> 00:10:41,040
Amazon isn't really trying to 
sell chips to the public, but 

222
00:10:41,160 --> 00:10:44,480
they are building custom silicon
for their own massive cloud 

223
00:10:44,480 --> 00:10:48,040
servers, and they've split the 
problem into two distinct parts.

224
00:10:48,760 --> 00:10:52,520
They have Tranium in Inferentia.
And that distinction is really 

225
00:10:52,520 --> 00:10:55,200
important for listeners to 
understand, right the whole life

226
00:10:55,200 --> 00:10:59,240
cycle of an AI model, what's the
difference between training and 

227
00:10:59,240 --> 00:11:01,720
inference? 
Think of training as the models 

228
00:11:01,720 --> 00:11:03,920
university education. 
That's when it's learning 

229
00:11:03,920 --> 00:11:05,880
everything it knows. 
Reading the whole Internet. 

230
00:11:05,880 --> 00:11:08,240
Literally, it's reading the 
entire Internet, finding 

231
00:11:08,240 --> 00:11:11,160
patterns, adjusting billions of 
tiny weights in its network. 

232
00:11:11,680 --> 00:11:13,720
This process takes weeks, 
sometimes months. 

233
00:11:13,720 --> 00:11:17,200
It requires enormous data sets, 
and it consumes just insane 

234
00:11:17,200 --> 00:11:19,040
amounts of power. 
That's what training was 

235
00:11:19,040 --> 00:11:21,160
designed for. 
Brute force learning, high 

236
00:11:21,160 --> 00:11:24,080
bandwidth, massive clusters of 
chips working together. 

237
00:11:24,600 --> 00:11:27,280
Inference is when the graduate 
finally gets a job. 

238
00:11:27,920 --> 00:11:29,640
It's when you actually use the 
model. 

239
00:11:29,800 --> 00:11:32,680
You ask it a question that gives
you an answer that has to happen

240
00:11:32,680 --> 00:11:35,760
in milliseconds. 
It requires less power per 

241
00:11:35,760 --> 00:11:40,000
transaction, but it needs to be 
incredibly fast and responsive, 

242
00:11:41,040 --> 00:11:43,680
so low latency. 
That's what infringe is for. 

243
00:11:43,880 --> 00:11:47,200
So AWS is trying to specialize 
the hardware for the specific 

244
00:11:47,200 --> 00:11:50,440
stage of the AI life cycle 
instead of using one general 

245
00:11:50,440 --> 00:11:52,320
purpose GPU for everything. 
Exactly. 

246
00:11:52,320 --> 00:11:55,040
It's all about efficiency and 
cost optimization at scale. 

247
00:11:55,640 --> 00:11:57,760
And then you have the startups 
that are just pushing the 

248
00:11:57,760 --> 00:12:00,760
boundaries of physics itself. 
The one that always, always 

249
00:12:00,760 --> 00:12:03,920
blows my mind is Cerebrous. 
All right, the way for scale 

250
00:12:03,920 --> 00:12:06,680
engine, I saw a picture of this 
thing, it looks, it looks 

251
00:12:06,680 --> 00:12:08,480
ridiculous. 
Looks like a prop from a sci-fi 

252
00:12:08,480 --> 00:12:10,880
movie. 
So usually to make a chip you 

253
00:12:10,880 --> 00:12:14,240
start with a big silicon wafer. 
Looks like a big round dinner 

254
00:12:14,240 --> 00:12:17,400
plate made of silicon and you 
etch hundreds of little square 

255
00:12:17,400 --> 00:12:18,960
chips onto. 
It then you cut them up, then 

256
00:12:19,000 --> 00:12:21,160
you cut it. 
Up Intel does this, NVIDIA does 

257
00:12:21,160 --> 00:12:24,480
this, everyone does this. 
Cerebras looked at that process 

258
00:12:24,480 --> 00:12:27,920
and said, what if we just didn't
cut it? 

259
00:12:28,000 --> 00:12:29,800
They use the whole plate as one 
chip. 

260
00:12:29,960 --> 00:12:31,880
The entire wafer is the 
processor. 

261
00:12:31,880 --> 00:12:36,440
It's the size of a keyboard. 
It has 2.6 trillion transistors 

262
00:12:36,640 --> 00:12:40,320
on a single piece of silicon. 
2.6 trillion. 

263
00:12:40,640 --> 00:12:43,280
That number is hard to even 
conceptualize. 

264
00:12:43,280 --> 00:12:46,000
Why on earth would you do? 
That to beat the speed of light.

265
00:12:46,120 --> 00:12:48,520
OK, explain that. 
One of the biggest slowdowns in 

266
00:12:48,640 --> 00:12:53,240
AI at scale is moving data from 
one chip to another across a 

267
00:12:53,240 --> 00:12:56,160
physical wire. 
Wires have resistance. 

268
00:12:56,480 --> 00:12:59,080
Signals take time to travel, 
even at lightspeed. 

269
00:12:59,600 --> 00:13:03,160
If the whole computer is 1 giant
chip, the data never has to 

270
00:13:03,160 --> 00:13:04,760
leave the silicon. 
It's all local. 

271
00:13:04,960 --> 00:13:08,400
It moves at effectively the 
speed of light across the wafer.

272
00:13:08,840 --> 00:13:11,720
It's an attempt to completely 
eliminate the communication 

273
00:13:11,720 --> 00:13:14,280
bottleneck between chips. 
So we have these massive chips. 

274
00:13:14,280 --> 00:13:17,200
We've got NVIDIA, Google, 
Cerebras, all pushing the 

275
00:13:17,200 --> 00:13:19,800
physical limits. 
But there's a practical side to 

276
00:13:19,800 --> 00:13:21,240
this, right? 
We can't just keep building 

277
00:13:21,240 --> 00:13:24,000
bigger and bigger brains if we 
can't fit them in our pockets or

278
00:13:24,000 --> 00:13:26,240
if they cost $10 million to run 
for an afternoon. 

279
00:13:26,240 --> 00:13:28,920
Exactly. 
Which brings us to optimization.

280
00:13:29,120 --> 00:13:30,520
Right. 
And this is where the rubber 

281
00:13:30,520 --> 00:13:32,240
really meets the road for 
engineers. 

282
00:13:32,680 --> 00:13:35,760
If you're building an AI product
today, you are constantly 

283
00:13:35,760 --> 00:13:39,480
fighting against memory limits. 
You have this massive, beautiful

284
00:13:39,480 --> 00:13:42,520
model, and you have to figure 
out how to fit it onto a GPU 

285
00:13:42,520 --> 00:13:44,800
that has a fixed amount of VRAM.
You can't just, you know, 

286
00:13:44,800 --> 00:13:46,480
download the Internet into your 
laptop. 

287
00:13:47,120 --> 00:13:49,600
And one of the key techniques 
mentioned in the Info queue 

288
00:13:49,600 --> 00:13:52,760
report is quantization. 
It sounds like something from 

289
00:13:52,760 --> 00:13:54,360
Star Trek, but let's break it 
down. 

290
00:13:54,520 --> 00:13:57,840
What are we actually doing when 
we quantize a model? 

291
00:13:58,160 --> 00:14:02,120
At its simplest, we are reducing
its precision to save space. 

292
00:14:03,080 --> 00:14:05,840
In a computer, numbers are 
typically stored as floating 

293
00:14:05,840 --> 00:14:09,120
point numbers, and the standard 
for scientific computing has 

294
00:14:09,120 --> 00:14:14,320
always been 32 bit or FP-32. 
That means every single weight, 

295
00:14:14,440 --> 00:14:17,440
every parameter in the neural 
network is represented by a 

296
00:14:17,440 --> 00:14:20,520
string of 32 zeros and ones. 
Which gives you incredibly high 

297
00:14:20,520 --> 00:14:22,560
precision. 
You can represent tiny, tiny 

298
00:14:22,560 --> 00:14:25,520
fractions out to dozens of 
decimal places if you need to. 

299
00:14:25,520 --> 00:14:28,320
Exactly. 
But engineers started asking a 

300
00:14:28,320 --> 00:14:31,320
really important question. 
Do we actually need that? 

301
00:14:31,520 --> 00:14:34,200
Think about it this way. 
If I ask you to drive a car 

302
00:14:34,200 --> 00:14:37,840
through a garage door, do you 
need to know the width of that 

303
00:14:37,840 --> 00:14:41,360
doorway down to the micrometer? 
No, not at all. 

304
00:14:41,360 --> 00:14:44,720
I just need to know if it's, you
know, roughly wider than my car.

305
00:14:44,760 --> 00:14:47,600
Precisely, neural networks are 
probabilistic. 

306
00:14:47,720 --> 00:14:51,160
They are fuzzy systems. 
They don't usually need 32 bit 

307
00:14:51,160 --> 00:14:52,800
precision to get the right 
answer. 

308
00:14:53,200 --> 00:14:55,200
So engineers started chopping 
the bits. 

309
00:14:55,240 --> 00:14:57,320
How far down? 
They went to 16 bit which is 

310
00:14:57,320 --> 00:15:00,120
half the size. 
Then they went even further to 8

311
00:15:00,120 --> 00:15:03,000
bit integers or in date. 
So you're literally cutting the 

312
00:15:03,000 --> 00:15:06,480
data size in half, and then in 
half again you're down to 1/4 of

313
00:15:06,480 --> 00:15:08,440
the original. 
Size and the benefits are 

314
00:15:08,440 --> 00:15:10,120
massive. 
First, memory. 

315
00:15:10,560 --> 00:15:14,720
An 8 bit model takes up 1/4 of 
the space of a 32 bit model. 

316
00:15:15,040 --> 00:15:18,880
Suddenly a model that required 4
expensive GPU's can run on just 

317
00:15:18,880 --> 00:15:20,400
one. 
And the second benefit? 

318
00:15:20,400 --> 00:15:22,680
Speed. 
Computers can multiply 8 bit 

319
00:15:22,680 --> 00:15:25,720
numbers way, way faster than 
they can multiply 32 bit 

320
00:15:25,720 --> 00:15:27,800
numbers. 
It's just less math for the 

321
00:15:27,800 --> 00:15:29,720
silicon to do. 
But there's no free lunch in 

322
00:15:29,720 --> 00:15:32,320
physics or engineering. 
You're throwing away information

323
00:15:32,320 --> 00:15:33,760
when you do this. 
What's the trade off? 

324
00:15:33,760 --> 00:15:36,520
What do you lose? 
You lose accuracy and you 

325
00:15:36,520 --> 00:15:40,560
introduce noise, but the biggest
and most dangerous risk is 

326
00:15:40,560 --> 00:15:42,920
something called overrun or 
underrun. 

327
00:15:43,160 --> 00:15:45,600
Explain overrun. 
I want to visualize this OK. 

328
00:15:45,600 --> 00:15:48,960
Think about the odometer on a 
really old car, the kind with 

329
00:15:48,960 --> 00:15:53,760
the physical spinning numbers. 
It only goes to say 99,999 

330
00:15:53,760 --> 00:15:56,200
miles, right? 
If you drive one more mile, what

331
00:15:56,200 --> 00:16:00,640
happens it? 
Rolls over to all zeros 0000000.

332
00:16:00,640 --> 00:16:03,000
Right, that's an overrun. 
You tried to store a number that

333
00:16:03,000 --> 00:16:04,800
was too big for the container 
you had. 

334
00:16:05,400 --> 00:16:08,720
Now in computing, an 8 bit 
integer is a container that can 

335
00:16:08,720 --> 00:16:11,960
only hold a number between 9:00 
is 128 and plus 127. 

336
00:16:12,440 --> 00:16:13,920
That's it. 
That's the hard limit. 

337
00:16:13,920 --> 00:16:16,120
That is a tiny, tiny range. 
It is. 

338
00:16:16,320 --> 00:16:18,840
So imagine you're doing a matrix
multiplication in your neural 

339
00:16:18,840 --> 00:16:21,520
net and the correct mathematical
result is 200. 

340
00:16:22,000 --> 00:16:23,520
The 8 bit container can't hold 
it. 

341
00:16:24,000 --> 00:16:25,200
The computer doesn't know what 
to do. 

342
00:16:25,320 --> 00:16:27,320
It might wrap around and call it
a negative number. 

343
00:16:27,400 --> 00:16:30,040
Which would completely trash the
logic of the model. 

344
00:16:30,040 --> 00:16:32,480
Instant garbage output. 
It's like your model confusing 

345
00:16:32,480 --> 00:16:35,320
very hot with very cold just 
because the thermometer broke at

346
00:16:35,320 --> 00:16:37,640
the top. 
So the art of quantization is 

347
00:16:37,640 --> 00:16:39,600
actually something called 
normalization. 

348
00:16:39,800 --> 00:16:43,080
Normalization, yes. 
Before you run the math, you 

349
00:16:43,080 --> 00:16:46,080
have to look at all of your data
and mathematically squash it. 

350
00:16:46,520 --> 00:16:49,720
You scale everything down so 
that the biggest and smallest 

351
00:16:49,720 --> 00:16:54,320
numbers fit perfectly inside 
that -128 to plus 127 range. 

352
00:16:54,440 --> 00:16:55,960
It's like trying to pack a 
suitcase. 

353
00:16:56,280 --> 00:16:58,120
You have to fold everything just
right. 

354
00:16:58,120 --> 00:17:00,480
If you just shove it all in, 
you'll break the zipper. 

355
00:17:00,720 --> 00:17:04,079
That is a perfect analogy. 
Quantization is the art of 

356
00:17:04,079 --> 00:17:06,960
folding the data so it fits in 
the hardware suitcase. 

357
00:17:07,359 --> 00:17:09,760
If you fold it too much, you 
wrinkle the close, you lose too 

358
00:17:09,760 --> 00:17:11,800
much accuracy. 
If you don't fold it enough, the

359
00:17:11,800 --> 00:17:13,839
suitcase doesn't close, you get 
an overrun. 

360
00:17:14,200 --> 00:17:16,599
That makes perfect sense. 
O we're otimizing the software 

361
00:17:16,599 --> 00:17:18,839
because the hardware has these 
very real limits. 

362
00:17:18,839 --> 00:17:20,640
And I want to talk more about 
those limits. 

363
00:17:20,960 --> 00:17:23,400
We always hear about Moore's law
that chips get twice as good 

364
00:17:23,400 --> 00:17:24,319
every two years. 
Is that? 

365
00:17:24,720 --> 00:17:27,560
Is that actually still true? 
This is a huge debate in the 

366
00:17:27,560 --> 00:17:30,080
hardware community, according to
the sources. 

367
00:17:30,200 --> 00:17:32,760
Strictly speaking, Moore's Law 
is alive. 

368
00:17:32,760 --> 00:17:35,480
It is. 
We are still somehow doubling 

369
00:17:35,480 --> 00:17:38,200
the number of transistors we can
pack into a square millimeter. 

370
00:17:38,240 --> 00:17:42,360
Yeah, Apple's M1 chip had 16 
billion transistors a few years 

371
00:17:42,360 --> 00:17:44,480
ago. 
The newer chips have way more. 

372
00:17:44,800 --> 00:17:48,880
We are getting denser. 
But I sense a but coming. 

373
00:17:48,880 --> 00:17:53,120
There's a huge But while Moore's
Law is alive and kicking, 

374
00:17:53,400 --> 00:17:57,760
another less famous law called 
Dennard scaling is dead, and its

375
00:17:57,760 --> 00:18:01,480
death is causing major headaches
for everyone from NVIDIA all the

376
00:18:01,480 --> 00:18:03,800
way up to the power grid. 
I have not heard of Dennard 

377
00:18:03,840 --> 00:18:04,840
scaling. 
What is it? 

378
00:18:05,280 --> 00:18:07,760
Dennard scaling was a beautiful 
observation from back in the 

379
00:18:07,760 --> 00:18:10,120
1970s. 
It basically said that as 

380
00:18:10,120 --> 00:18:13,400
transistors got smaller and more
dense, they also used less 

381
00:18:13,400 --> 00:18:16,160
power. 
So for a long, long time, we 

382
00:18:16,160 --> 00:18:18,960
could pack twice as many chips 
into the same space and the 

383
00:18:18,960 --> 00:18:21,160
total power usage stayed pretty 
much flat. 

384
00:18:21,600 --> 00:18:23,360
It was magical. 
We got free performance. 

385
00:18:23,360 --> 00:18:26,600
And that stopped working. 
It crashed and burned about 15 

386
00:18:26,600 --> 00:18:27,920
years ago. 
The physics just stopped 

387
00:18:27,920 --> 00:18:30,320
cooperating. 
We hit a quantum leakage limit. 

388
00:18:30,960 --> 00:18:33,840
So now when we pack more 
transistors onto a chip, the 

389
00:18:33,840 --> 00:18:35,840
power usage doesn't stay flat. 
It goes. 

390
00:18:36,200 --> 00:18:38,560
Up it skyrockets and power 
equals heat. 

391
00:18:38,640 --> 00:18:41,040
And we're right back to the 
nuclear engine room analogy we 

392
00:18:41,040 --> 00:18:41,920
started with. 
We are. 

393
00:18:42,240 --> 00:18:46,280
A modern NVIDIA H-100 GPU can 
draw 700 watts of power. 

394
00:18:47,000 --> 00:18:49,360
By itself, just one chip. 
Just one chip. 

395
00:18:49,680 --> 00:18:52,600
Now put eight of them in a 
server rack, which is a standard

396
00:18:52,600 --> 00:18:56,000
configuration, and you have 
thousands and thousands of watts

397
00:18:56,000 --> 00:18:59,440
of heat being generated in a box
the size of a mini fridge. 

398
00:18:59,520 --> 00:19:01,920
That's like running a couple of 
space heaters inside your 

399
00:19:01,920 --> 00:19:03,600
computer case. 
It's insane. 

400
00:19:03,600 --> 00:19:06,640
It's enough to melt silicon. 
This is why you don't really see

401
00:19:06,640 --> 00:19:09,240
fans in these high end AI 
servers anymore. 

402
00:19:09,480 --> 00:19:13,400
Fans aren't enough. 
We are literally pumping fluid 

403
00:19:13,400 --> 00:19:17,400
special dielectric coolant 
directly through the servers to 

404
00:19:17,400 --> 00:19:20,200
carry the heat away. 
And this is the hard physical 

405
00:19:20,200 --> 00:19:22,200
limit we are now hitting. 
It is the limit. 

406
00:19:22,320 --> 00:19:26,240
We can't make the chips run 
faster in terms of clock speed, 

407
00:19:26,440 --> 00:19:29,120
you know, gigahertz because they
would literally catch fire. 

408
00:19:29,600 --> 00:19:31,800
Have you noticed that computers 
haven't actually gotten that 

409
00:19:31,800 --> 00:19:34,120
much faster in terms of 
gigahertz in the last decade? 

410
00:19:34,360 --> 00:19:36,880
We've been stuck around 3 to 5 
gigahertz for a long time. 

411
00:19:36,920 --> 00:19:38,360
That's. 
Wild, I hadn't really thought 

412
00:19:38,360 --> 00:19:40,440
about that, but you're right, my
processor speed is about the 

413
00:19:40,440 --> 00:19:42,760
same as it was on my computer 
from 2015. 

414
00:19:42,800 --> 00:19:44,600
Exactly. 
The only way we get more 

415
00:19:44,600 --> 00:19:46,360
performance now is by going 
wider. 

416
00:19:46,520 --> 00:19:49,240
More cores, more parallelism, 
bigger chips. 

417
00:19:49,680 --> 00:19:52,640
But bigger means more power and 
more heat. 

418
00:19:53,240 --> 00:19:56,320
It's a vicious cycle. 
This physical reality, the heat,

419
00:19:56,440 --> 00:20:00,160
the power, the sheer cost, it 
seems to be forcing a change in 

420
00:20:00,160 --> 00:20:04,000
the entire industry's direction.
The Info Q report really 

421
00:20:04,000 --> 00:20:08,000
highlights a move towards small 
language models or SLMS. Yes, 

422
00:20:08,000 --> 00:20:11,040
this is a direct logical 
response to the hardware 

423
00:20:11,040 --> 00:20:13,120
constraints for the last few 
years. 

424
00:20:13,120 --> 00:20:15,080
The trend was just bigger is 
better. 

425
00:20:15,160 --> 00:20:18,440
GPT 3, GPT 4 are llama 3. 
Exactly. 

426
00:20:18,600 --> 00:20:21,840
Trillions of parameters, but 
those models require massive 

427
00:20:21,840 --> 00:20:24,720
clusters of GPU's that cost 
hundreds of millions of dollars 

428
00:20:24,920 --> 00:20:27,240
and sucked down megawatts of 
power from the grid. 

429
00:20:27,240 --> 00:20:29,480
Which is just unsustainable for 
most companies. 

430
00:20:29,480 --> 00:20:32,200
If I'm a mid sized business, I 
can't afford my own H-100 

431
00:20:32,200 --> 00:20:33,920
cluster You. 
Can't even rank one easily. 

432
00:20:34,280 --> 00:20:35,800
So now we are seeing a 
correction. 

433
00:20:36,040 --> 00:20:39,480
Companies are leasing models 
like Microsoft's FI 3 or Mistral

434
00:20:39,480 --> 00:20:41,560
or Tiny Llama. 
These are models that are 

435
00:20:41,560 --> 00:20:43,840
specifically designed to be 
small and efficient. 

436
00:20:43,840 --> 00:20:46,320
And the whole pitch for these is
that they are good enough for 

437
00:20:46,320 --> 00:20:47,960
most tasks. 
Exactly. 

438
00:20:48,560 --> 00:20:50,800
They might not be able to write 
a Symphony or code an entire 

439
00:20:50,800 --> 00:20:54,480
application from scratch like 
GBT 4 can, but if you just need 

440
00:20:54,480 --> 00:20:57,440
to summarize an e-mail or 
organize a spreadsheet or answer

441
00:20:57,440 --> 00:21:00,360
a common customer support query,
they're perfect. 

442
00:21:00,480 --> 00:21:03,480
And, crucially, they can run on 
the edge. 

443
00:21:03,520 --> 00:21:06,360
The edge, which just means 
locally right on my own device. 

444
00:21:06,400 --> 00:21:08,160
Right. 
Instead of sending your data to 

445
00:21:08,160 --> 00:21:11,520
some massive server farm in 
Virginia to be processed by a 

446
00:21:11,520 --> 00:21:15,720
700 Watt GPU, you process it 
right there on your laptop or 

447
00:21:15,720 --> 00:21:17,920
even your phone. 
I've noticed this with all the 

448
00:21:17,920 --> 00:21:22,960
new hardware releases, the Apple
M4 chip, the new NVIDIA RTX 

449
00:21:22,960 --> 00:21:26,640
cards for laptops, the whole 
AIPC push from Microsoft. 

450
00:21:26,920 --> 00:21:30,080
They are all marketing their NPU
or AI performance now. 

451
00:21:30,160 --> 00:21:32,840
They're optimizing that consumer
grade hardware to run these 

452
00:21:32,840 --> 00:21:34,240
small language models 
efficiently. 

453
00:21:34,680 --> 00:21:37,360
And this isn't just about saving
money on cloud compute bills. 

454
00:21:37,560 --> 00:21:41,280
This is a huge, huge deal for 
privacy and security, which the 

455
00:21:41,280 --> 00:21:43,480
report highlights as a major 
enterprise trend. 

456
00:21:43,600 --> 00:21:45,120
Right. 
Because if I'm a bank or 

457
00:21:45,120 --> 00:21:48,480
hospital or a law firm. 
You do not want to upload your 

458
00:21:48,480 --> 00:21:52,360
patient data or your client's 
sensitive legal strategy to open

459
00:21:52,360 --> 00:21:54,200
AI's cloud. 
You just don't. 

460
00:21:54,400 --> 00:21:57,480
The risk is too high. 
But if you can, run a small 

461
00:21:57,480 --> 00:22:02,240
language model entirely on your 
own secure laptop completely 

462
00:22:02,240 --> 00:22:04,800
disconnected from the Internet. 
That solves the problem. 

463
00:22:04,800 --> 00:22:07,920
It solves a massive problem. 
So the hardware constraint, the 

464
00:22:07,920 --> 00:22:11,320
fact that these big ships are 
scarce, hot and expensive, is 

465
00:22:11,320 --> 00:22:14,160
actually driving better privacy 
and more efficient software 

466
00:22:14,160 --> 00:22:16,440
design. 
Scarcity breeds innovation. 

467
00:22:16,520 --> 00:22:18,800
It has to. 
The sources mentioned that the 

468
00:22:18,800 --> 00:22:22,600
demand for AI hardware outstrips
the supply by a factor of 10. 

469
00:22:22,600 --> 00:22:24,880
Right now. 
You literally cannot buy your 

470
00:22:24,880 --> 00:22:27,400
way out of the problem. 
You have to engineer your way 

471
00:22:27,400 --> 00:22:29,240
out of it. 
So you have to use quantization.

472
00:22:29,240 --> 00:22:31,520
You have to use quantization. 
You have to use SLMS. You have 

473
00:22:31,520 --> 00:22:33,760
to be clever. 
OK, so let's synthesize all of 

474
00:22:33,760 --> 00:22:35,640
this for the listener. 
We covered a lot of ground. 

475
00:22:35,880 --> 00:22:42,480
We started with the the happy 
accident of the gaming PC, how 

476
00:22:42,480 --> 00:22:45,200
graphics cards turned out to be 
the perfect engine for parallel 

477
00:22:45,200 --> 00:22:47,760
processing because of how they 
render pixels. 

478
00:22:47,760 --> 00:22:52,000
Which all runs on matrix math, 
or more generally, tensors. 

479
00:22:52,000 --> 00:22:54,000
That's the native language of 
the modern AI chip. 

480
00:22:54,160 --> 00:22:57,760
Right, but because that hardware
is physically constrained by 

481
00:22:57,760 --> 00:23:01,320
heat and the death of Dennard 
scaling, we can't just brute 

482
00:23:01,320 --> 00:23:03,720
force it forever by making 
things bigger. 

483
00:23:03,720 --> 00:23:05,440
We have to be clever. 
We have to be clever. 

484
00:23:05,680 --> 00:23:09,520
We use optimization tricks like 
quantization to shrink the data 

485
00:23:09,520 --> 00:23:13,280
from 32 bit down to 8 bit while 
carefully managing that risk of 

486
00:23:13,280 --> 00:23:15,600
overrun. 
And the whole industry is now 

487
00:23:15,600 --> 00:23:18,400
moving towards small language 
models to run things efficiently

488
00:23:18,400 --> 00:23:22,120
on the edge, which solves for 
both power and privacy. 

489
00:23:22,120 --> 00:23:24,720
It's a full stack view. 
You can't really understand the 

490
00:23:24,800 --> 00:23:28,160
AI software revolution without 
understanding the silicon it 

491
00:23:28,160 --> 00:23:30,240
lives on. 
The hardware dictates the 

492
00:23:30,240 --> 00:23:32,880
architecture, and the 
architecture is now dictating 

493
00:23:32,880 --> 00:23:35,400
the future of the software. 
And I want to leave the listener

494
00:23:35,400 --> 00:23:37,800
with a final thought about that 
future because everything we've 

495
00:23:37,800 --> 00:23:41,120
discussed today, the liquid 
cooling, the 700 Watt chips, the

496
00:23:41,120 --> 00:23:44,920
massive data centers, it all 
sounds incredibly energy 

497
00:23:44,920 --> 00:23:45,920
intensive. 
It is. 

498
00:23:45,920 --> 00:23:48,880
It's arguably the single biggest
bottleneck we face as an 

499
00:23:48,880 --> 00:23:51,080
industry. 
And it raises a really profound 

500
00:23:51,080 --> 00:23:53,880
question. 
Are we approaching a wall? 

501
00:23:53,960 --> 00:23:56,200
A physical wall. 
A hard physical wall. 

502
00:23:56,880 --> 00:24:00,720
If we can't make chips run 
faster because they'll melt, and

503
00:24:00,720 --> 00:24:02,880
we can't just make them 
infinitely bigger because they 

504
00:24:02,880 --> 00:24:05,640
use too much power, what happens
next? 

505
00:24:05,640 --> 00:24:08,120
What's the next paradigm? 
I think the future might not be 

506
00:24:08,120 --> 00:24:13,160
bigger AI, the future might have
to be denser AI, and for that 

507
00:24:13,160 --> 00:24:15,320
you have to look at biology. 
The human brain. 

508
00:24:15,360 --> 00:24:17,440
The human brain. 
It's the most advanced neural 

509
00:24:17,440 --> 00:24:21,920
network in the known universe. 
It has 86 billion neurons, 

510
00:24:21,920 --> 00:24:25,440
trillions of synapses and do you
have any idea how much power it 

511
00:24:25,440 --> 00:24:27,840
runs on? 
I'm going to assume it's not 700

512
00:24:27,840 --> 00:24:30,040
watts. 
It runs on about 20 watts. 20 

513
00:24:30,040 --> 00:24:31,480
watts. 
It's like a dim light bulb. 

514
00:24:31,480 --> 00:24:33,000
Exactly. 
A dim light bulb. 

515
00:24:33,440 --> 00:24:36,600
Meanwhile, to simulate anything 
even remotely close to that 

516
00:24:36,600 --> 00:24:40,240
level of connectivity, a cluster
of H1 hundreds needs its own 

517
00:24:40,240 --> 00:24:43,360
dedicated power plant. 
The gap between biology and 

518
00:24:43,360 --> 00:24:47,040
silicon, that massive efficiency
gap, That's the next frontier. 

519
00:24:47,560 --> 00:24:49,920
We have to figure out how to get
that level of intelligence 

520
00:24:50,240 --> 00:24:53,840
without boiling the oceans to do
it. 20 watts versus a MW. 

521
00:24:53,840 --> 00:24:55,680
That really puts the whole thing
in perspective. 

522
00:24:55,840 --> 00:24:58,000
We are still essentially using 
brute force. 

523
00:24:58,200 --> 00:25:00,600
We are in the steam engine era 
of AI. 

524
00:25:01,120 --> 00:25:05,800
It works, it's powerful, but 
it's loud, hot, and wildly 

525
00:25:05,800 --> 00:25:08,560
inefficient. 
The electric motor era hasn't 

526
00:25:08,560 --> 00:25:10,800
even started yet. 
So next time you look at your 

527
00:25:10,800 --> 00:25:14,120
laptop or your gaming console, 
don't just see a screen and a 

528
00:25:14,120 --> 00:25:16,600
keyboard. 
Look at it as a parallel math 

529
00:25:16,600 --> 00:25:18,960
engine. 
It's a baby supercomputer, and 

530
00:25:18,960 --> 00:25:22,520
we're only just now figuring out
how to really turn it on. 

531
00:25:22,760 --> 00:25:25,160
And if you're a developer out 
there, stop writing loops. 

532
00:25:25,440 --> 00:25:27,880
Start thinking in tensors, 
because that's where the entire 

533
00:25:27,880 --> 00:25:30,120
world is going. 
Thanks for diving deep into the 

534
00:25:30,120 --> 00:25:32,640
engine room with us today. 
It's absolutely fascinating to 

535
00:25:32,640 --> 00:25:34,360
see what's really under the hood
of all this. 

536
00:25:34,360 --> 00:25:36,280
Always a pleasure. 
We'll catch you on the next deep

537
00:25:36,280 --> 00:25:37,760
dive. 
Keep learning.

