1
00:00:00,040 --> 00:00:02,480
I want to start today with a 
mental image. 

2
00:00:03,080 --> 00:00:07,520
I want you to picture the 
absolute cutting edge of 

3
00:00:07,520 --> 00:00:11,520
artificial intelligence. 
Right now, most people, when I 

4
00:00:11,520 --> 00:00:14,720
ask them to do that, they 
picture a massive data center. 

5
00:00:14,720 --> 00:00:17,480
You know, rows and rows of 
servers blinking lights. 

6
00:00:17,480 --> 00:00:19,760
Cooling fans roaring like jet 
engines. 

7
00:00:19,760 --> 00:00:22,440
Exactly. 
Consuming enough electricity to 

8
00:00:22,440 --> 00:00:26,600
power a small city. 
They're picturing NVIDIA H1 

9
00:00:26,600 --> 00:00:28,880
hundreds stacked right up to the
ceiling. 

10
00:00:29,360 --> 00:00:31,400
And they wouldn't be wrong. 
I mean, that is certainly where 

11
00:00:31,400 --> 00:00:33,800
the large language models live. 
That's sort of the big AI 

12
00:00:33,800 --> 00:00:37,640
paradigm, right? 
But today we are going to go in 

13
00:00:37,640 --> 00:00:40,720
the, well, the complete opposite
direction. 

14
00:00:40,920 --> 00:00:43,320
I want to shrink our world down.
Way down. 

15
00:00:43,360 --> 00:00:45,440
We're leaving the cloud, we're 
leaving the server farm. 

16
00:00:45,440 --> 00:00:47,520
We were even leaving the 
smartphone. 

17
00:00:47,520 --> 00:00:48,920
It's probably in your pocket 
right now. 

18
00:00:48,920 --> 00:00:52,240
We're headed to the edge. 
Yeah, the the extreme edge. 

19
00:00:52,240 --> 00:00:54,000
We are talking about 
microcontrollers. 

20
00:00:54,000 --> 00:00:57,040
We're talking chips that cost 
less than $5, chips that can run

21
00:00:57,040 --> 00:01:00,920
on a coin cell battery. 
And the mission today for this 

22
00:01:00,920 --> 00:01:05,519
deep dive is to figure out how 
you can take that massive, you 

23
00:01:05,519 --> 00:01:09,360
know, brain melting AI 
capability and shove it into a 

24
00:01:09,360 --> 00:01:12,160
piece of silicon smaller than 
your fingernail. 

25
00:01:12,400 --> 00:01:15,360
It's a fascinating area. 
It's often called tiny ML and 

26
00:01:15,360 --> 00:01:17,760
honestly I think it's a harder 
problem than the big stuff. 

27
00:01:17,800 --> 00:01:19,680
Really harder. 
Oh for sure. 

28
00:01:19,800 --> 00:01:22,960
When you have basically infinite
computing power, you can get a 

29
00:01:22,960 --> 00:01:25,640
little easy. 
When you have $5 and a few 

30
00:01:25,640 --> 00:01:28,440
kilobytes of memory, you have to
be a genius. 

31
00:01:28,600 --> 00:01:31,880
And to really set the stage for 
why this is so hard, I want to 

32
00:01:31,880 --> 00:01:36,920
take us back in time a bit to 
the early 1980s, Carnegie Mellon

33
00:01:36,920 --> 00:01:38,960
University. 
Right, the birth place of so 

34
00:01:38,960 --> 00:01:40,920
much computer science history. 
It really is. 

35
00:01:41,040 --> 00:01:42,680
So picture a group of grad 
students. 

36
00:01:42,680 --> 00:01:45,400
It's the 80s, so you can imagine
the big hair, the chunky 

37
00:01:45,400 --> 00:01:47,080
glasses. 
They're working in their office 

38
00:01:47,080 --> 00:01:49,600
and they have a problem. 
A very, very serious problem. 

39
00:01:49,600 --> 00:01:51,680
Let me guess, Yeah, they needed 
coffee. 

40
00:01:51,800 --> 00:01:53,280
Close. 
They wanted a Coca-Cola. 

41
00:01:53,520 --> 00:01:55,840
The fuel of all night coding 
sessions. 

42
00:01:56,040 --> 00:01:58,480
I get it. 
But the vending machine is all 

43
00:01:58,480 --> 00:02:01,600
the way down the hall, maybe 
even in another building, and 

44
00:02:01,600 --> 00:02:04,840
there was just nothing worse 
than trekking all the way there,

45
00:02:05,120 --> 00:02:07,400
putting your quarters in and 
getting a warm. 

46
00:02:07,400 --> 00:02:10,360
Coat oh the worst. 
Or finding out it's empty. 

47
00:02:10,520 --> 00:02:12,720
Right. 
A classic engineer's dilemma. 

48
00:02:12,880 --> 00:02:15,400
Laziness really is the mother of
invention. 

49
00:02:15,400 --> 00:02:19,080
It truly is. 
So they hacked the machine, they

50
00:02:19,080 --> 00:02:22,400
installed these little micro 
switches inside the columns of 

51
00:02:22,400 --> 00:02:25,400
the Coke machine to track how 
many bottles were left. 

52
00:02:25,480 --> 00:02:27,200
OK. 
And this is the key part. 

53
00:02:27,560 --> 00:02:30,720
How long they had been there so 
they could estimate if they were

54
00:02:30,720 --> 00:02:32,000
cold? 
That's brilliant. 

55
00:02:32,120 --> 00:02:35,240
Then they wired it all up to the
departmental server. 

56
00:02:35,240 --> 00:02:37,360
Which meant they could just sit 
at their desk terminals and 

57
00:02:37,360 --> 00:02:40,160
check the stock. 
They invented the first Internet

58
00:02:40,160 --> 00:02:42,360
of Things device because they 
didn't want to walk down the 

59
00:02:42,360 --> 00:02:46,080
hall for a warm soda. 
It's a great story, but it also 

60
00:02:46,080 --> 00:02:48,520
highlights a very specific kind 
of computing right? 

61
00:02:48,920 --> 00:02:52,360
That Coke machine was doing 
something, well, very simple. 

62
00:02:52,360 --> 00:02:54,080
It was just logic. 
If then. 

63
00:02:54,240 --> 00:02:56,680
Exactly. 
If the bottle count is greater 

64
00:02:56,680 --> 00:02:59,800
than 0, return true. 
If the time in the machine is 

65
00:02:59,800 --> 00:03:01,760
greater than 3 hours, return 
cold. 

66
00:03:02,240 --> 00:03:04,880
It's a lookup table. 
It's basic and Fast forward to 

67
00:03:04,880 --> 00:03:10,120
today, we have something like 
250 billion microcontrollers 

68
00:03:10,120 --> 00:03:12,600
MCU's out there in the world. 
Billion with AB. 

69
00:03:12,960 --> 00:03:17,120
They're in your car's braking 
system, your microwave, your TV 

70
00:03:17,120 --> 00:03:19,920
remote, soil sensors in farm 
fields. 

71
00:03:20,080 --> 00:03:21,840
Everywhere. 
And for the most part, they're 

72
00:03:21,840 --> 00:03:23,760
still kind of stuck in that Coke
machine era. 

73
00:03:23,760 --> 00:03:25,880
They're reactive. 
They just follow simple 

74
00:03:25,880 --> 00:03:28,000
preprogrammed rules. 
But we're greedy now. 

75
00:03:28,000 --> 00:03:29,920
We don't want them to just 
follow rules anymore. 

76
00:03:29,920 --> 00:03:31,920
No, we want them to see. 
We want them to hear. 

77
00:03:31,920 --> 00:03:34,840
We want to understand context. 
We want deep learning on them. 

78
00:03:34,880 --> 00:03:37,720
And that right there is where we
hit the wall. 

79
00:03:38,560 --> 00:03:41,320
The memory wall. 
Let's unpack that term because I

80
00:03:41,320 --> 00:03:43,880
think people hear chip and they 
just assume it's like a mini 

81
00:03:43,880 --> 00:03:46,000
computer. 
I mean, my phone is essentially 

82
00:03:46,000 --> 00:03:49,040
one big chip and it has 
gigabytes of RAM. 

83
00:03:49,280 --> 00:03:51,560
And that is such a dangerous 
assumption. 

84
00:03:52,160 --> 00:03:55,000
We need to be very, very precise
about the hardware we're talking

85
00:03:55,000 --> 00:03:57,760
about today. 
The source material for this 

86
00:03:57,760 --> 00:04:02,280
dive, it focuses on a specific, 
really popular commercial 

87
00:04:02,280 --> 00:04:06,040
microcontroller, the STM 32F746.
What's? 

88
00:04:06,280 --> 00:04:08,080
Your name. 
Oh, engineers love their 

89
00:04:08,080 --> 00:04:10,800
alphanumeric codes, but here are
the specs that matter. 

90
00:04:11,000 --> 00:04:14,200
A typical smartphone, like you 
said, has maybe 8 gigabytes of 

91
00:04:14,200 --> 00:04:17,120
RAM, right? 
This microcontroller has 320 

92
00:04:17,120 --> 00:04:19,040
kilobytes of RAM. 
Wait, wait. 

93
00:04:19,040 --> 00:04:20,240
Kilobytes. 
Kilobytes. 

94
00:04:20,519 --> 00:04:23,680
That is what, roughly 25,000 
times less working memory than 

95
00:04:23,680 --> 00:04:26,080
your phone. 
I have single e-mail attachments

96
00:04:26,080 --> 00:04:27,520
that are larger. 
Than that definitely do. 

97
00:04:27,680 --> 00:04:29,880
And then there's the storage, 
the flash memory. 

98
00:04:30,120 --> 00:04:32,000
This is where you keep the 
program itself. 

99
00:04:32,400 --> 00:04:36,000
Your phone might have 128 gigs, 
256. 

100
00:04:36,480 --> 00:04:38,680
This chip has one MB. 
One MB. 

101
00:04:38,680 --> 00:04:40,920
So we are trying to fit a neural
network which can take up 

102
00:04:40,920 --> 00:04:44,560
hundreds of megabytes just to 
store its weights into one MB of

103
00:04:44,560 --> 00:04:46,520
program space. 
And then run it in a working 

104
00:04:46,520 --> 00:04:49,680
memory space smaller than a 
floppy disk from 1995. 

105
00:04:49,680 --> 00:04:52,160
That sounds. 
It is if you try to take a 

106
00:04:52,160 --> 00:04:56,080
standard small deep learning 
model, something like Mobile Net

107
00:04:56,080 --> 00:04:58,840
V2, which is widely used on 
phones, and you try to load it 

108
00:04:58,840 --> 00:05:01,600
onto this chip. 
Yeah, it's not that it runs 

109
00:05:01,600 --> 00:05:03,800
slowly. 
It physically cannot exist 

110
00:05:03,800 --> 00:05:04,640
there. 
It just doesn't. 

111
00:05:04,680 --> 00:05:07,480
Fit the Arameters alone need 68 
megabytes. 

112
00:05:07,720 --> 00:05:10,760
You have one MB of total space. 
It's like trying to park a 

113
00:05:10,760 --> 00:05:12,920
Greyhound bus inside a coat 
closet. 

114
00:05:13,040 --> 00:05:14,680
OK, so it was considered 
impossible. 

115
00:05:14,760 --> 00:05:17,880
It pretty much was, until the 
research we're looking at today.

116
00:05:17,880 --> 00:05:21,760
We're doing a deep dive into a 
system called MCU Net, which was

117
00:05:21,760 --> 00:05:25,880
developed by researchers at MIT,
specifically A-Team led by Ji 

118
00:05:25,880 --> 00:05:28,320
Lin and Song Han. 
And looking at the results they 

119
00:05:28,320 --> 00:05:31,320
sent over, they didn't just get 
a model to run, they achieved 

120
00:05:31,320 --> 00:05:35,000
over 70% accuracy on Imagenet. 
Which is the absolute gold 

121
00:05:35,000 --> 00:05:38,440
standard benchmark we're talking
about recognizing 1000 different

122
00:05:38,440 --> 00:05:41,520
categories of objects. 
Golden retrievers, Fire trucks, 

123
00:05:41,560 --> 00:05:43,760
Espresso machines. 
All on a $5 chip. 

124
00:05:43,800 --> 00:05:46,320
With 70% accuracy, that is a 
massive deal. 

125
00:05:46,480 --> 00:05:49,120
The previous best attempts were 
stuck down around, what, 54 

126
00:05:49,120 --> 00:05:51,840
percent or they discouraged? 
So how did they do it? 

127
00:05:51,840 --> 00:05:55,360
I mean, did they just find a 
better way to compress the file?

128
00:05:55,360 --> 00:05:57,200
A better zip algorithm? 
No. 

129
00:05:57,760 --> 00:06:00,160
And that's the old way of 
thinking about this problem. 

130
00:06:00,520 --> 00:06:04,120
To really get why MCU net is so 
revolutionary, we have to look 

131
00:06:04,120 --> 00:06:06,040
at how this kind of AI is 
usually built. 

132
00:06:06,480 --> 00:06:10,800
It's typically a two stage 
process with, well, with two 

133
00:06:10,800 --> 00:06:12,640
different tribes of people. 
The tribes. 

134
00:06:12,720 --> 00:06:16,640
OK, I like this analogy. 
So you have tribe one, the model

135
00:06:16,640 --> 00:06:18,040
designers. 
These are the architects. 

136
00:06:18,400 --> 00:06:21,240
They work in Python using 
libraries like Pytorch or 

137
00:06:21,240 --> 00:06:23,760
Tensorflow. 
Their whole world is designing 

138
00:06:23,760 --> 00:06:26,120
the neural network structure. 
They care about accuracy. 

139
00:06:26,200 --> 00:06:29,760
And they assume hardware is just
this infinite resource pool 

140
00:06:29,760 --> 00:06:31,440
somewhere in the cloud. 
Exactly. 

141
00:06:31,760 --> 00:06:34,680
They're drawing the blueprints 
for a beautiful massive 

142
00:06:34,680 --> 00:06:36,800
cathedral. 
Then they hand those blueprints 

143
00:06:36,800 --> 00:06:39,280
over to Tribe 2, the system 
designers. 

144
00:06:39,480 --> 00:06:41,760
These are the folks right in the
compilers, the drivers, the low 

145
00:06:41,760 --> 00:06:44,040
level code. 
Their job is to take that 

146
00:06:44,040 --> 00:06:46,400
blueprint and actually make it 
run on the real silicon. 

147
00:06:46,440 --> 00:06:48,320
They're the construction crew. 
Precisely. 

148
00:06:48,960 --> 00:06:51,880
But in the world of 
microcontrollers, this hand off 

149
00:06:51,920 --> 00:06:56,360
is a complete disaster. 
The architect hands over a 

150
00:06:56,360 --> 00:06:59,240
blueprint for a cathedral. 
The builder looks at the 

151
00:06:59,240 --> 00:07:03,200
physical site and says, but we 
only have a 10 foot by 10 foot 

152
00:07:03,200 --> 00:07:07,160
lot and our total budget is $5. 
So the builder just yells back 

153
00:07:07,160 --> 00:07:08,800
at the architect. 
Make it smaller. 

154
00:07:08,840 --> 00:07:12,240
Right, and the architect who 
doesn't understand construction 

155
00:07:12,240 --> 00:07:14,160
just starts chopping things off 
randomly. 

156
00:07:14,480 --> 00:07:17,560
OK, let's remove the Spire. 
Get rid of the lobby. 

157
00:07:18,120 --> 00:07:20,400
They make the model Dumber and 
Dumber and Dumber until it 

158
00:07:20,400 --> 00:07:22,400
finally fits. 
But there's a catch, I'm 

159
00:07:22,400 --> 00:07:24,000
guessing. 
Here's the kicker. 

160
00:07:24,400 --> 00:07:27,200
The builder is also using really
inefficient tools. 

161
00:07:27,320 --> 00:07:30,120
Standard libraries like 
Tensorflow, light, micro, they 

162
00:07:30,120 --> 00:07:32,720
have a lot of overhead. 
They waste space by their very 

163
00:07:32,720 --> 00:07:34,200
nature. 
So the builder is wasting space 

164
00:07:34,200 --> 00:07:36,560
on the tiny lot, which forces 
the architect to make the 

165
00:07:36,560 --> 00:07:39,080
building even smaller and Dumber
than it actually needs to be. 

166
00:07:39,080 --> 00:07:40,720
It's a vicious cycle of 
inefficiency. 

167
00:07:40,720 --> 00:07:43,160
A vicious cycle. 
OK, and this is where MCNET 

168
00:07:43,160 --> 00:07:46,720
proposes a new way system 
algorithm Co design. 

169
00:07:46,760 --> 00:07:50,520
Which means what exactly? 
It means the architect and the 

170
00:07:50,520 --> 00:07:54,320
builder are the same person, or 
at the very least they are in 

171
00:07:54,320 --> 00:07:58,120
the same room looking at the 
same constraints and designing 

172
00:07:58,120 --> 00:08:00,400
the engine and the car 
simultaneously. 

173
00:08:00,440 --> 00:08:03,320
OK, that makes intuitive sense. 
So let's break this down. 

174
00:08:03,480 --> 00:08:06,640
The paper splits MCU net into 
two main components. 

175
00:08:07,000 --> 00:08:09,440
There's tiny Nas on the 
algorithm side. 

176
00:08:09,440 --> 00:08:11,040
And tiny engine on the system 
side. 

177
00:08:11,240 --> 00:08:14,040
Let's start with tiny Nas. 
Nas stands for Neural 

178
00:08:14,040 --> 00:08:15,640
Architecture Search. 
Correct. 

179
00:08:16,120 --> 00:08:18,040
I've heard of this. 
This is basically the concept of

180
00:08:18,200 --> 00:08:20,280
AI designing AI, right? 
That's a great way to put it. 

181
00:08:20,280 --> 00:08:23,720
Instead of a human engineer 
manually picking all the layers,

182
00:08:23,920 --> 00:08:26,000
I'll put a convolution here, a 
pooling layer there. 

183
00:08:26,280 --> 00:08:28,760
You write an algorithm that 
automatically searches through 

184
00:08:28,760 --> 00:08:32,760
millions or even billions of 
possible architectures to find 

185
00:08:32,760 --> 00:08:34,880
the best one for your task. 
But wait a second. 

186
00:08:34,919 --> 00:08:39,600
NASA is notoriously 
computationally expensive. 

187
00:08:39,640 --> 00:08:42,480
I've read about Google running 
these searches on thousands of 

188
00:08:42,480 --> 00:08:45,440
TP us for weeks just to find one
good model. 

189
00:08:45,720 --> 00:08:47,880
A great point. 
If our whole goal here is to 

190
00:08:47,880 --> 00:08:51,760
save resources for a tiny chip, 
isn't running a massive power 

191
00:08:51,760 --> 00:08:53,920
hungry search totally counter 
intuitive? 

192
00:08:54,160 --> 00:08:56,280
It would be if he used a 
standard Nas. 

193
00:08:56,600 --> 00:08:59,760
Standard Nas is expensive, and 
it's usually just optimizing for

194
00:08:59,760 --> 00:09:03,680
one thing, accuracy. 
But tiny Nas has to solve a 

195
00:09:03,680 --> 00:09:07,320
much, much harder problem, which
is it has to find a model that 

196
00:09:07,320 --> 00:09:11,760
is not only accurate, but also 
strictly obeys the hard physical

197
00:09:11,760 --> 00:09:13,520
constraints of the 
microcontroller. 

198
00:09:13,880 --> 00:09:17,920
It can't use one single byte 
more than 320 kilobytes of RAM, 

199
00:09:18,320 --> 00:09:20,840
not one. 
So it's constrained optimization

200
00:09:20,840 --> 00:09:23,040
on steroids. 
And there's another wrinkle. 

201
00:09:23,120 --> 00:09:25,280
Microcontrollers are incredibly 
diverse. 

202
00:09:25,720 --> 00:09:27,800
There are thousands of different
chips out there with slightly 

203
00:09:27,800 --> 00:09:29,520
different memory sizes, 
different features. 

204
00:09:29,520 --> 00:09:31,720
So you can't just find one 
perfect model. 

205
00:09:31,720 --> 00:09:34,120
Exactly. 
If you have to rerun a massive 

206
00:09:34,120 --> 00:09:37,120
expensive search for every 
single specific chip you might 

207
00:09:37,120 --> 00:09:39,360
want to use, well, it's just not
scalable. 

208
00:09:39,360 --> 00:09:41,560
It's not practical. 
So how does Tiny Nest solve 

209
00:09:41,560 --> 00:09:42,280
that? 
Problem. 

210
00:09:42,400 --> 00:09:45,080
It starts by being very, very 
clever about where it looks for 

211
00:09:45,080 --> 00:09:47,000
a solution. 
This is the first big 

212
00:09:47,000 --> 00:09:50,120
innovation, the concept of 
automated search space 

213
00:09:50,120 --> 00:09:52,520
optimization. 
OK, the outline mentions 

214
00:09:52,520 --> 00:09:54,400
something here that sounds 
pretty technical. 

215
00:09:54,400 --> 00:09:57,920
The CDF of FLOPS. 
I'll be honest, that sounds like

216
00:09:57,920 --> 00:10:00,120
a tax form. 
Can you help us visualize what 

217
00:10:00,120 --> 00:10:02,280
this actually means? 
I can try. 

218
00:10:02,600 --> 00:10:05,200
Let's stick with an analogy. 
Imagine you're trying to pack a 

219
00:10:05,200 --> 00:10:07,120
suitcase for a trip. 
OK, I'm with you. 

220
00:10:07,280 --> 00:10:11,760
You have a very small suitcase 
that's your 320 KB memory limit 

221
00:10:12,280 --> 00:10:14,800
and you want to pack as much 
utility into that suitcase as 

222
00:10:14,800 --> 00:10:18,520
you possibly can't close 
gadgets, shoes right now. 

223
00:10:18,520 --> 00:10:21,640
FLOPS stands for floating point 
operations. 

224
00:10:22,160 --> 00:10:25,400
In our little analogy, FLOPS are
like the utility or the 

225
00:10:25,400 --> 00:10:27,120
smartness of the items you're 
packing. 

226
00:10:27,520 --> 00:10:31,440
Generally an AI, the more math, 
the more FLOPSA model performs, 

227
00:10:31,800 --> 00:10:32,800
the smarter it is. 
So. 

228
00:10:32,800 --> 00:10:36,360
I want to pack as many FLOPS, as
much smartness into my little 

229
00:10:36,360 --> 00:10:38,240
suitcase as I can. 
Exactly. 

230
00:10:38,240 --> 00:10:40,360
Now imagine you have different 
wardrobes you can choose your 

231
00:10:40,360 --> 00:10:41,880
items from before you start 
packing. 

232
00:10:42,320 --> 00:10:45,000
Some wardrobes are full of 
bulky, inefficient clothes. 

233
00:10:45,280 --> 00:10:47,760
You know, a giant wool coat that
takes up half the suitcase but 

234
00:10:47,760 --> 00:10:49,920
isn't actually that warm. 
That's a bad search space. 

235
00:10:50,320 --> 00:10:53,640
That's a bad search space. 
Other wardrobes are full of high

236
00:10:53,640 --> 00:10:56,960
tech thin thermal gear. 
You can fit a lot more warmth, a

237
00:10:56,960 --> 00:10:59,840
lot more FLOPS into the very 
same suitcase. 

238
00:10:59,920 --> 00:11:04,120
So before Tiny Nest even starts 
picking the specific items, the 

239
00:11:04,120 --> 00:11:07,680
specific model it first analyzes
the wardrobes. 

240
00:11:07,760 --> 00:11:10,720
That is exactly what it does. 
That's the CDF part, the 

241
00:11:10,720 --> 00:11:12,160
cumulative distribution 
function. 

242
00:11:12,440 --> 00:11:15,480
It's a statistical method. 
It basically samples the 

243
00:11:15,480 --> 00:11:17,520
wardrobes. 
It looks at the probability 

244
00:11:17,520 --> 00:11:21,160
distribution of what's inside. 
It asks if I pick random items 

245
00:11:21,160 --> 00:11:24,720
from wardrobe A, how many FLOPS 
can I usually fit in my 

246
00:11:24,720 --> 00:11:26,560
suitcase? 
Now what about wardrobe B? 

247
00:11:26,720 --> 00:11:29,440
And it picks the wardrobe that 
offers the highest density of 

248
00:11:29,440 --> 00:11:32,000
math per byte, the most 
efficient set of options. 

249
00:11:32,000 --> 00:11:34,320
Spot on, and the results were 
staggering. 

250
00:11:34,600 --> 00:11:37,280
They found that by simply 
picking the right search space, 

251
00:11:37,560 --> 00:11:40,080
the right wardrobe, before they 
even started the fine grained 

252
00:11:40,080 --> 00:11:43,600
search, they could boost the 
final models accuracy by over 

253
00:11:43,600 --> 00:11:46,440
4%. 4% is huge in this field. 
It's. 

254
00:11:46,440 --> 00:11:48,040
Massive. 
And that's before they've even 

255
00:11:48,040 --> 00:11:50,040
chosen a single layer. 
That's fascinating. 

256
00:11:50,040 --> 00:11:52,680
So it's not just about the final
model, it's about the potential 

257
00:11:52,680 --> 00:11:54,000
of the model family you start 
with. 

258
00:11:54,320 --> 00:11:56,120
Right. 
So once it's picked the best 

259
00:11:56,120 --> 00:11:59,960
search space then he uses 
something called one shot Nas to

260
00:11:59,960 --> 00:12:02,320
pick the specific final 
architecture. 

261
00:12:02,320 --> 00:12:04,040
Why one shot? 
What does that mean? 

262
00:12:04,240 --> 00:12:07,080
Because training a neural 
network from scratch is the slow

263
00:12:07,080 --> 00:12:08,960
part. 
Training millions of them to 

264
00:12:08,960 --> 00:12:10,560
test each one would take 100 
years. 

265
00:12:10,600 --> 00:12:12,720
So instead they train a super 
network. 

266
00:12:12,720 --> 00:12:14,680
A super network. 
They give it like a giant 

267
00:12:14,680 --> 00:12:17,840
network that contains every 
possible variation of the 

268
00:12:17,840 --> 00:12:20,240
smaller networks inside it all 
at once. 

269
00:12:20,280 --> 00:12:23,680
So like a Swiss army knife that 
is every possible tool already 

270
00:12:23,680 --> 00:12:25,720
attached to it. 
That's a perfect analogy. 

271
00:12:25,840 --> 00:12:28,520
They train that big Swiss army 
knife knife just one time. 

272
00:12:28,920 --> 00:12:33,400
Then they use a clever fast 
evolutionary algorithm to select

273
00:12:33,400 --> 00:12:37,000
the specific subset of tools, 
the specific subnetwork that 

274
00:12:37,000 --> 00:12:40,520
works best for a specific chip. 
So it can tailor the model. 

275
00:12:40,560 --> 00:12:43,080
It custom tailors the suit to 
the exact measurements of the 

276
00:12:43,080 --> 00:12:46,600
hardware. 
It asks OK for this STM 32 chip 

277
00:12:46,600 --> 00:12:50,160
with 320 kilobytes of RAM, is it
better to have a high resolution

278
00:12:50,160 --> 00:12:53,960
image with fewer channels or a 
lower resolution image with more

279
00:12:53,960 --> 00:12:56,320
channels and more layers? 
And for a different chip with 

280
00:12:56,320 --> 00:12:58,800
say 512 kilobe, the answer might
be different. 

281
00:12:59,000 --> 00:13:00,920
The answer will almost certainly
be different. 

282
00:13:01,560 --> 00:13:05,240
It automatically finds that 
perfect trade off to squeeze 

283
00:13:05,280 --> 00:13:08,360
every last drop of performance 
out of that specific memory 

284
00:13:08,360 --> 00:13:11,280
budget. 
OK, so we have our custom design

285
00:13:11,280 --> 00:13:13,160
model. 
The architect has done their job

286
00:13:13,160 --> 00:13:16,600
perfectly, but now we have to 
actually build it and run it. 

287
00:13:17,560 --> 00:13:19,000
And that brings us to the other 
half. 

288
00:13:19,560 --> 00:13:22,800
Tiny engine. 
This is the system side and for 

289
00:13:22,800 --> 00:13:25,320
the engineers listening, this is
where you should really lean in 

290
00:13:25,320 --> 00:13:29,000
because this part directly 
challenges the industry standard

291
00:13:29,160 --> 00:13:31,920
way of doing things. 
And the standard way would be 

292
00:13:31,920 --> 00:13:36,280
using libraries like Tensorflow,
Light micro or CMSIS and and 

293
00:13:36,280 --> 00:13:38,440
yes. 
Most of these existing 

294
00:13:38,440 --> 00:13:41,080
frameworks use what's called an 
interpreter based approach. 

295
00:13:41,080 --> 00:13:43,640
Can you explain that? 
Why is an interpreter a bad 

296
00:13:43,640 --> 00:13:46,160
thing for a microcontroller? 
Let's go back to the kitchen. 

297
00:13:46,160 --> 00:13:49,440
An interpreter is like a cook 
who has to read the recipe line 

298
00:13:49,440 --> 00:13:51,720
by line while they are actively 
cooking. 

299
00:13:51,800 --> 00:13:55,880
So they read the first line, 
chop one onion, they stop, they 

300
00:13:55,880 --> 00:13:58,440
put the book down, they find the
knife, they find the cutting 

301
00:13:58,440 --> 00:14:01,280
board, they chop the onion, and 
they pick the book back up, find

302
00:14:01,280 --> 00:14:03,040
their place, and read the next 
line. 

303
00:14:03,440 --> 00:14:07,080
Saute the onion in a pan. 
They put the book down again. 

304
00:14:07,360 --> 00:14:09,640
There's a lot of switching back 
and forth, a lot of overhead. 

305
00:14:09,640 --> 00:14:12,480
Tons of overhead, yeah. 
But even more importantly, you 

306
00:14:12,480 --> 00:14:15,320
need space on your counter for 
the cookbook itself. 

307
00:14:15,800 --> 00:14:19,480
And that cookbook here is the 
model graph, the metadata that 

308
00:14:19,480 --> 00:14:22,440
tells the system what layer 
comes next, what its parameters 

309
00:14:22,440 --> 00:14:24,360
are. 
I see, on a server with 32 

310
00:14:24,360 --> 00:14:27,320
gigabytes of RAM, you just don't
care if the cookbook takes up 

311
00:14:27,320 --> 00:14:30,440
100 kilobytes. 
But on our little chip that 

312
00:14:30,440 --> 00:14:35,000
metadata, that cookbook can 
consume 20% or even 30% of your 

313
00:14:35,000 --> 00:14:37,240
total precious memory. 
Just to know what to do next, 

314
00:14:37,240 --> 00:14:39,960
that feels incredibly wasteful. 
It is so tiny. 

315
00:14:39,960 --> 00:14:41,520
Engine takes a completely 
different approach. 

316
00:14:41,800 --> 00:14:43,960
It is a compiler. 
So instead of the cook with the 

317
00:14:43,960 --> 00:14:45,920
open book. 
It's like a robot arm in a 

318
00:14:45,920 --> 00:14:49,680
factory that has been hardwired 
to perform only that one 

319
00:14:49,680 --> 00:14:52,600
specific recipe. 
There's no book Tiny Engine 

320
00:14:52,600 --> 00:14:56,320
looks at the model that tiny ass
designed and it generates 

321
00:14:56,320 --> 00:15:01,040
specific custom binary code for 
just that model says I see you 

322
00:15:01,040 --> 00:15:03,920
need a three by three depth wise
convolution here. 

323
00:15:04,320 --> 00:15:07,200
I will write the exact machine 
code for that. 

324
00:15:07,560 --> 00:15:10,080
And I see you don't use any 
pooling layers anywhere in this 

325
00:15:10,080 --> 00:15:12,160
model. 
OK, I will just delete all the 

326
00:15:12,160 --> 00:15:13,640
pooling code. 
So there's no cookbook at 

327
00:15:13,640 --> 00:15:15,320
runtime, no interpreter. 
None. 

328
00:15:15,720 --> 00:15:19,040
The logic is baked directly into
the executable file, and this 

329
00:15:19,040 --> 00:15:23,760
simple change slashes the final 
binary size by a factor of 5. 

330
00:15:23,760 --> 00:15:26,680
Compared to Tensorflow light, it
frees up all that flash sorts. 

331
00:15:26,680 --> 00:15:28,800
But what about the SRAM, the 
counter space? 

332
00:15:28,800 --> 00:15:30,840
You said earlier that was 
usually the hardest constraint. 

333
00:15:30,880 --> 00:15:33,640
And this is where a tiny engine 
gets really, really clever with 

334
00:15:33,640 --> 00:15:36,160
something they call model 
Adaptive Memory scheduling. 

335
00:15:36,240 --> 00:15:38,840
I saw a note about this, they 
compared it to Tetris. 

336
00:15:39,000 --> 00:15:42,080
It is exactly like Tetris in a 
standard library. 

337
00:15:42,200 --> 00:15:44,640
Memory allocation is often lazy 
and inefficient. 

338
00:15:45,200 --> 00:15:47,800
The system looks at layer one 
and says, OK, I need 50 

339
00:15:47,800 --> 00:15:50,480
kilobytes for this, it grabs a 
chunk of memory. 

340
00:15:50,920 --> 00:15:53,640
Then layer 2 needs space, it 
grabs another chunk. 

341
00:15:53,880 --> 00:15:57,040
Often they just allocate a huge 
buffer for the worst case layer 

342
00:15:57,040 --> 00:15:59,720
and let it sit there unused for 
most of the time. 

343
00:15:59,800 --> 00:16:03,320
So you end up with all these 
gaps, wasted space like a bad 

344
00:16:03,320 --> 00:16:04,800
game of Tetris. 
Exactly. 

345
00:16:04,960 --> 00:16:07,960
But tiny engine, because it's a 
compiler, knows the future. 

346
00:16:08,120 --> 00:16:10,640
You can see the entire life 
cycle of the network before it 

347
00:16:10,640 --> 00:16:12,640
ever runs. 
It knows that tensor A is 

348
00:16:12,640 --> 00:16:16,600
created at Step 5 and is never 
ever used again after step 6. 

349
00:16:16,720 --> 00:16:19,800
So it's step 6.001. 
That memory is free to be used 

350
00:16:19,800 --> 00:16:23,280
for something else immediately. 
And tiny engine scheduler slides

351
00:16:23,280 --> 00:16:26,240
tensor V right into that same 
physical memory address. 

352
00:16:26,600 --> 00:16:28,760
It creates a perfect memory 
schedule where the buffers 

353
00:16:28,760 --> 00:16:31,320
overlap in time with no wasted 
gaps. 

354
00:16:31,360 --> 00:16:33,240
It's like extreme hot dusting 
for data. 

355
00:16:33,240 --> 00:16:35,680
As soon as one number leaves the
seat, another one sits right 

356
00:16:35,680 --> 00:16:37,040
down. 
That's a great way to put it, 

357
00:16:37,360 --> 00:16:40,480
and because it knows the whole 
topology ahead of time, it can 

358
00:16:40,480 --> 00:16:42,440
make smarter decisions about 
things like tiling. 

359
00:16:42,760 --> 00:16:45,880
If a layer is too big to fit in 
memory all at once, you have to 

360
00:16:45,880 --> 00:16:49,040
chop it into smaller tiles. 
Tiny Engine calculates the 

361
00:16:49,040 --> 00:16:51,840
perfect tile size based on what 
else is on the desk at that 

362
00:16:51,840 --> 00:16:54,840
specific moment in time. 
That seems incredibly efficient,

363
00:16:55,040 --> 00:16:57,320
but there's one more trick in 
tiny engine that the paper 

364
00:16:57,320 --> 00:16:58,840
highlighted and I really want to
understand it. 

365
00:16:59,280 --> 00:17:02,760
The in place depth wise 
convolution. 

366
00:17:03,240 --> 00:17:07,839
Yes, this is a very low level 
optimization, but its impact is 

367
00:17:07,839 --> 00:17:09,160
huge. 
It's brilliant. 

368
00:17:09,359 --> 00:17:12,280
OK, walk us through it. 
So in a standard convolution, 

369
00:17:12,440 --> 00:17:15,079
you're taking an input image or 
a feature map from a previous 

370
00:17:15,079 --> 00:17:18,200
layer, and you're applying a 
filter to it to create an output

371
00:17:18,200 --> 00:17:20,720
map, right? 
To do this you almost always 

372
00:17:20,720 --> 00:17:24,720
need 2 separate memory buffers, 
one to hold the input so you can

373
00:17:24,720 --> 00:17:27,520
read from it, and a second 
separate one to write the output

374
00:17:27,520 --> 00:17:30,240
into. 
So if the image data is size N, 

375
00:17:30,320 --> 00:17:32,920
you need at least 2 N of memory 
at that moment. 

376
00:17:33,120 --> 00:17:38,080
Yes, your peak memory usage is 2
N You can't just overwrite the 

377
00:17:38,080 --> 00:17:40,400
input while you're working, 
because you might need those 

378
00:17:40,400 --> 00:17:42,640
same input pixels for the next 
calculation. 

379
00:17:43,160 --> 00:17:45,680
However. 
Many modern efficient networks 

380
00:17:45,840 --> 00:17:48,680
use a special type of layer 
called depth wise convolution. 

381
00:17:49,440 --> 00:17:53,080
The key detail here is that the 
filter only looks at 1 channel 

382
00:17:53,080 --> 00:17:57,240
at a time, so the output for 
channel 1 depends only on the 

383
00:17:57,240 --> 00:17:59,880
input from channel 1, doesn't 
need to look at Channel 2 or 

384
00:17:59,880 --> 00:18:01,720
channel 3. 
OK, so the channels are 

385
00:18:01,720 --> 00:18:03,960
independent. 
How does that help with memory? 

386
00:18:04,000 --> 00:18:07,560
Because the researchers have 
this key insight, as soon as you

387
00:18:07,560 --> 00:18:12,720
have calculated the output for 
pixel 000 in channel 1, you will

388
00:18:12,720 --> 00:18:16,480
never ever need the input for 
pixel 00 in channel one ever 

389
00:18:16,480 --> 00:18:18,720
again. 
So you can just overrate it. 

390
00:18:18,800 --> 00:18:21,040
You can write the answer 
directly on top of the question.

391
00:18:21,040 --> 00:18:23,720
That's like destructive editing.
Exactly. 

392
00:18:23,720 --> 00:18:25,160
It's like painting over a 
canvas. 

393
00:18:25,160 --> 00:18:27,440
You don't need a second canvas 
to copy your work to. 

394
00:18:27,440 --> 00:18:29,800
You just transform the image 
right there in place. 

395
00:18:29,800 --> 00:18:31,440
And what does that do to the 
memory requirement? 

396
00:18:31,440 --> 00:18:34,200
It drops from 2 N. 
To north plus a tiny tiny 

397
00:18:34,200 --> 00:18:36,480
buffer. 
So basically it drops to north, 

398
00:18:37,000 --> 00:18:40,600
it cuts the peak memory usage 
for those specific layers almost

399
00:18:40,600 --> 00:18:43,040
perfectly in half and. 
When you are fighting for every 

400
00:18:43,040 --> 00:18:46,040
single KB, cutting your peak 
usage in half is. 

401
00:18:46,520 --> 00:18:48,360
That's miraculous. 
It really is. 

402
00:18:48,360 --> 00:18:50,640
And this brings us right back to
the Co design principle. 

403
00:18:51,080 --> 00:18:55,480
The Architect Tiny Nass knows 
that the Builder Tiny Engine has

404
00:18:55,480 --> 00:18:57,920
this special in place trick up 
its sleeve. 

405
00:18:58,000 --> 00:19:00,480
So it changes what the architect
designs. 

406
00:19:00,480 --> 00:19:03,960
Of course Tiny Nass is now more 
likely to design networks that 

407
00:19:03,960 --> 00:19:07,920
use a lot of these depth wise 
convolutions because it they are

408
00:19:07,920 --> 00:19:10,040
super cheap on memory when tiny 
engine runs. 

409
00:19:10,080 --> 00:19:12,680
They feed off each other. 
The capabilities of the system 

410
00:19:12,680 --> 00:19:16,440
software directly influence the 
design of the AI model itself. 

411
00:19:16,440 --> 00:19:18,640
Precisely. 
So we've built the car, we've 

412
00:19:18,640 --> 00:19:21,600
custom tuned the engine, we've 
stripped out all the unnecessary

413
00:19:21,600 --> 00:19:23,040
way. 
Let's take it to the track. 

414
00:19:23,040 --> 00:19:25,040
What are the actual performance 
results? 

415
00:19:25,440 --> 00:19:28,920
The headline number as we 
mentioned at the top is 70.7% 

416
00:19:29,160 --> 00:19:32,480
top one accuracy on Imagedet. 
Can you contextualize that for 

417
00:19:32,480 --> 00:19:34,720
me? 
Again, is 70% considered good in

418
00:19:34,720 --> 00:19:38,000
the grand scheme of things? 
70% is widely considered the 

419
00:19:38,000 --> 00:19:42,200
threshold of real world utility.
You know 50% is basically a coin

420
00:19:42,200 --> 00:19:44,520
toss. 
It's useless. 60% is OK for a 

421
00:19:44,520 --> 00:19:48,120
cool demo, but 70% is where you 
start to see things that could 

422
00:19:48,120 --> 00:19:51,080
be commercial products. 
It means the system is reliable 

423
00:19:51,080 --> 00:19:53,840
enough to actually be useful. 
And how did this compare to the 

424
00:19:53,840 --> 00:19:55,960
previous state-of-the-art for 
this kind of chip? 

425
00:19:56,360 --> 00:19:59,440
O the previous best solution 
they benchmarked against which 

426
00:19:59,440 --> 00:20:01,960
was using a quantized mobile net
V2 model. 

427
00:20:02,480 --> 00:20:07,920
This standard CMSISNN library it
only got 54% accuracy. 

428
00:20:08,040 --> 00:20:10,720
That is a massive gap, 16 
points. 

429
00:20:10,720 --> 00:20:13,200
That's the difference between a 
prototype and a real product. 

430
00:20:13,400 --> 00:20:14,880
And that's not even the whole 
story. 

431
00:20:15,320 --> 00:20:17,680
In many of their tests, the 
mobile Net V2 model simply 

432
00:20:17,680 --> 00:20:19,480
wouldn't run at all. 
It would just crash with an out 

433
00:20:19,480 --> 00:20:21,520
of memory error. 
So MC Net doesn't just make it 

434
00:20:21,520 --> 00:20:23,840
better, it actually makes it 
possible in the first place. 

435
00:20:23,840 --> 00:20:26,680
OK, what about speed? 
Usually when you compress things

436
00:20:26,680 --> 00:20:29,560
this much or get this clever, 
you have to pay a penalty and 

437
00:20:29,560 --> 00:20:32,040
latency. 
Not here because of that 

438
00:20:32,040 --> 00:20:35,120
compiler smaller approach, no 
interpreter overhead, Tiny 

439
00:20:35,120 --> 00:20:39,520
engine was consistently 1.7 X to
3.3 X faster than running a 

440
00:20:39,520 --> 00:20:42,560
model on tensorflow light micro.
Faster, smaller and more 

441
00:20:42,560 --> 00:20:44,640
accurate. 
It's the Holy Trinity of 

442
00:20:44,640 --> 00:20:47,520
embedded systems engineering, 
and they proved it on other 

443
00:20:47,520 --> 00:20:49,920
tasks too. 
They benchmarked it on Visual 

444
00:20:49,920 --> 00:20:53,120
Wake words. 
That's like Hey Siri or OK 

445
00:20:53,120 --> 00:20:55,560
Google but for video, right? 
Exactly. 

446
00:20:56,040 --> 00:20:59,000
The task is just detecting if a 
person is present in the frame. 

447
00:20:59,560 --> 00:21:02,560
It's used for things like smart 
doorbells or security cameras to

448
00:21:02,560 --> 00:21:06,480
wake them up. 
MC Net ran 2.4 times faster than

449
00:21:06,480 --> 00:21:09,160
the existing solutions. 
And was it more accurate? 

450
00:21:09,320 --> 00:21:12,280
Yes, and this is where the 
trade-offs get interesting. 

451
00:21:12,520 --> 00:21:15,480
The previous winner of the 
Visual Wake Words challenge had 

452
00:21:15,480 --> 00:21:18,320
optimized really hard for low 
memory, but it was very slow. 

453
00:21:19,080 --> 00:21:22,520
MC Net managed to optimize for 
both small memory and low 

454
00:21:22,520 --> 00:21:24,120
latency. 
They also mentioned they tried 

455
00:21:24,120 --> 00:21:27,760
it on object detection. 
They did using a tiny version of

456
00:21:27,760 --> 00:21:30,760
a Yolo model which stands for 
You Only Look Once. 

457
00:21:31,120 --> 00:21:33,920
This is a much harder task than 
just classification because you 

458
00:21:33,920 --> 00:21:36,240
have to draw a bounding box 
around the object you find 

459
00:21:36,280 --> 00:21:39,720
right, And on that task they 
achieved a 20% improvement in 

460
00:21:39,720 --> 00:21:43,560
mean average precision while 
staying under a strict 512 

461
00:21:43,560 --> 00:21:46,760
kilobytes RAM limit. 
A 20% improvement on a chip this

462
00:21:46,760 --> 00:21:49,240
constrained is huge. 
It's the difference between a 

463
00:21:49,240 --> 00:21:52,640
Smart car seeing a pedestrian 
and not seeing a pedestrian. 

464
00:21:53,120 --> 00:21:55,680
It's a game changer. 
So this works. 

465
00:21:55,680 --> 00:21:59,280
The technology is real. 
We have successfully shrunk the 

466
00:21:59,440 --> 00:22:01,640
AI brain. 
I want to spend our last few 

467
00:22:01,640 --> 00:22:03,600
minutes talking about the So 
what? 

468
00:22:04,280 --> 00:22:06,320
Why should you, the listener, 
care? 

469
00:22:06,400 --> 00:22:08,680
What does this unlock for the 
world? 

470
00:22:09,040 --> 00:22:11,320
Well, I think there are three 
major implications that jump out

471
00:22:11,320 --> 00:22:12,960
immediately. 
The first one is privacy. 

472
00:22:12,960 --> 00:22:15,000
You have that great line in the 
prep material. 

473
00:22:15,120 --> 00:22:17,320
What happens on the toaster 
stays on the toaster. 

474
00:22:17,320 --> 00:22:20,080
Exactly. 
Right now, to have a truly smart

475
00:22:20,080 --> 00:22:23,120
home, you basically have to 
accept a certain level of 

476
00:22:23,120 --> 00:22:25,560
surveillance. 
Your smart speaker, your 

477
00:22:25,560 --> 00:22:29,200
doorbell camera, they are 
constantly streaming raw audio 

478
00:22:29,200 --> 00:22:31,440
and video data to the cloud to 
be processed. 

479
00:22:31,440 --> 00:22:34,160
Which a lot of people are, you 
know, increasingly uncomfortable

480
00:22:34,160 --> 00:22:36,320
with, and for good reason. 
For very good reason. 

481
00:22:36,600 --> 00:22:39,720
With a system like MCU Net, we 
can move that processing to the 

482
00:22:39,720 --> 00:22:42,440
extreme edge. 
The audio analysis, the person 

483
00:22:42,440 --> 00:22:45,960
detection, the face recognition,
it all happened on that $5 chip 

484
00:22:45,960 --> 00:22:48,920
inside the device itself. 
The raw data never leaves your 

485
00:22:48,920 --> 00:22:50,600
house. 
So the device just sends a 

486
00:22:50,600 --> 00:22:56,120
simple signal, a flag, door 
opened or person detected, not 

487
00:22:56,120 --> 00:22:57,880
the actual video of you opening 
the door. 

488
00:22:57,920 --> 00:22:59,800
Right. 
And that is a massive selling 

489
00:22:59,800 --> 00:23:01,880
point for consumers who value 
their privacy. 

490
00:23:01,920 --> 00:23:03,680
OK, what's the second big 
implication? 

491
00:23:03,840 --> 00:23:05,440
The second one is 
democratization. 

492
00:23:05,480 --> 00:23:09,320
Just lowering the cost of entry.
Drastically if you need $1000 

493
00:23:09,320 --> 00:23:15,040
GPU or an $800 smartphone to run
useful AI, that automatically 

494
00:23:15,040 --> 00:23:18,080
excludes a huge part of the 
world and a huge number of 

495
00:23:18,080 --> 00:23:21,120
potential use cases. 
But if you can, run it on a $5 

496
00:23:21,120 --> 00:23:22,640
chip. 
You can put it in literally 

497
00:23:22,640 --> 00:23:24,200
everything. 
You can put crop disease 

498
00:23:24,200 --> 00:23:27,360
detection sensors in the fields 
of developing nations and they 

499
00:23:27,360 --> 00:23:31,080
can run on a small solar cell. 
You can put smart monitoring 

500
00:23:31,080 --> 00:23:33,280
into basic low cost medical 
devices. 

501
00:23:33,280 --> 00:23:36,560
You can put it in toys, 
appliances, infrastructure. 

502
00:23:36,880 --> 00:23:40,080
It makes AI truly ubiquitous. 
And the third one, I think the 

503
00:23:40,080 --> 00:23:42,160
third one is green AI. 
We talked a lot about the 

504
00:23:42,520 --> 00:23:45,440
massive carbon footprint of 
training these huge models like 

505
00:23:45,480 --> 00:23:48,840
PPG 4, but running the models 
inference consumes a ton of 

506
00:23:48,840 --> 00:23:52,280
energy too. 
Every single time I ask my smart

507
00:23:52,280 --> 00:23:55,920
assistant a question, a GPU has 
to spin up somewhere in a data 

508
00:23:55,920 --> 00:23:56,960
center. 
That's a good point. 

509
00:23:56,960 --> 00:24:00,760
Running a model locally on a 
microcontroller consumes a tiny,

510
00:24:00,760 --> 00:24:04,040
tiny fraction of the energy of 
sending data to the cloud, 

511
00:24:04,240 --> 00:24:06,600
processing it there, and sending
the result back. 

512
00:24:06,680 --> 00:24:09,680
Yeah, it enables always on 
intelligence without constantly 

513
00:24:09,680 --> 00:24:11,320
draining the battery or the 
power grid. 

514
00:24:11,440 --> 00:24:15,600
This brings me to a final, maybe
slightly provocative thought to 

515
00:24:15,600 --> 00:24:18,600
leave our listeners with. 
For the past 20 years, we've 

516
00:24:18,600 --> 00:24:20,800
gotten used to the Internet 
being cloud centric. 

517
00:24:21,040 --> 00:24:24,080
The brains are all in the server
farms, the endpoints, our 

518
00:24:24,080 --> 00:24:26,800
phones, our laptops are 
basically just fancy screens. 

519
00:24:26,800 --> 00:24:30,920
Right, they're thin clients. 
But if these 250 billion devices

520
00:24:30,920 --> 00:24:34,480
suddenly wake up, if every light
bulb, every toaster, every shoe 

521
00:24:34,480 --> 00:24:37,760
has a tiny but effective neural 
network inside it that actually 

522
00:24:37,760 --> 00:24:41,560
works, are we moving from that 
cloud centric world to an edge 

523
00:24:41,560 --> 00:24:44,240
heavy world? 
That is the big question, isn't 

524
00:24:44,240 --> 00:24:46,320
it? 
If the edge becomes truly 

525
00:24:46,320 --> 00:24:49,320
intelligent, the entire 
architecture of the Internet 

526
00:24:49,320 --> 00:24:52,400
could change. 
We might stop moving raw data 

527
00:24:52,400 --> 00:24:55,240
around. 
We would start moving knowledge 

528
00:24:55,920 --> 00:24:57,680
insights. 
We don't send the video of the 

529
00:24:57,680 --> 00:24:59,360
living room, we just send the 
insight. 

530
00:24:59,880 --> 00:25:03,160
The cat is on the couch. 
The bandwidth requirements 

531
00:25:03,160 --> 00:25:06,320
plummet, but the distributed 
computing power of the whole 

532
00:25:06,320 --> 00:25:09,840
network skyrockets. 
We might be accidentally 

533
00:25:09,840 --> 00:25:14,080
building a global distributed 
supercomputer, a digital nervous

534
00:25:14,080 --> 00:25:17,360
system for the physical world. 
And MCU net feels like it could 

535
00:25:17,360 --> 00:25:19,360
be the synapse that makes it all
connect. 

536
00:25:19,440 --> 00:25:22,640
It's certainly a key enabling 
technology, a huge step in that 

537
00:25:22,640 --> 00:25:23,720
direction. 
Too. 

538
00:25:24,120 --> 00:25:26,720
So for the developers and the 
engineers listening out there, 

539
00:25:27,000 --> 00:25:28,440
maybe take a look at your own 
stack. 

540
00:25:28,440 --> 00:25:30,600
If you're just grabbing a 
standard library because it's 

541
00:25:30,600 --> 00:25:33,600
easy, you might be leaving a 
huge amount of performance on 

542
00:25:33,600 --> 00:25:35,360
the table. 
Don't just buy the suit off the 

543
00:25:35,360 --> 00:25:37,120
rack. 
It might be time to tailor it, 

544
00:25:37,640 --> 00:25:40,440
or better yet, learn how to 
codesign it from the fabric U. 

545
00:25:40,800 --> 00:25:43,480
Tailor it, I like that. 
We will leave it there. 

546
00:25:43,520 --> 00:25:46,200
Thank you for helping us shrink 
the world and get the synthesis 

547
00:25:46,200 --> 00:25:47,160
today. 
It was a pleasure. 

548
00:25:47,360 --> 00:25:50,080
And thank you for listening to 
the deep dive. 

549
00:25:50,200 --> 00:25:50,920
We'll see you next time.
