1
00:00:00,040 --> 00:00:03,240
Welcome back to the Deep Dive. 
So we're shifting gears a little

2
00:00:03,240 --> 00:00:05,080
bit today. 
Yeah, a new format. 

3
00:00:05,160 --> 00:00:08,039
A new format. 
Usually we look at, you know, 

4
00:00:08,039 --> 00:00:11,280
the results of technology, how 
it changes society, the market 

5
00:00:11,280 --> 00:00:12,880
impact. 
Right, the big picture stuff. 

6
00:00:12,880 --> 00:00:16,880
But today we are taking a 
strictly engineering 

7
00:00:16,880 --> 00:00:18,520
perspective. 
We're launching what we're 

8
00:00:18,520 --> 00:00:22,320
calling the Practical AI Digest.
I love this. 

9
00:00:22,320 --> 00:00:25,040
So we're going to take one 
specific architecture, just one 

10
00:00:25,640 --> 00:00:30,080
strip away all the marketing 
hype and explain exactly how it 

11
00:00:30,080 --> 00:00:32,479
works so you can actually 
understand the mechanism under 

12
00:00:32,479 --> 00:00:33,720
the hood. 
Exactly. 

13
00:00:33,760 --> 00:00:37,680
And the topic for today is 
something that I think a lot of 

14
00:00:37,680 --> 00:00:40,600
people feel they understand on a
surface level, but the actual 

15
00:00:40,600 --> 00:00:43,520
mechanics are quite surprising. 
Yeah, we are looking at the 

16
00:00:43,520 --> 00:00:47,480
engine behind the massive, 
massive explosion in AI art 

17
00:00:47,480 --> 00:00:49,160
we've seen over the last couple 
of years. 

18
00:00:49,160 --> 00:00:52,840
We're talking about latent 
diffusion models, LDMS. This is 

19
00:00:52,840 --> 00:00:55,640
the thing that powers Stable 
Diffusion is basically the 

20
00:00:55,640 --> 00:00:58,760
reason you can now generate 
these incredible high fidelity 

21
00:00:58,760 --> 00:01:01,240
images on a home computer. 
Instead of needing a giant 

22
00:01:01,240 --> 00:01:04,280
server farm. 
And that is the key hook here. 

23
00:01:04,280 --> 00:01:07,920
I mean everyone has seen the 
output you know mid journey 

24
00:01:07,920 --> 00:01:12,440
deli, stable diffusion, you type
a cyberpunk city in the style of

25
00:01:12,440 --> 00:01:16,120
Van Gogh and boom 5 seconds 
later you have a masterpiece. 

26
00:01:16,480 --> 00:01:20,600
But very few people ask the next
question, which is what is the 

27
00:01:20,640 --> 00:01:22,120
engine driving this? 
Yeah. 

28
00:01:22,360 --> 00:01:25,640
And specifically, how did we get
from models that needed, like 

29
00:01:25,640 --> 00:01:29,000
you said, a super powder to 
something that can run on a 

30
00:01:29,000 --> 00:01:31,440
gaming laptop? 
And that's the core conflict. 

31
00:01:31,520 --> 00:01:34,240
You know the problem we need to 
address before this paper? 

32
00:01:34,240 --> 00:01:36,800
The high resolution image 
synthesis with late diffusion 

33
00:01:36,800 --> 00:01:39,880
models paper. 
This whole field was incredibly 

34
00:01:39,880 --> 00:01:41,640
expensive. 
It was a witch man's game, 

35
00:01:41,760 --> 00:01:43,920
totally. 
Let's put a number on expensive 

36
00:01:43,920 --> 00:01:46,960
because I think people just 
assume AI lives in the cloud and

37
00:01:46,960 --> 00:01:50,040
it's, you know, basically free. 
Oh, it's so far from free. 

38
00:01:50,040 --> 00:01:51,000
Yeah. 
I mean, if you look at the 

39
00:01:51,000 --> 00:01:53,600
source material, they highlight 
that training one of these 

40
00:01:53,600 --> 00:01:55,840
powerful pixel based diffusion 
models. 

41
00:01:55,840 --> 00:01:58,680
Yeah, it took hundreds of GPU 
days. 

42
00:01:58,760 --> 00:02:00,520
Hundreds. 
Yeah, they cite examples from 

43
00:02:00,520 --> 00:02:05,080
like 150 up to 1000 V 100 days. 
OK, let's just clarify that for 

44
00:02:05,080 --> 00:02:08,759
you listening. 
AV 100 is a top tier enterprise 

45
00:02:08,759 --> 00:02:12,840
grade graphics card. 
So 1000 V 100 days means if you 

46
00:02:12,840 --> 00:02:16,640
had one of these very expensive 
cards, it would take you almost 

47
00:02:16,640 --> 00:02:19,600
three years to train the model 
once, just once. 

48
00:02:19,600 --> 00:02:20,840
Exactly. 
And that's just the training 

49
00:02:20,840 --> 00:02:23,400
part, the inference. 
So actually making an image 

50
00:02:23,400 --> 00:02:26,440
after you type a prompt, that 
was also painfully slow. 

51
00:02:26,560 --> 00:02:28,200
Why? 
Because the neural network had 

52
00:02:28,200 --> 00:02:32,040
to sequentially evaluate a 
massive, massive amount of data.

53
00:02:32,040 --> 00:02:34,960
So the mission of this deep dive
is to unpack the architecture 

54
00:02:34,960 --> 00:02:38,360
that solved this bottleneck. 
We're not just making images. 

55
00:02:38,360 --> 00:02:40,600
We are, I guess, making them 
efficiently. 

56
00:02:41,040 --> 00:02:44,160
We're going to break down the 
latent part of latent diffusion 

57
00:02:44,160 --> 00:02:46,160
models. 
It's what democratize high 

58
00:02:46,160 --> 00:02:49,400
resolution image generation and 
the whole trick. 

59
00:02:49,400 --> 00:02:52,880
The core idea was separating the
compression phase from the 

60
00:02:52,880 --> 00:02:55,000
generative phase. 
OK, so let's get into that 

61
00:02:55,000 --> 00:02:57,480
Section 1, The intuition and the
problem with pixels. 

62
00:02:57,960 --> 00:03:00,720
So before we can understand why 
latent diffusion is so special, 

63
00:03:00,840 --> 00:03:03,800
we need to understand what 
standard diffusion even is. 

64
00:03:03,840 --> 00:03:06,160
Right, the baseline. 
If I'm a developer looking at 

65
00:03:06,160 --> 00:03:08,360
this for the first time, what's 
the basic mechanism? 

66
00:03:08,520 --> 00:03:11,240
The basic idea is something 
called iterative denoising. 

67
00:03:11,400 --> 00:03:13,720
It's. 
You can think of it as a model 

68
00:03:13,720 --> 00:03:15,720
that's designed to learn a data 
distribution. 

69
00:03:15,720 --> 00:03:18,760
Bye. 
By gradually removing noise. 

70
00:03:18,760 --> 00:03:22,200
It's a two way St. the forward. 
Process and the reverse process.

71
00:03:22,240 --> 00:03:24,040
Right, so the forward process is
fixed. 

72
00:03:24,040 --> 00:03:27,480
It's easy. 
You take a clean image, let's 

73
00:03:27,480 --> 00:03:31,160
say a photo of the cat, and you 
just gradually add Gaussian 

74
00:03:31,160 --> 00:03:32,840
noise to it. 
A little bit at a time. 

75
00:03:32,840 --> 00:03:36,080
Step by step, so you know, step 
one, it's a little grainy. 

76
00:03:36,360 --> 00:03:40,680
Step 50 it's very grainy. 
By step 1000 is just Ure static,

77
00:03:40,720 --> 00:03:43,200
the image is completely gone. 
OK, that seems simle enough. 

78
00:03:43,200 --> 00:03:46,440
Ruining a iPhoto is easy. 
It is the magic is the reverse 

79
00:03:46,440 --> 00:03:48,360
process. 
This is where you train a neural

80
00:03:48,360 --> 00:03:52,520
network to look at that Ure 
static and predict the noise 

81
00:03:52,520 --> 00:03:54,160
that was added at the previous 
step. 

82
00:03:54,320 --> 00:03:56,000
Ah, so it's learning to go 
backwards. 

83
00:03:56,000 --> 00:03:57,440
Exactly. 
Because if you can predict the 

84
00:03:57,440 --> 00:03:58,680
noise, you can subtract it. 
Yeah. 

85
00:03:58,680 --> 00:04:01,240
And if you do that over and 
over, step by step backwards 

86
00:04:01,240 --> 00:04:05,240
from total randomness, you 
eventually reveal a clear image.

87
00:04:05,360 --> 00:04:07,840
So the model is kind of like a 
professional art restorer. 

88
00:04:08,000 --> 00:04:12,080
It's looking at this messy noisy
signal and saying, OK, I know 

89
00:04:12,080 --> 00:04:15,000
what part of this is just static
and what part is actual 

90
00:04:15,000 --> 00:04:16,800
structure. 
And it just wipes the static 

91
00:04:16,800 --> 00:04:18,200
away. 
That's a perfect way to put it. 

92
00:04:18,240 --> 00:04:21,480
And ideally you end up with a 
beautiful high fidelity image. 

93
00:04:21,480 --> 00:04:26,080
You do, but here's the problem, 
and it's a huge problem with 

94
00:04:26,080 --> 00:04:30,800
standard models like Dally 2 or 
the original guided diffusion 

95
00:04:30,800 --> 00:04:33,400
models. 
They do all of this in pixel 

96
00:04:33,400 --> 00:04:35,400
space. 
Meaning it's literally 

97
00:04:35,400 --> 00:04:38,240
calculating every single pixel. 
Every single pic. 

98
00:04:38,400 --> 00:04:41,040
Think about it, you have a 10/24
by 1024 image. 

99
00:04:41,160 --> 00:04:42,960
That's over 1,000,000 pixels and
each. 

100
00:04:42,960 --> 00:04:45,360
One has three color channels. 
Red, green and blue. 

101
00:04:45,680 --> 00:04:49,400
So it's more like 3,000,000 data
points, and the model has to 

102
00:04:49,400 --> 00:04:53,680
process that entire massive grid
for every single step of the 

103
00:04:53,680 --> 00:04:55,840
denoise. 
Every step which can be hundreds

104
00:04:55,840 --> 00:04:57,960
or thousands of steps. 
OK, I see it now. 

105
00:04:57,960 --> 00:05:00,880
And this is where we hit what 
the paper calls the pixel space 

106
00:05:00,880 --> 00:05:03,480
bottleneck, right? 
And this right here is probably 

107
00:05:03,480 --> 00:05:06,000
the most critical conceptual 
take away from this whole thing.

108
00:05:06,200 --> 00:05:09,280
Digital images contain a massive
amount of information, right? 

109
00:05:09,800 --> 00:05:12,440
But not all of that information 
is actually important. 

110
00:05:13,040 --> 00:05:15,520
Semantically important. 
What do you mean by semantically

111
00:05:15,520 --> 00:05:18,120
important? 
OK, well imagine a photo of a 

112
00:05:18,120 --> 00:05:21,200
brick wall. 
The semantic information is this

113
00:05:21,200 --> 00:05:22,440
is a brick wall. 
It's red. 

114
00:05:22,600 --> 00:05:24,360
The mortar is Gray. 
The concept. 

115
00:05:24,360 --> 00:05:28,120
The concept exactly, but the 
actual digital image also 

116
00:05:28,120 --> 00:05:31,720
contains all this high frequency
detail. 

117
00:05:31,720 --> 00:05:35,240
High frequency meaning like the 
microscopic grain on the brick. 

118
00:05:35,240 --> 00:05:39,840
Yes, imperceptible noise. 
Tiny, tiny shifts in the 

119
00:05:39,840 --> 00:05:42,400
gradient of a shadow on one 
single brick. 

120
00:05:43,000 --> 00:05:46,480
The exact texture of the mortar 
if you zoomed in 1000 times. 

121
00:05:46,640 --> 00:05:50,320
All these little details take up
a huge amount of data to encode.

122
00:05:50,400 --> 00:05:53,320
But do they matter to me as the 
person looking at the picture? 

123
00:05:53,400 --> 00:05:55,320
Not really. 
I mean, if I showed you a 

124
00:05:55,320 --> 00:05:57,280
picture of that wall and then I 
showed you another one where I 

125
00:05:57,280 --> 00:05:59,840
just slightly blurred the 
texture or shifted the noise 

126
00:05:59,840 --> 00:06:02,520
pattern, you still just say 
that's a brick wall. 

127
00:06:02,680 --> 00:06:04,400
Right, the wallness of it hasn't
changed. 

128
00:06:04,400 --> 00:06:06,040
Exactly. 
But the standard diffusion 

129
00:06:06,040 --> 00:06:07,280
models, they don't know that 
they. 

130
00:06:07,280 --> 00:06:09,320
Can't tell the difference 
between the concept and the tiny

131
00:06:09,320 --> 00:06:11,280
details. 
No, because they're what are 

132
00:06:11,280 --> 00:06:14,840
called likelihood based models. 
Mathematically, they're 

133
00:06:14,840 --> 00:06:19,400
obligated to model everything, 
so they spend a completely 

134
00:06:19,400 --> 00:06:23,280
disproportionate amount of their
compute power modeling these 

135
00:06:23,280 --> 00:06:25,840
imperceptible high frequency 
details. 

136
00:06:25,840 --> 00:06:28,720
So it's like trying to write a 
book summary by, I don't know, 

137
00:06:29,000 --> 00:06:32,360
analyzing the microscopic fibers
of the paper on every single 

138
00:06:32,360 --> 00:06:33,840
page. 
Instead of just reading the 

139
00:06:33,840 --> 00:06:35,160
plot. 
Instead of just reading the plot

140
00:06:35,160 --> 00:06:37,520
points, yeah. 
That is the perfect analogy. 

141
00:06:37,640 --> 00:06:42,720
You are burning these incredibly
expensive GPU cycles just to 

142
00:06:42,720 --> 00:06:45,080
figure figure out the texture of
the paper, but all you really 

143
00:06:45,080 --> 00:06:47,480
want is the story. 
That's the difference between 

144
00:06:47,480 --> 00:06:49,680
perceptual compression and 
semantic compression. 

145
00:06:49,840 --> 00:06:53,000
So the goal of the latent 
diffusion model is to stop 

146
00:06:53,000 --> 00:06:55,960
looking at the paper fibers. 
Yes, the goal is to find a 

147
00:06:55,960 --> 00:06:59,440
representation of the image, a 
different space that is 

148
00:06:59,440 --> 00:07:03,160
computationally way way cheaper,
but is perceptually equivalent 

149
00:07:03,160 --> 00:07:05,200
to the original. 
We want to find a space where we

150
00:07:05,200 --> 00:07:07,880
can do the heavy lifting of 
generation without being bogged 

151
00:07:07,880 --> 00:07:10,000
down by millions of useless 
pixels. 

152
00:07:10,000 --> 00:07:13,040
Precisely. 
OK, this brings us to Section 2,

153
00:07:13,720 --> 00:07:16,800
the system architecture. 
The paper splits this whole 

154
00:07:16,800 --> 00:07:20,480
system into two very distinct 
phases, and phase one you said 

155
00:07:20,480 --> 00:07:24,560
is strictly about compression. 
Right, first we need to get out 

156
00:07:24,560 --> 00:07:27,320
of pixel space and to do that 
they introduce a component 

157
00:07:27,320 --> 00:07:30,480
called an auto encoder. 
OK, for the engineers listening,

158
00:07:30,480 --> 00:07:32,880
let's breakdown the specs of 
this auto encoder. 

159
00:07:32,880 --> 00:07:34,480
So an auto encoder has two 
parts. 

160
00:07:35,000 --> 00:07:37,240
They have an encoder, which we 
can call egoder, which we'll 

161
00:07:37,240 --> 00:07:39,720
call encoder and encoder. 
The encoder takes your high 

162
00:07:39,720 --> 00:07:42,920
dimensional image 6 LR. 
That's your massive grid of RGB 

163
00:07:42,920 --> 00:07:46,160
pixels, and it crunches it down 
to something called a latent 

164
00:07:46,160 --> 00:07:48,200
representation. 
We call that D getter. 

165
00:07:48,200 --> 00:07:52,440
So $6 is the big original image 
and Zeidol is the tiny 

166
00:07:52,440 --> 00:07:55,640
compressed version, correct? 
And Zelda lives in what we call 

167
00:07:55,640 --> 00:07:58,240
latent space. 
It's a much, much lower 

168
00:07:58,240 --> 00:08:01,280
dimensional space. 
The paper talks about a specific

169
00:08:01,280 --> 00:08:03,880
parameter for this, a down 
sampling factor, which they call

170
00:08:03,880 --> 00:08:05,440
Philidol. 
And what is that control? 

171
00:08:05,440 --> 00:08:07,240
It just controls the ratio of 
compression. 

172
00:08:07,320 --> 00:08:11,280
So if Philidol is equal to 4, 
the latent representations has a

173
00:08:11,280 --> 00:08:13,640
height and width that's 1/4 of 
the original image. 

174
00:08:13,920 --> 00:08:18,520
And if $5 is 8, it's 18. 
So we're shrinking the map. 

175
00:08:18,600 --> 00:08:20,760
But I mean, that brings me back 
to my earlier concern. 

176
00:08:20,920 --> 00:08:22,960
When you shrink an image, you 
lose quality. 

177
00:08:23,120 --> 00:08:25,800
If I take a big photo in 
Photoshop and I shrink it down 

178
00:08:25,800 --> 00:08:29,560
to 10% and then try to blow it 
back up, it looks awful. 

179
00:08:29,560 --> 00:08:33,840
It's all pixelated and blurry. 
How does the LDM avoid that? 

180
00:08:34,480 --> 00:08:36,679
This is such a critical design 
choice in the paper. 

181
00:08:37,240 --> 00:08:40,799
They don't just use a standard 
loss function like an L1 or L2 

182
00:08:40,799 --> 00:08:42,480
loss. 
Mean squared error things like. 

183
00:08:42,559 --> 00:08:46,280
That right, those standard loss 
functions tend to produce really

184
00:08:46,280 --> 00:08:50,280
blurry results because the model
is trying to average out all the

185
00:08:50,280 --> 00:08:52,840
uncertainty. 
Right, the mean of all possible 

186
00:08:52,840 --> 00:08:55,400
sharp edges ends up being just a
blurry gradient. 

187
00:08:55,440 --> 00:08:57,840
Exactly. 
So to keep things sharp, they do

188
00:08:57,840 --> 00:09:00,320
something clever. 
They use a perceptual loss 

189
00:09:00,400 --> 00:09:03,640
combined with a patch based 
adversarial objective. 

190
00:09:03,720 --> 00:09:05,520
Adversary. 
OK, that sounds like we're in 

191
00:09:05,520 --> 00:09:07,920
Jan territory. 
It's borrowed directly from Jans

192
00:09:07,920 --> 00:09:11,080
Generative Adversarial Networks.
They essentially have a second 

193
00:09:11,080 --> 00:09:14,320
network, a discriminator, that 
looks at the reconstructed image

194
00:09:14,320 --> 00:09:16,240
that comes out of the decoder. 
And it critiques it. 

195
00:09:16,440 --> 00:09:19,840
It critiques it. 
It basically asks does this 

196
00:09:19,840 --> 00:09:22,400
little patch of the image look 
like a real texture? 

197
00:09:22,400 --> 00:09:25,040
Is this sharp or is this blurry 
mush? 

198
00:09:25,200 --> 00:09:28,240
So it forces the auto encoder to
be honest, it can't just get 

199
00:09:28,240 --> 00:09:30,240
away with a blurry 
approximation. 

200
00:09:30,360 --> 00:09:33,800
No, it has to maintain what they
call local realism. 

201
00:09:33,800 --> 00:09:36,640
Local realism I like. 
That, and that's what ensures 

202
00:09:36,640 --> 00:09:40,040
that even though Siller is 
highly compressed, when you run 

203
00:09:40,040 --> 00:09:43,560
it back through the decoder 
dollar, the output image looks 

204
00:09:43,560 --> 00:09:46,680
crisp and detailed. 
Now, there's a technical nuance 

205
00:09:46,680 --> 00:09:49,760
here about regularization. 
The paper mentions two ways to 

206
00:09:49,760 --> 00:09:51,720
keep this latent space well 
behaved. 

207
00:09:52,200 --> 00:09:54,720
Why do we care if the latent 
space is well behaved? 

208
00:09:54,720 --> 00:09:57,520
Well, imagine if you just mapped
a bunch of images to a latent 

209
00:09:57,520 --> 00:10:00,960
space and the points were just 
scattered completely randomly. 

210
00:10:00,960 --> 00:10:03,960
So a cat is here, and a 
millimeter away is a toaster. 

211
00:10:04,360 --> 00:10:06,520
In in between it's just garbage 
static. 

212
00:10:06,640 --> 00:10:08,800
That would be a high variance 
chaotic space. 

213
00:10:08,800 --> 00:10:11,960
You can't navigate that. 
You can't, and you certainly 

214
00:10:11,960 --> 00:10:13,480
can't train a diffusion model on
it. 

215
00:10:14,000 --> 00:10:17,320
Diffusion models need a smooth 
continuous space that they can 

216
00:10:17,320 --> 00:10:21,040
walk through, so regularization 
is about forcing the latent 

217
00:10:21,040 --> 00:10:24,760
vectors to follow a standard 
predictable distribution. 

218
00:10:25,400 --> 00:10:26,880
And what are the two options 
they talk about? 

219
00:10:27,240 --> 00:10:29,760
Option A is called KL 
regularization. 

220
00:10:30,000 --> 00:10:33,200
This adds A slight penalty that 
kind of nudges the latent 

221
00:10:33,200 --> 00:10:36,520
distribution towards a standard 
normal distribution is very 

222
00:10:36,520 --> 00:10:40,080
similar to how variational auto 
encoders or VA ES work. 

223
00:10:40,080 --> 00:10:41,320
OK. 
And option B. 

224
00:10:41,560 --> 00:10:45,680
Option B is VQ regularization, 
which stands for Vector 

225
00:10:45,680 --> 00:10:48,480
quantization. 
This uses a kind of code book of

226
00:10:48,480 --> 00:10:51,400
learned vectors. 
It quantizes the latent space 

227
00:10:51,400 --> 00:10:54,320
into discrete separate steps. 
This is more like what you see 

228
00:10:54,320 --> 00:10:56,360
in VQGAN. 
So which one did they end up 

229
00:10:56,360 --> 00:10:58,760
going with? 
They experimented with both, but

230
00:10:58,760 --> 00:11:01,840
the kale regularization version 
is what became the standard for 

231
00:11:01,840 --> 00:11:04,200
Stable Diffusion. 
It just offers a really nice 

232
00:11:04,200 --> 00:11:05,600
balance of quality and 
stability. 

233
00:11:05,640 --> 00:11:07,400
OK, so let's just recap phase 
one. 

234
00:11:07,560 --> 00:11:10,360
We take an image, we use an 
encoder to crunch it down into a

235
00:11:10,360 --> 00:11:14,040
tiny efficient latent vector 7. 
We then use a decoder to prove 

236
00:11:14,040 --> 00:11:16,960
that we can turn Zeiger's back 
into a sharp looking image, but 

237
00:11:16,960 --> 00:11:18,440
we haven't generated anything 
new yet. 

238
00:11:18,440 --> 00:11:20,080
Not a thing. 
We're just compressing and 

239
00:11:20,080 --> 00:11:21,840
decompressing. 
It's just building the stage. 

240
00:11:21,840 --> 00:11:24,200
Exactly. 
Phase 2 is where the performance

241
00:11:24,200 --> 00:11:26,480
actually happens. 
This is latent diffusion. 

242
00:11:26,480 --> 00:11:29,160
This is the core shift. 
We are no longer diffusing 

243
00:11:29,160 --> 00:11:32,280
pixels, we are diffusing the 
latent vectors. 

244
00:11:32,280 --> 00:11:33,680
Right. 
We take that compressed 

245
00:11:33,680 --> 00:11:37,800
representation 70 hours, we add 
noise to that, and then we train

246
00:11:37,800 --> 00:11:39,840
a model to remove the noise from
Ziller. 

247
00:11:40,040 --> 00:11:43,120
So we are essentially 
hallucinating the compressed 

248
00:11:43,120 --> 00:11:46,040
concept of an image rather than 
the image itself. 

249
00:11:46,160 --> 00:11:49,760
That's a great way to put it. 
And because ZZS is so, so much 

250
00:11:49,760 --> 00:11:53,400
smaller than the original Image 
60's, the computational cost 

251
00:11:53,640 --> 00:11:56,120
just drops off a Cliff. 
OK, let's talk about the 

252
00:11:56,120 --> 00:11:57,840
backbone. 
What's the neural network that's

253
00:11:57,840 --> 00:12:00,160
actually doing this work? 
Is it a transformer? 

254
00:12:00,160 --> 00:12:02,320
It feels like everything is a 
transformer these days. 

255
00:12:02,360 --> 00:12:04,000
Right, that's the default 
assumption. 

256
00:12:04,120 --> 00:12:07,200
Daily one used a transformer. 
It did, but no. 

257
00:12:07,200 --> 00:12:10,160
And this is a really specific 
and I think very smart 

258
00:12:10,160 --> 00:12:13,280
engineering decision. 
They use a time conditional U 

259
00:12:13,280 --> 00:12:14,240
net AU. 
Net. 

260
00:12:14,240 --> 00:12:15,040
Why AU? 
Net. 

261
00:12:15,280 --> 00:12:18,600
So the paper argues that 
Transformers are naturally 

262
00:12:18,600 --> 00:12:21,720
sequence based. 
They treat data like a stream of

263
00:12:21,720 --> 00:12:23,840
words in a sentence. 
One after another. 

264
00:12:24,080 --> 00:12:28,000
But images, even latent images, 
they aren't sequences, they're 

265
00:12:28,000 --> 00:12:30,320
2D spatial grids. 
They have structure. 

266
00:12:30,520 --> 00:12:33,800
The top left corner is related 
to the top right corner in a way

267
00:12:33,800 --> 00:12:36,280
that the first word of a 
sentence isn't necessarily 

268
00:12:36,280 --> 00:12:38,080
related to the last. 
Exactly. 

269
00:12:38,240 --> 00:12:41,400
A unit is built from 
convolutional layers, and 

270
00:12:41,400 --> 00:12:44,520
convolutions are designed 
specifically to process spatial 

271
00:12:44,520 --> 00:12:45,880
grids. 
They have what's called an 

272
00:12:45,880 --> 00:12:50,080
inductive bias towards images. 
They naturally understand things

273
00:12:50,080 --> 00:12:53,080
like locality and, you know, 
translation and variance, so. 

274
00:12:53,080 --> 00:12:56,280
By using a unit, they're working
with the structure of the data 

275
00:12:56,280 --> 00:12:58,960
and not trying to force it into 
a different shape, and I assume 

276
00:12:58,960 --> 00:13:00,720
this is faster. 
Much faster. 

277
00:13:00,960 --> 00:13:03,240
Remember our down sampling 
factor $5? 

278
00:13:03,640 --> 00:13:08,400
If $5 is 8, the latent space has
64 times fewer data points than 

279
00:13:08,400 --> 00:13:12,120
the original pixel space, so the
UNET is processing a tiny, tiny 

280
00:13:12,120 --> 00:13:14,720
fraction of the data that a 
standard diffusion model would 

281
00:13:14,720 --> 00:13:16,800
have to. 
And that explains how we can run

282
00:13:16,800 --> 00:13:20,240
this on a single consumer GPU. 
That's the whole ball game. 

283
00:13:20,440 --> 00:13:24,640
This design allows for training 
on limited resources and, 

284
00:13:24,840 --> 00:13:27,080
crucially, very efficient 
inference. 

285
00:13:27,440 --> 00:13:29,000
Let's look at the math for just 
a second. 

286
00:13:29,040 --> 00:13:31,760
We don't need to get bogged 
down, but for the engineers 

287
00:13:31,760 --> 00:13:34,440
trying to implement this, what's
the objective function? 

288
00:13:34,440 --> 00:13:36,760
What is the model actually 
trying to minimize? 

289
00:13:36,760 --> 00:13:39,440
It's a simplified version of the
standard diffusion loss. 

290
00:13:39,960 --> 00:13:44,440
The formula is the LLDM equals 
the expectation of the squared 

291
00:13:44,440 --> 00:13:46,920
difference between epsilon and 
epsilon Theta. 

292
00:13:47,080 --> 00:13:49,680
OK, let's translate that into 
plain English for sure, yeah. 

293
00:13:50,280 --> 00:13:53,320
You have the actual noise that's
epsilon which you just sampled 

294
00:13:53,320 --> 00:13:56,720
from a normal distribution. 
OK, And you have the predicted 

295
00:13:56,720 --> 00:13:59,400
noise that's epsilon Theta, 
which comes from your neural 

296
00:13:59,400 --> 00:14:02,840
network, EU net, right? 
EU net looks at the noisy latent

297
00:14:02,840 --> 00:14:07,840
vector Z dollar at a specific 
time step T dollars, and it 

298
00:14:07,840 --> 00:14:10,280
makes a prediction. 
The objective is simply to make 

299
00:14:10,280 --> 00:14:13,440
the predicted noise as close as 
possible to the actual noise. 

300
00:14:13,440 --> 00:14:15,440
So the model is just playing a 
guessing game. 

301
00:14:15,440 --> 00:14:17,520
It's saying I see this messy 
static. 

302
00:14:17,720 --> 00:14:19,960
I think the noise you added 
looks like this specific 

303
00:14:19,960 --> 00:14:21,880
pattern. 
And we grade it on how close 

304
00:14:21,880 --> 00:14:24,280
it's guess was. 
If it guesses the noise 

305
00:14:24,280 --> 00:14:27,800
correctly, we can abstract it 
and take one step closer to a 

306
00:14:27,800 --> 00:14:29,120
clean image. 
That's it. 

307
00:14:29,120 --> 00:14:32,280
So now we have a system that can
generate random images from 

308
00:14:32,280 --> 00:14:34,600
noise, and it can do it really 
efficiently. 

309
00:14:35,040 --> 00:14:38,280
But we don't want random images.
No, we want specific images. 

310
00:14:38,280 --> 00:14:41,800
We want a cat eating pizza. 
Right, and this brings us to 

311
00:14:41,800 --> 00:14:45,840
Section 4, conditioning. 
This is the mechanism that 

312
00:14:45,840 --> 00:14:47,960
enables text to image 
generation. 

313
00:14:48,160 --> 00:14:51,080
How do we tell the model what to
generate? 

314
00:14:51,080 --> 00:14:55,480
I mean EU net speaks this 
spatial image language, but our 

315
00:14:55,480 --> 00:14:57,640
prompts are in language sequence
language. 

316
00:14:57,640 --> 00:14:59,000
We need a bridge between the 
two. 

317
00:14:59,000 --> 00:15:01,960
We do and the solution they 
implemented is brilliant. 

318
00:15:01,960 --> 00:15:04,920
It's called cross attention. 
This is a mechanism that's 

319
00:15:04,920 --> 00:15:06,560
borrowed from the transformer 
world, isn't? 

320
00:15:06,560 --> 00:15:09,320
It it is, yeah. 
They effectively graphed these 

321
00:15:09,320 --> 00:15:12,800
cross attention layers onto the 
unit backbone, but they make the

322
00:15:12,800 --> 00:15:14,920
really smart architectural move 
here. 

323
00:15:15,160 --> 00:15:17,480
They separate the text 
processing from the image 

324
00:15:17,480 --> 00:15:19,560
generation. 
They introduce what they call a 

325
00:15:19,560 --> 00:15:23,800
domain specific encoder, which 
they label as taffeta. 

326
00:15:24,240 --> 00:15:26,200
Taffeta. 
It's a totally separate model. 

327
00:15:26,200 --> 00:15:29,280
It could be bird, it could be a 
colleague model that takes your 

328
00:15:29,280 --> 00:15:32,960
input, say your text prompt 
right, and IT projects it into 

329
00:15:32,960 --> 00:15:36,200
an intermediate representation. 
So the text goes into its own 

330
00:15:36,200 --> 00:15:39,600
dedicated processor first, 
before it ever touches the image

331
00:15:39,600 --> 00:15:42,760
generator. 
Yes, and only then is that 

332
00:15:42,760 --> 00:15:46,440
process text representation fed 
into the unit through the 

333
00:15:46,440 --> 00:15:49,400
attention mechanism. 
If you think of the attention 

334
00:15:49,400 --> 00:15:54,320
formula, attention to QKV. 
Query key and value. 

335
00:15:54,440 --> 00:15:56,440
Right, let's map those terms to 
this system. 

336
00:15:56,440 --> 00:15:58,680
OK, who's the query? 
The query comes from the unit. 

337
00:15:58,680 --> 00:16:01,120
It's coming from the 
intermediate layers of the image

338
00:16:01,120 --> 00:16:02,800
in progress. 
You can think of it as the image

339
00:16:02,800 --> 00:16:05,000
asking questions. 
And the keys and values the. 

340
00:16:05,000 --> 00:16:08,520
Keys and values come from our 
conditioning input, so the text 

341
00:16:08,520 --> 00:16:11,400
prompt after it's been processed
by that tough head encoder. 

342
00:16:11,560 --> 00:16:14,360
So visually, can you walk me 
through what's happening here? 

343
00:16:14,520 --> 00:16:18,640
OK, imagine the unit is trying 
to denoise a little patch of the

344
00:16:18,640 --> 00:16:21,440
latent image. 
It generates a query that is 

345
00:16:21,440 --> 00:16:24,440
effectively asking what is 
supposed to be in this top left 

346
00:16:24,440 --> 00:16:27,240
corner right now. 
OK, the crossattention layer 

347
00:16:27,360 --> 00:16:30,240
then takes that query and 
compares it to the keys from the

348
00:16:30,240 --> 00:16:33,120
text prompt. 
If the prompt is a red cat on a 

349
00:16:33,120 --> 00:16:37,000
blue table, the keys for red and
cat might light up for that 

350
00:16:37,000 --> 00:16:37,840
area. 
They become. 

351
00:16:37,840 --> 00:16:39,200
Relevant. 
They become relevant and then 

352
00:16:39,200 --> 00:16:41,520
the values provide the actual 
information. 

353
00:16:42,040 --> 00:16:45,040
This area should have red furry 
texture. 

354
00:16:45,040 --> 00:16:47,680
That's a fantastic explanation 
that So the image generation 

355
00:16:47,680 --> 00:16:51,040
process is literally querying 
the text prompt for guidance at 

356
00:16:51,040 --> 00:16:52,840
every single step of the 
denoising. 

357
00:16:52,840 --> 00:16:56,440
And this design is so incredibly
flexible because the unit only 

358
00:16:56,440 --> 00:16:58,880
interacts with the conditioning 
through this clean cross 

359
00:16:58,880 --> 00:17:01,640
attention interface. 
The conditioning input can be 

360
00:17:02,160 --> 00:17:03,840
anything. 
Doesn't have to be text. 

361
00:17:04,040 --> 00:17:05,599
Nope. 
As long as you have a domain 

362
00:17:05,599 --> 00:17:08,880
specific encoder atop of data 
that can turn your input into 

363
00:17:08,880 --> 00:17:11,400
the right shape for the key and 
value pairs, you can. 

364
00:17:11,400 --> 00:17:13,839
Plug anything in. 
You can condition on anything. 

365
00:17:13,839 --> 00:17:16,599
You can use semantic maps, so 
like a color-coded layout. 

366
00:17:16,880 --> 00:17:20,160
You can condition on class 
labels or even other images. 

367
00:17:20,880 --> 00:17:24,319
And this explains why the whole 
Stable Diffusion ecosystem is so

368
00:17:24,319 --> 00:17:26,440
modular. 
People are constantly swapping 

369
00:17:26,440 --> 00:17:29,560
out these conditioning encoders 
to build things like control net

370
00:17:29,560 --> 00:17:31,880
or IP adapter. 
Without having to retrain the 

371
00:17:31,880 --> 00:17:33,880
core unit from scratch. 
It's a plug and play 

372
00:17:33,880 --> 00:17:35,400
architecture for control 
signals. 

373
00:17:35,400 --> 00:17:38,000
Exactly. 
OK, so we've covered the auto 

374
00:17:38,000 --> 00:17:41,400
encoder, the latent diffusion 
process itself, and the 

375
00:17:41,400 --> 00:17:44,440
conditioning. 
Now let's get practical Section 

376
00:17:44,440 --> 00:17:48,920
5 trade-offs in tuning. 
If I'm an engineer trying to 

377
00:17:48,920 --> 00:17:51,720
implement this, or even just a 
power user trying to optimize my

378
00:17:51,720 --> 00:17:54,040
workflow, what are the levers I 
need to know about? 

379
00:17:54,440 --> 00:17:57,400
The paper spends a lot of time 
on that down sampling factor, 

380
00:17:57,480 --> 00:17:59,400
$5. 
Yes, this is the sweet spot 

381
00:17:59,400 --> 00:18:01,600
analysis. 
They tested a whole range of 

382
00:18:01,600 --> 00:18:07,040
factors and five five of 124816 
all the way up to 32. 

383
00:18:07,040 --> 00:18:09,960
And just to remind everyone, 
$5.00 is how much we compress 

384
00:18:09,960 --> 00:18:12,240
the image, right? 
So let's look at the extremes. 

385
00:18:12,240 --> 00:18:16,360
What happens at five $302.00? 
That's a massive amount of 

386
00:18:16,360 --> 00:18:20,320
compression. 
It is at feet three, 302 hours. 

387
00:18:20,560 --> 00:18:25,480
You're squashing a 256 by 256 
image all the way down to an 8 

388
00:18:25,480 --> 00:18:27,400
by 8 grid. 
It's tiny. 

389
00:18:27,520 --> 00:18:29,080
And the benefit is speed I 
assume. 

390
00:18:29,360 --> 00:18:31,800
The benefit is the diffusion 
model is lightning fast. 

391
00:18:31,800 --> 00:18:33,840
It has almost no data to 
process. 

392
00:18:34,240 --> 00:18:37,280
But the downside is the auto 
encoder just fails. 

393
00:18:37,680 --> 00:18:39,880
It abstracts away way too much 
information. 

394
00:18:39,880 --> 00:18:41,560
And that becomes too simple to 
be useful. 

395
00:18:41,560 --> 00:18:43,440
Exactly. 
The perceptual quality just 

396
00:18:43,440 --> 00:18:45,280
stagnates. 
You lose all the texture, you 

397
00:18:45,280 --> 00:18:48,120
lose text legibility, the 
reconstruction looks. 

398
00:18:48,360 --> 00:18:51,520
I don't know, sketchy because 
the latent space just doesn't 

399
00:18:51,520 --> 00:18:53,600
have the capacity to store the 
fine details anymore. 

400
00:18:53,640 --> 00:18:54,800
OK. 
And what about the other 

401
00:18:54,800 --> 00:18:57,320
extreme? 
And four of one or two? 

402
00:18:57,320 --> 00:19:00,000
Well, 4111 is just pixel space. 
You're right back to the 

403
00:19:00,000 --> 00:19:02,720
original problem. 
So it's slow high memory usage. 

404
00:19:02,720 --> 00:19:05,120
You're not gaining any of the 
efficiency benefits that make 

405
00:19:05,120 --> 00:19:07,640
the LDM architecture so great. 
So what's the verdict? 

406
00:19:07,640 --> 00:19:10,480
Where's the Goldilocks zone? 
So the paper concludes that the 

407
00:19:10,480 --> 00:19:12,920
audible balance is somewhere 
between in five of four and in 

408
00:19:12,920 --> 00:19:15,000
five of eight. 
This is where they achieved the 

409
00:19:15,000 --> 00:19:16,840
best FID scores. 
And FI. 

410
00:19:17,120 --> 00:19:20,720
First yet inception distance. 
It's basically the standard 

411
00:19:20,720 --> 00:19:24,200
metric for measuring how real 
generated images look compared 

412
00:19:24,200 --> 00:19:27,120
to a data set of real images. 
Lower is better. 

413
00:19:27,120 --> 00:19:30,040
OK. 
And the models they call LDM 4 

414
00:19:30,040 --> 00:19:33,760
and LDM 8 down sampling by 4:00 
and 8:00, they achieved 

415
00:19:33,760 --> 00:19:37,040
state-of-the-art results on day 
sets like Celebi HQ and 

416
00:19:37,040 --> 00:19:38,480
Imagenet. 
And they were faster. 

417
00:19:38,520 --> 00:19:42,120
They significantly outperformed 
the pixel based models in sample

418
00:19:42,120 --> 00:19:45,080
throughput, which is speed, 
while also getting lower FID 

419
00:19:45,080 --> 00:19:47,960
scores which is better quality. 
It is so rare in engineering to 

420
00:19:47,960 --> 00:19:50,360
get faster and better at the 
same time. 

421
00:19:50,360 --> 00:19:53,000
Usually you have to pick one. 
It's very rare, and that's 

422
00:19:53,000 --> 00:19:54,840
really why this paper was such a
breakthrough. 

423
00:19:55,160 --> 00:19:57,240
But there's another practical 
feature that I think is even 

424
00:19:57,240 --> 00:19:58,920
cooler. 
It's called convolutional 

425
00:19:58,920 --> 00:20:00,200
sampling. 
And what does that? 

426
00:20:00,200 --> 00:20:01,720
Enable. 
It enables resolution solution 

427
00:20:01,720 --> 00:20:04,160
independence. 
Most transformer based models 

428
00:20:04,160 --> 00:20:07,840
are totally stuck generating 
images at the exact resolution 

429
00:20:07,840 --> 00:20:08,840
they were trained on. 
Right. 

430
00:20:08,880 --> 00:20:13,120
If you trained on 256 by 256, 
you get 256 by 256. 

431
00:20:13,280 --> 00:20:15,080
That's it. 
But LDMS are different. 

432
00:20:15,320 --> 00:20:17,200
Because the model is fully 
convolutional. 

433
00:20:17,360 --> 00:20:20,280
It learns features rather than 
fixed positions. 

434
00:20:20,280 --> 00:20:22,640
It learns what a tree looks like
or what sky looks like. 

435
00:20:23,200 --> 00:20:27,280
Which means you can take a model
that was trained on 256 by 266 

436
00:20:27,280 --> 00:20:31,280
images and just apply it to a 
10/24 by the 424 canvas. 

437
00:20:31,280 --> 00:20:33,640
Yeah, it doesn't just break. 
It doesn't break, it just 

438
00:20:33,640 --> 00:20:36,360
generates more stuff. 
If you ask for a landscape, 

439
00:20:36,360 --> 00:20:39,080
it'll just generate A wider 
landscape with more trees and 

440
00:20:39,080 --> 00:20:41,920
more mountains. 
The paper shows these amazing 

441
00:20:41,920 --> 00:20:46,280
examples of generating huge, 
consistent panoramas that are 

442
00:20:46,280 --> 00:20:48,440
way larger than any of the 
training images. 

443
00:20:48,440 --> 00:20:50,640
That is. 
Huge for creative applications. 

444
00:20:50,640 --> 00:20:53,840
It implies the model understands
textures and relationships on a 

445
00:20:53,840 --> 00:20:56,440
local level, so it can just tile
them indefinitely. 

446
00:20:56,440 --> 00:20:59,280
It also enables these really 
powerful editing workflows like 

447
00:20:59,280 --> 00:21:01,600
in painting. 
So filling in a missing part of 

448
00:21:01,600 --> 00:21:02,720
an image. 
Exactly. 

449
00:21:02,720 --> 00:21:05,160
You can mask out a part of an 
image and just ask the model to 

450
00:21:05,160 --> 00:21:07,560
fill it in. 
The LDM handles this completely 

451
00:21:07,560 --> 00:21:10,040
naturally. 
It just treats the unmasked part

452
00:21:10,040 --> 00:21:13,880
as known latent codes and then 
diffuses the mask part to match 

453
00:21:13,880 --> 00:21:16,480
the context. 
How does it stack up up against 

454
00:21:16,480 --> 00:21:19,000
specialized tools built just for
that? 

455
00:21:19,240 --> 00:21:22,280
So the paper compares it against
specialized in painting 

456
00:21:22,280 --> 00:21:25,640
architectures like a model 
called Llama, and the LDM 

457
00:21:25,640 --> 00:21:27,280
achieves competitive 
performance. 

458
00:21:27,520 --> 00:21:30,120
Which is pretty remarkable for a
general purpose. 

459
00:21:30,120 --> 00:21:31,200
Model. 
It's amazing. 

460
00:21:31,200 --> 00:21:33,920
It wasn't built just for in 
painting, but it happens to be 

461
00:21:33,920 --> 00:21:36,840
great at repair work too. 
And what about super resolution 

462
00:21:37,240 --> 00:21:40,600
upscaling? 
Same story, you can use LDMS to 

463
00:21:40,600 --> 00:21:45,160
upscale low res images and 
because they're likelihood based

464
00:21:45,400 --> 00:21:47,720
they avoid something called mode
collapse. 

465
00:21:47,800 --> 00:21:49,880
That's a classic jam problem 
right? 

466
00:21:49,880 --> 00:21:52,840
Where the model just learns to 
output the same generic good 

467
00:21:52,840 --> 00:21:55,120
looking face over and over. 
Again, exactly. 

468
00:21:55,200 --> 00:21:59,200
LDM's preserve diversity, and 
unlike the big autoaggressive 

469
00:21:59,200 --> 00:22:02,440
models, they don't need billions
and billions of parameters to do

470
00:22:02,440 --> 00:22:04,200
it. 
They're just efficient, diverse,

471
00:22:04,200 --> 00:22:06,400
and sharp. 
So if we step back and look at 

472
00:22:06,400 --> 00:22:09,800
the big picture here, the LDM 
architecture really is a kind of

473
00:22:09,880 --> 00:22:11,840
best of all worlds approach. 
I think so. 

474
00:22:12,360 --> 00:22:15,200
They took the perceptual quality
of Jans through that auto 

475
00:22:15,200 --> 00:22:18,560
encoders discriminator, they 
took the stability and diversity

476
00:22:18,560 --> 00:22:21,800
of diffusion models, they took 
the control of Transformers 

477
00:22:21,800 --> 00:22:25,360
using cross attention, and they 
combined all of it into a system

478
00:22:25,360 --> 00:22:27,760
that's efficient enough to run 
on consumer hardware. 

479
00:22:27,760 --> 00:22:30,800
That is the summary. 
They identified that the pixel 

480
00:22:30,920 --> 00:22:33,800
was the wrong unit of currency 
for generation. 

481
00:22:34,280 --> 00:22:36,840
The latent vector is the right 
unit. 

482
00:22:37,280 --> 00:22:39,800
Let's wrap up with the aha 
moment. 

483
00:22:40,760 --> 00:22:44,000
For me, the genius of latent 
diffusion is that separation of 

484
00:22:44,000 --> 00:22:46,240
concerns. 
You don't need to paint every 

485
00:22:46,240 --> 00:22:50,720
single pixel from scratch. 
You can dream in this compressed

486
00:22:50,720 --> 00:22:54,000
abstract space, the latent 
space, and then you have a 

487
00:22:54,000 --> 00:22:57,760
separate dedicated translator, 
the decoder, to turn that dream 

488
00:22:57,760 --> 00:23:00,720
into a high fidelity reality. 
I like to think of it like a 

489
00:23:00,720 --> 00:23:04,480
chef versus a chemist. 
The diffusion model is the chef,

490
00:23:04,640 --> 00:23:07,080
It's the artist. 
It decides what ingredients go 

491
00:23:07,080 --> 00:23:08,520
together to make a delicious 
meal. 

492
00:23:09,040 --> 00:23:11,920
The auto encoder is the chemist.
It worries about the precise 

493
00:23:11,920 --> 00:23:15,200
molecular structure of the food.
But the chef doesn't need to 

494
00:23:15,200 --> 00:23:17,760
know chemistry to cook, right? 
They just need to know what 

495
00:23:17,760 --> 00:23:20,280
flavors work together. 
And by separating those roles. 

496
00:23:20,280 --> 00:23:22,000
The chef can work much much 
faster. 

497
00:23:22,040 --> 00:23:24,240
And the impact of the separation
has been huge. 

498
00:23:24,520 --> 00:23:27,680
It completely reduced the 
barrier to entry from tech giant

499
00:23:27,680 --> 00:23:30,360
supercomputer down to single 
GPU. 

500
00:23:30,640 --> 00:23:35,000
It effectively democratized high
resolution image synthesis. 

501
00:23:35,080 --> 00:23:38,360
Yeah, and that's where we're 
seeing this incredible explosion

502
00:23:38,720 --> 00:23:41,920
of creativity and tools in the 
open source community right now.

503
00:23:42,040 --> 00:23:44,840
So here's a final provocative 
thought for you to Mull over. 

504
00:23:45,120 --> 00:23:47,920
We've just spent this whole time
talking about how compressing 

505
00:23:47,920 --> 00:23:51,520
images into latent space let's 
us generate visuals efficiently,

506
00:23:51,560 --> 00:23:53,160
right? 
We're already seeing similar 

507
00:23:53,160 --> 00:23:56,840
logic being applied to audio 
generating music by diffusing 

508
00:23:56,840 --> 00:24:00,360
spectrograms in a latent space. 
Yeah. 

509
00:24:00,440 --> 00:24:03,360
What happens when we start 
applying this latent diffusion 

510
00:24:03,360 --> 00:24:07,840
logic to video day or to entire 
3D environments? 

511
00:24:07,840 --> 00:24:09,600
We are already standing on that 
precipice. 

512
00:24:09,600 --> 00:24:12,560
I mean, if we can dream up a 
static snapshot on a home 

513
00:24:12,560 --> 00:24:15,560
computer today, are we moving 
toward a future where we can 

514
00:24:15,560 --> 00:24:18,560
generate entire realities? 
You know, movies, video games, 

515
00:24:18,560 --> 00:24:21,760
immersive worlds, all of it all 
optimized to run on the hardware

516
00:24:21,760 --> 00:24:23,480
you already on? 
We're moving from generating 

517
00:24:23,480 --> 00:24:25,200
snapshots to generating 
realities. 

518
00:24:25,240 --> 00:24:27,520
And the engine under the hood is
exactly what we talked about 

519
00:24:27,520 --> 00:24:28,880
today. 
It's going to be a wild. 

520
00:24:29,440 --> 00:24:32,520
It certainly is. 
I hope this deep dive gave you a

521
00:24:32,520 --> 00:24:35,920
solid mental model of how the 
magic trick actually works. 

522
00:24:36,240 --> 00:24:38,520
Thanks for listening and we'll 
see you in the next deep dive.

