1
00:00:00,040 --> 00:00:03,280
Welcome back to the Deep Dive. 
Today we are wading into waters 

2
00:00:03,280 --> 00:00:06,080
that feel, honestly, a little 
dangerous. 

3
00:00:06,080 --> 00:00:07,840
Dangerous. 
That's a strong way to start. 

4
00:00:07,840 --> 00:00:09,640
Well, think about it. 
We're talking about high stakes 

5
00:00:09,640 --> 00:00:11,840
medicine. 
We're talking about diagnosing 

6
00:00:11,840 --> 00:00:17,560
rare legal diseases, situations 
where accuracy is literally life

7
00:00:17,560 --> 00:00:20,040
or death. 
And the premise of the research 

8
00:00:20,040 --> 00:00:22,600
we're looking at today is that 
the best way to save those lives

9
00:00:22,720 --> 00:00:27,120
is to use, well, fake data. 
But that way it does sound a bit

10
00:00:27,120 --> 00:00:28,760
like malpractice, doesn't it? 
Right. 

11
00:00:29,040 --> 00:00:32,080
I mean, if I'm a patient and I 
hear my doctor is training their

12
00:00:32,080 --> 00:00:35,960
diagnostic AI on patients that 
you know don't even exist, I'm 

13
00:00:35,960 --> 00:00:38,440
getting a second opinion 
immediately. 

14
00:00:38,440 --> 00:00:40,880
And you'd. 
Be right to be skeptical, but 

15
00:00:41,280 --> 00:00:44,760
this is the big butt for today. 
What if I told you that relying 

16
00:00:44,760 --> 00:00:47,960
only on real patient data is 
actually what's making the AI 

17
00:00:47,960 --> 00:00:48,920
fail? 
How so? 

18
00:00:49,080 --> 00:00:53,120
That the real data is, in a way,
mathematically broken when it 

19
00:00:53,120 --> 00:00:55,800
comes to the sickest people. 
That is the paradox we are 

20
00:00:55,800 --> 00:00:58,400
unpacking today. 
This is our Practical AI Digest 

21
00:00:58,400 --> 00:01:00,200
edition. 
So for all of you out there who 

22
00:01:00,200 --> 00:01:02,040
are building these systems, this
one's for you. 

23
00:01:02,360 --> 00:01:06,040
We are looking at a stack of 
research, primarily a new 

24
00:01:06,280 --> 00:01:09,960
technical paper from Scientific 
Reports that was validated on 

25
00:01:11,360 --> 00:01:14,840
COVID-19, kidney disease and 
Dengo datasets. 

26
00:01:14,880 --> 00:01:16,560
A pretty serious lineup of 
diseases. 

27
00:01:16,800 --> 00:01:18,720
Exactly. 
And we're also pulling in some 

28
00:01:18,720 --> 00:01:22,400
industry specs from Blue Gen. 
AI and an explainer from IBM 

29
00:01:22,400 --> 00:01:25,160
Technology. 
But we aren't just talking 

30
00:01:25,160 --> 00:01:27,840
theory here. 
We are breaking down a specific 

31
00:01:27,920 --> 00:01:30,200
buildable framework. 
That's the key. 

32
00:01:30,200 --> 00:01:31,720
We're looking at a hybrid 
pipeline. 

33
00:01:31,720 --> 00:01:35,080
We're going to talk about 
combining classical statistical 

34
00:01:35,080 --> 00:01:38,280
tricks with some some really 
heavy hitting deep learning. 

35
00:01:38,400 --> 00:01:41,400
Specifically. 
Specifically Deep CT, Jan and 

36
00:01:41,400 --> 00:01:43,480
something called residual 
networks or resnets. 

37
00:01:43,480 --> 00:01:46,680
And just to really set the stage
for you builders listening, by 

38
00:01:46,680 --> 00:01:48,720
the end of this, we want you to 
walk away with the recipe. 

39
00:01:49,080 --> 00:01:51,920
We're going to cover the 
architecture, the specific hyper

40
00:01:51,920 --> 00:01:53,680
parameters. 
I'm talking learning rates, 

41
00:01:53,760 --> 00:01:57,200
batch sizes, the works, and how 
to validate it so you don't end 

42
00:01:57,200 --> 00:01:59,480
up, you know, just hallucinating
patient data. 

43
00:01:59,520 --> 00:02:02,400
It's basically a master class in
handling the imbalance data set 

44
00:02:02,400 --> 00:02:04,080
problem. 
So let's start there. 

45
00:02:04,240 --> 00:02:09,280
The intuition You mentioned that
real data is broken or biased. 

46
00:02:09,600 --> 00:02:11,640
I think most people assume data 
is neutral. 

47
00:02:11,960 --> 00:02:13,640
It's just numbers in a 
spreadsheet. 

48
00:02:14,480 --> 00:02:17,200
How can a spreadsheet have a 
bias against sick people? 

49
00:02:17,320 --> 00:02:19,200
It's not a bias in the human 
sense. 

50
00:02:19,480 --> 00:02:21,560
It comes down to the math of 
optimization. 

51
00:02:21,800 --> 00:02:25,000
We call it majority class bias. 
OK, think about a typical 

52
00:02:25,000 --> 00:02:28,160
hospital database. 
Let's say you pull records for, 

53
00:02:28,320 --> 00:02:32,120
I don't know, 1000 patients, OK.
1000 rows in my CSE file. 

54
00:02:32,120 --> 00:02:34,600
I've got it open. 
Now, how many of those thousand 

55
00:02:34,600 --> 00:02:38,520
people actually have a rare, 
specific form of kidney disease?

56
00:02:38,520 --> 00:02:39,680
Very. 
Few I'd imagine. 

57
00:02:39,680 --> 00:02:41,680
Maybe 50, maybe even fewer. 
Right. 

58
00:02:41,800 --> 00:02:44,640
The vast majority are there for 
check UPS or broken arms or the 

59
00:02:44,640 --> 00:02:46,600
flu. 
They don't have this specific 

60
00:02:46,600 --> 00:02:48,000
critical condition. 
Exactly. 

61
00:02:48,000 --> 00:02:52,720
So you have a 95 to 5 split, 950
healthy people, 50 sick people. 

62
00:02:53,440 --> 00:02:56,680
Now you feed this to a standard 
machine learning classifier, say

63
00:02:57,160 --> 00:02:59,520
a random forest or a basic 
neural net. 

64
00:03:00,240 --> 00:03:03,280
The models only goal is to 
maximize its accuracy. 

65
00:03:03,360 --> 00:03:05,120
It just wants to get the highest
score possible. 

66
00:03:05,120 --> 00:03:06,760
That's it. 
And that's where the trap is. 

67
00:03:06,960 --> 00:03:11,000
The model quickly realizes, wait
a minute, if I just predict 

68
00:03:11,360 --> 00:03:16,320
healthy for every single person,
I'm right 95% of the time. 

69
00:03:16,400 --> 00:03:19,920
It's the lazy student strategy. 
It crams for the easiest part of

70
00:03:19,920 --> 00:03:22,520
the test and aces it, but it 
doesn't learn anything. 

71
00:03:22,520 --> 00:03:25,960
Precisely. 
In a computer science class, 95%

72
00:03:25,960 --> 00:03:27,640
accuracy is amazing. 
You get an A. 

73
00:03:28,120 --> 00:03:31,560
But in an oncology ward, that 
model is a catastrophe. 

74
00:03:31,560 --> 00:03:35,400
It's worse than useless. 
It has 0% sensitivity. 

75
00:03:35,520 --> 00:03:37,960
It missed every single cancer 
case. 

76
00:03:38,360 --> 00:03:42,400
It maximized its score by 
completely ignoring the minority

77
00:03:42,400 --> 00:03:45,200
class. 
So the data itself isn't 

78
00:03:45,240 --> 00:03:48,400
prejudiced in a way we'd 
recognize, but the distribution 

79
00:03:48,400 --> 00:03:51,760
of the data forces the algorithm
to ignore the edge cases. 

80
00:03:51,760 --> 00:03:54,120
And in medicine, the edge cases 
are usually the ones that need 

81
00:03:54,120 --> 00:03:55,960
saving. 
This is why we need synthetic 

82
00:03:55,960 --> 00:03:58,080
data. 
We have to artificially inflate 

83
00:03:58,080 --> 00:04:00,520
that minority class so the model
can't just brush it aside 

84
00:04:00,520 --> 00:04:02,640
anymore. 
The IBM source had a really good

85
00:04:02,640 --> 00:04:04,720
analogy for this. 
I thought they used football. 

86
00:04:04,960 --> 00:04:07,360
The Premier League one, right? 
Yeah, that's a great way to 

87
00:04:07,360 --> 00:04:08,720
think about. 
It walk us through it. 

88
00:04:08,960 --> 00:04:12,360
OK, so if you trained an AI on 
the last 20 or 30 years of 

89
00:04:12,360 --> 00:04:15,320
English football and and asked 
it to predict the next winner, 

90
00:04:15,400 --> 00:04:18,200
what would it say? 
Manchester City, Manchester 

91
00:04:18,200 --> 00:04:20,920
United, Chelsea. 
One of the big rich clubs, 

92
00:04:21,000 --> 00:04:23,240
right? 
It creates a simple rule in its 

93
00:04:23,240 --> 00:04:25,400
head. 
Rich big teams win. 

94
00:04:25,960 --> 00:04:27,920
It's seen that pattern over and 
over. 

95
00:04:28,160 --> 00:04:31,160
It's the majority class. 
And then 2016 happens. 

96
00:04:31,320 --> 00:04:32,800
And Leicester City wins the 
whole thing. 

97
00:04:32,800 --> 00:04:36,080
The ultimate outlier, a total 
statistical anomaly. 

98
00:04:36,160 --> 00:04:38,800
Exactly. 
If your model is only looking at

99
00:04:38,800 --> 00:04:42,680
the majority history, it says 
that event is impossible, but it

100
00:04:42,680 --> 00:04:45,200
happened. 
Synthetic data lets us simulate 

101
00:04:45,200 --> 00:04:47,240
thousands of Leicester City 
seasons. 

102
00:04:47,480 --> 00:04:50,960
It lets us create these what if 
scenarios that force the model 

103
00:04:51,080 --> 00:04:53,040
to learn the characteristics of 
the underdog. 

104
00:04:53,040 --> 00:04:55,120
Or, in our case, the 
characteristics of the rare 

105
00:04:55,120 --> 00:04:56,880
disease. 
That's the perfect parallel. 

106
00:04:57,200 --> 00:05:00,200
We aren't just copying data, we 
are creating new examples that 

107
00:05:00,200 --> 00:05:03,520
follow the same underlying rules
of physics and biology, but they

108
00:05:03,520 --> 00:05:06,040
fill in the gaps where a real 
data is thin. 

109
00:05:06,400 --> 00:05:09,480
So it's crucial to understand we
aren't just typing in random 

110
00:05:09,480 --> 00:05:11,160
numbers. 
Oh, absolutely not. 

111
00:05:11,160 --> 00:05:13,720
That would be useless. 
We're learning the statistical 

112
00:05:13,720 --> 00:05:16,760
distribution, the hidden shape 
of the data, and then we're 

113
00:05:16,760 --> 00:05:20,600
sampling from that shape to 
create new plausible points. 

114
00:05:20,680 --> 00:05:23,520
Now before we get to the how, I 
want to clarify the use case 

115
00:05:23,520 --> 00:05:26,680
because the blue Gen. 
AI material mentions 2 main 

116
00:05:26,680 --> 00:05:28,440
reasons for using synthetic 
data. 

117
00:05:28,680 --> 00:05:31,720
One is privacy, which is 
massive. 

118
00:05:31,720 --> 00:05:32,360
Huge. 
Yeah. 

119
00:05:32,720 --> 00:05:36,080
GDPRAPA, right? 
If the patient is completely 

120
00:05:36,080 --> 00:05:38,680
fake, you can't get sued for 
leaking their data. 

121
00:05:38,680 --> 00:05:40,560
You can share it with 
researchers freely. 

122
00:05:40,600 --> 00:05:42,920
True, that's what we call the 
replacement use case. 

123
00:05:43,600 --> 00:05:46,400
But for this deep dive, for what
this paper is about, we are 

124
00:05:46,400 --> 00:05:49,200
focused on augmentation. 
OK, what's the difference? 

125
00:05:49,200 --> 00:05:50,680
We aren't throwing away the real
data. 

126
00:05:50,680 --> 00:05:53,920
We are supplementing it. 
We are taking those 50 real sick

127
00:05:53,920 --> 00:05:58,640
patients and generating say 450 
new synthetic sick patients. 

128
00:05:58,920 --> 00:06:03,200
So you turn the 50 into 500. 
And suddenly the lazy student 

129
00:06:03,200 --> 00:06:06,080
can't get an A by ignoring them.
They have to actually study. 

130
00:06:06,080 --> 00:06:07,640
OK, let's get into the 
engineering. 

131
00:06:07,800 --> 00:06:09,760
This hasn't been a straight line
to success. 

132
00:06:09,760 --> 00:06:13,680
There's been an evolution here, 
a couple of tiers of technology.

133
00:06:14,000 --> 00:06:16,360
A pretty clear one actually. 
It's really been a journey from 

134
00:06:16,520 --> 00:06:20,800
simple geometry to complex deep 
learning. 

135
00:06:21,040 --> 00:06:23,200
We can call tier one classical 
resampling. 

136
00:06:23,840 --> 00:06:26,400
This is stuff you could run on a
laptop back in 2010. 

137
00:06:26,400 --> 00:06:29,600
Pretty much the big name here, 
the one everyone learns, is 

138
00:06:29,600 --> 00:06:33,600
SOBT. 
SMOTE Synthetic minority 

139
00:06:33,600 --> 00:06:36,800
oversampling technique. 
It's a staple in data science 

140
00:06:36,800 --> 00:06:39,120
interviews. 
It is, and it's intuitively 

141
00:06:39,120 --> 00:06:41,480
very, very simple. 
Imagine your data points are 

142
00:06:41,480 --> 00:06:44,840
plotted on a 2D graph. 
You have your big cluster of 

143
00:06:44,840 --> 00:06:47,920
healthy patients and then you 
have this little cluster of 50 

144
00:06:47,920 --> 00:06:50,160
sick patients floating off in a 
corner. 

145
00:06:50,200 --> 00:06:53,040
OK, I can picture it. 
SMODI picks one of those sick 

146
00:06:53,040 --> 00:06:56,360
patients at random, then it 
finds its K nearest neighbors, 

147
00:06:56,800 --> 00:06:59,240
so the other sick patients that 
are closest to it on the grass. 

148
00:06:59,240 --> 00:07:00,720
And then it just connects the 
dots. 

149
00:07:00,880 --> 00:07:03,480
Literally it draws a straight 
line between the patient it 

150
00:07:03,480 --> 00:07:06,720
picked and one of its neighbors,
and then it just picks a random 

151
00:07:06,720 --> 00:07:08,840
spot on that line and says new 
patient here. 

152
00:07:08,840 --> 00:07:10,600
So it's just linear 
interpolation. 

153
00:07:10,600 --> 00:07:13,320
It's all it is. 
So if patient A has high blood 

154
00:07:13,320 --> 00:07:17,480
pressure and patient B who's 
nearby has slightly higher blood

155
00:07:17,480 --> 00:07:21,200
pressure, smell may just assumes
A valid patient C exists 

156
00:07:21,200 --> 00:07:22,640
somewhere in the middle. 
Precisely. 

157
00:07:22,880 --> 00:07:27,440
The math is just X nu plus 
Lambda X neighbor a she. 

158
00:07:27,760 --> 00:07:30,480
It's simple, and for some 
problems it works surprisingly 

159
00:07:30,480 --> 00:07:32,400
well, but there's a massive 
flaw. 

160
00:07:32,640 --> 00:07:36,520
Biology isn't a straight. 
Line lingo biology is messy. 

161
00:07:36,520 --> 00:07:39,280
It's nonlinear. 
Complex diseases don't always 

162
00:07:39,280 --> 00:07:41,960
follow a straight, predictable 
path between two patients. 

163
00:07:42,320 --> 00:07:44,680
If you just draw lines, you 
might be creating patients that 

164
00:07:44,680 --> 00:07:47,920
are biologically impossible. 
Or just not very representative 

165
00:07:47,920 --> 00:07:51,760
of the true complexity, right? 
Plus, COT is kind of dumb in how

166
00:07:51,760 --> 00:07:54,280
it samples. 
It treats every minority point 

167
00:07:54,280 --> 00:07:55,960
the same. 
It doesn't care if a data point 

168
00:07:55,960 --> 00:07:58,680
is easy to classify or really 
confusing and sitting right on 

169
00:07:58,680 --> 00:08:00,120
the fence. 
And that's where the upgrade 

170
00:08:00,120 --> 00:08:01,560
came in, right? 
Ada sign. 

171
00:08:01,680 --> 00:08:04,120
ADA sign hole adaptive synthetic
sampling. 

172
00:08:04,200 --> 00:08:05,920
This is like SMOTE scene with a 
bit of a brain. 

173
00:08:06,520 --> 00:08:08,760
It looks at the data set and 
asks where's the model 

174
00:08:08,760 --> 00:08:11,240
struggling the most. 
It hunts for the hard to learn 

175
00:08:11,240 --> 00:08:12,520
examples. 
Yes. 

176
00:08:13,120 --> 00:08:17,520
It looks at each minority point,
each sick patient, and counts 

177
00:08:17,520 --> 00:08:20,200
how many of its neighbors are 
from the majority class, the 

178
00:08:20,200 --> 00:08:23,840
healthy patients. 
So if a sick patient is totally 

179
00:08:23,840 --> 00:08:25,680
surrounded by healthy patients, 
it's. 

180
00:08:25,840 --> 00:08:28,400
In what we call a hostile 
environment, it's an ambiguous 

181
00:08:28,400 --> 00:08:32,360
case. 
Adis YN sees that and assigns a 

182
00:08:32,360 --> 00:08:35,919
higher weight to that point. 
It says we need more troops here

183
00:08:36,159 --> 00:08:38,600
on the frontline. 
So it generates more synthetic 

184
00:08:38,600 --> 00:08:42,120
data specifically along that 
messy decision boundary to try 

185
00:08:42,120 --> 00:08:44,200
and clear up the confusion. 
Exactly. 

186
00:08:44,200 --> 00:08:46,000
It's much smarter, but it's 
still geometric. 

187
00:08:46,000 --> 00:08:48,600
It's still just drawing lines. 
And that brings us to the limit 

188
00:08:48,640 --> 00:08:51,120
of Tier 1. 
It does when you have high 

189
00:08:51,120 --> 00:08:52,800
dimensional data. 
We're talking hundreds of 

190
00:08:52,800 --> 00:08:56,200
columns with gene expressions, 
complex symptom interactions, 

191
00:08:56,200 --> 00:08:58,760
lab results. 
Simple geometry just breaks down

192
00:08:58,760 --> 00:09:00,960
completely. 
You can't draw a line in 500 

193
00:09:00,960 --> 00:09:02,920
dimensional space and expect it 
to mean anything. 

194
00:09:02,920 --> 00:09:04,920
You need something that can 
learn the manifold. 

195
00:09:04,920 --> 00:09:08,280
The curved, twisted shape of the
data's probability distribution.

196
00:09:08,280 --> 00:09:10,840
You need deep. 
Learning and enter the Jans 

197
00:09:11,360 --> 00:09:16,760
generative adversarial networks.
Tier 2, the famous cat and mouse

198
00:09:16,760 --> 00:09:19,280
game. 
Or maybe artist and critic is a 

199
00:09:19,280 --> 00:09:20,560
better way to put it. 
Break it down. 

200
00:09:20,560 --> 00:09:22,640
You have two neural networks 
fighting each other. 

201
00:09:23,240 --> 00:09:26,440
The generator is the artist. 
It tries to create a fake 

202
00:09:26,440 --> 00:09:28,360
sample, like a fake patient 
record. 

203
00:09:28,720 --> 00:09:31,880
The discriminator is the critic.
It looks at the fake record 

204
00:09:31,880 --> 00:09:34,760
alongside a real 1 and has to 
guess which is which. 

205
00:09:34,760 --> 00:09:36,800
And if the discriminator spots 
the fake. 

206
00:09:36,800 --> 00:09:40,240
The generator gets punished, 
gets a bad signal and has to go 

207
00:09:40,240 --> 00:09:43,280
back and try harder next time it
learns from its mistake. 

208
00:09:43,360 --> 00:09:46,280
And this goes on and on for. 
Thousands of rounds until 

209
00:09:46,280 --> 00:09:49,080
eventually the generator is so 
good at creating fakes that the 

210
00:09:49,080 --> 00:09:52,640
discriminator is no better than 
guessing its accuracy is 50%. 

211
00:09:52,720 --> 00:09:55,680
And at that point you have a 
master forger, you have high 

212
00:09:55,680 --> 00:10:00,160
fidelity synthetic data. 
That's the theory, but here's 

213
00:10:00,160 --> 00:10:01,960
the big snag. 
Standard jams are terrible at 

214
00:10:01,960 --> 00:10:03,200
sreadsheets. 
Why? 

215
00:10:03,200 --> 00:10:06,600
I mean, they can generate 
photorealistic faces of people 

216
00:10:06,600 --> 00:10:09,360
who don't exist. 
Why is a spreadsheet of patient 

217
00:10:09,360 --> 00:10:11,920
data harder than a human face? 
It's a great question. 

218
00:10:12,040 --> 00:10:15,960
It's because of the data types. 
Pixels in an image are all more 

219
00:10:15,960 --> 00:10:18,320
or less the same kind of thing. 
They're continuous values from 

220
00:10:18,320 --> 00:10:21,080
zero to 255. 
They flow into each other 

221
00:10:21,080 --> 00:10:23,640
smoothly. 
But a spreadsheet, it's a total 

222
00:10:23,640 --> 00:10:25,480
mix. 
You have a column for blood 

223
00:10:25,480 --> 00:10:27,240
pressure, which is a continuous 
number. 

224
00:10:27,560 --> 00:10:30,480
From right next to it is gender,
which is a discrete category. 

225
00:10:30,840 --> 00:10:32,560
Then blood type another 
category. 

226
00:10:33,360 --> 00:10:36,400
This mix of discrete and 
continuous data completely 

227
00:10:36,400 --> 00:10:39,520
confuses standard Jans. 
And the distribution of those 

228
00:10:39,520 --> 00:10:41,320
numbers is usually really weird,
right? 

229
00:10:41,320 --> 00:10:44,320
Like blood pressure isn't just 
one simple bell curve, it might 

230
00:10:44,320 --> 00:10:46,960
have a couple of different peaks
for different patient groups. 

231
00:10:46,960 --> 00:10:49,080
Exactly. 
It's multimodal, yeah, and 

232
00:10:49,080 --> 00:10:52,120
standard Jans suffer from 
something called mode collapse. 

233
00:10:52,240 --> 00:10:54,880
Mode collapse. 
They find one type of easy to 

234
00:10:54,880 --> 00:10:58,280
fake patient one mode in the 
data and they just keep 

235
00:10:58,280 --> 00:11:00,440
generating that same kind of 
patient over and over again. 

236
00:11:00,920 --> 00:11:03,880
They ignore all the variety and 
especially the rare cases. 

237
00:11:03,880 --> 00:11:06,360
Which defeats the entire purpose
of what we're trying to do. 

238
00:11:06,360 --> 00:11:09,360
We need the rare stuff. 
That is precisely why the 

239
00:11:09,360 --> 00:11:12,920
Scientific Reports paper didn't 
just use a vanilla Jan, they 

240
00:11:12,920 --> 00:11:14,960
used something called Deep CT 
Jan. 

241
00:11:15,040 --> 00:11:16,560
OK, let's breakdown that 
acronym. 

242
00:11:16,920 --> 00:11:20,720
CT Jan. 
Right, so the T is for tabular 

243
00:11:20,800 --> 00:11:22,280
because it's designed for 
spreadsheets. 

244
00:11:22,760 --> 00:11:26,360
The C is the real magic here. 
It stands for conditional. 

245
00:11:26,440 --> 00:11:29,000
Conditional. 
And that lets us explicitly tell

246
00:11:29,000 --> 00:11:31,520
the generator what we want. 
We can give it a condition, we 

247
00:11:31,520 --> 00:11:34,000
can say I want you to generate a
sample, but only for the 

248
00:11:34,000 --> 00:11:36,240
minority class. 
So instead of just hoping it 

249
00:11:36,240 --> 00:11:39,760
stumbles upon making a sick 
patient, you can actually force 

250
00:11:39,760 --> 00:11:42,200
its hand. 
You condition the input, you say

251
00:11:42,200 --> 00:11:46,480
give me a plausible breakdown of
a patient with dengue fever, and

252
00:11:46,480 --> 00:11:50,120
this forces the model to learn 
the specific complex 

253
00:11:50,120 --> 00:11:53,120
distribution of that class, 
which is what prevents mode 

254
00:11:53,120 --> 00:11:55,320
collapse. 
So that's the engine, the deep 

255
00:11:55,320 --> 00:11:58,760
C2 Gan, but the paper we're 
discussing today proposes a 

256
00:11:58,760 --> 00:12:01,360
hybrid framework. 
They didn't just turn on the Jan

257
00:12:01,360 --> 00:12:03,320
walkway, they built a whole 
pipeline around it. 

258
00:12:03,400 --> 00:12:06,360
This is the real how to section 
for the engineers listening. 

259
00:12:06,400 --> 00:12:09,080
If you're building this, you 
need to follow these three steps

260
00:12:09,080 --> 00:12:11,560
they laid out. 
OK, Step 1, preprocessing. 

261
00:12:11,800 --> 00:12:15,680
This is standard data hygiene. 
You can't train a model on empty

262
00:12:15,680 --> 00:12:19,640
cells for any missing numbers 
that used mean imputation. 

263
00:12:20,160 --> 00:12:23,360
For missing categories they used
mode imputation. 

264
00:12:23,800 --> 00:12:26,680
So fill in the blanks with the 
average or the most common 

265
00:12:26,680 --> 00:12:29,360
value. 
Nothing fancy, but essential. 

266
00:12:29,360 --> 00:12:31,000
Absolutely essential garbage in 
garbage. 

267
00:12:31,000 --> 00:12:34,360
Out Step 2 is interesting. 
They call it initial balancing. 

268
00:12:35,320 --> 00:12:38,800
This is a really clever move. 
Before they even touched the big

269
00:12:38,800 --> 00:12:42,000
expensive deep learning model, 
they used the classical methods 

270
00:12:42,000 --> 00:12:45,680
we just talked about, Seals, 
matter, SMOTE to create a 

271
00:12:45,680 --> 00:12:48,760
roughly balanced data set. 
Wait, so they use the old tech 

272
00:12:48,920 --> 00:12:50,720
to prep the data for the new 
tech? 

273
00:12:50,720 --> 00:12:52,720
Why do that? 
Think of it as priming the 

274
00:12:52,720 --> 00:12:56,200
canvas before you paint. 
If you feed a drastically 

275
00:12:56,200 --> 00:13:00,560
imbalanced data set like 95 to 5
into again, even a conditional 

276
00:13:00,560 --> 00:13:02,520
Jan might struggle to converge 
at the beginning. 

277
00:13:02,520 --> 00:13:05,200
It's a hard problem. 
OK by balancing it roughly with 

278
00:13:05,240 --> 00:13:07,800
eight is 1 first maybe to like a
6040 split? 

279
00:13:08,440 --> 00:13:11,520
They gave the deep CT Yan a much
easier starting point, a head 

280
00:13:11,520 --> 00:13:12,480
start. 
That makes sense. 

281
00:13:12,480 --> 00:13:14,600
And then Step 3, advanced 
augmentation. 

282
00:13:14,880 --> 00:13:17,520
Right. 
They feed this primed, roughly 

283
00:13:17,520 --> 00:13:21,960
balanced data into the deep CT 
Jan but they modified the 

284
00:13:21,960 --> 00:13:24,280
architecture. 
This is the core innovation. 

285
00:13:24,280 --> 00:13:26,480
I think they added a resnet 
backbone. 

286
00:13:26,480 --> 00:13:28,560
This is the most technical part 
of the paper and probably the 

287
00:13:28,560 --> 00:13:31,160
most important. 
Resnet stands for residual 

288
00:13:31,160 --> 00:13:33,040
networks. 
We usually hear about this in 

289
00:13:33,040 --> 00:13:34,880
the context of image 
recognition. 

290
00:13:34,880 --> 00:13:38,840
Resnet 50, Resnet 101. 
So why are we jamming an image 

291
00:13:38,840 --> 00:13:43,560
recognition architecture into a 
generator for tabular data? 

292
00:13:43,560 --> 00:13:46,000
It seems like a weird. 
Fit It's all about solving a 

293
00:13:46,000 --> 00:13:48,800
classic deep learning problem 
called the vanishing gradient. 

294
00:13:48,800 --> 00:13:51,400
OK, Eli, five time. 
Give me the simple version. 

295
00:13:51,400 --> 00:13:54,480
What is the vanishing gradient? 
Imagine you're playing a game of

296
00:13:54,480 --> 00:13:57,160
telephone, but the line of 
people is a mile long. 

297
00:13:57,640 --> 00:13:59,960
That's a deep neural network 
with many layers. 

298
00:13:59,960 --> 00:14:01,800
OK, you whisper a message at the
start. 

299
00:14:01,800 --> 00:14:04,040
That's your input data. 
By the time it gets to the 

300
00:14:04,040 --> 00:14:07,520
person at the very end, the 
output layer, the message is so 

301
00:14:07,520 --> 00:14:09,880
distorted and quiet and 
basically gibberish. 

302
00:14:10,480 --> 00:14:12,960
The person at the end has no 
idea what the original signal 

303
00:14:12,960 --> 00:14:14,640
was. 
The signal just dies out as it 

304
00:14:14,640 --> 00:14:16,480
passes through all the layers, 
correct? 

305
00:14:17,000 --> 00:14:19,800
And in a neural network, that 
signal is the gradient. 

306
00:14:19,800 --> 00:14:22,200
It's the error signal that tells
the network how to update its 

307
00:14:22,200 --> 00:14:25,080
weights, how to learn. 
If that signal dies, the model 

308
00:14:25,080 --> 00:14:26,960
can't learn. 
It stops updating. 

309
00:14:27,000 --> 00:14:31,400
It gets stuck. 
And resnet fixes this with skip 

310
00:14:31,400 --> 00:14:32,880
connections. 
The shortcut. 

311
00:14:33,080 --> 00:14:34,280
That's the perfect way to think 
of it. 

312
00:14:35,000 --> 00:14:38,120
Instead of passing the message 
only to the very next person in 

313
00:14:38,120 --> 00:14:41,040
line, you also get to shout it 
to the person two or three spots

314
00:14:41,040 --> 00:14:43,200
down the line. 
So the original information gets

315
00:14:43,200 --> 00:14:45,720
to jump ahead. 
Mathematically, instead of a 

316
00:14:45,720 --> 00:14:50,360
layer trying to learn the full 
complex function FX, it only has

317
00:14:50,360 --> 00:14:53,240
to learn the residual or the 
difference FX plus X. 

318
00:14:53,520 --> 00:14:55,640
That plus X is the skip 
connection. 

319
00:14:55,880 --> 00:14:58,560
It carries the original 
unmodified input forward. 

320
00:14:58,560 --> 00:15:01,240
So it never forgets where it 
started, it always has the 

321
00:15:01,240 --> 00:15:02,840
original context. 
Precisely. 

322
00:15:02,880 --> 00:15:05,840
And for tabular data this is 
absolutely be crucial because 

323
00:15:05,840 --> 00:15:08,920
the correlations can be very 
subtle and exist at different 

324
00:15:08,920 --> 00:15:10,320
levels. 
You have simple patterns like 

325
00:15:10,320 --> 00:15:13,440
older people have higher risk, 
but also very complex high level

326
00:15:13,440 --> 00:15:16,480
patterns like this specific 
enzyme level combined with this 

327
00:15:16,480 --> 00:15:19,200
blood type and this genetic 
marker equals danger. 

328
00:15:19,360 --> 00:15:22,120
And the skip connections help 
the model capture both. 

329
00:15:22,240 --> 00:15:23,760
Exactly. 
It allows the generator to 

330
00:15:23,760 --> 00:15:27,320
capture both the simple low 
level features and the complex 

331
00:15:27,320 --> 00:15:30,880
high level dependencies without 
getting lost in the noise of a 

332
00:15:30,880 --> 00:15:33,480
very deep network. 
It just stabilizes the whole 

333
00:15:33,480 --> 00:15:35,960
training process. 
OK, I want to get granular on 

334
00:15:35,960 --> 00:15:38,920
the hyper parameters here. 
The paper actually lists the 

335
00:15:38,920 --> 00:15:41,320
settings they used. 
So if I'm a data scientist 

336
00:15:41,320 --> 00:15:44,080
sitting at my GPU right now 
trying to reproduce this for a 

337
00:15:44,080 --> 00:15:46,640
project, what am I typing into 
my code? 

338
00:15:46,640 --> 00:15:49,440
Get your notebooks out. 
Yeah, for the core deep CT Jan 

339
00:15:49,440 --> 00:15:54,080
training, they kept the learning
rate very, very low .0002. 

340
00:15:54,160 --> 00:15:56,680
That is tiny. 
The default in most frameworks 

341
00:15:56,680 --> 00:16:00,000
is what .0001? 
Why go so slow? 

342
00:16:00,000 --> 00:16:03,400
Stability. 
GNS are notoriously unstable. 

343
00:16:03,640 --> 00:16:06,800
It's a delicate tug of war. 
If one side, the generator or 

344
00:16:06,800 --> 00:16:09,920
the discriminator, learns too 
fast, it overpowers the other 

345
00:16:09,920 --> 00:16:11,120
and the whole system falls 
apart. 

346
00:16:11,120 --> 00:16:13,160
The rope snaps. 
The model just oscillates and 

347
00:16:13,160 --> 00:16:14,600
never finds a good solution. 
Right. 

348
00:16:14,600 --> 00:16:16,240
So you have to creep up on the 
solution. 

349
00:16:16,360 --> 00:16:18,240
Slow and steady adversarial 
training. 

350
00:16:18,360 --> 00:16:19,880
Makes sense. 
What about batch size? 

351
00:16:20,080 --> 00:16:22,640
128 That's a pretty classic 
Goldilocks number. 

352
00:16:22,920 --> 00:16:25,440
It's small enough that it fits 
in memory and provides a bit of 

353
00:16:25,440 --> 00:16:28,320
stochastic noise, which helps 
with generalization, but it's 

354
00:16:28,320 --> 00:16:31,680
also large enough to give you a 
stable, reliable estimate of the

355
00:16:31,680 --> 00:16:34,000
gradient. 
And the embedding dimension for 

356
00:16:34,000 --> 00:16:37,760
a categorical features. 
We use between 128 and 256 and 

357
00:16:37,760 --> 00:16:40,080
they said it depends on the 
complexity of the data set. 

358
00:16:40,800 --> 00:16:44,520
This is for converting those 
text based categories like city 

359
00:16:44,520 --> 00:16:48,480
or hospital ID into dense 
numerical vectors that the math 

360
00:16:48,480 --> 00:16:49,840
can actually work with. 
Got it. 

361
00:16:49,920 --> 00:16:53,160
Now for the Resonant backbone 
specifically, what did they use 

362
00:16:53,160 --> 00:16:54,960
there? 
They used a Resonant 18 

363
00:16:54,960 --> 00:16:56,600
architecture. 
That's one of the lighter, 

364
00:16:56,600 --> 00:16:59,800
faster versions of Resonant. 
But the real secret sauce I 

365
00:16:59,800 --> 00:17:01,600
think was the regularization 
they applied. 

366
00:17:02,040 --> 00:17:03,920
They use dropout at a rate of 
.3. 

367
00:17:03,920 --> 00:17:07,880
Which means they are randomly 
turning off 30% of the neurons 

368
00:17:07,920 --> 00:17:10,880
during each training step. 
Yes, exactly. 

369
00:17:11,200 --> 00:17:13,319
And they also use weight decay 
at 1 E 5. 

370
00:17:14,119 --> 00:17:17,240
Both of these are absolutely 
crucial because remember, our 

371
00:17:17,240 --> 00:17:21,480
starting data set of sick 
patients is tiny, maybe 50 or 60

372
00:17:21,480 --> 00:17:23,119
people. 
Deep learning models are 

373
00:17:23,119 --> 00:17:25,040
effectively massive memory 
machines. 

374
00:17:25,440 --> 00:17:28,760
If you let them, they will just 
memorize those 50 patients 

375
00:17:28,760 --> 00:17:30,080
perfectly. 
That's overfitting. 

376
00:17:30,160 --> 00:17:33,840
Extreme overfitting. 
The model becomes a parrot, not 

377
00:17:33,840 --> 00:17:37,440
a thinker. 
Dropout forces it to be robust. 

378
00:17:37,720 --> 00:17:40,200
You can't rely on any single 
neuron, because that neuron 

379
00:17:40,200 --> 00:17:41,760
might be switched off at any 
time. 

380
00:17:42,040 --> 00:17:45,080
It forces the model to learn the
underlying concept of the 

381
00:17:45,080 --> 00:17:48,200
disease, not just the specific 
patient records in the training 

382
00:17:48,200 --> 00:17:49,160
set. 
OK. 

383
00:17:49,160 --> 00:17:51,880
So we've built the generator, 
we've primed it with ATIS YN, 

384
00:17:51,880 --> 00:17:54,360
we've stabilized it with Resnet 
and we've tuned it with these 

385
00:17:54,360 --> 00:17:56,760
very specific slow hyper 
parameters. 

386
00:17:57,000 --> 00:17:58,800
We generate our new synthetic 
data. 

387
00:17:58,800 --> 00:18:01,240
Now we have a nice big balanced 
data set. 

388
00:18:01,240 --> 00:18:03,920
We still need a classifier to 
actually diagnose the patients. 

389
00:18:04,480 --> 00:18:06,840
Did they just use something 
simple like random forest? 

390
00:18:06,840 --> 00:18:09,000
They compared a few. 
They looked at Random Forest, 

391
00:18:09,000 --> 00:18:12,200
XG, Boost, KNN, but the hands 
down winner, the one they 

392
00:18:12,200 --> 00:18:15,400
focused on was Tabnet. 
Tabnet, another deep learning 

393
00:18:15,400 --> 00:18:17,200
model. 
Why use a deep learning 

394
00:18:17,200 --> 00:18:20,960
sledgehammer when a screwdriver 
like Random Forest usually works

395
00:18:20,960 --> 00:18:23,920
just fine for tabular data? 
Because Tabnet is a deep 

396
00:18:23,920 --> 00:18:26,920
learning model that's 
specifically designed to beat 

397
00:18:26,920 --> 00:18:30,560
decision trees at their own 
game, it has an architecture 

398
00:18:30,560 --> 00:18:33,880
that really tries to mimic the 
way a decision tree, or even a 

399
00:18:33,880 --> 00:18:37,280
human doctor thinks. 
It uses a mechanism called 

400
00:18:37,280 --> 00:18:40,640
sequential attention. 
Sequential attention. 

401
00:18:40,640 --> 00:18:42,440
What does that mean in this 
context? 

402
00:18:43,000 --> 00:18:44,720
Think about how you diagnose a 
car problem. 

403
00:18:45,160 --> 00:18:46,840
You don't check everything all 
at once. 

404
00:18:46,880 --> 00:18:49,160
You pop the hood, You check the 
battery first. 

405
00:18:49,920 --> 00:18:51,080
Is it dead? 
No. 

406
00:18:51,480 --> 00:18:53,680
OK, now check the starter. 
Is that making a clicking sound?

407
00:18:53,720 --> 00:18:56,040
It's a step by step process of 
elimination. 

408
00:18:56,040 --> 00:18:57,640
Exactly. 
Tabinet does this. 

409
00:18:58,080 --> 00:19:01,680
It uses a learnable mask at each
step to focus on a specific 

410
00:19:01,680 --> 00:19:03,880
subset of features. 
For the first step, it might 

411
00:19:03,880 --> 00:19:05,400
only look at age and blood 
pressure. 

412
00:19:05,800 --> 00:19:08,880
Then based on that, for the 
second step, it decides to look 

413
00:19:08,880 --> 00:19:11,920
at hemoglobin and leukocytes. 
So it learns not just what's 

414
00:19:11,920 --> 00:19:14,520
important, but in what order to 
look at things. 

415
00:19:14,520 --> 00:19:15,920
And it learns what to ignore 
this. 

416
00:19:15,920 --> 00:19:19,160
The other key part uses a very 
specific activation function 

417
00:19:19,160 --> 00:19:20,840
called antmax. 
Antmax. 

418
00:19:20,840 --> 00:19:23,680
I've heard of softmax. 
Softmax is what's typically 

419
00:19:23,680 --> 00:19:25,840
used. 
It gives a probability score to 

420
00:19:25,840 --> 00:19:29,320
everything, but nothing is ever 
truly 0A feature might be 

421
00:19:29,440 --> 00:19:35,600
.00001% important at Max, allows
for true sparsity. 

422
00:19:35,960 --> 00:19:38,400
It can actually say you know 
what, these 10 columns in the 

423
00:19:38,400 --> 00:19:40,280
spreadsheet. 
They're completely irrelevant 

424
00:19:40,280 --> 00:19:42,560
for this decision probability. 
Zero. 

425
00:19:42,560 --> 00:19:44,960
It effectively deletes the noise
from the calculation. 

426
00:19:44,960 --> 00:19:46,960
Which is huge. 
It makes the model much more 

427
00:19:46,960 --> 00:19:49,360
efficient and critically, more 
interpretable. 

428
00:19:49,360 --> 00:19:51,840
It tells you exactly which 
features it chose to ignore. 

429
00:19:52,240 --> 00:19:55,240
So that's the full stack, a two 
element baiting feeding into a 

430
00:19:55,240 --> 00:19:59,280
deep CTGN with a resnet backbone
which generates data to train a 

431
00:19:59,280 --> 00:20:01,440
tab net classifier. 
That's the recipe. 

432
00:20:01,760 --> 00:20:04,040
Now the big question, does it 
actually work? 

433
00:20:04,320 --> 00:20:06,440
Or is this just a cool 
engineering science project that

434
00:20:06,440 --> 00:20:08,760
looks good on paper? 
The validation results are 

435
00:20:08,760 --> 00:20:12,000
honestly startling, but first we
have to define how they tested 

436
00:20:12,000 --> 00:20:13,800
it. 
This is a hell I will die on. 

437
00:20:13,960 --> 00:20:16,800
You have to use TSTR. 
Trade on synthetic, test on 

438
00:20:16,800 --> 00:20:18,600
real. 
It is the golden rule of 

439
00:20:18,600 --> 00:20:22,560
synthetic data validation. 
You train your model only on the

440
00:20:22,560 --> 00:20:26,600
augmented or fully synthetic 
data, but you must must must 

441
00:20:26,600 --> 00:20:30,480
test it on a holdout set of real
flesh and blood patients that 

442
00:20:30,480 --> 00:20:33,280
the model has never seen before.
Because if you test on your own 

443
00:20:33,280 --> 00:20:35,680
synthetic data, you're just 
grading your own homework. 

444
00:20:35,680 --> 00:20:37,960
Exactly. 
It's a meaningless self 

445
00:20:37,960 --> 00:20:41,920
congratulatory loop, but if a 
model that was trained entirely 

446
00:20:41,920 --> 00:20:46,480
on fake data can accurately 
diagnose a real patient it has 

447
00:20:46,480 --> 00:20:48,320
never met. 
And you know, the synthetic data

448
00:20:48,320 --> 00:20:51,560
successfully captured the 
fundamental underlying rules of 

449
00:20:51,560 --> 00:20:53,040
the disease. 
That's the proof. 

450
00:20:53,080 --> 00:20:55,000
So what were the scores? 
What did they find? 

451
00:20:55,000 --> 00:20:56,640
OK. 
For the chronic kidney disease 

452
00:20:56,640 --> 00:21:01,000
data set using the full hybrid 
framework, they achieved 99.4% 

453
00:21:01,000 --> 00:21:03,520
accuracy. 
Wow, that is that's basically 

454
00:21:03,520 --> 00:21:04,720
perfect. 
What about the others? 

455
00:21:04,720 --> 00:21:10,200
For COVID-19 they hit 99.2% and 
for Ding Fever 99.5%. 

456
00:21:10,200 --> 00:21:12,040
That's incredible. 
How did the more traditional 

457
00:21:12,040 --> 00:21:13,600
models compare on the same 
tasks? 

458
00:21:13,600 --> 00:21:16,160
They were good, don't get me 
wrong, but Tabnet consistently 

459
00:21:16,320 --> 00:21:19,600
beat Random Forest and XG boost 
across the board on these 

460
00:21:19,600 --> 00:21:22,280
augmented data sets. 
But the most interesting failure

461
00:21:22,280 --> 00:21:25,280
to me was KNN. 
Key nearest neighbors, the one 

462
00:21:25,280 --> 00:21:27,760
that works based on distance. 
Right, it's a very simple 

463
00:21:27,760 --> 00:21:30,920
algorithm and in some of their 
experiments KNN actually 

464
00:21:30,920 --> 00:21:33,400
performed worse when they added 
the synthetic data. 

465
00:21:33,440 --> 00:21:36,280
How is that possible? 
How does more data make a model 

466
00:21:36,280 --> 00:21:39,440
worse? 
Because KNN relies purely on 

467
00:21:39,440 --> 00:21:43,480
geometric distance and synthetic
data, even very high quality 

468
00:21:43,480 --> 00:21:46,960
data like this introduces a 
tiny, tiny bit of noise. 

469
00:21:47,320 --> 00:21:50,760
It's not perfect, it fuzzes the 
decision boundaries just 

470
00:21:50,760 --> 00:21:53,200
slightly. 
And KNN is super sensitive to 

471
00:21:53,200 --> 00:21:54,280
that. 
Very sensitive. 

472
00:21:54,440 --> 00:21:57,400
It got confused by these new 
fake neighbors that weren't 

473
00:21:57,400 --> 00:21:59,320
placed perfectly in the 
geometric space. 

474
00:21:59,480 --> 00:22:01,560
It started making mistakes it 
wouldn't have made before. 

475
00:22:01,920 --> 00:22:04,520
That is a huge practical take 
away for anyone listening. 

476
00:22:04,560 --> 00:22:07,360
If you are going to implement 
this kind of synthetic strategy,

477
00:22:07,520 --> 00:22:09,760
you might need to upgrade your 
classifier too. 

478
00:22:10,400 --> 00:22:13,680
You can't just plug this into a 
legacy KNN system and expect 

479
00:22:13,680 --> 00:22:15,520
magic to happen. 
You really can't. 

480
00:22:15,520 --> 00:22:18,160
You need a more robust 
classifier like tab net that can

481
00:22:18,160 --> 00:22:21,640
use an attention mechanism to 
filter out that slight noise and

482
00:22:21,640 --> 00:22:23,840
find the real signal. 
There was one other metric in 

483
00:22:23,840 --> 00:22:26,760
the paper I wanted to ask about,
the similarity score. 

484
00:22:27,040 --> 00:22:30,920
They said they achieved between 
84 and 87% right? 

485
00:22:31,200 --> 00:22:34,960
Which sounds like AB Plus why 
wouldn't we want 100% 

486
00:22:34,960 --> 00:22:37,080
similarity? 
That's a great question because 

487
00:22:37,080 --> 00:22:40,160
it's counterintuitive. 
If your synthetic data is 100% 

488
00:22:40,160 --> 00:22:42,880
similar to your real data, 
you've just cloned it. 

489
00:22:42,880 --> 00:22:44,800
You've perfectly memorized the 
training set. 

490
00:22:45,280 --> 00:22:48,040
Which is a privacy violation and
it's terrible for 

491
00:22:48,040 --> 00:22:49,240
generalization. 
Exactly. 

492
00:22:49,240 --> 00:22:50,520
You haven't learned anything 
new. 

493
00:22:50,680 --> 00:22:54,200
You want the synthetic data to 
be close enough to capture the 

494
00:22:54,200 --> 00:22:57,400
statistical properties the 
physics of the system, but 

495
00:22:57,400 --> 00:23:01,640
different enough to be novel and
to cover new ground. 87% is 

496
00:23:01,640 --> 00:23:04,680
actually a very healthy score. 
Means you're in the sweet spot. 

497
00:23:04,840 --> 00:23:08,040
So the model is incredibly 
accurate, but in medicine, 

498
00:23:08,040 --> 00:23:10,600
accuracy isn't enough. 
I can't just walk into a 

499
00:23:10,600 --> 00:23:12,760
patient's room and say, well, 
the computer says you're sick, 

500
00:23:12,760 --> 00:23:14,800
so bye. 
I mean to know why? 

501
00:23:14,800 --> 00:23:16,360
The black box problem. 
Yes. 

502
00:23:17,200 --> 00:23:19,360
This brings us to the fifth 
section of the paper, 

503
00:23:19,480 --> 00:23:22,720
Interpretability. 
To open up that black box, they 

504
00:23:22,720 --> 00:23:27,200
used SHAP values. 
SHAP Shapley additive 

505
00:23:27,200 --> 00:23:28,760
explanations. 
It comes from game theory, 

506
00:23:28,760 --> 00:23:29,920
doesn't it? 
It does, yeah. 

507
00:23:30,280 --> 00:23:33,400
It basically treats the model's 
prediction like a cooperative 

508
00:23:33,400 --> 00:23:38,160
game where every feature, every 
symptom, every lab result is a 

509
00:23:38,160 --> 00:23:39,640
player on a team. 
OK? 

510
00:23:40,000 --> 00:23:43,680
And it calculates how much each 
individual player contributed to

511
00:23:43,680 --> 00:23:45,880
the final score, to the final 
diagnosis. 

512
00:23:46,160 --> 00:23:48,600
It tells you who the MVP was for
that prediction. 

513
00:23:48,800 --> 00:23:51,000
So it gives us a clinical 
reality check. 

514
00:23:51,080 --> 00:23:55,160
The ultimate reality check If 
your model is 99% accurate, but 

515
00:23:55,160 --> 00:23:57,720
it tells you the most important 
feature for diagnosing cancer is

516
00:23:57,960 --> 00:24:01,000
the patient ID number, you know 
something is fundamentally 

517
00:24:01,000 --> 00:24:02,680
broken. 
You know it's garbage. 

518
00:24:02,800 --> 00:24:05,240
So let's look at what the model 
actually learned for the 

519
00:24:05,240 --> 00:24:07,920
COVID-19 data set. 
What was the top predictor 

520
00:24:07,920 --> 00:24:10,840
according to SHAP? 
The top feature was leukocytes. 

521
00:24:10,840 --> 00:24:12,520
White blood cells. 
Exactly. 

522
00:24:12,680 --> 00:24:15,080
A high white blood cell count 
indicates an infection. 

523
00:24:15,080 --> 00:24:17,000
It's one of the most basic 
tenets of medicine. 

524
00:24:17,000 --> 00:24:19,560
So that's biology one O 1 the 
model learned the right. 

525
00:24:19,560 --> 00:24:22,600
Thing it validates that the 
synthetic data preserved the 

526
00:24:22,600 --> 00:24:26,160
true underlying biological 
signal for kidney disease. 

527
00:24:26,160 --> 00:24:29,760
The top 2 features it flagged 
were hemoglobin and hematocrit. 

528
00:24:30,000 --> 00:24:34,640
OK, so low hemoglobin means 
anemia, and kidneys produce a 

529
00:24:34,640 --> 00:24:38,280
hormone erythropoietin that 
tells the body to make red blood

530
00:24:38,280 --> 00:24:40,400
cells. 
So if your kidneys fail, you 

531
00:24:40,400 --> 00:24:42,200
can't make that hormone and you 
become anemic. 

532
00:24:42,200 --> 00:24:46,240
See, the model figured out the 
entire renal anemia connection 

533
00:24:46,240 --> 00:24:47,840
just by looking at the synthetic
data. 

534
00:24:48,200 --> 00:24:52,000
It learned real medicine. 
But the dengue data set, that 

535
00:24:52,000 --> 00:24:53,880
one had a really interesting 
curveball. 

536
00:24:54,000 --> 00:24:55,440
Oh. 
What was it? 

537
00:24:55,440 --> 00:24:57,400
The top predictor wasn't a blood
test. 

538
00:24:57,720 --> 00:25:01,120
It wasn't a symptom. 
It was the column labeled year 

539
00:25:01,280 --> 00:25:03,320
the. 
Year like the calendar year the 

540
00:25:03,320 --> 00:25:07,400
patient was seen 2018-2019. 
And at first you might think, 

541
00:25:07,400 --> 00:25:09,920
OK, that's a data leak, that's a
bias, that's a mistake. 

542
00:25:10,080 --> 00:25:12,240
But then you should think about 
what dengue fever is. 

543
00:25:12,320 --> 00:25:14,560
It's an outbreak disease. 
It's carried by mosquitoes. 

544
00:25:14,560 --> 00:25:16,760
It's epidemiological, it comes 
in waves. 

545
00:25:16,760 --> 00:25:19,480
It's seasonal encyclical. 
The model learned that if a 

546
00:25:19,480 --> 00:25:22,240
patient walks into the clinic 
during 2019, which is a big 

547
00:25:22,240 --> 00:25:25,400
outbreak year, the data set 
their prior probability of 

548
00:25:25,400 --> 00:25:28,360
having dengue is way, way higher
than if they walked in during 

549
00:25:28,360 --> 00:25:30,320
2018. 
That is fascinating. 

550
00:25:30,320 --> 00:25:34,320
It didn't just learn Physiology 
from the synthetic data, it 

551
00:25:34,320 --> 00:25:37,720
learned epidemiology. 
It learned about public health 

552
00:25:37,720 --> 00:25:39,200
trends. 
And it also found that the 

553
00:25:39,200 --> 00:25:42,280
remarks column, which was this 
unstructured text notes for the 

554
00:25:42,280 --> 00:25:46,200
doctors, was highly predictive. 
Which Tabnet is well suited to 

555
00:25:46,200 --> 00:25:48,440
handle because of the embedding 
layers we talked about. 

556
00:25:48,640 --> 00:25:52,920
It shows that the deep CTGN was 
able to encode these really 

557
00:25:52,920 --> 00:25:58,200
complex nonlinear relationships 
between calendar year, the free 

558
00:25:58,200 --> 00:26:00,920
form text notes and the final 
diagnosis. 

559
00:26:00,920 --> 00:26:03,920
It's incredible, really. 
So we have a working system, 

560
00:26:03,960 --> 00:26:06,160
it's accurate, it's 
interpretable, but I want to 

561
00:26:06,160 --> 00:26:09,560
bring this back down to earth. 
This is the practical AI digest.

562
00:26:09,560 --> 00:26:11,720
There is no free lunch in 
engineering. 

563
00:26:12,080 --> 00:26:14,080
What is the cost? 
What's the trade off for this 

564
00:26:14,080 --> 00:26:17,800
fancy hybrid framework? 
The cost is time and compute and

565
00:26:17,800 --> 00:26:20,040
the paper is very transparent 
about this which is great to 

566
00:26:20,040 --> 00:26:20,880
see. 
They timed it. 

567
00:26:21,040 --> 00:26:24,400
Training a standard Smoke plus a
Tabnet model took them less than

568
00:26:24,400 --> 00:26:25,960
10 minutes. 
So, quick coffee break and 

569
00:26:25,960 --> 00:26:28,280
you're done, right? 
But training the full hybrid 

570
00:26:28,280 --> 00:26:31,920
model, the ADCNN pre balancing, 
the deep CDD GAN with a resnet 

571
00:26:31,920 --> 00:26:34,880
backbone and then the Tabnet 
classifier, that took 42 

572
00:26:34,880 --> 00:26:36,680
minutes. 
So that's more than a four fold 

573
00:26:36,680 --> 00:26:39,440
increase in training time. 
It is now. 

574
00:26:39,480 --> 00:26:42,720
In the grand scheme of things, 
42 minutes isn't long if you're,

575
00:26:42,720 --> 00:26:44,840
say, training ChatGPT from 
scratch. 

576
00:26:45,520 --> 00:26:48,240
But in a clinical production 
environment where you might be 

577
00:26:48,240 --> 00:26:52,640
retraining these models every 
night on new data, that adds up.

578
00:26:52,640 --> 00:26:56,080
So you are paying in GPU hours 
for this extra couple of 

579
00:26:56,080 --> 00:26:59,320
percentage points of accuracy. 
You are, and you have to ask if 

580
00:26:59,320 --> 00:27:02,000
it's worth it. 
For life threatening diseases 

581
00:27:02,000 --> 00:27:05,520
like these, I would argue 
absolutely yes, that extra 1 or 

582
00:27:05,520 --> 00:27:09,120
2% in accuracy translates to 
real people not getting 

583
00:27:09,120 --> 00:27:11,040
misdiagnosed. 
But if you're using this to 

584
00:27:11,040 --> 00:27:15,200
predict, say, customer churn, 
maybe the simpler, faster SOD is

585
00:27:15,200 --> 00:27:16,640
good enough. 
The Blue Gen. 

586
00:27:16,720 --> 00:27:19,400
AI source had some great 
implementation advice here too. 

587
00:27:19,400 --> 00:27:21,920
They talked about the point of 
diminishing returns. 

588
00:27:21,920 --> 00:27:23,880
It's. 
So crucial for the builders out 

589
00:27:23,880 --> 00:27:25,520
there. 
Don't just turn on the generator

590
00:27:25,520 --> 00:27:27,440
and create a billion new rows of
data. 

591
00:27:27,680 --> 00:27:30,280
It doesn't work that way. 
Blue Gen. suggests starting by 

592
00:27:30,280 --> 00:27:34,000
generating 2 to five times your 
original minority data set size.

593
00:27:34,560 --> 00:27:37,640
I have 1000 real rows of the 
rare disease I should aim to 

594
00:27:37,640 --> 00:27:40,760
generate, maybe 2000 to 5000 
synthetic ones. 

595
00:27:40,880 --> 00:27:43,080
Exactly. 
The study showed that if you go 

596
00:27:43,080 --> 00:27:45,840
much beyond that, you aren't 
really adding new information 

597
00:27:45,840 --> 00:27:48,120
anymore. 
You're just adding noise and 

598
00:27:48,120 --> 00:27:49,920
increasing your training time 
for no benefit. 

599
00:27:50,120 --> 00:27:53,240
The accuracy starts to plateau. 
You have to find that sweet 

600
00:27:53,240 --> 00:27:55,240
spot. 
And what about their hybrid 

601
00:27:55,240 --> 00:27:57,880
workflow idea? 
This is just best practice for 

602
00:27:57,880 --> 00:28:00,800
safety and reliability. 
Use the synthetic data for 

603
00:28:00,800 --> 00:28:03,200
training. 
Use it to cover all those weird 

604
00:28:03,200 --> 00:28:05,960
edge cases, the Leicester City 
moments we talked about, but 

605
00:28:05,960 --> 00:28:09,120
never let synthetic data leak 
into your final validation set. 

606
00:28:09,160 --> 00:28:12,320
Your final exam must be on real 
world problems. 

607
00:28:12,320 --> 00:28:15,560
Always, and you have to 
constantly monitor for 

608
00:28:15,560 --> 00:28:18,720
hallucinations. 
If your similarity score starts 

609
00:28:18,720 --> 00:28:22,360
to drop or your TFTR performance
tanks, it might mean your 

610
00:28:22,360 --> 00:28:25,200
generator has drifted and is 
inventing relationships that 

611
00:28:25,200 --> 00:28:28,600
don't exist in reality. 
You need those automated checks 

612
00:28:28,600 --> 00:28:30,720
before you deploy anything to a 
real clinic. 

613
00:28:30,800 --> 00:28:33,880
So to summarize the entire stack
we've uncovered today, we have a

614
00:28:33,880 --> 00:28:37,560
fundamental problem. 
Imbalance causes lazy AI that 

615
00:28:37,560 --> 00:28:41,200
misses the most critical cases. 
We have a solution, the hybrid 

616
00:28:41,200 --> 00:28:43,760
framework. 
We prep the data with a classic 

617
00:28:43,760 --> 00:28:47,120
tool like Adis Wyan. 
We generate new data with a 

618
00:28:47,120 --> 00:28:51,280
powerful conditional deep CT jam
that's been stabilized by Resnet

619
00:28:51,280 --> 00:28:54,840
skip connections, and then we 
classify with a smart, attentive

620
00:28:54,840 --> 00:28:57,720
model like Tabnet. 
And we validate everything with 

621
00:28:57,720 --> 00:29:00,680
TSTR. 
It's a robust, modern and 

622
00:29:01,040 --> 00:29:05,240
frankly very effective pipeline.
As we wrap up, I want to leave 

623
00:29:05,240 --> 00:29:07,320
our listeners with the thought 
that came from the conclusion of

624
00:29:07,320 --> 00:29:09,000
the paper. 
They mentioned real time 

625
00:29:09,000 --> 00:29:10,920
streams. 
This felt like a genuine glimpse

626
00:29:10,920 --> 00:29:12,800
since the near future. 
This is where it gets really 

627
00:29:12,800 --> 00:29:14,760
exciting. 
Beyond just a single static data

628
00:29:14,760 --> 00:29:16,840
set. 
And imagine a synthetic data 

629
00:29:16,840 --> 00:29:19,120
generator that isn't a script 
you run once a year. 

630
00:29:19,640 --> 00:29:21,400
Imagine it's a live running 
service. 

631
00:29:21,400 --> 00:29:24,360
So it's connected directly to 
the hospital's intake feed. 

632
00:29:24,360 --> 00:29:26,440
Exactly. 
Every single night it 

633
00:29:26,440 --> 00:29:29,120
automatically ingests the day's 
new patient cases. 

634
00:29:29,440 --> 00:29:32,080
And if a new strain of a virus 
starts to appear, like we all 

635
00:29:32,080 --> 00:29:34,960
saw happen with the COVID 
variance, or if the common 

636
00:29:34,960 --> 00:29:37,520
symptoms of Denguing start to 
shift slightly because of a 

637
00:29:37,520 --> 00:29:40,440
climate change factor. 
The generator would see that new

638
00:29:40,440 --> 00:29:42,560
pattern emerging in the real 
time data. 

639
00:29:42,720 --> 00:29:45,680
And it would immediately 
generate thousands of new 

640
00:29:45,720 --> 00:29:48,400
simulated cases of this new 
pattern. 

641
00:29:48,800 --> 00:29:51,800
It would then use that data to 
retrain the main diagnostic 

642
00:29:51,800 --> 00:29:53,840
model all overnight 
automatically. 

643
00:29:54,000 --> 00:29:57,280
So by the time the doctors clock
in for the morning shift, the AI

644
00:29:57,280 --> 00:29:59,880
has already practiced on 
thousands of cases places of the

645
00:29:59,880 --> 00:30:03,240
new strain before the human 
doctors have even seen their 

646
00:30:03,240 --> 00:30:05,720
second real patient of the day. 
Precisely. 

647
00:30:06,080 --> 00:30:09,320
It turns the AI from a 
historian, something that only 

648
00:30:09,320 --> 00:30:13,200
looks at the past, into a scout.
It can adapt to a changing 

649
00:30:13,200 --> 00:30:16,280
disease faster than a human run 
public health system can. 

650
00:30:16,280 --> 00:30:20,360
That is slightly terrifying, but
also incredibly hopeful. 

651
00:30:20,360 --> 00:30:24,000
It's the future of adaptive real
time diagnostics is where this 

652
00:30:24,000 --> 00:30:26,280
is all heading. 
So here is your call to action 

653
00:30:26,280 --> 00:30:28,320
for the week. 
If you're an engineer sitting on

654
00:30:28,320 --> 00:30:31,200
an imbalance data set, and I 
know you are because basically 

655
00:30:31,200 --> 00:30:34,960
all real world data sets are, 
try implementing Deepctgen. 

656
00:30:35,600 --> 00:30:38,480
Don't just settle for the old 
SEM OT standby. 

657
00:30:38,480 --> 00:30:41,000
But please don't forget the 
resnet skip connections or your 

658
00:30:41,000 --> 00:30:43,560
gradients will vanish into thin 
air and you'll spend a week 

659
00:30:43,560 --> 00:30:45,800
debugging. 
It and always always check your 

660
00:30:45,800 --> 00:30:48,880
feature importance if patient ID
is your top predictor. 

661
00:30:49,120 --> 00:30:52,160
Delete your model, delete your 
features and start over from 

662
00:30:52,160 --> 00:30:54,240
scratch. 
Wise words. 

663
00:30:54,960 --> 00:30:56,400
Thanks for listening to this 
deep dive. 

664
00:30:56,440 --> 00:30:59,160
We'll catch you on the next one.
Goodbye and Abby Building.

