1
00:00:03,080 --> 00:00:06,640
Greetings and welcome to EHA 
Unplug, the official podcast 

2
00:00:06,640 --> 00:00:10,240
channel of the European 
Hematology Association EHA. 

3
00:00:10,600 --> 00:00:13,600
Hello, I am Isabella Rivera, a 
medical writer for EHA. 

4
00:00:14,400 --> 00:00:18,680
Today we're thrilled to have 
Doctor Jan Moritz Medeka, a 

5
00:00:18,840 --> 00:00:21,600
distinguished hematologist and 
AI specialist. 

6
00:00:22,880 --> 00:00:26,680
We're going to be talking today 
about AI to accelerate clinical 

7
00:00:26,680 --> 00:00:28,480
trials. 
Welcome, Doctor. 

8
00:00:28,480 --> 00:00:30,920
Medeka, can you please introduce
yourself? 

9
00:00:31,800 --> 00:00:35,760
Hi, yeah, thanks for having me. 
I'm, I'm Moritz Malika from the 

10
00:00:35,760 --> 00:00:38,880
University Hospital in Dresden. 
I'm working there as a 

11
00:00:38,880 --> 00:00:40,960
hematologist and have a. 
Research Group. 

12
00:00:41,600 --> 00:00:47,840
On AI in hematology. 
So we want to understand how you

13
00:00:47,840 --> 00:00:51,840
can accelerate clinical trials 
by generating data with AI. 

14
00:00:52,680 --> 00:00:55,720
Can you please explain first 
what is synthetic data? 

15
00:00:56,400 --> 00:00:58,960
Synthetic data is as the name 
says. 

16
00:00:59,240 --> 00:01:01,960
Data which is. 
Made-up and. 

17
00:01:02,200 --> 00:01:06,520
The idea is that we face a. 
Big problem with clinical trials

18
00:01:06,640 --> 00:01:11,880
in general, but especially in 
hematology. 

19
00:01:11,880 --> 00:01:14,400
Where a lot of the diseases are 
rare diseases for. 

20
00:01:14,400 --> 00:01:16,520
Example acute myeloid leukemia 
is a rare. 

21
00:01:16,520 --> 00:01:20,360
Disease per seat. 
And we already know that. 

22
00:01:20,360 --> 00:01:24,560
There are several subgroups with
targeted therapies, so the. 

23
00:01:24,560 --> 00:01:28,360
Number of patients we can 
include in clinical trials are 

24
00:01:28,360 --> 00:01:34,040
even lower and therefore we need
to think about new ways how we 

25
00:01:34,040 --> 00:01:37,160
can. 
Accelerate how we. 

26
00:01:37,160 --> 00:01:43,200
Perform clinical trials in order
to get sufficient data to get 

27
00:01:43,200 --> 00:01:46,320
the knowledge about novel. 
Therapies. 

28
00:01:47,640 --> 00:01:52,120
So what types of synthetic data 
are you generating that could be

29
00:01:52,120 --> 00:01:55,480
used for clinical trials? 
So the idea. 

30
00:01:55,480 --> 00:01:57,280
Is that you have synthetic? 
Data which is. 

31
00:01:57,560 --> 00:02:00,680
Reassembling the original cohort
in a A. 

32
00:02:00,720 --> 00:02:05,320
Most perfect way. 
But without being able. 

33
00:02:05,320 --> 00:02:07,920
To get back. 
To the original data. 

34
00:02:09,320 --> 00:02:12,400
So you're generating images. 
This can be used. 

35
00:02:12,480 --> 00:02:14,000
Used both for. 
Images. 

36
00:02:14,000 --> 00:02:17,560
But also for tabular. 
Data and what we did in recent 

37
00:02:17,560 --> 00:02:22,160
study and also other groups have
shown is that you can reassemble

38
00:02:22,600 --> 00:02:27,600
real clinical trial data in 
terms of baseline data, but also

39
00:02:27,600 --> 00:02:34,520
in terms of outcome results 
based on cohorts of of. 

40
00:02:34,520 --> 00:02:36,720
Real. 
Patients and the synthetic data 

41
00:02:36,720 --> 00:02:40,560
is reassembling all the 
information about the original 

42
00:02:40,560 --> 00:02:42,600
patient which is like. 
Age. 

43
00:02:42,600 --> 00:02:46,200
Sex. 
But also molecular changes and 

44
00:02:46,200 --> 00:02:49,840
then the outcome after a given 
intervention. 

45
00:02:51,200 --> 00:02:53,480
And how do you generate the 
synthetic data? 

46
00:02:53,800 --> 00:02:57,600
So for this. 
We used guns, Generative 

47
00:02:57,600 --> 00:03:04,280
Adversarial Networks, which is a
known form of AI. 

48
00:03:04,280 --> 00:03:06,600
Which has been. 
Used for image generation. 

49
00:03:06,720 --> 00:03:08,160
Quite successful. 
Already. 

50
00:03:08,440 --> 00:03:12,600
But can also. 
Be used for tabular data and in 

51
00:03:12,600 --> 00:03:15,600
our specific work we use two 
types of these, so. 

52
00:03:15,920 --> 00:03:21,200
Nflow and C Top Gun. 
So these are types of deep 

53
00:03:21,200 --> 00:03:23,480
learning and it's generative 
modelling. 

54
00:03:23,560 --> 00:03:28,040
You have described 4 scenarios 
in which synthetic data 

55
00:03:28,360 --> 00:03:33,840
generation can be useful and I 
would like to walk through this 

56
00:03:33,960 --> 00:03:39,720
for the audience. 
So first, AI generated data 

57
00:03:40,560 --> 00:03:48,000
could help to prevent privacy 
concerns about clinical trial 

58
00:03:48,000 --> 00:03:50,480
date data of the patients. 
Exactly so. 

59
00:03:51,840 --> 00:03:55,560
The generated synthetic data can
be shared without any. 

60
00:03:55,560 --> 00:03:57,840
Privacy issues, so we checked 
that. 

61
00:03:58,240 --> 00:04:02,120
There was no single patient in 
the synthetic cohort. 

62
00:04:02,120 --> 00:04:05,560
Which was exactly. 
Matching the original patient. 

63
00:04:06,760 --> 00:04:10,800
So there's no way that you can 
go back to the original patient 

64
00:04:10,800 --> 00:04:12,920
which. 
Of course is the case if you. 

65
00:04:13,000 --> 00:04:17,360
Just have an anonymized data, so
this makes. 

66
00:04:17,360 --> 00:04:20,399
It really easy. 
You can just upload these data 

67
00:04:20,399 --> 00:04:22,880
and make them available. 
Also for other research. 

68
00:04:22,880 --> 00:04:27,760
Group and see public. 
So how does this compare to the 

69
00:04:27,760 --> 00:04:31,960
de identification or 
anonymisation of patient data? 

70
00:04:31,960 --> 00:04:37,880
So anonymisation of these data 
isn't 100% sure, and with the 

71
00:04:37,880 --> 00:04:41,720
synthetic data it's 100% sure. 
So there's no way that you can 

72
00:04:41,720 --> 00:04:44,640
go back to the original patient.
Because see, there's. 

73
00:04:44,640 --> 00:04:49,840
No original data coming from one
single patient. 

74
00:04:50,360 --> 00:04:54,320
So it's basically 100% sure that
you can't go back to the. 

75
00:04:56,200 --> 00:04:58,600
Characteristics of. 
One single patient. 

76
00:04:59,280 --> 00:05:02,600
So this is a big step forward 
for this kind of data. 

77
00:05:02,720 --> 00:05:07,560
Also, you would create new 
patients and this would lead to 

78
00:05:07,600 --> 00:05:12,120
either cohort augmentation, so 
some synthetic patients among 

79
00:05:12,120 --> 00:05:17,440
patients that are real or 
completely a substitution and 

80
00:05:17,440 --> 00:05:23,600
you have a complete new cohort. 
How robust is this process? 

81
00:05:25,040 --> 00:05:28,000
So for the. 
Generation it's. 

82
00:05:28,000 --> 00:05:30,960
Very robust. 
We showed and also other groups.

83
00:05:30,960 --> 00:05:33,120
Matteo de la Porta did this for 
MD's. 

84
00:05:33,400 --> 00:05:37,280
Showed that it's. 
Reassembling the whole. 

85
00:05:37,280 --> 00:05:40,440
Crestrix of the. 
Patients for a given disease. 

86
00:05:40,440 --> 00:05:45,400
So it's not only age and sex in 
terms of the distribution, but 

87
00:05:45,400 --> 00:05:49,520
it also comes down to the 
molecular changes, which is very

88
00:05:49,520 --> 00:05:52,640
important because we know that 
there are certain relationships 

89
00:05:52,640 --> 00:05:55,800
between certain mutations. 
So for example, if you have an 

90
00:05:55,920 --> 00:05:59,480
NPM 1 mutation, you're more 
likely to also have inflat 3 

91
00:05:59,480 --> 00:06:02,720
mutation. 
And these correlations are 

92
00:06:02,720 --> 00:06:04,800
preserved in the synthetic. 
Data set. 

93
00:06:05,120 --> 00:06:06,840
And this. 
Holds also true. 

94
00:06:06,840 --> 00:06:10,160
For correlation between age and 
certain cytogenetic 

95
00:06:10,160 --> 00:06:13,120
abnormalities like complex. 
Karyotypes with with. 

96
00:06:13,120 --> 00:06:16,120
Higher age and so on so. 
Biologically. 

97
00:06:16,120 --> 00:06:19,040
Important correlations are. 
Preserved. 

98
00:06:19,040 --> 00:06:22,760
Within this data and even more 
important, this also. 

99
00:06:22,760 --> 00:06:26,680
Holds true for the. 
Correlation to the outcome data,

100
00:06:26,680 --> 00:06:29,760
so known factors for. 
Poor or good? 

101
00:06:29,760 --> 00:06:33,080
Prognosis are also preserved in 
the synthetic. 

102
00:06:33,080 --> 00:06:35,040
Data so if. 
You have a patient with some 

103
00:06:37,200 --> 00:06:39,200
complex karyotype. 
He. 

104
00:06:39,480 --> 00:06:41,360
Also has a worse prognosis in 
the. 

105
00:06:41,360 --> 00:06:45,240
Synthetic group of patients, and
this is super important because 

106
00:06:46,120 --> 00:06:51,680
if we want to use him as a real 
control group, we have to make 

107
00:06:51,680 --> 00:06:53,400
sure that. 
This is also. 

108
00:06:53,400 --> 00:06:55,000
Like in the. 
Real world. 

109
00:06:56,560 --> 00:06:59,720
So there's also A use for 
exploratory analysis. 

110
00:07:00,000 --> 00:07:04,840
So synthetically enlarged data 
sets, what is the interest? 

111
00:07:04,880 --> 00:07:09,080
Of this, we know that there are 
several important subgroups when

112
00:07:09,080 --> 00:07:12,520
we stay in the case of acute 
myeloid leukemia, certain 

113
00:07:12,520 --> 00:07:14,960
molecular subgroups which are 
recurrent. 

114
00:07:14,960 --> 00:07:16,720
But are rare. 
And. 

115
00:07:16,960 --> 00:07:20,200
So for these group. 
Of patients, it's of course very

116
00:07:20,200 --> 00:07:23,720
important to know. 
How they fare with a certain. 

117
00:07:23,720 --> 00:07:27,480
Therapy and. 
With a synthetic. 

118
00:07:28,600 --> 00:07:31,720
Generation of data you can. 
Substitute or. 

119
00:07:31,720 --> 00:07:34,720
Enlarge. 
These rare subgroups. 

120
00:07:36,560 --> 00:07:40,160
How reliable is this process? 
We have all heard about 

121
00:07:40,280 --> 00:07:47,400
hallucinations in AI. 
Can this generative AI generate 

122
00:07:48,200 --> 00:07:50,360
misleading data? 
Well. 

123
00:07:51,520 --> 00:07:53,880
Of course, you always have to 
show it. 

124
00:07:53,880 --> 00:07:57,600
For a given group, but for the 
examples I mentioned. 

125
00:07:57,840 --> 00:08:02,600
This is very robust, so. 
There are some small 

126
00:08:02,600 --> 00:08:05,920
differences, but in general. 
This is. 

127
00:08:05,920 --> 00:08:11,560
Very robust in terms of how the 
baseline distribution of. 

128
00:08:11,560 --> 00:08:16,200
Characteristics is but also. 
For the outcome data, however, 

129
00:08:16,840 --> 00:08:20,840
especially when the outcome 
events get lower in the training

130
00:08:21,240 --> 00:08:24,680
group. 
This leads to a lower. 

131
00:08:24,680 --> 00:08:28,040
Robustness of the model. 
So there is still open question 

132
00:08:28,040 --> 00:08:32,880
where we have to work on. 
And the last case you mentioned 

133
00:08:33,000 --> 00:08:36,360
where this kind of synthetic 
data can be used is in model 

134
00:08:36,360 --> 00:08:39,840
training or benchmarking. 
Can you explain this further? 

135
00:08:40,360 --> 00:08:42,720
So the. 
Idea is that. 

136
00:08:42,720 --> 00:08:46,960
We also have to generate. 
Data for rare subgroups. 

137
00:08:47,320 --> 00:08:50,720
And for that we can use 
synthetic data to enlarge the 

138
00:08:50,720 --> 00:08:52,440
group of. 
Patients, basically. 

139
00:08:53,880 --> 00:08:59,440
For in clinical trials this 
generated patients are called 

140
00:08:59,440 --> 00:09:02,360
now synthetic patients. 
Can you summarize what the 

141
00:09:02,360 --> 00:09:06,720
advantages to use this kind of 
synthetic patients? 

142
00:09:06,720 --> 00:09:09,880
We have to point out that. 
So far synthetic. 

143
00:09:10,280 --> 00:09:13,400
Patients have not been used 
within clinical trials and. 

144
00:09:13,760 --> 00:09:16,800
I believe that there's still a 
lot of work to do, also from a 

145
00:09:16,800 --> 00:09:19,360
regulatory. 
Aspect to. 

146
00:09:19,360 --> 00:09:21,280
Show that this is really 
feasible. 

147
00:09:21,520 --> 00:09:25,880
On the other hand, we know that 
the generation works and the. 

148
00:09:25,880 --> 00:09:29,200
Problem is so big that. 
We have to think of new ways how

149
00:09:29,200 --> 00:09:31,960
we can accelerate clinical 
testing. 

150
00:09:32,400 --> 00:09:37,320
So for this example we used. 
Only intensively treated acute 

151
00:09:37,320 --> 00:09:40,960
myeloid leukemia patients almost
with. 

152
00:09:41,040 --> 00:09:43,120
Our targeted therapy, so this 
has to. 

153
00:09:43,120 --> 00:09:45,920
Be shown with. 
Targeted therapies this has. 

154
00:09:45,920 --> 00:09:48,160
To be shown with. 
Real life data. 

155
00:09:48,520 --> 00:09:51,280
And with certain intervention 
and there's a lot of. 

156
00:09:51,280 --> 00:09:53,440
Work to do. 
Before we can really. 

157
00:09:53,440 --> 00:09:57,480
Say OK. 
So we one group of patients 

158
00:09:57,480 --> 00:10:02,680
within a clinical trial for in 
the end leading to an approval 

159
00:10:02,680 --> 00:10:07,440
or identification of certain 
efficacy of a of a given drug. 

160
00:10:08,840 --> 00:10:13,080
So really is work in progress, 
how closely does these cohorts 

161
00:10:13,080 --> 00:10:17,800
match real world cohorts? 
OK. 

162
00:10:17,800 --> 00:10:19,560
That's a difficult question 
because. 

163
00:10:19,560 --> 00:10:24,080
It hasn't been shown yet. 
So for the moment you don't know

164
00:10:24,080 --> 00:10:26,200
really how well and it's being 
tested. 

165
00:10:26,400 --> 00:10:28,920
We know how well it matches. 
For. 

166
00:10:29,160 --> 00:10:31,880
A given group of patients we 
studied. 

167
00:10:32,040 --> 00:10:34,400
Or other groups. 
Studied and there it is working 

168
00:10:34,400 --> 00:10:36,360
well. 
However, we have to. 

169
00:10:36,720 --> 00:10:39,080
Yeah. 
Prove this for real world 

170
00:10:39,080 --> 00:10:42,600
cohorts. 
And and certain interventions. 

171
00:10:43,480 --> 00:10:47,960
So the idea would be to use 
these cohorts mostly in the 

172
00:10:47,960 --> 00:10:51,960
control group in a clinical 
trial at the beginning at least.

173
00:10:52,040 --> 00:10:54,000
Yes, I, I. 
I think this is the most 

174
00:10:54,000 --> 00:10:57,680
realistic and and closest 
scenario that we use this as a 

175
00:10:57,680 --> 00:11:01,240
control group, an enlarged 
control group. 

176
00:11:01,800 --> 00:11:03,640
Yes. 
With a certain intervention. 

177
00:11:04,840 --> 00:11:07,720
Which would be interesting 
because then you don't have to 

178
00:11:07,720 --> 00:11:10,400
assign patients to a control 
group. 

179
00:11:10,640 --> 00:11:13,520
So you mentioned the regulatory 
considerations. 

180
00:11:13,720 --> 00:11:17,280
We have some tools at the 
moment, like GDPR or HIPAA. 

181
00:11:18,920 --> 00:11:22,840
Are they enough to address the 
potential issues that this kind 

182
00:11:22,840 --> 00:11:26,880
of synthetic patients would 
bring? 

183
00:11:30,440 --> 00:11:32,320
Well, I think this. 
Requires. 

184
00:11:33,240 --> 00:11:37,400
Complete new thinking about the 
approval process and the. 

185
00:11:37,440 --> 00:11:40,080
The way we look at. 
Clinical trials. 

186
00:11:40,640 --> 00:11:43,080
We need a. 
Way to prove that this is. 

187
00:11:43,080 --> 00:11:46,040
Really working and and I 
envision this in the 1st. 

188
00:11:46,040 --> 00:11:48,680
Place to go. 
Alongside clinical. 

189
00:11:48,680 --> 00:11:51,160
Trials and this is also. 
What we are currently doing, so 

190
00:11:51,160 --> 00:11:54,480
we're going back to clinical. 
Trials we already. 

191
00:11:54,480 --> 00:12:01,400
Did and we are going to see how 
many patients we can substitute 

192
00:12:01,560 --> 00:12:06,360
by synthetic patients. 
To get the results, we. 

193
00:12:06,360 --> 00:12:10,560
Already know, so for me this is 
an important step. 

194
00:12:10,880 --> 00:12:15,520
To show how. 
We can incorporate A synthetic 

195
00:12:15,520 --> 00:12:20,720
patients in real clinical trials
and I think this step has to be 

196
00:12:20,960 --> 00:12:24,840
done in close contact. 
With the authorities. 

197
00:12:24,840 --> 00:12:28,680
That we together define 
criteria. 

198
00:12:28,680 --> 00:12:33,560
Where we say OK this is. 
Working and we can move on 

199
00:12:33,560 --> 00:12:40,040
further. 
So if you had a crystal ball and

200
00:12:40,040 --> 00:12:44,320
could see what the future holds 
in five years time, do you think

201
00:12:44,320 --> 00:12:46,800
it's going to be used a lot? 
Yeah, what 5? 

202
00:12:46,800 --> 00:12:51,520
Years is close in in clinical 
testing, but I I know. 

203
00:12:51,840 --> 00:12:56,200
That the way we. 
Did clinical trials in rare 

204
00:12:56,200 --> 00:12:58,240
disease like acute myeloid 
leukemia in the. 

205
00:12:58,240 --> 00:13:01,720
Past is over. 
So there won't be clinical. 

206
00:13:01,720 --> 00:13:07,880
Trials with like. 1000 patient 
randomized with a certain drug 

207
00:13:07,880 --> 00:13:12,520
because in the meanwhile other 
drugs are coming up we we gain 

208
00:13:12,600 --> 00:13:16,520
further knowledge. 
So this is getting harder and 

209
00:13:16,520 --> 00:13:18,720
harder to perform these large 
trials. 

210
00:13:18,720 --> 00:13:21,840
So we have to. 
Adopt new methods and I think 

211
00:13:21,840 --> 00:13:24,880
synthetic data is 1 very 
promising way. 

212
00:13:25,120 --> 00:13:28,720
But there are several steps. 
Before we can really say, OK, 

213
00:13:28,880 --> 00:13:32,480
that this is working and 
enhancing our clinical testing. 

214
00:13:32,960 --> 00:13:39,080
Are there other uses of this 
data that you are now developing

215
00:13:39,440 --> 00:13:47,040
for prognostic capabilities in 
hematological malignancies? 

216
00:13:47,880 --> 00:13:53,840
So coming back to images, 
there's another interesting 

217
00:13:53,840 --> 00:13:57,800
field where we can enhance the 
amount of images. 

218
00:13:57,800 --> 00:14:02,600
For rare subgroups. 
Of patients, so basically doing 

219
00:14:02,600 --> 00:14:05,800
the same thing for small 
subgroups within clinical. 

220
00:14:05,800 --> 00:14:07,760
Trials. 
We can also do this for. 

221
00:14:07,760 --> 00:14:10,600
Images of of bone marrow smears,
for example. 

222
00:14:11,360 --> 00:14:15,440
This is already, it's been used.
Well, I wouldn't say it's been 

223
00:14:15,440 --> 00:14:19,320
used, but it's like we and 
others have shown that this is 

224
00:14:19,320 --> 00:14:25,240
working and that you can 
generate synthetic images based 

225
00:14:25,240 --> 00:14:29,200
on real bone marrow smears and 
that you can't really tell which

226
00:14:29,200 --> 00:14:33,040
is a synthetic or. 
A real bone marrow smear 

227
00:14:33,040 --> 00:14:36,080
picture. 
And these have been used to 

228
00:14:36,560 --> 00:14:40,000
train. 
Well, it hasn't been used, but 

229
00:14:40,000 --> 00:14:40,840
that's. 
The idea. 

230
00:14:40,840 --> 00:14:43,280
So the idea. 
Is that you can generate more 

231
00:14:43,280 --> 00:14:47,280
data and then increase model 
development. 

232
00:14:49,520 --> 00:14:53,880
So in the future you think that 
you will go beyond the control 

233
00:14:53,880 --> 00:14:58,040
group, so you would use the 
synthetic data also in the test 

234
00:14:58,040 --> 00:15:00,720
group. 
The improvement of the 

235
00:15:01,200 --> 00:15:06,520
techniques is so fast that we 
can really, like simulate what 

236
00:15:06,520 --> 00:15:09,600
is going on within a single 
cancer. 

237
00:15:09,600 --> 00:15:13,440
Cell with a. 
Certain intervention and with 

238
00:15:13,440 --> 00:15:18,200
that I believe that in I don't 
know how many years, but in the 

239
00:15:18,200 --> 00:15:19,880
future. 
We will be able to. 

240
00:15:19,880 --> 00:15:24,480
Also simulate what will happen 
to a certain group of patients 

241
00:15:24,480 --> 00:15:27,440
when we. 
Give the drug X. 

242
00:15:28,480 --> 00:15:34,080
Yes, I'm I'm a biologist by 
training and I find difficult to

243
00:15:34,080 --> 00:15:40,520
imagine that you will be able to
reproduce this complex system in

244
00:15:40,520 --> 00:15:43,520
its totality, but I look forward
to it. 

245
00:15:43,600 --> 00:15:46,240
It would be fantastic. 
Yeah. 

246
00:15:46,400 --> 00:15:51,000
Agree on that but like. 
Five years back, you probably 

247
00:15:51,000 --> 00:15:52,960
would. 
Have said that in predicting 

248
00:15:52,960 --> 00:15:56,720
protein structure which is now 
working almost. 

249
00:15:56,720 --> 00:16:00,280
Perfectly. 
So yeah. 

250
00:16:01,320 --> 00:16:04,080
I think it's really promising 
and I really hope this works 

251
00:16:04,080 --> 00:16:10,920
because it would help enormously
in accelerating clinical trials 

252
00:16:10,920 --> 00:16:15,640
and treating the other patients.
So thank you very much, Doctor 

253
00:16:15,640 --> 00:16:18,800
Mideke, for sharing your 
insights and your experience 

254
00:16:18,800 --> 00:16:22,160
with us. 
Thank you to everybody for 

255
00:16:22,160 --> 00:16:24,640
listening. 
And if you like this episode, 

256
00:16:25,040 --> 00:16:27,600
don't forget to like it and 
share it with your colleagues. 

257
00:16:28,040 --> 00:16:30,560
And stay tuned for our next 
episode. 

258
00:16:31,040 --> 00:16:32,320
Thank you. 
Thank you.

