1
00:00:03,080 --> 00:00:06,640
Greetings and welcome to EHA 
Unplug, the official podcast 

2
00:00:06,640 --> 00:00:10,240
channel of the European 
Hematology Association EHA. 

3
00:00:10,600 --> 00:00:12,560
Hello. 
I am Isabella Rivera, a medical 

4
00:00:12,560 --> 00:00:16,520
writer for EHA. 
Today we're thrilled to have 

5
00:00:16,520 --> 00:00:20,680
Doctor Jan Moritz Mideka, a 
distinguished hematologist and 

6
00:00:20,720 --> 00:00:24,240
AI specialist. 
We're going to be talking today 

7
00:00:24,240 --> 00:00:27,160
about AI to accelerate clinical 
trials. 

8
00:00:27,560 --> 00:00:30,920
Welcome, Doctor Mideke, can you 
please introduce yourself? 

9
00:00:31,800 --> 00:00:35,760
Hi, yes, thanks for having me. 
I'm I'm Moritz Malika from the 

10
00:00:35,760 --> 00:00:38,920
University Hospital in Dresden. 
I'm working there as a 

11
00:00:38,920 --> 00:00:43,360
hematologist and have a research
group on AI in hematology. 

12
00:00:43,920 --> 00:00:49,600
So we want to understand how you
can accelerate clinical trials 

13
00:00:49,600 --> 00:00:54,320
by generating data with AI. 
Can you please explain first 

14
00:00:54,320 --> 00:00:58,440
what is synthetic data? 
Synthetic data is, as the name 

15
00:00:58,440 --> 00:01:04,840
says, data which is made-up, and
the idea is that we face a big 

16
00:01:04,840 --> 00:01:10,920
problem with clinical trials in 
general, but especially in 

17
00:01:10,920 --> 00:01:14,200
hematology where a lot of the 
diseases are rare diseases. 

18
00:01:14,200 --> 00:01:17,160
For example, acute myeloid 
leukemia is a rare disease per 

19
00:01:17,160 --> 00:01:22,440
see and we already know that 
there are several subgroups with

20
00:01:22,760 --> 00:01:26,200
targeted therapies. 
So the number of patients we can

21
00:01:26,200 --> 00:01:29,280
include in clinical trials are 
even lower. 

22
00:01:30,240 --> 00:01:34,520
And therefore we need to think 
about new ways how we can 

23
00:01:35,200 --> 00:01:40,960
accelerate how we perform 
clinical trials in order to get 

24
00:01:41,200 --> 00:01:46,320
sufficient data to get the 
knowledge about novel therapies.

25
00:01:47,640 --> 00:01:52,120
So what types of synthetic data 
are you generating that could be

26
00:01:52,120 --> 00:01:56,000
used for clinical trials? 
So the idea is that you have 

27
00:01:56,000 --> 00:01:59,840
synthetic data which is 
reassembling the original cohort

28
00:02:00,120 --> 00:02:06,720
in a a most perfect way, but 
without being able to get back 

29
00:02:06,720 --> 00:02:11,320
to the original data. 
So you're generating images. 

30
00:02:11,480 --> 00:02:15,400
This can be used used both for 
images but also for tabular 

31
00:02:15,400 --> 00:02:18,040
data. 
And what we did in recent study 

32
00:02:18,040 --> 00:02:23,160
and also other groups have shown
is that you can reassemble real 

33
00:02:23,400 --> 00:02:28,480
clinical trial data in terms of 
baseline data, but also in terms

34
00:02:28,480 --> 00:02:35,520
of outcome results based on 
cohorts of of real patients. 

35
00:02:35,520 --> 00:02:39,480
And the synthetic data is 
reassembling all the information

36
00:02:39,480 --> 00:02:44,200
about the original patient, 
which is like age, sex, but also

37
00:02:44,200 --> 00:02:49,000
molecular changes and then the 
outcome after a given 

38
00:02:49,000 --> 00:02:52,640
intervention. 
And how do you generate the 

39
00:02:52,640 --> 00:02:56,720
synthetic data? 
So for this we used guns, 

40
00:02:57,000 --> 00:03:04,280
Generative Adversarial Networks 
which is a known form of AI 

41
00:03:04,280 --> 00:03:07,640
which has been used for image 
generation, quite successful 

42
00:03:07,640 --> 00:03:10,480
already, but can also be used 
for tabular data. 

43
00:03:12,240 --> 00:03:16,680
And in our specific work, we use
two types of these, so and flow 

44
00:03:16,680 --> 00:03:21,200
and C Top Gun. 
So these are types of deep 

45
00:03:21,200 --> 00:03:23,480
learning and it's generative 
modelling. 

46
00:03:23,560 --> 00:03:28,040
You have described 4 scenarios 
in which synthetic data 

47
00:03:28,360 --> 00:03:33,840
generation can be useful and I 
would like to walk through this 

48
00:03:33,960 --> 00:03:39,720
for the audience. 
So first, AI generated data 

49
00:03:40,600 --> 00:03:48,000
could help to prevent privacy 
concerns about clinical trial 

50
00:03:48,000 --> 00:03:50,120
date data of the patients. 
Exactly. 

51
00:03:50,120 --> 00:03:55,560
So the generated synthetic data 
can be shared without any 

52
00:03:55,560 --> 00:03:59,160
privacy issues. 
So we checked that there was no 

53
00:03:59,240 --> 00:04:03,760
single patient in the synthetic 
cohort which was exactly 

54
00:04:03,880 --> 00:04:08,800
matching the original patient. 
So there's no way that you can 

55
00:04:08,800 --> 00:04:12,760
go back to the original patient,
which of course is the case if 

56
00:04:12,760 --> 00:04:15,120
you just have an anonymized 
data. 

57
00:04:16,440 --> 00:04:20,399
So this makes it really easy. 
You can just upload these data 

58
00:04:20,399 --> 00:04:23,600
and make them available also for
other research group and see 

59
00:04:24,240 --> 00:04:27,760
public. 
So how does this compare to the 

60
00:04:27,760 --> 00:04:31,960
DI identification or 
anonymisation of patient data? 

61
00:04:31,960 --> 00:04:36,800
So anonymisation of these data 
isn't 100% sure. 

62
00:04:37,040 --> 00:04:40,120
And with the synthetic data it's
100% sure. 

63
00:04:40,120 --> 00:04:43,160
So there's no way that you can 
go back to the original patient 

64
00:04:43,160 --> 00:04:49,240
because see, there's no original
data coming from one single 

65
00:04:49,240 --> 00:04:52,240
patient. 
So it's basically 100% sure that

66
00:04:52,240 --> 00:04:58,040
you can't go back to the 
characteristics of one single 

67
00:04:58,040 --> 00:05:01,320
patient. 
So this is a big step forward 

68
00:05:01,320 --> 00:05:04,880
for this kind of data. 
Also, you would create new 

69
00:05:04,880 --> 00:05:09,800
patients and this would lead to 
either cohort augmentation, so 

70
00:05:10,560 --> 00:05:14,360
some synthetic patients among 
patients that are real or 

71
00:05:15,320 --> 00:05:19,640
completely a substitution and 
you have a complete new cohort. 

72
00:05:22,080 --> 00:05:28,360
How robust is this process? 
So for the generation, it's very

73
00:05:28,360 --> 00:05:30,960
robust. 
We showed and also other groups,

74
00:05:30,960 --> 00:05:34,320
Matteo de la Porta did this for 
MD's showed that it's 

75
00:05:35,120 --> 00:05:40,000
reassembling the whole Crestrix 
of the patients forgiven 

76
00:05:40,000 --> 00:05:43,400
disease. 
So it's not only age and sex in 

77
00:05:43,400 --> 00:05:47,360
terms of the distribution, but 
it also comes down to the 

78
00:05:47,360 --> 00:05:51,040
molecular changes, which is very
important because we know that 

79
00:05:51,040 --> 00:05:54,080
there are certain relationships 
between certain mutations. 

80
00:05:54,080 --> 00:05:57,800
So for example, if you have an 
NPM 1 mutation, you are more 

81
00:05:57,800 --> 00:06:00,160
likely to also have inflate 3 
mutation. 

82
00:06:00,480 --> 00:06:04,480
And these correlations are 
preserved in the synthetic data 

83
00:06:04,480 --> 00:06:07,120
set. 
And this holds also true for 

84
00:06:07,120 --> 00:06:10,160
correlation between age and 
certain cytogenetic 

85
00:06:10,160 --> 00:06:13,800
abnormalities like complex 
karyotypes with with higher age 

86
00:06:13,800 --> 00:06:17,040
and so on. 
So biologically important 

87
00:06:17,200 --> 00:06:20,080
correlations are preserved 
within this data. 

88
00:06:20,440 --> 00:06:23,800
And even more important, this 
also holds true for the 

89
00:06:24,400 --> 00:06:29,480
correlation to the outcome data.
So known factors for poor or 

90
00:06:29,480 --> 00:06:33,400
good prognosis are also 
preserved in the synthetic data.

91
00:06:33,400 --> 00:06:39,840
So if you have a patient with 
some complex karyotype, he also 

92
00:06:39,840 --> 00:06:43,400
has a worse prognosis in the 
synthetic group of patients. 

93
00:06:43,400 --> 00:06:49,000
And this is super important 
because if we want to use them 

94
00:06:49,000 --> 00:06:53,400
as a real control group, we have
to make sure that this is also 

95
00:06:53,400 --> 00:06:58,080
like in the real world. 
So there's also A use for 

96
00:06:58,080 --> 00:07:02,440
exploratory analysis. 
So synthetically enlarged data 

97
00:07:02,440 --> 00:07:05,280
sets. 
What is the interest of this? 

98
00:07:05,320 --> 00:07:09,640
We know that there are several 
important subgroups when we stay

99
00:07:09,640 --> 00:07:13,120
in the case of acute myeloid 
leukemia, certain molecular 

100
00:07:13,120 --> 00:07:16,200
subgroups which are recurrent 
but are rare. 

101
00:07:16,320 --> 00:07:20,200
And so for these group of 
patients, it's of course very 

102
00:07:20,200 --> 00:07:26,720
important to know how they fare 
with a certain therapy and with 

103
00:07:26,720 --> 00:07:32,640
a synthetic generation of data, 
you can substitute or enlarge 

104
00:07:32,680 --> 00:07:38,520
these rare subgroups. 
How reliable is this process? 

105
00:07:38,960 --> 00:07:42,400
We have all heard about 
hallucinations in AI. 

106
00:07:43,920 --> 00:07:49,240
Can this generative AI generate 
misleading data? 

107
00:07:50,120 --> 00:07:55,280
Well, of course, you always have
to show it for a given group, 

108
00:07:55,600 --> 00:07:59,160
but for the examples I 
mentioned, this is very robust. 

109
00:07:59,160 --> 00:08:05,800
So there are some small 
differences, but in general this

110
00:08:05,800 --> 00:08:11,560
is very robust in terms of how 
the baseline distribution of 

111
00:08:11,560 --> 00:08:15,520
characteristics is, but also for
the outcome data. 

112
00:08:15,520 --> 00:08:20,120
However, especially when the 
outcome events get lower in the 

113
00:08:20,120 --> 00:08:26,160
training group, this leads to a 
lower robustness of the model. 

114
00:08:26,160 --> 00:08:29,240
So there is still open questions
where we have to work on. 

115
00:08:31,120 --> 00:08:34,880
And the last case you mentioned 
where this kind of synthetic 

116
00:08:34,880 --> 00:08:37,960
data can be used is in model 
training or benchmarking. 

117
00:08:38,520 --> 00:08:43,960
Can you explain this further? 
So the idea is that we also have

118
00:08:43,960 --> 00:08:48,400
to generate data for rare 
subgroups and for that we can 

119
00:08:48,400 --> 00:08:52,400
use synthetic data to enlarge 
the group of patients basically.

120
00:08:53,880 --> 00:08:59,400
For in clinical trials this 
generated patients are called 

121
00:08:59,400 --> 00:09:02,320
now synthetic patients. 
Can you summarize what the 

122
00:09:02,320 --> 00:09:06,680
advantages to use this kind of 
synthetic patients? 

123
00:09:06,680 --> 00:09:11,560
We have to point out that so far
synthetic patients have not been

124
00:09:11,560 --> 00:09:15,000
used within clinical trials. 
And I believe that there's still

125
00:09:15,000 --> 00:09:20,040
a lot of work to do also from a 
regulatory aspect to show that 

126
00:09:20,040 --> 00:09:23,080
this is really feasible. 
On the other hand, we know that 

127
00:09:23,080 --> 00:09:27,560
the generation works and the 
problem is so big that we have 

128
00:09:27,560 --> 00:09:31,960
to think of new ways how we can 
accelerate clinical testing. 

129
00:09:32,400 --> 00:09:37,320
So for this example, we used 
only intensively treated acute 

130
00:09:37,320 --> 00:09:42,200
myeloid leukemia patients almost
with our targeted therapy. 

131
00:09:42,200 --> 00:09:45,560
So this has to be shown with 
targeted therapies. 

132
00:09:45,560 --> 00:09:49,760
This has to be shown with real 
life data and with certain 

133
00:09:49,760 --> 00:09:52,080
intervention. 
And there's a lot of work to do 

134
00:09:52,080 --> 00:09:57,760
before we can really say, OK, so
we one group of patients within 

135
00:09:57,760 --> 00:10:03,200
a clinical trial for in the end 
leading to an approval or 

136
00:10:03,200 --> 00:10:07,440
identification of certain 
efficacy of a of a given drug. 

137
00:10:08,840 --> 00:10:13,080
So really is work in progress, 
how closely does these cohorts 

138
00:10:13,080 --> 00:10:19,160
match real world cohorts? 
OK, that's a difficult question 

139
00:10:19,160 --> 00:10:21,720
because it hasn't been shown 
yet. 

140
00:10:22,680 --> 00:10:25,720
So for the moment you don't know
really how well and it's being 

141
00:10:25,720 --> 00:10:28,920
tested. 
We know how well it matches for 

142
00:10:29,160 --> 00:10:33,200
a given group of patients we 
studied or other groups studied 

143
00:10:33,240 --> 00:10:37,840
and there it is working well. 
However, we have to, yeah, prove

144
00:10:37,840 --> 00:10:42,600
this for real world cohorts and 
and certain interventions. 

145
00:10:43,480 --> 00:10:47,960
So the idea would be to use 
these cohorts mostly in the 

146
00:10:47,960 --> 00:10:51,960
control group in a clinical 
trial at the beginning at least.

147
00:10:52,040 --> 00:10:56,000
Yes, I, I, I think this is the 
most realistic and and closest 

148
00:10:56,000 --> 00:11:00,200
scenario that we use this as a 
control group, an enlarged 

149
00:11:00,200 --> 00:11:03,640
control group, yes, with a 
certain intervention. 

150
00:11:04,840 --> 00:11:07,720
Which would be interesting 
because then you don't have to 

151
00:11:07,720 --> 00:11:10,400
assign patients to a control 
group. 

152
00:11:10,640 --> 00:11:13,520
So you mentioned the regulatory 
considerations. 

153
00:11:13,720 --> 00:11:17,280
We have some tools at the 
moment, like GDPR or HIPAA. 

154
00:11:18,920 --> 00:11:22,840
Are they enough to address the 
potential issues that this kind 

155
00:11:22,840 --> 00:11:26,880
of synthetic patients would 
bring? 

156
00:11:30,440 --> 00:11:35,480
Well, I think this requires 
complete new thinking about the 

157
00:11:35,480 --> 00:11:40,080
approval process and the, the 
way we look at clinical trials. 

158
00:11:40,640 --> 00:11:43,880
We need a way to prove that this
is really working. 

159
00:11:43,880 --> 00:11:48,200
And and I envision this in the 
1st place to go alongside 

160
00:11:48,200 --> 00:11:50,120
clinical trials. 
And this is also what we are 

161
00:11:50,120 --> 00:11:53,280
currently doing. 
So we're going back to clinical 

162
00:11:53,280 --> 00:11:59,160
trials, we already did and we 
are going to see how many 

163
00:11:59,160 --> 00:12:05,640
patients we can substitute by 
synthetic patients to get the 

164
00:12:05,640 --> 00:12:10,000
results we already know. 
So for me this is an important 

165
00:12:10,000 --> 00:12:16,200
step to show how we can 
incorporate A synthetic patients

166
00:12:16,800 --> 00:12:20,720
in real clinical trials. 
And I think this step has to be 

167
00:12:20,960 --> 00:12:27,160
done in close contact with the 
authorities that we together 

168
00:12:27,160 --> 00:12:32,840
define criteria where we say, 
OK, this is working and we can 

169
00:12:32,920 --> 00:12:40,040
move on further. 
So if you had a crystal ball and

170
00:12:40,040 --> 00:12:44,320
could see what the future holds 
in five years time, do you think

171
00:12:44,320 --> 00:12:47,640
it's going to be used a lot? 
Yeah, what five years is close 

172
00:12:47,640 --> 00:12:54,080
in in clinical testing, but I, I
know that the way we did 

173
00:12:54,520 --> 00:12:58,160
clinical trials in rare disease 
like acute myeloid leukemia in 

174
00:12:58,160 --> 00:13:01,720
the past is over. 
So there won't be clinical 

175
00:13:01,720 --> 00:13:07,880
trials with like 1000 patient 
randomized with a certain drug 

176
00:13:07,880 --> 00:13:12,560
because in the meanwhile other 
drugs are coming up we we gain 

177
00:13:12,600 --> 00:13:16,480
further knowledge. 
So this is getting harder and 

178
00:13:16,480 --> 00:13:18,720
harder to perform these large 
trials. 

179
00:13:18,720 --> 00:13:23,240
So we have to adopt new methods.
And I think synthetic data is 1 

180
00:13:23,440 --> 00:13:27,120
very promising way. 
But there are several steps 

181
00:13:27,120 --> 00:13:30,400
before we can really say, OK, 
that this is working and 

182
00:13:30,400 --> 00:13:35,560
enhancing our clinical testing. 
Are there other uses of this 

183
00:13:35,560 --> 00:13:45,480
data that you are now developing
for prognostic capabilities in 

184
00:13:45,480 --> 00:13:51,760
hematological malignancies? 
So coming back to images, 

185
00:13:52,680 --> 00:13:56,280
there's another interesting 
field where we can enhance the 

186
00:13:56,320 --> 00:14:00,640
amount of images for rare 
subgroups of patients. 

187
00:14:00,640 --> 00:14:05,360
So basically doing the same 
thing for small subgroups within

188
00:14:05,360 --> 00:14:08,400
clinical trials. 
We can also do this for images 

189
00:14:08,400 --> 00:14:10,600
of of bone marrow smears, for 
example. 

190
00:14:11,360 --> 00:14:15,440
This is already, it's been used.
Well, I wouldn't say it's been 

191
00:14:15,440 --> 00:14:19,320
used, but it's like we and 
others have shown that this is 

192
00:14:19,320 --> 00:14:25,240
working and that you can 
generate synthetic images based 

193
00:14:25,240 --> 00:14:29,200
on real bone marrow smears and 
that you can't really tell which

194
00:14:29,200 --> 00:14:33,520
is a synthetic or a real bone 
marrow smear picture. 

195
00:14:34,760 --> 00:14:37,240
And these have been used to. 
Train. 

196
00:14:37,280 --> 00:14:40,840
Well, it hasn't been used, but 
that's the idea. 

197
00:14:40,840 --> 00:14:44,400
So the idea is that you can 
generate more data and then 

198
00:14:44,400 --> 00:14:52,280
increase model development. 
So in the future you think that 

199
00:14:52,280 --> 00:14:54,840
you will go beyond the control 
group, so you would use the 

200
00:14:54,840 --> 00:14:58,440
synthetic data also in the test 
group. 

201
00:14:59,440 --> 00:15:03,480
The improvement of the 
techniques is so fast that we 

202
00:15:03,480 --> 00:15:09,600
can really like simulate what is
going on within a single cancer 

203
00:15:09,600 --> 00:15:12,040
cell with a certain 
intervention. 

204
00:15:12,960 --> 00:15:17,720
And with that, I believe that 
in, I don't know how many years,

205
00:15:17,720 --> 00:15:21,640
but in the future we will be 
able to also simulate what will 

206
00:15:21,640 --> 00:15:26,680
happen to a certain group of 
patients when we give the drug 

207
00:15:27,000 --> 00:15:30,120
X. 
Yes, I'm I'm a biologist by 

208
00:15:30,120 --> 00:15:35,920
training and I find difficult to
imagine that you will be able to

209
00:15:37,080 --> 00:15:43,200
reproduce this complex system in
its totality, but I look forward

210
00:15:43,200 --> 00:15:44,760
to it. 
It would be fantastic. 

211
00:15:45,880 --> 00:15:50,600
Yeah, agree on that. 
But like five years back, you 

212
00:15:50,600 --> 00:15:54,040
probably would have said that in
predicting protein structure, 

213
00:15:54,040 --> 00:15:57,440
which is now working almost 
perfectly. 

214
00:15:58,560 --> 00:16:02,600
So yeah. 
I think it's really promising 

215
00:16:02,600 --> 00:16:07,600
and I really hope this works 
because it would help enormously

216
00:16:07,600 --> 00:16:12,480
in accelerating clinical trials 
and treating the other patients.

217
00:16:13,120 --> 00:16:17,240
So thank you very much, Doctor 
Mideke, for sharing your 

218
00:16:17,240 --> 00:16:19,240
insights and your experience 
with us. 

219
00:16:20,360 --> 00:16:22,800
Thank you to everybody for 
listening. 

220
00:16:22,840 --> 00:16:26,400
And if you like this episode, 
don't forget to like it and 

221
00:16:26,400 --> 00:16:30,040
share it with your colleagues. 
And stay tuned for our next 

222
00:16:30,080 --> 00:16:31,440
episode. 
Thank you. 

223
00:16:31,840 --> 00:16:32,280
Thank you.
