1
00:00:00,240 --> 00:00:02,360
Hi, and welcome to the Neil 
 
Ashton Podcast. 

2
00:00:03,000 --> 00:00:05,880
In each episode, we explained 
 
some of the fascinating ways 

3
00:00:05,880 --> 00:00:09,080
that science and engineering are

 changing the world around us. 

4
00:00:09,760 --> 00:00:12,720
We talked to leading engineers 

from elite level sports like 

5
00:00:12,800 --> 00:00:16,840
cycling and Formula One to some 
 of the world's top academics to

6
00:00:16,840 --> 00:00:20,960
understand how fluid dynamics, 

machine learning, supercomputing

7
00:00:21,360 --> 00:00:22,960
are bringing in a new era of 
 
discovery. 

8
00:00:23,920 --> 00:00:27,040
We also hear some of their life 
 stories, their career advice, 

9
00:00:27,640 --> 00:00:30,040
the lessons they've learned on 

the way that I hope will be 

10
00:00:30,040 --> 00:00:33,800
helpful to you too. 
 
So sit back and enjoy this 

11
00:00:33,800 --> 00:00:40,840
episode. 
 
Hi and welcome back to the Neil 

12
00:00:40,840 --> 00:00:43,800
Ashton podcast. 
 
So today's guest is Professor 

13
00:00:43,800 --> 00:00:47,000
Nils Thuerey. 
 
He's an associate professor at 

14
00:00:47,000 --> 00:00:51,000
the Technical University of 
 
Munich TUM and has been a 

15
00:00:51,000 --> 00:00:53,920
pioneer in physics-based deep 
 
learning. 

16
00:00:54,240 --> 00:01:00,160
We had a really interesting 
 
discussion today on first of 

17
00:01:00,160 --> 00:01:03,200
all, some of his background, 
 
Very interesting that he 

18
00:01:03,600 --> 00:01:08,720
actually won an, an Oscar 
 
working in the visual effects 

19
00:01:08,720 --> 00:01:14,560
industry for, for movies after 

his postdoc and PhD before he 

20
00:01:14,560 --> 00:01:20,160
then turned to academia full 
 
time, which is kind of 

21
00:01:20,360 --> 00:01:23,880
incredible actually, and is 
 
interesting because I've noticed

22
00:01:24,640 --> 00:01:28,080
the lessons learned from the 
 
visual effects industry in terms

23
00:01:28,080 --> 00:01:31,720
of photorealistic 
 
representations of, of, of the 

24
00:01:31,720 --> 00:01:36,920
world is actually something we, 
 we then talked right the end of

25
00:01:36,920 --> 00:01:39,160
the, the podcast. 
 
So make sure you listen to the 

26
00:01:39,160 --> 00:01:43,920
end about the link between 
 
foundation models and surrogates

27
00:01:43,920 --> 00:01:46,400
to world models. 
 
These world models that are 

28
00:01:46,800 --> 00:01:51,080
being created essentially as a 

synthetic environment to train 

29
00:01:51,160 --> 00:01:52,440
autonomous vehicles and robots. 
 

30
00:01:52,448 --> 00:01:55,932
And we actually talked a lot 
about whether there needs to be 

31
00:01:55,932 --> 00:01:58,530
 more physics in these world 
models, which I actually thought

32
00:01:58,530 --> 00:01:59,760

 was a very interesting 
discussion. 
 

33
00:01:59,768 --> 00:02:01,855
And we only got to it at the 
end. 
 

34
00:02:01,863 --> 00:02:08,380
But you know, he he's been one 
of these people who was doing 
 

35
00:02:08,388 --> 00:02:13,682
deep learning in the mid twenty 
10's and early twenty 20s before

36
00:02:13,682 --> 00:02:16,252

 ChatGPT and the sort of big 
rise. 
 

37
00:02:16,260 --> 00:02:20,620
And it was interesting to him 
talk about how his work was 
 

38
00:02:20,628 --> 00:02:24,259
looked in those times compared 
to now and some of the barriers,

39
00:02:24,259 --> 00:02:26,715

 you know, and, and skepticism 
maybe people had. 
 

40
00:02:26,723 --> 00:02:31,980
He's also been someone who's 
really pioneered open source and

41
00:02:31,980 --> 00:02:35,191

 pushing the boundaries of 
differentiable physics and some 

42
00:02:35,191 --> 00:02:39,650
 of the codes that him and his 
team has developed and now have 

43
00:02:39,650 --> 00:02:43,860
 moved into both generating data
sets like the SuperWing data 
 

44
00:02:43,868 --> 00:02:47,508
set, but also addressing the 
challenge of foundation models 


45
00:02:47,516 --> 00:02:52,987
also from a sort of PDE point of
view in the latest Tadpole 
 

46
00:02:52,995 --> 00:02:56,320
paper. 
So we, we, we kind of go through

47
00:02:56,320 --> 00:03:00,112

 some of those topics around 
those different strategies of 
 

48
00:03:00,120 --> 00:03:04,460
sort of PDE based methods, you 
know, pre training on those 
 

49
00:03:04,468 --> 00:03:07,620
versus doing more of the 
traditional, I guess, building 


50
00:03:07,628 --> 00:03:10,992
out large data sets. 
We talk about the role of 
 

51
00:03:11,000 --> 00:03:12,632
startups and industry and 
academia. 
 

52
00:03:12,640 --> 00:03:16,102
Interesting to get his 
perspectives on that and how all

53
00:03:16,102 --> 00:03:18,932

 three can work together to 
solve some of these problems. 
 

54
00:03:18,940 --> 00:03:22,000
And we also talk about the open 
source, why he's so pro open 
 

55
00:03:22,008 --> 00:03:25,795
source and how we think that can
help progress the, the, the 
 

56
00:03:25,803 --> 00:03:29,885
field along. 
He's, he's actually contributed 

57
00:03:29,885 --> 00:03:34,474
 together with his team, many 
important papers covering not 
 

58
00:03:34,482 --> 00:03:38,458
just on fluids, but also on the,
on the weather side, interesting

59
00:03:38,458 --> 00:03:41,435

 like WeatherBench. 
So I, I put a whole link list of

60
00:03:41,435 --> 00:03:44,284

 papers in the chat notes that 
you can in the show notes that 


61
00:03:44,292 --> 00:03:47,498
you can have a look at. 
But I, I really enjoyed this 
 

62
00:03:47,506 --> 00:03:49,464
conversation. 
He's someone who is very modest 

63
00:03:49,464 --> 00:03:52,222
 in, in what he's done. 
But if you actually read his 
 

64
00:03:52,230 --> 00:03:55,260
papers and look through, you can
see he's being quite influential

65
00:03:55,260 --> 00:03:59,190

 in steering these topics and 
always seems to be one step 
 

66
00:03:59,198 --> 00:04:02,128
ahead of where most of the field
is at. 
 

67
00:04:02,136 --> 00:04:06,420
And so if you look at his 2025 
and 2026 papers, you'll know 
 

68
00:04:06,428 --> 00:04:11,192
what I mean by that. 
So hope you enjoy this episode 


69
00:04:11,200 --> 00:04:15,340
with Professor Nils Thuerey. 
Well, yeah, thanks for thanks 
 

70
00:04:15,348 --> 00:04:18,958
for joining today. 
I, I've been very keen to speak 

71
00:04:18,958 --> 00:04:22,695
 to you because as I've been 
getting up to speed myself over 

72
00:04:22,695 --> 00:04:25,510
 the past years on machine 
learning and differentiable 
 

73
00:04:25,518 --> 00:04:28,142
side. 
Your name is a, a constant, you 

74
00:04:28,142 --> 00:04:32,226
 know, top of the list in terms 
of influential papers, in terms 

75
00:04:32,226 --> 00:04:36,395
 of doing things what seemed to 
be earlier than most others have

76
00:04:36,395 --> 00:04:39,900

 done, you know, setting the 
scene, which suggests to me that

77
00:04:39,900 --> 00:04:44,102

 you have quite a good outlook 
and, and, and thought on it. 
 

78
00:04:44,110 --> 00:04:47,735
So, yeah, thanks very much for 
for for joining. 
 

79
00:04:47,743 --> 00:04:53,200
And yeah, maybe as a starting 
question, you took a diversion. 

80
00:04:53,200 --> 00:04:55,520
 
Well, not diversion, but you 

81
00:04:55,520 --> 00:04:59,400
spent some time doing visual 
 
effects after your PhD and 

82
00:04:59,400 --> 00:05:02,160
postdoc maybe. 
 
What did you do your PhD? 

83
00:05:02,160 --> 00:05:04,920
What was the postdoc? 
 
And then I'm I'm intrigued on on

84
00:05:04,920 --> 00:05:09,320
the movie stuff. 
 
Thanks for the invite. 

85
00:05:09,480 --> 00:05:11,880
First of all, Neil. 
 
Yeah, I'd be happy to tell about

86
00:05:11,880 --> 00:05:15,520
this. 
 
My background, right or over my 

87
00:05:15,520 --> 00:05:18,040
career, I have switched fields a

 couple of times. 

88
00:05:18,040 --> 00:05:23,400
My PhD was in a in a 
 
computational numerics 

89
00:05:23,600 --> 00:05:28,480
computational physics lab, 
 
basically targeting multigrid 

90
00:05:28,520 --> 00:05:32,240
methods and and fluids. 
 
And then I switched to computer 

91
00:05:32,240 --> 00:05:34,360
graphics for a while. 
 
I guess that that actually was 

92
00:05:34,360 --> 00:05:37,080
one of my driving my motivations

 from the start. 

93
00:05:37,280 --> 00:05:39,520
I guess I'm a visual person. 
 
I'd like to see how things 

94
00:05:39,520 --> 00:05:41,720
evolve. 
 
And I think that's that's still 

95
00:05:41,720 --> 00:05:44,240
actually one of the main things 
 that's fascinate me about 

96
00:05:44,240 --> 00:05:47,840
fluids, these swirling motions 

and the things that actually do 

97
00:05:47,840 --> 00:05:53,000
come out once it's working. 
 
So we also worked on computer 

98
00:05:53,000 --> 00:05:55,100
graphics, basically, for 
 
computer-based animation for 

99
00:05:55,100 --> 00:05:58,000
quite a while. 
 
My postdoc at ETH and then 

100
00:05:58,000 --> 00:06:01,200
afterwards also did actually 
 
work in industry for a while 

101
00:06:01,200 --> 00:06:04,840
because I was curious how things

 would actually work out in 

102
00:06:04,840 --> 00:06:06,760
industry in a practical 
 
environment. 

103
00:06:06,760 --> 00:06:08,560
And partially also because I 
 
couldn't really get a good 

104
00:06:09,560 --> 00:06:14,840
faculty position at the time. 
 
So right, not not completely 

105
00:06:14,840 --> 00:06:19,160
voluntary in a while in a way, 

but right, it was good to see in

106
00:06:19,160 --> 00:06:21,400
the end. 
 
I noticed for me, research is 

107
00:06:21,400 --> 00:06:24,320
the more I think the better 
 
field. 

108
00:06:24,760 --> 00:06:28,080
But yes, that's why I did my 
 
postdoc, basically. 

109
00:06:28,080 --> 00:06:32,426
And then I think that if you 10 
 years ago, we started looking 

110
00:06:32,426 --> 00:06:36,920
at machine learning. 
 
So previously we did in this ETH

111
00:06:36,920 --> 00:06:39,880
group what was called 
 
data-driven approaches in 

112
00:06:39,880 --> 00:06:42,520
computer graphics, essentially 

also trying to work with data 

113
00:06:42,520 --> 00:06:47,160
and right, somewhat aligned, but

 we didn't have the tools back 

114
00:06:47,160 --> 00:06:48,880
then. 
 
So it was, it didn't really 

115
00:06:48,880 --> 00:06:51,360
work, to be honest, but we 
 
tried. 

116
00:06:51,360 --> 00:06:54,840
And then when AlphaGo came 
 
along, I thought, oh, this has 

117
00:06:54,840 --> 00:06:57,040
to be something that works. 
 
I didn't understand it at the 

118
00:06:57,040 --> 00:06:59,440
time, but I thought this, 
 
there's got to be something 

119
00:06:59,440 --> 00:07:02,160
there. 
 
And then, yeah, since 10 years, 

120
00:07:02,160 --> 00:07:06,240
I think we've been working on, 

on basically finding these 

121
00:07:06,240 --> 00:07:08,360
applying these techniques to 
 
numerical simulations and fluid 

122
00:07:08,360 --> 00:07:11,760
mechanics specifically. 
 
And I think, in a way, it's a 

123
00:07:12,240 --> 00:07:14,280
good time at the moment. 
 
It's finally starting to work. 

124
00:07:14,400 --> 00:07:19,720
But right, it took a while. 
 
So you then joined TUM and 

125
00:07:19,720 --> 00:07:24,520
you've been there ever since. 
 
That was your your period. 

126
00:07:24,520 --> 00:07:30,560
And so where did you? 
 
What was it like? 

127
00:07:30,760 --> 00:07:33,320
I'm trying to put it into the 
 
picture because it's very hard 

128
00:07:33,320 --> 00:07:38,480
to the people who were working 

on these things before now, you 

129
00:07:38,480 --> 00:07:41,920
know, before ChatGPT, what 
 was
the field like? 

130
00:07:41,920 --> 00:07:44,120
What was the community like? 
 
What was the reception like when

131
00:07:44,120 --> 00:07:47,680
you were bringing up ideas of 
 
sort of data-driven and machine 

132
00:07:47,680 --> 00:07:49,480
learning? 
 
Was there a lot of skepticism 

133
00:07:49,480 --> 00:07:52,000
and push back? 
 
Yes, actually there was. 

134
00:07:52,480 --> 00:07:56,800
So I mean, initially, especially

 in the early phases, all these

135
00:07:56,800 --> 00:08:00,280
AI-based—back then called deep 
learning— 
 techniques were 

136
00:08:01,520 --> 00:08:03,280
focusing on, on learning 
 
descriptors. 

137
00:08:03,280 --> 00:08:05,360
I remember our very first papers

 and graphics, they, they 

138
00:08:05,360 --> 00:08:07,160
basically couldn't synthesize 
 
anything. 

139
00:08:07,160 --> 00:08:09,560
You learned some reduced 
 
representation. 

140
00:08:09,800 --> 00:08:12,840
They told you about what happens

 in high dimensional data, but 

141
00:08:12,840 --> 00:08:14,960
right, it, it couldn't really 
 
generate anything. 

142
00:08:14,960 --> 00:08:18,080
Also, if you think about 
 
AlphaGo, the early success 

143
00:08:18,080 --> 00:08:21,400
stories, it was basically 8 by 8

 with, with three states. 

144
00:08:21,400 --> 00:08:25,720
That's a tiny domain and tiny 
 
state space. 

145
00:08:26,360 --> 00:08:28,240
And back then that was 
 
challenging. 

146
00:08:28,680 --> 00:08:30,240
So in a way, it really couldn't 
 do much. 

147
00:08:30,760 --> 00:08:34,400
And I, I remember quite well 
 
talking to colleagues and 

148
00:08:34,400 --> 00:08:36,080
telling them, look, we're 
 
interested in this deep learning

149
00:08:36,080 --> 00:08:38,679
stuff. 
 
We, we're doing simulations. 

150
00:08:38,679 --> 00:08:41,280
And then some of my colleagues 

literally laugh me in the face 

151
00:08:41,280 --> 00:08:43,558
like, you're doing learning for 
 PDEs. 

152
00:08:43,558 --> 00:08:45,760
What, what could be more 
 
ridiculous? 

153
00:08:46,520 --> 00:08:49,280
If you have a PDE and you can 
 
solve it, why, why, why on earth

154
00:08:49,280 --> 00:08:54,800
would you do anything right 
 
learning based And I think so 

155
00:08:54,800 --> 00:08:57,480
there was a lot of scepticism 
 
for years. 

156
00:08:57,480 --> 00:09:00,600
I always had to defend in the 
 
very first slides of my talks 

157
00:09:01,160 --> 00:09:04,280
that this makes sense at all, at

 least makes sense to consider 

158
00:09:04,280 --> 00:09:07,880
this combination. 
 
And luckily, interestingly, I 

159
00:09:07,880 --> 00:09:11,788
think for the wide 
 community, 
I think the picture really 

160
00:09:11,788 --> 00:09:14,892
changed with ChatGPT. 
I 
 think in computer science 

161
00:09:14,892 --> 00:09:18,640
and and computational fields. 
 
People have been aware earlier, 

162
00:09:18,640 --> 00:09:24,192
but since ChatGPT, it's 
 really
on everybody's plate and in 

163
00:09:24,192 --> 00:09:26,640
every view but before. 
 
But you but you were 

164
00:09:26,640 --> 00:09:29,760
interestingly was looking 
 
through some of the papers and 

165
00:09:29,760 --> 00:09:33,400
you you had slightly earlier on 
 probably some of the first ones

166
00:09:33,400 --> 00:09:36,480
are sort of like in the AIAA, 
 
the deep learning of like some 

167
00:09:36,680 --> 00:09:43,720
RANS of airfoils, and so quite a

 few in the 20 twenties before 

168
00:09:44,120 --> 00:09:47,760
ChatGPT. 
 
So what was it like sort of pre 

169
00:09:47,760 --> 00:09:50,600
and post? 
 
Was it was the issue partly the 

170
00:09:50,600 --> 00:09:54,840
compute side? 
 
Was it the the architectures? 

171
00:09:54,840 --> 00:09:56,680
What? 
 
What was holding back? 

172
00:09:56,760 --> 00:09:59,280
Exactly the the architectures 
 
were a fundamental issue. 

173
00:09:59,680 --> 00:10:03,600
I think the whole infrastructure

 right in the beginning, CNNs 

174
00:10:03,600 --> 00:10:06,280
were were a big breakthrough, 
 
yes, before you had MLPs. 

175
00:10:06,280 --> 00:10:09,768
So CNNs definitely for 
field-like 
 data worked quite 

176
00:10:09,768 --> 00:10:13,104
nicely, but also right the, the,
the shift 
 from basically 

177
00:10:13,104 --> 00:10:15,400
computer graphics to things like
aerodynamics. 
 

178
00:10:15,408 --> 00:10:19,565
The a paper you mentioned, we're
a bit motivated by the 
 

179
00:10:19,573 --> 00:10:22,480
observation that it's, it's 
really tough to synthesize stuff

180
00:10:22,480 --> 00:10:25,210

 in 3D plus time. 
You essentially have 4D fields. 

181
00:10:25,210 --> 00:10:27,240
 
Even if you simplify it and and 

182
00:10:27,240 --> 00:10:31,520
treat regular geometry as 
 
regular grids, this is really 

183
00:10:31,520 --> 00:10:34,680
difficult to do. 
 
It's still challenging these 

184
00:10:34,680 --> 00:10:36,320
days actually with, with neural 
 networks. 

185
00:10:36,560 --> 00:10:39,643
So back then it was right bridge

 out of the question to do 

186
00:10:39,643 --> 00:10:42,880
this. 
And in graphics, you actually 
 

187
00:10:42,888 --> 00:10:45,400
have extremely good 
approximations that aim for the 

188
00:10:45,400 --> 00:10:50,248
 visual aspect of, of things. 
So it, it was extremely tough to

189
00:10:50,248 --> 00:10:52,940

 compete with that. 
And it was interestingly in the 

190
00:10:52,940 --> 00:10:55,705
 computer graphics field also, 
if you cannot directly apply it 

191
00:10:55,705 --> 00:10:59,632
on 
 a scale and, and with a 
fidelity that's fit for movies, 

192
00:10:59,632 --> 00:11:02,200
then your 
 papers typically get
rejected. 

193
00:11:02,760 --> 00:11:04,840
It's really, it's computer 
 
graphics makes sense, but it 

194
00:11:04,840 --> 00:11:09,080
needs to look good in the end. 

So write this 2D proof of 

195
00:11:09,080 --> 00:11:12,040
concept simulations. 
 
Even if they were fairly 

196
00:11:12,040 --> 00:11:13,400
accurate, you you couldn't do 
 
much. 

197
00:11:14,200 --> 00:11:16,760
It had had to be right movie 
 
quality basically. 

198
00:11:16,760 --> 00:11:19,000
And and that just was extremely 
 difficult. 

199
00:11:19,040 --> 00:11:21,880
Only also for niche applications

 like super resolution. 

200
00:11:21,880 --> 00:11:24,080
This was was one that worked 
 
very nicely. 

201
00:11:24,080 --> 00:11:28,560
But they can look at repeating 

structures and very local fields

202
00:11:29,480 --> 00:11:33,760
for synthesis on on larger 
 
scales, it was basically not 

203
00:11:33,760 --> 00:11:35,680
possible. 
 
And that's why we also started 

204
00:11:35,680 --> 00:11:40,200
actually doing these proof-of- 

concept papers like the RANS 

205
00:11:40,200 --> 00:11:45,960
predictions by AI back 
 then 
also these were tiny 

206
00:11:45,960 --> 00:11:49,920
resolutions, right? 
 
128 by 128, basically. 

207
00:11:49,920 --> 00:11:54,560
So very small domains. 
 
And also I remember back then 

208
00:11:54,560 --> 00:11:58,327
actually, I'm glad it got into 

AIAA, but there were also 

209
00:11:58,327 --> 00:11:59,760
critical questions even 
 back 
then. 

210
00:12:00,120 --> 00:12:01,760
Why would you even do this at 
 
all? 

211
00:12:01,760 --> 00:12:04,960
This is crude approximations. 
 
Does this make sense? 

212
00:12:07,840 --> 00:12:11,600
But yeah, basically we noticed 

in in aerodynamics and in 

213
00:12:11,600 --> 00:12:14,600
engineering applications, 
 
there's much more demand for 

214
00:12:14,760 --> 00:12:17,800
getting things right and really 
 going forward converged and 

215
00:12:17,840 --> 00:12:21,200
solutions that are accurate and 
 can be evaluated. 

216
00:12:21,960 --> 00:12:25,720
And there are many problems 
 
there that that are like really 

217
00:12:25,720 --> 00:12:29,440
important basically for for a 
 
variety of industries. 

218
00:12:29,440 --> 00:12:32,960
So we basically switched from 
 
from graphics trying to apply 

219
00:12:32,960 --> 00:12:34,680
these for engineering 
 
applications. 

220
00:12:37,440 --> 00:12:40,880
And where did the, you know, 
 
PhiFlow and the sort of 

221
00:12:40,880 --> 00:12:45,040
differentiable physics, where 
 
did that start to come into the 

222
00:12:45,040 --> 00:12:46,520
grid? 
 
Because that seems to be now 

223
00:12:46,520 --> 00:12:49,840
quite a hot topic. 
 
But you were publishing in 2020 

224
00:12:49,840 --> 00:12:52,760
on these, so where did that 
 
first come from? 

225
00:12:52,760 --> 00:12:53,840
I. 
 
Don't even remember, it seemed 

226
00:12:53,840 --> 00:12:57,520
like a very obvious thing. 
 
We were working on computational

227
00:12:57,560 --> 00:13:02,080
and numerical methods, the the 

networks couldn't really 

228
00:13:02,080 --> 00:13:05,640
generate a whole simulation. 
 
So why not start with the 

229
00:13:05,640 --> 00:13:08,647
simulator, combining it with the

 network and then if you want 

230
00:13:08,647 --> 00:13:12,480
to get a gradient through it 
 
actually need differentiability 

231
00:13:12,480 --> 00:13:17,240
on on the solver side and right 
 the the available ones couldn't

232
00:13:17,240 --> 00:13:19,520
do this. 
 
So we just started experimenting

233
00:13:19,520 --> 00:13:23,546
with our own solvers. 
PhiFlow 
 became a good testing 

234
00:13:23,546 --> 00:13:28,280
ground for very basic methods 
across 
 different APIs. 

235
00:13:28,280 --> 00:13:33,680
Back then, right, there was 
TensorFlow, 
 PyTorch and JAX 

236
00:13:33,680 --> 00:13:37,560
and the whole bandwidth of of 
 
different methods. 

237
00:13:38,720 --> 00:13:42,120
So I think it was good for 
 
experimentation and flexibility.

238
00:13:42,640 --> 00:13:46,680
And ultimately we notice here 
 
it's a certain trade off. 

239
00:13:46,680 --> 00:13:49,600
Flexibility and generality don't

 go too well together. 

240
00:13:50,800 --> 00:13:53,480
And so now we're specialising a 
 bit more. 

241
00:13:53,480 --> 00:13:56,280
But PhiFlow was one of these 
 
early testing grounds. 

242
00:13:56,280 --> 00:14:00,400
But for us, it really seemed 
 
like in a way obvious, obvious 

243
00:14:00,400 --> 00:14:02,000
way to go. 
 
And then later on, we notice, 

244
00:14:02,000 --> 00:14:04,440
yeah, there's actually quite 
 
some work to do and also 

245
00:14:04,440 --> 00:14:06,400
interesting questions. 
 
How do you get gradients through

246
00:14:07,120 --> 00:14:10,880
long chains of solver 
 
operations, which is still a 

247
00:14:10,880 --> 00:14:16,320
topic? 
 
So maybe shifting a little bit 

248
00:14:16,320 --> 00:14:19,640
towards the what's going on 
 
right now. 

249
00:14:19,880 --> 00:14:25,200
I mean, there's 22 fundamental 

questions that I see you've also

250
00:14:25,200 --> 00:14:27,680
been trying to tackle and be 
 
interested to get your thoughts 

251
00:14:27,680 --> 00:14:29,560
on it. 
 
And that is around, I guess the 

252
00:14:29,560 --> 00:14:32,440
foundation model topic. 
 
The idea of I guess it's 

253
00:14:32,440 --> 00:14:34,840
logical, isn't it? 
 
You look at ChatGPT and you see 

254
00:14:34,840 --> 00:14:39,120
it as essentially a data-driven 
 problem that you know, applying

255
00:14:39,120 --> 00:14:42,240
up data, have a decent 
 
architecture and you solve it. 

256
00:14:44,120 --> 00:14:50,760
But you well, two questions. 
 
One was on the idea of the 

257
00:14:50,760 --> 00:14:53,760
ground truth and can you ever 
 
surpass the ground truth? 

258
00:14:53,800 --> 00:14:57,400
And you know, one of your recent

 papers looked at can the model

259
00:14:57,400 --> 00:15:00,520
do better than essentially the 

quality of the training data? 

260
00:15:01,080 --> 00:15:05,800
And the second one is more just 
 a general question on where you

261
00:15:05,800 --> 00:15:08,320
see the likelihood of foundation

 models going. 

262
00:15:08,320 --> 00:15:11,000
How, how realistic is it? 
 
But on the first one, maybe you 

263
00:15:11,000 --> 00:15:13,560
could explain that paper because

 I that really intrigued me 

264
00:15:13,560 --> 00:15:15,920
actually, right? 
 
There was actually a great 

265
00:15:16,120 --> 00:15:18,920
observation by Felix, one of 
 
one of the PhD students from my 

266
00:15:18,920 --> 00:15:22,280
group, that he noticed in some 

situations. 

267
00:15:22,280 --> 00:15:25,360
He also looked at basic 
 
combinations of solvers and 

268
00:15:25,360 --> 00:15:29,560
networks and that's the networks

 really seem to give extremely 

269
00:15:29,560 --> 00:15:33,240
good situations in certain 
 
situations extremely good 

270
00:15:33,960 --> 00:15:37,680
results that seem to be even 
 
better than the the training 

271
00:15:37,680 --> 00:15:39,640
data we essentially used for 
 
these cases. 

272
00:15:40,200 --> 00:15:46,080
And looking into it a bit more, 
 it's unfortunately not a topic 

273
00:15:46,080 --> 00:15:49,920
that I think can be generally 
 
applied and when analysing it in

274
00:15:49,920 --> 00:15:53,160
terms of numerical errors and, 

and how these pieces fit 

275
00:15:53,160 --> 00:15:55,720
together. 
 
It's basically a combination of 

276
00:15:55,720 --> 00:15:59,520
the priors imposed by neural 
 
networks in terms of smoothness,

277
00:15:59,520 --> 00:16:03,730
in terms of right, being able to

 over sharp more reproduce and 

278
00:16:03,730 --> 00:16:07,760
be trained for producing certain

 parts of the solutions that 

279
00:16:07,760 --> 00:16:13,280
benefits the the outcome. 
 
So in a way, if you can leverage

280
00:16:13,280 --> 00:16:16,560
or if you understand your 
 
solutions enough to leverage 

281
00:16:16,680 --> 00:16:21,440
these particular behaviours of 

neural networks, then basically 

282
00:16:21,440 --> 00:16:25,200
the outputs can be better than 

what you're trained with. 

283
00:16:25,880 --> 00:16:29,600
But it's really this interplay 

of, of what neural networks do 

284
00:16:29,680 --> 00:16:32,640
as essentially also numerical 
 
methods to compute a solution 

285
00:16:33,160 --> 00:16:36,600
and what, what's in your data 
 
and what you like to get out. 

286
00:16:37,120 --> 00:16:40,760
So it also it raises some 
 
interesting questions on how to 

287
00:16:40,760 --> 00:16:44,040
evaluate it because if 
 
especially in the numerical 

288
00:16:44,040 --> 00:16:46,488
field, we're trusting the ground

 truth, there are also 

289
00:16:46,488 --> 00:16:48,776
typically approximation errors 
in there. 
 

290
00:16:48,784 --> 00:16:53,675
And it it can actually happen 
that solutions output by the 
 

291
00:16:53,683 --> 00:16:58,255
network might be even better if 
you consider other solves or or 

292
00:16:58,255 --> 00:17:01,294
 other ways to approximate the 
same problems. 
 

293
00:17:01,302 --> 00:17:06,866
But I think it's we also for 
this paper thought about it. 
 

294
00:17:06,874 --> 00:17:10,382
It is very difficult to really 
do this in a targeted way. 
 

295
00:17:10,390 --> 00:17:13,615
So we had some new cases where 
it worked for some basic 
 

296
00:17:13,623 --> 00:17:16,160
advection-diffusion and Burgers.
Burgers equation was a good 
 

297
00:17:16,167 --> 00:17:20,800
example actually because the the
shock representation and 
 

298
00:17:20,808 --> 00:17:25,640
smoothness—Burgers equation is 
non-trivial, but 
 seemed to be 

299
00:17:25,640 --> 00:17:30,200
a very good showcase for this. 

Directly applying it to other 

300
00:17:30,200 --> 00:17:34,000
scenarios is is challenging so 

we haven't unfortunately we 

301
00:17:34,000 --> 00:17:37,200
haven't gotten around to finding

 a good way to leverage this 

302
00:17:37,200 --> 00:17:40,720
for real-world Navier–Stokes 
 
problems or so. 

303
00:17:40,720 --> 00:17:47,310
But what about 
 foundation 
models in general? 
 

304
00:17:47,318 --> 00:17:49,067
Like, well, how was your group 
tackling this? 
 

305
00:17:49,075 --> 00:17:51,460
What? 
What's your strategy in a sense,

306
00:17:51,460 --> 00:17:55,375

 you know where, where, where 
do you see this going? 
 

307
00:17:55,383 --> 00:17:58,720
Recently we switched pretty much
from these differentiable 

308
00:17:58,720 --> 00:18:02,608
solvers 
 to foundation models. 
I was quite skeptical of it 
 

309
00:18:02,616 --> 00:18:06,232
for quite a while. 
I think it's also there's also 


310
00:18:06,240 --> 00:18:08,960
just quite some skepticism 
talking to other people in the 


311
00:18:08,968 --> 00:18:11,321
field with foundation models. 
I think partly it's because of 


312
00:18:11,329 --> 00:18:13,514
the name, right? 
Foundation model sounds super 
 

313
00:18:13,522 --> 00:18:19,072
general and I'm sure you and and
most people know these fields of

314
00:18:19,072 --> 00:18:22,688

 PDEs are extremely diverse. 
The thought of having one model 

315
00:18:22,688 --> 00:18:27,400
 that just is everything 
out-of-the-box is yeah, but is 


316
00:18:27,408 --> 00:18:33,692
is a bit unbelievable and I 
think also still out of quite a 

317
00:18:33,692 --> 00:18:37,132
 bit out of reach. 
Nonetheless, if you just right 


318
00:18:37,140 --> 00:18:40,909
foundation models also by now I 
think it's actually not so well 

319
00:18:40,909 --> 00:18:42,881
 defined. 
People use it in in all various 

320
00:18:42,881 --> 00:18:46,362
 forms and where I think it 
makes a lot of sense and in a 

321
00:18:46,362 --> 00:18:49,372
way it's 
 also the direction we
are typically tech typically 
 

322
00:18:49,380 --> 00:18:53,900
tackling is to see it as a model
that's pre trained on on 
 

323
00:18:53,908 --> 00:18:58,224
something and then applicable to
other tasks, other problems. 
 

324
00:18:58,232 --> 00:19:03,580
So there's just some disparity 
between training data and the 
 

325
00:19:03,588 --> 00:19:06,300
actual application we saw. 
We want models to generalise. 
 

326
00:19:06,308 --> 00:19:09,360
Also, generalisation is not a 
well defined topic, right? 
 

327
00:19:09,368 --> 00:19:12,906
It's in distribution is clear, 
but how far out of distribution 

328
00:19:12,906 --> 00:19:14,820
 you need to go to have proper 
generalization. 
 

329
00:19:14,828 --> 00:19:18,760
It's completely open. 
But in the past that's been 
 

330
00:19:18,768 --> 00:19:20,812
really tricky. 
Anyway, it's a classical 
 

331
00:19:20,820 --> 00:19:23,480
transfer learning problem. 
Now, I think with these 
 

332
00:19:23,488 --> 00:19:25,785
foundation model tools borrowing
a lot from large language 

333
00:19:25,785 --> 00:19:29,208
models, 
 we can actually do 
this if we pre train on on one 

334
00:19:29,208 --> 00:19:31,520
part and 
 then apply it to 
something else. 

335
00:19:31,760 --> 00:19:38,200
And how else how how different 

is is up for very specific to 

336
00:19:38,200 --> 00:19:40,920
problems and and application 
 
domains. 

337
00:19:41,320 --> 00:19:43,680
But that seems to work. 
 
And that's specifically what 

338
00:19:43,680 --> 00:19:48,160
what I'm quite excited about is 
 that we notice you can actually

339
00:19:48,160 --> 00:19:51,320
pre train on very cheaply 
 
generated data, offer us 

340
00:19:52,200 --> 00:19:54,960
training foundation models in in

 3D plus time. 

341
00:19:55,080 --> 00:19:59,760
The amount of data needed seems 
 very scary and also somewhat 

342
00:19:59,760 --> 00:20:03,160
impractical given given the 
 
current cost of of doing these 

343
00:20:03,160 --> 00:20:07,440
training runs. 
 
But we noticed that in a way you

344
00:20:07,440 --> 00:20:09,960
don't really need these huge 
 
amounts of data. 

345
00:20:09,960 --> 00:20:11,840
And I think that that really 
 
changes the picture, that 

346
00:20:11,840 --> 00:20:14,400
changes how we can approach 
 
foundation model training. 

347
00:20:15,240 --> 00:20:19,120
And I think in a way it makes it

 much more attractive as as a 

348
00:20:19,120 --> 00:20:21,600
starting point. 
 
Could you maybe go into a little

349
00:20:21,600 --> 00:20:25,120
bit more detail when you when 
 
you talk about not needing as as

350
00:20:25,120 --> 00:20:27,960
much data or all the high- 
 
quality data? 

351
00:20:28,680 --> 00:20:31,840
Right. 
 
So you also thought about this 

352
00:20:31,840 --> 00:20:35,000
quite a bit, right, as you 
 
recent paper discussing the the 

353
00:20:35,000 --> 00:20:39,680
scaling of of of data and and 
 
computer requirements for fluids

354
00:20:40,240 --> 00:20:43,160
is is a good starting point to 

think about it. 

355
00:20:43,760 --> 00:20:48,280
What we noticed in this Tadpole 
 work, we we called it typically

356
00:20:48,520 --> 00:20:54,000
was basically that if you have a

 good solver. 

357
00:20:54,000 --> 00:20:59,440
So it is actually also based in 
 a way on on previous work on on

358
00:20:59,440 --> 00:21:03,320
this APEBench benchmark with 
 
very synthetic data. 

359
00:21:03,320 --> 00:21:07,200
So if you and if you take the 
 
right solver, these spectral 

360
00:21:07,200 --> 00:21:12,080
solvers, then you can extremely 
 efficiently compute super good 

361
00:21:12,280 --> 00:21:15,143
solutions for a class of basic 

PDEs. 

362
00:21:15,143 --> 00:21:19,035
For advection-diffusion, these 
ETDRK solvers—in a way, 
 

363
00:21:19,043 --> 00:21:23,002
they're really perfect. 
So also doing the learning on 
 

364
00:21:23,010 --> 00:21:26,104
this on this level does not 
really make sense because the 
 

365
00:21:26,112 --> 00:21:29,965
solvers are so good, the errors 
are so low, you can prove that 


366
00:21:29,973 --> 00:21:34,715
they are basically as accurate 
as it gets for for these basic 


367
00:21:34,723 --> 00:21:37,580
solutions. 
So we can basically generate 
 

368
00:21:37,588 --> 00:21:42,026
this data extremely quickly and 
very broadly if you vary the 
 

369
00:21:42,034 --> 00:21:45,875
parameters and we notice that 
right, you can then pre train a 

370
00:21:45,875 --> 00:21:50,650
 model with this just generating
a lot of different solutions for

371
00:21:50,650 --> 00:21:55,698

 for varying parameter ranges 
for these different PDEs, from 


372
00:21:55,706 --> 00:22:01,640
advection-diffusion to transport
to some higher-order terms with 

373
00:22:01,640 --> 00:22:04,320
 Kuramoto–Sivashinsky and some 
some polynomial 
 chemical 

374
00:22:04,320 --> 00:22:09,360
reaction like constraints. 
 
And basically pre training this 

375
00:22:09,360 --> 00:22:15,480
was on this very broad class of 
 very basic synthetic solutions.

376
00:22:15,960 --> 00:22:19,720
Nonetheless, it's actually not 

not a very easy task and 

377
00:22:19,720 --> 00:22:23,920
benefits or seems to benefit in 
 in quite a bit of downstream 

378
00:22:23,920 --> 00:22:26,400
tasks. 
 
And the nice thing then is 

379
00:22:26,400 --> 00:22:28,440
right, these solutions to 
 
these synthetic problems are 

380
00:22:28,440 --> 00:22:30,760
really not interested in or 
 
interesting in a way you can 

381
00:22:30,760 --> 00:22:33,880
compute them almost as quickly 

with the solver from from 

382
00:22:33,880 --> 00:22:36,160
scratch basically. 
 
So you can train with it, you 

383
00:22:36,160 --> 00:22:38,840
can run it through a model to 
 
get a gradient basically, but 

384
00:22:38,840 --> 00:22:41,200
then you can throw it away. 
 
So it's basically just running 

385
00:22:41,200 --> 00:22:44,360
the solver alongside. 
You 
 typically have a node with

386
00:22:44,360 --> 00:22:46,360
four GPUs. 
 
So one GPU does the data 

387
00:22:46,360 --> 00:22:49,520
generation, the other 3 train 
 
and just generates data on the 

388
00:22:49,520 --> 00:22:52,920
fly, throws it away right after 
 it produces more data. 

389
00:22:53,120 --> 00:22:56,680
But the nice thing is it's 
 
produced on the fly and it's 

390
00:22:56,680 --> 00:22:58,960
really just a continuous stream 
 of data to train with. 

391
00:22:59,880 --> 00:23:03,800
You don't, you don't have 
 
limitations on the bandwidth 

392
00:23:03,800 --> 00:23:08,194
side on the hardware disk space 
 to manage 10s of terabytes of 

393
00:23:08,194 --> 00:23:12,270
or hundreds of terabytes to 
upload 
 to HPC centres for for 

394
00:23:12,270 --> 00:23:14,736
anyone who's tried this, it's, 
it's 
 very painful actually to 

395
00:23:14,736 --> 00:23:20,080
do this and it's not trivial. 
 
So right, having this online 

396
00:23:20,080 --> 00:23:24,360
data generation with the 
 
training is is really quite neat

397
00:23:24,360 --> 00:23:27,000
and worked quite well in our 
 
experiments. 

398
00:23:27,520 --> 00:23:31,480
But where? 
 
Do you, how do you see that 

399
00:23:31,920 --> 00:23:33,720
progressing? 
 
I mean, I think this has often 

400
00:23:33,720 --> 00:23:39,880
been my challenges and you've 
 
been on both sides of it. 

401
00:23:39,880 --> 00:23:44,960
You know, one side of it is, 
 
let's take realistic in well, in

402
00:23:45,280 --> 00:23:47,840
if we assume that the model 
 
should work for industrial scale

403
00:23:47,840 --> 00:23:53,160
problems, then you know, the the

 idea is you generate data like

404
00:23:53,160 --> 00:23:56,680
you did with the SuperWing or, 

or now this HiLiftAeroML or any 

405
00:23:56,680 --> 00:23:59,080
or company, you know, 
 
generating data. 

406
00:23:59,080 --> 00:24:01,644
And then you you train on that. 
 

407
00:24:01,652 --> 00:24:04,046
And you know that is one 
argument. 
 

408
00:24:04,054 --> 00:24:07,308
And the other one is you're 
starting from the far left side 

409
00:24:07,308 --> 00:24:11,770
 of basic PDEs. 
But where do you see how can 
 

410
00:24:11,778 --> 00:24:16,480
they go in the middle? 
The basic PDEs are—can you 
 go 

411
00:24:16,480 --> 00:24:19,115
to Navier–Stokes? 
How does that scale to more 
 

412
00:24:19,123 --> 00:24:20,240
complexity? 
Yeah. 
 

413
00:24:20,248 --> 00:24:23,700
So what we basically notice is 
that even if we pre trained also

414
00:24:23,700 --> 00:24:26,875

 one of the one of the tricks 
basically that I didn't mention 

415
00:24:26,875 --> 00:24:30,295
 just now was also that we don't
train for a time evolution. 
 

416
00:24:30,303 --> 00:24:33,855
It's really more of a of a 
latent space representation that

417
00:24:33,855 --> 00:24:37,058

 is learned for for single 
spatial fields. 
 

418
00:24:37,066 --> 00:24:42,826
So there's no time in the pre 
training and then the the basic 

419
00:24:42,826 --> 00:24:46,722
 architecture is basically so so
general that you can just put 
 

420
00:24:46,730 --> 00:24:49,160
time on top of that, right. 
So we learn in this latent 
 

421
00:24:49,168 --> 00:24:52,616
space, one step to the next 
seems to work nicely. 
 

422
00:24:52,624 --> 00:24:57,544
And also time in a way is also 
specific application already I 


423
00:24:57,552 --> 00:25:00,790
would argue, right. 
There are steady state problems 

424
00:25:00,790 --> 00:25:04,360
 that that essentially go for an
equilibrium already or an 
 

425
00:25:04,368 --> 00:25:07,502
average solution where time does
not really the the evolution of 

426
00:25:07,502 --> 00:25:09,040
 time doesn't really play a 
role. 

427
00:25:09,040 --> 00:25:10,880
It's more a mapping of some 
 
initial conditions in the 

428
00:25:10,880 --> 00:25:13,040
geometry or so to the steady 
 
state. 

429
00:25:13,920 --> 00:25:17,560
You might need time and other 
 
solutions, the occasions where 

430
00:25:17,560 --> 00:25:19,320
you really want to resolve what 
 happens over time. 

431
00:25:19,320 --> 00:25:25,044
But in a way taking that out of 
 pre training already makes 

432
00:25:25,044 --> 00:25:27,920
sense in in retrospect I think. 
 

433
00:25:27,928 --> 00:25:32,440
And then with the synthetic 
PDEs, probably you also might 

434
00:25:32,440 --> 00:25:37,520
not 
 get the out-of-the-box 
best accuracy for right high 

435
00:25:37,520 --> 00:25:39,440
accuracy 
 aerodynamic 
predictions. 

436
00:25:40,000 --> 00:25:42,240
But I think it, it seems to be a

 very good starting point also 

437
00:25:42,240 --> 00:25:45,560
for Navier–Stokes. 
 
So what I see as a practical 

438
00:25:46,160 --> 00:25:49,160
kind of approach in in the 
 
future would be to pre train 

439
00:25:49,160 --> 00:25:51,800
some very generic models as a 
 
starting point. 

440
00:25:51,800 --> 00:25:55,210
And then step by step was first 
 further pre training stages 

441
00:25:55,210 --> 00:25:59,880
move towards more application 
 
specific models, aerodynamics, 

442
00:25:59,880 --> 00:26:03,520
maybe other one, maybe aero 
 
acoustics, other one, maybe 

443
00:26:03,600 --> 00:26:08,240
structural structural problems 

or branch off different 

444
00:26:08,240 --> 00:26:10,560
directions. 
 
But pre training doesn't need to

445
00:26:10,560 --> 00:26:13,560
be restricted to these canonical

 PDEs, but could go in certain 

446
00:26:13,560 --> 00:26:16,720
stages and then at a company you

 might for a certain use case 

447
00:26:16,720 --> 00:26:21,840
then have 10 final models to 
 
fine tune the the last one for a

448
00:26:21,840 --> 00:26:25,560
certain application. 
 
I just think just this this 

449
00:26:25,560 --> 00:26:29,440
outlook is basically not 
 
starting from scratch but pre 

450
00:26:29,440 --> 00:26:32,600
training and then downstream 
 
needing less data. 

451
00:26:33,600 --> 00:26:35,984
I think it's potentially a good.

 

452
00:26:35,992 --> 00:26:41,890
So do you so because of that, do
you feel that there can be, I 
 

453
00:26:41,898 --> 00:26:44,965
mean how general can these 
models go do you think? 
 

454
00:26:44,973 --> 00:26:47,548
I mean if you if you put a 
looking glass. 
 

455
00:26:47,556 --> 00:26:52,280
Admittedly, that's an open 
question and in in a way also we

456
00:26:52,280 --> 00:26:57,155

 were surprised that the models
can learn something useful from 

457
00:26:57,155 --> 00:27:00,470
 these admittedly relatively 
useless synthetic solutions. 
 

458
00:27:00,478 --> 00:27:04,664
But it's a really interesting 
question to analyse what they 
 

459
00:27:04,672 --> 00:27:06,996
actually learned. 
So there are we had we had some 

460
00:27:06,996 --> 00:27:09,180
 hypothesis. 
Basically, do they really just 


461
00:27:09,188 --> 00:27:13,220
learn or less random abstract 
fields or physical structures 
 

462
00:27:13,228 --> 00:27:16,815
like the key modes? 
Or do they learn differential 
 

463
00:27:16,823 --> 00:27:19,690
operators? 
Also, these PDEs are built from 

464
00:27:19,690 --> 00:27:23,766
certain 
 derivatives and it's 
difficult to disentangle. 
 

465
00:27:23,774 --> 00:27:28,334
But for example, the initial 
conditions are relatively easy 


466
00:27:28,342 --> 00:27:30,490
to test. 
So the PDEs are initialized. 
 

467
00:27:30,498 --> 00:27:33,780
If you only train with these 
random fields, the models 
 

468
00:27:33,788 --> 00:27:38,205
actually also learn something, 
but it clearly does less well on

469
00:27:38,205 --> 00:27:41,110

 Navier–Stokes than if you 
pre-train with the PDEs. 
 

470
00:27:41,118 --> 00:27:44,084
So basically a transformation of
initial conditions into actual 


471
00:27:44,092 --> 00:27:46,675
states of a PDE seems to be 
beneficial. 
 

472
00:27:46,683 --> 00:27:48,914
I think this is just the first 
step. 
 

473
00:27:48,922 --> 00:27:53,260
It's it's actually not right 
right now difficult to 
 

474
00:27:53,268 --> 00:27:56,095
disentangle what they learned. 
And intuitively, I think the 
 

475
00:27:56,103 --> 00:27:59,695
network network currently don't 
has any reason to clearly 
 

476
00:27:59,703 --> 00:28:02,525
disentangled modes or 
differential operators or so 
 

477
00:28:02,533 --> 00:28:06,214
probably just mixes the most 
prominent structures that that 


478
00:28:06,222 --> 00:28:10,420
come up. 
But in a way also most of the 
 

479
00:28:10,428 --> 00:28:13,816
problems we're dealing with in 
in engineering and in real world

480
00:28:13,816 --> 00:28:16,480

 applications are typically 
built from certain like building

481
00:28:16,480 --> 00:28:20,116

 blocks to differential 
equations in a way have a have a

482
00:28:20,116 --> 00:28:23,490
very 
 limited vocabulary. 
So my guess is that the 
 

483
00:28:23,498 --> 00:28:28,264
structures that come from these 
actually in a way can be pre 
 

484
00:28:28,272 --> 00:28:33,000
trended and reused. 
Yeah. 
 

485
00:28:33,008 --> 00:28:41,505
And I guess the the other, I 
mean, so maybe to put it in a 
 

486
00:28:41,513 --> 00:28:45,305
nutshell, in terms of your 
research and your thinking, do 


487
00:28:45,313 --> 00:28:51,160
you think it is more better to 
go down this sort of PDE route 


488
00:28:51,168 --> 00:28:56,070
with an integrated solver in the
loop rather than coming from the

489
00:28:56,070 --> 00:29:00,142

 other direction, which is just
generating lots and lots of, you

490
00:29:00,142 --> 00:29:05,032

 know, if if we're trying to 
ultimately get to predict a 
 

491
00:29:05,040 --> 00:29:09,060
wing, practically speaking, do 
you think it's better just to 
 

492
00:29:09,068 --> 00:29:12,920
simulate a million wings in all 
different conditions or to sort 

493
00:29:12,920 --> 00:29:17,880
 of build up from the ground 
from a sort of PDE based that 

494
00:29:17,880 --> 00:29:21,770
may be 
 pre trained? 
And then you go to more complex,

495
00:29:21,770 --> 00:29:25,040

 more complex, because they 
still do feel like different 
 

496
00:29:25,048 --> 00:29:29,489
directions to go, right? 
The PDEs definitely aim for 
 

497
00:29:29,497 --> 00:29:34,556
generality in a way I think also
based on previous work, if your 

498
00:29:34,556 --> 00:29:37,360
 data set is large enough, if 
you have enough data to really 

499
00:29:37,360 --> 00:29:40,770
train 
 the large scale model 
specifically for the task you're

500
00:29:40,770 --> 00:29:45,885

 interested in, I don't think 
you actually need this this much

501
00:29:45,885 --> 00:29:53,580

 more generic starting points. 
So if you can do this, I think 


502
00:29:53,588 --> 00:29:58,040
probably it's not not necessary.
It's really more motivated by 
 

503
00:29:58,048 --> 00:30:00,962
the practical constraints. 
I think at the moment that's 
 

504
00:30:00,970 --> 00:30:05,125
especially in in 3D for large 
scale PDEs, it's actually 
 

505
00:30:05,133 --> 00:30:09,467
difficult to gather these really
large collective, Yeah. 
 

506
00:30:09,475 --> 00:30:14,278
I, I'm, I'm, I'm kind of 
intrigued on that point just 
 

507
00:30:14,286 --> 00:30:17,488
because, you know, as that fluid
intelligence paper tried to 
 

508
00:30:17,496 --> 00:30:24,320
point out, there is a certain 
economic challenge, which is 
 

509
00:30:24,328 --> 00:30:30,624
feels to be different than 
ChatGPT where, you know, 
 

510
00:30:30,632 --> 00:30:36,287
ChatGPT's cost does not really 
include the data—the token 
 

511
00:30:36,295 --> 00:30:40,550
cost per se. 
And therefore it's all on 
 the 

512
00:30:40,550 --> 00:30:45,776
training side. 
And I, I still, I still 
 

513
00:30:45,784 --> 00:30:48,578
fundamentally wonder, is it 
massively inefficient? 
 

514
00:30:48,586 --> 00:30:52,780
Because someone said to me, 
well, if you run cases, maybe 
 

515
00:30:52,788 --> 00:30:58,070
I'll pose this question to you. 
So someone said, OK, so if you 


516
00:30:58,078 --> 00:31:04,800
need to generate 2,000,000 cases
in order to fully scope out what

517
00:31:04,800 --> 00:31:08,860

 is a reasonable definition of 
problems within fluids, So all 


518
00:31:08,868 --> 00:31:12,944
possible planes and cars and 
data centers and all the rest of

519
00:31:12,944 --> 00:31:16,425

 it, and then you use all that 
to train a model. 
 

520
00:31:16,433 --> 00:31:19,708
Haven't you already got the 
solution from CFD for most of 
 

521
00:31:19,716 --> 00:31:21,875
these problems anyway? 
Because you probably run it 
 

522
00:31:21,883 --> 00:31:25,599
like, what? 
Where is it just a, you know, 
 

523
00:31:25,607 --> 00:31:29,048
you're almost so massively in 
distribution that you almost 
 

524
00:31:29,056 --> 00:31:31,574
could just map that solution to 
a new. 
 

525
00:31:31,582 --> 00:31:36,030
Yeah. 
But you're right that that that 

526
00:31:36,030 --> 00:31:40,082
 might happen. 
So 2 million seems seems a 
 

527
00:31:40,090 --> 00:31:44,260
little bit scarce for for actual
geometry aside. 
 

528
00:31:44,268 --> 00:31:48,598
But if if you could pretend 
enough of them, I guess you 
 

529
00:31:48,606 --> 00:31:51,786
could hope that any new one you 
run would actually just hit 
 

530
00:31:51,794 --> 00:31:54,550
close to one of the existing 
samples basically. 
 

531
00:31:54,558 --> 00:31:57,710
Then I think a practical problem
would still be that essentially 

532
00:31:57,710 --> 00:31:59,990
 you need to look up this data 
somehow, right? 
 

533
00:31:59,998 --> 00:32:03,346
So these two all right, millions
of simulations would probably be

534
00:32:03,346 --> 00:32:08,394

 I don't know how how many 
petabytes or gigantic lots of 
 

535
00:32:08,402 --> 00:32:10,870
storage. 
So it could be beneficial to pre

536
00:32:10,870 --> 00:32:13,770

 train just to have a reduced 
compressed representation to be 

537
00:32:13,770 --> 00:32:17,810
 able to efficiently look up 
these data points and maybe also

538
00:32:17,810 --> 00:32:20,166

 give give some smoothness in 
between, right. 
 

539
00:32:20,174 --> 00:32:23,998
That's even if your exact 
geometry is not in there, you 
 

540
00:32:24,006 --> 00:32:28,002
get in in between solutions. 
So it, it could actually still 


541
00:32:28,010 --> 00:32:31,680
be beneficial from a from a 
practical standpoint. 
 

542
00:32:31,688 --> 00:32:39,514
But yeah, given how difficult it
is and, and the amount of data 


543
00:32:39,522 --> 00:32:43,430
involved with these online 
generated data sets, it seems 
 

544
00:32:43,438 --> 00:32:47,763
that you can actually get away 
with much less data in a way. 
 

545
00:32:47,771 --> 00:32:52,888
The the hope would be that you 
do this as a first step and then

546
00:32:52,888 --> 00:32:58,275

 right maybe you can get away 
with 1,000,000 off for CFD 
 

547
00:32:58,283 --> 00:33:01,720
simulations or significantly 
less at least to to get to the 


548
00:33:01,728 --> 00:33:04,660
same level of accuracy. 
Practically in your maybe one of

549
00:33:04,660 --> 00:33:09,577

 the topics that also interests
people and be interested to see 

550
00:33:09,577 --> 00:33:13,224
 how from a coding point of view
you've seen this progress. 
 

551
00:33:13,232 --> 00:33:17,615
So you raise the point of 
training the data, sorry, 
 

552
00:33:17,623 --> 00:33:20,908
generating the data, throwing it
away because you're ultimately 


553
00:33:20,916 --> 00:33:23,362
using it. 
So maybe you could talk a little

554
00:33:23,362 --> 00:33:26,116

 bit more what you saw at the 
challenges from an 
 

555
00:33:26,124 --> 00:33:28,950
implementation from a coding 
point of view, because that 
 

556
00:33:28,958 --> 00:33:31,869
that's I guess also part of the 
problem if everybody's 
 

557
00:33:31,877 --> 00:33:35,224
generating these hundreds of 
thousands of cases offline and 


558
00:33:35,232 --> 00:33:39,600
then trying to bring into a 
model that also seems quite 
 

559
00:33:39,608 --> 00:33:44,340
inefficient, so. 
The infrastructure work is quite

560
00:33:44,340 --> 00:33:48,488

 substantial and I think it's 
getting better also largely 
 

561
00:33:48,496 --> 00:33:54,908
thanks to all the methodology is
converging on this LLM 
 

562
00:33:54,916 --> 00:34:00,520
transformer-style processing, 
which probably is is one of the 

563
00:34:00,520 --> 00:34:03,861
 reasons why now finally we're 
we're in a stage where it starts

564
00:34:03,861 --> 00:34:07,900

 to work because also we have 
the the tools and outlooks for 

565
00:34:07,900 --> 00:34:11,667
how 
 to scale things up to 
these larger resolutions. 
 

566
00:34:11,675 --> 00:34:16,610
It's nonetheless still quite 
tricky in practice. 
 

567
00:34:16,619 --> 00:34:21,340
So also in my group took quite a
while to to figure out how to 
 

568
00:34:21,348 --> 00:34:24,070
put these pieces together. 
And there are numerous caveats 


569
00:34:24,078 --> 00:34:26,580
and, and things that can go 
wrong. 
 

570
00:34:26,588 --> 00:34:31,445
So actually, admittedly, one of 
the things we're still fighting 

571
00:34:31,445 --> 00:34:34,860
 with is the non linear scaling.
So even with Transformers, which

572
00:34:34,860 --> 00:34:37,476

 which are demonstrated to 
scale up to billions of 

573
00:34:37,476 --> 00:34:40,719
parameters, if 
 you take one 
model and you just try to 

574
00:34:40,719 --> 00:34:44,464
increase the size of of 
 layers
and the overall capacity, it's 

575
00:34:44,464 --> 00:34:46,840
highly non linear. 
 
So it's not guaranteed to to 

576
00:34:46,840 --> 00:34:51,280
work if you rerun this, just 
 
longer with an increased size. 

577
00:34:51,600 --> 00:34:55,960
Most likely at least need to 
 
adjust the learning rate, the 

578
00:34:55,960 --> 00:34:59,400
additional hyper parameters like

 smoothing of of network states

579
00:34:59,880 --> 00:35:02,800
over time. 
 
How to actually adjust how to 

580
00:35:02,800 --> 00:35:05,570
scale up the size of the network

 in terms of this embedding 

581
00:35:05,570 --> 00:35:08,440
space that the Transformers have
or 
 the the weights for the 

582
00:35:08,440 --> 00:35:12,560
additional components that 
 
unfortunately or it seems it's 

583
00:35:12,560 --> 00:35:17,640
necessary to revisit for, for 
 
every new case again and also 

584
00:35:17,640 --> 00:35:20,080
for the PDE case, 
 
unfortunately. 

585
00:35:21,400 --> 00:35:26,920
So it's still not not right, not

 not really trivial to. 

586
00:35:27,880 --> 00:35:30,240
Do this. 
 
And then right, also just 

587
00:35:31,360 --> 00:35:33,800
ability, we fought quite a bit 

just with with the data 

588
00:35:33,800 --> 00:35:37,520
management. 
 
So we rely on these official 

589
00:35:37,520 --> 00:35:41,080
compute infrastructures here 
 
from the very end. 

590
00:35:41,080 --> 00:35:43,720
And in Germany we now have 
 
supercomputers with a fair 

591
00:35:43,720 --> 00:35:46,520
number of GPUs. 
 
But your someone need to get the

592
00:35:46,520 --> 00:35:49,200
data over there and then you can

 just store it on on temporary 

593
00:35:49,200 --> 00:35:51,480
drives. 
 
And if you don't pay attention 

594
00:35:51,480 --> 00:35:55,240
then suddenly your trainer data 
 has gone if you don't train 

595
00:35:55,240 --> 00:35:58,760
often enough and issues issues 

like this basically. 

596
00:35:58,800 --> 00:36:02,240
But how did you solve that with 
 some of the just wanted to 

597
00:36:02,240 --> 00:36:06,640
double click on the you know one

 GPU to generate the data, 3 to

598
00:36:06,640 --> 00:36:07,440
train. 
 
How? 

599
00:36:07,600 --> 00:36:11,760
How were you getting around the 
 passing the data between? 

600
00:36:13,560 --> 00:36:16,034
Yeah, I'm just trying to learn a

 little bit more how you 

601
00:36:16,034 --> 00:36:19,040
approach that that topic of the 
online 
 training. 

602
00:36:19,080 --> 00:36:21,720
With the with the online 
 
training, so right, the classic 

603
00:36:21,720 --> 00:36:24,240
approach would be right there. 

You, you have your huge data 

604
00:36:24,240 --> 00:36:28,480
sets on disk and eventually if 

you don't want to overfit to 

605
00:36:28,480 --> 00:36:30,960
what fits into memory, you 
 
somehow need to get it from the 

606
00:36:30,960 --> 00:36:36,396
disk, get it to the GPUs 
 with 
this large data set that that 

607
00:36:36,396 --> 00:36:41,368
becomes a bottleneck, at 
 least
for right, actually for even if 

608
00:36:41,368 --> 00:36:44,800
you've hundreds of 
 millions of
parameters on HPC systems, I 

609
00:36:44,800 --> 00:36:48,379
think we notice 
 especially and
the the hardware, the 

610
00:36:48,379 --> 00:36:51,760
interconnects are good, but 
 
not fast enough to really get 

611
00:36:51,760 --> 00:36:54,000
the data in quickly enough for 

training. 

612
00:36:54,280 --> 00:36:55,600
So the standard set up would be 
 right. 

613
00:36:55,600 --> 00:36:58,626
You have your GPUs; they 
 need 
to load the data from this, then

614
00:36:58,626 --> 00:37:01,090
shuffle it through the 
 network
to get a gradient and they need 

615
00:37:01,090 --> 00:37:02,920
to load the next 
 sample 
basically. 

616
00:37:03,400 --> 00:37:07,291
And for this online training, we

 are already basically forced 

617
00:37:07,291 --> 00:37:10,440
to deal with a separate process 
 that does the generation. 

618
00:37:10,880 --> 00:37:12,600
So it's natural to put a buffer 
 in between. 

619
00:37:12,800 --> 00:37:16,080
So you basically have a, a 
 
buffer that is filled up by the 

620
00:37:16,080 --> 00:37:20,840
simulator and then the training 
 sets just pull the data from 

621
00:37:20,840 --> 00:37:24,800
there or you take a random 
 
sample that's available, then 

622
00:37:24,800 --> 00:37:27,000
train with it. 
 
And we try to replace it as 

623
00:37:27,000 --> 00:37:29,960
quickly as possible. 
 
So in practice this this doesn't

624
00:37:29,960 --> 00:37:34,640
guarantee every sample is used 

only once, but you can measure 

625
00:37:34,840 --> 00:37:37,640
right the the throughput of your

 simulator versus the training 

626
00:37:37,640 --> 00:37:41,000
and it's typically at least 
 
below 2, so somewhere between 

627
00:37:41,400 --> 00:37:44,160
11.5. 
 
The South descendants get reused

628
00:37:45,960 --> 00:37:48,480
before they get replaced. 
 
So did you buffering? 

629
00:37:49,120 --> 00:37:50,640
That's interesting. 
 
I mean, I need to look a little 

630
00:37:50,640 --> 00:37:54,600
bit more at the the paper and 
 
the code, but that was, that was

631
00:37:54,600 --> 00:37:59,560
the bit that I guess I was 
 
assuming maybe for your PDE 

632
00:37:59,560 --> 00:38:03,080
problem is easier. 
 
But if we're doing what is, 

633
00:38:03,480 --> 00:38:06,720
let's say if we want to have 
 
something really accurate and 

634
00:38:06,720 --> 00:38:11,080
we're doing an LES simulation, 

the actual LES simulation might 

635
00:38:11,080 --> 00:38:15,640
need to run on 64 GPUs for like 
 8 hours, but the training is 

636
00:38:15,640 --> 00:38:17,390
obviously way faster than that. 
 

637
00:38:17,398 --> 00:38:20,075
So do you have any thoughts on 
that? 
 

638
00:38:20,083 --> 00:38:24,189
How you would balance from a 
time from a loading point of 
 

639
00:38:24,197 --> 00:38:27,709
view? 
So my guess is that that's why 


640
00:38:27,717 --> 00:38:31,420
this online training so far 
hasn't really taken off before 


641
00:38:31,428 --> 00:38:34,648
because like you mentioned for 
all classic simulations, exactly

642
00:38:34,648 --> 00:38:38,280

 you have the the warm up time 
until you get some equilibrium 


643
00:38:38,288 --> 00:38:42,485
that's physically valid and that
that could be used and it might 

644
00:38:42,485 --> 00:38:46,202
 take for a realistic real-world
case, it might take hours until 

645
00:38:46,202 --> 00:38:50,356
 you're in that regime and need 
fair number of CPUs at least, 
 

646
00:38:50,364 --> 00:38:53,112
or GPUs if you have a modern 
solver. 
 

647
00:38:53,120 --> 00:38:56,168
So that's, that's really 
unattractive because the, the 
 

648
00:38:56,176 --> 00:39:00,242
speed of generating the data is 
way below what you need for 
 

649
00:39:00,250 --> 00:39:03,920
training, which is why I think 
it's actually so interesting to 

650
00:39:03,920 --> 00:39:07,424
 use these canonical PDEs with 
these spectral solvers because 


651
00:39:07,432 --> 00:39:12,050
it, it's really, the solver is 
basically as fast as, as a 
 

652
00:39:12,058 --> 00:39:15,539
network or typically we run 
actually the, the training we 
 

653
00:39:15,547 --> 00:39:20,100
run on little regions like 64 ^3
and the solver typically produce

654
00:39:20,100 --> 00:39:23,068

 large, produces larger ones 
like 256 or so. 
 

655
00:39:23,076 --> 00:39:27,682
And even those are pretty close 
to the, to the training speed. 


656
00:39:27,690 --> 00:39:31,214
And then you can basically cut 
out different pieces for data 
 

657
00:39:31,222 --> 00:39:35,204
augmentation. 
But it, it matches quite nicely.

658
00:39:35,204 --> 00:39:37,680

 
So I think without such a solver

659
00:39:38,080 --> 00:39:41,520
doing a large scale pre training

 is is infeasible. 

660
00:39:42,080 --> 00:39:44,840
I think there are some 
 
approaches doing this, but you 

661
00:39:45,080 --> 00:39:48,920
if you want to match the 
 
generation capacity with the 

662
00:39:48,920 --> 00:39:51,692
training capacity you would need

 for a supercomputer to 

663
00:39:51,692 --> 00:39:56,160
generate data on the fly just to
feed a 
 decent sized model. 

664
00:39:56,800 --> 00:40:00,400
Yeah, that though then my other 
 thought was actually whilst it 

665
00:40:00,400 --> 00:40:05,520
takes 8 hours on 64 GPUs, let's 
 say, to do it, that is the 

666
00:40:05,520 --> 00:40:08,720
entire simulation. 
 
But actually to generate one 

667
00:40:08,720 --> 00:40:12,960
time step is probably only a 
 
second or two seconds. 

668
00:40:13,840 --> 00:40:19,640
So part of my interest, and 
 
maybe some groups already doing 

669
00:40:19,640 --> 00:40:23,322
this is training per time step. 
 

670
00:40:23,330 --> 00:40:28,234
Now, the variation from one time
step to another is pretty 
 

671
00:40:28,242 --> 00:40:30,875
minimal. 
So you could argue that you're 


672
00:40:30,883 --> 00:40:34,780
not really learning that much 
between, but you are. 
 

673
00:40:34,788 --> 00:40:39,575
That's the only way I could see 
that the time scales being 
 

674
00:40:39,583 --> 00:40:44,180
similar, You know, which I guess
ultimately addresses one of the 

675
00:40:44,180 --> 00:40:48,360
 challenges that I think many 
people have, which is almost all

676
00:40:48,360 --> 00:40:52,260

 of the standard data sets and 
standard machine learning 
 

677
00:40:52,268 --> 00:40:55,650
approaches do take some time 
averaged solution. 
 

678
00:40:55,658 --> 00:41:02,140
You know, which if you're trying
to get to real problems and real

679
00:41:02,140 --> 00:41:05,975

 accuracy, turbulence is 
transient, you know, and giving 

680
00:41:05,975 --> 00:41:11,860
 just a steady state answer or 
time average answer is, you 
 

681
00:41:11,868 --> 00:41:16,460
know, not, not really 
representing the ultimate goal 


682
00:41:16,468 --> 00:41:21,208
of like weather forecasting. 
I guess you would have the 
 

683
00:41:21,216 --> 00:41:25,400
temporal evolution of it. 
So I don't know how you how we 


684
00:41:25,408 --> 00:41:26,898
can get over that. 
Issue. 
 

685
00:41:26,906 --> 00:41:29,828
I think it's a great order to to
really target time. 
 

686
00:41:29,836 --> 00:41:35,228
I think there are some probably 
what I see there are quite some 

687
00:41:35,228 --> 00:41:38,080
 hurdles for for practitioners 
because we have all these 
 

688
00:41:38,088 --> 00:41:40,338
pipelines set up for for 
averaged quantities. 
 

689
00:41:40,346 --> 00:41:43,936
All right, potentially I think 
this would be would be neat to 


690
00:41:43,944 --> 00:41:46,805
have in place. 
Regarding your first point 
 

691
00:41:46,813 --> 00:41:51,640
though, with the kind of 
exploiting the, the fast solves 

692
00:41:51,640 --> 00:41:54,985
 over time for data generation. 
Unfortunately, here with 

693
00:41:54,985 --> 00:42:00,068
Tadpole, 
 with this online 
training, We have some, some, 

694
00:42:00,068 --> 00:42:03,423
some data 
 that is pointing to 
to problems there. 
 

695
00:42:03,431 --> 00:42:06,880
So what we noticed even with 
this online generation if we 
 

696
00:42:06,888 --> 00:42:09,596
don't replace the data fast 
enough. 
 

697
00:42:09,604 --> 00:42:13,787
So if this reuse is too high of 
our pre-training buffer 
 

698
00:42:13,795 --> 00:42:17,048
basically the models do start to
overfit. 
 

699
00:42:17,056 --> 00:42:20,876
So typically for foundation 
models you're dealing with 
 

700
00:42:20,884 --> 00:42:23,337
fairly large models. 
So ours are not even extreme, 
 

701
00:42:23,345 --> 00:42:25,770
but on the order of maybe 100 
million parameters. 
 

702
00:42:25,778 --> 00:42:30,264
And if the samples are reused 
too often, there was already—we 

703
00:42:30,264 --> 00:42:33,636
 saw some signs of of 
performance deteriorating if the

704
00:42:33,636 --> 00:42:38,782
generator 
 didn't catch up. 
And basically this was just out 

705
00:42:38,782 --> 00:42:42,280
 of pure luck because of of 
scheduling on these high 
 

706
00:42:42,288 --> 00:42:45,583
performance systems of right. 
If there's some one of the GPUs 

707
00:42:45,583 --> 00:42:49,215
 is is for some reason slower 
and suddenly on one training run

708
00:42:49,215 --> 00:42:54,090

 you reuse the data more often.
We saw a drop in performance and

709
00:42:54,090 --> 00:42:57,355

 even there the the correlation
between the samples was not 
 

710
00:42:57,363 --> 00:42:59,854
overly strong. 
So it still generated a fair 
 

711
00:42:59,862 --> 00:43:03,348
amount of data. 
But I would be worried if you 
 

712
00:43:03,356 --> 00:43:05,292
have these strongly correlated 
samples over time. 
 

713
00:43:05,300 --> 00:43:09,800
Even if you're able to swap them
out and reload reload different 

714
00:43:09,800 --> 00:43:14,595
 configurations for a solver, I 
think this would be would be 
 

715
00:43:14,603 --> 00:43:18,382
difficult to ensure that the 
variability in the data is large

716
00:43:18,382 --> 00:43:21,190

 enough to right not overfit a 
large model. 
 

717
00:43:21,198 --> 00:43:25,828
Yeah, that's a very good point 
actually, that that's, that's 
 

718
00:43:25,836 --> 00:43:27,727
the gang. 
Yeah. 
 

719
00:43:27,735 --> 00:43:33,125
The drawback that we, the 
temporal evolution of the PDE 
 

720
00:43:33,133 --> 00:43:37,205
needs to be small in for 
numerical reasons, you know, to,

721
00:43:37,205 --> 00:43:41,190

 to avoid blowing up, you know,
for because of the CFL 
 

722
00:43:41,198 --> 00:43:44,318
constraints, etcetera. 
So you you're kind of forced to 

723
00:43:44,318 --> 00:43:47,795
 slowly March in time, whereas I
guess you're right, if you feed 

724
00:43:47,795 --> 00:43:51,520
 the machine learning model that
it's just going to keep seeing 


725
00:43:51,528 --> 00:43:53,920
essentially the same solution, 
very similar. 
 

726
00:43:53,928 --> 00:43:57,624
Ones that. 
And and yeah, that's, but that's

727
00:43:57,624 --> 00:44:01,360

 where I still see a bit of a 
fundamental. 
 

728
00:44:01,368 --> 00:44:04,080
I just don't see. 
That's why I see almost as 
 

729
00:44:04,088 --> 00:44:06,180
that's why I was interested in 
your differentiable physics, 

730
00:44:06,180 --> 00:44:10,624
your 
 your code writing 
essentially because I feel some 

731
00:44:10,624 --> 00:44:16,520
of the 
 breakthroughs in this 
probably will come through a 

732
00:44:16,520 --> 00:44:21,282
clever, 
 clever use of the 
training being linked to the 

733
00:44:21,282 --> 00:44:26,925
data generation in 
 a way that.
I think right now our focus is 


734
00:44:26,933 --> 00:44:29,172
on foundation models. 
I'm also confident that at some 

735
00:44:29,172 --> 00:44:33,699
point 
 it's going to make sense
to bring the solvers back in the

736
00:44:33,699 --> 00:44:39,490

 our our tests back back then a
few years ago basically did did 

737
00:44:39,490 --> 00:44:41,888
 pretty short. 
If you have a solver, there's 
 

738
00:44:41,896 --> 00:44:44,444
anything decent, the learning 
task is just simpler. 
 

739
00:44:44,452 --> 00:44:49,850
So once once we have kind of 
reached the decent reasonable 
 

740
00:44:49,858 --> 00:44:55,040
capacity for this pre training, 
I think it's might make a lot of

741
00:44:55,040 --> 00:44:58,140

 sense to put a solver back 
into the loop to just make it 

742
00:44:58,140 --> 00:45:04,040
more 
 accurate in the end. 
Also, they are it's it is 
 

743
00:45:04,048 --> 00:45:07,952
challenging at the moment at the
at the scope given the current 


744
00:45:07,960 --> 00:45:10,920
infrastructures and and large 
scale models. 
 

745
00:45:10,928 --> 00:45:15,221
It's typically challenging to 
just train a single model and in

746
00:45:15,221 --> 00:45:18,982

 a decent amount of time and 
then trying to get a solver into

747
00:45:18,982 --> 00:45:21,627
the 
 picture and maybe going 
over multiple time steps. 
 

748
00:45:21,635 --> 00:45:27,760
Things like this are are tricky,
but it's also going to change 
 

749
00:45:27,768 --> 00:45:29,824
next year's. 
Yeah, yeah. 
 

750
00:45:29,832 --> 00:45:35,244
No, I I would agree. 
And maybe the the other topic is

751
00:45:35,244 --> 00:45:41,220

 around what were some of your 
lessons learned from doing, you 

752
00:45:41,220 --> 00:45:45,152
 know, WeatherBench and APEBench
and, and some of these more 
 

753
00:45:45,160 --> 00:45:48,764
benchmarking efforts because I 
guess that seems to have helped 

754
00:45:48,764 --> 00:45:52,380
 quite a bit on the weather and 
climate side. 
 

755
00:45:52,388 --> 00:45:55,640
So what? 
What was some of the genesis 
 

756
00:45:55,648 --> 00:46:00,300
behind those efforts? 
So the WeatherBench effort, 
 

757
00:46:00,308 --> 00:46:04,766
yeah, in retrospect, it's, it's 
great that it's, it's doing so 


758
00:46:04,774 --> 00:46:09,142
well. 
I, I think it's especially, I 
 

759
00:46:09,150 --> 00:46:11,932
guess I should think that so, 
right. 
 

760
00:46:11,940 --> 00:46:14,400
What I'm trying to say is 
basically the, these benchmarks 

761
00:46:14,400 --> 00:46:18,390
 have a big impact in the field.
I think by now that is widely 
 

762
00:46:18,398 --> 00:46:22,102
accepted across all the fields. 
Back then it was probably more 


763
00:46:22,110 --> 00:46:26,190
clear in the vision area where 
fair number of of benchmarks 
 

764
00:46:26,198 --> 00:46:27,847
have been around for quite a 
while. 
 

765
00:46:27,855 --> 00:46:32,705
But in other fields, like 
weather, there was very little, 

766
00:46:32,705 --> 00:46:36,056
 basically having established 
benchmarks and reliable 

767
00:46:36,056 --> 00:46:39,828
evaluations 
 is really 
important for for any discipline

768
00:46:39,828 --> 00:46:43,680
in the field. 
 
I think that's a nice pointer 

769
00:46:43,680 --> 00:46:46,160
and right, it's civic. 
 
It's it's quite some work 

770
00:46:46,560 --> 00:46:48,680
collecting the data, right and 

thinking about what should be in

771
00:46:48,680 --> 00:46:51,840
there and how to evaluate it. 
 
So the way those efforts are 

772
00:46:52,320 --> 00:46:57,440
really important for any 
 
subfield within any data-driven 

773
00:46:57,840 --> 00:46:59,920
discipline. 
 
And I think by now it's great 

774
00:46:59,920 --> 00:47:05,040
also in this scientific Yeah, I 
 feel that this is noticed that 

775
00:47:05,040 --> 00:47:07,720
by now more and more benchmarks 
 are coming out. 

776
00:47:08,480 --> 00:47:10,200
I think that's. 
 
Really important you you 

777
00:47:10,200 --> 00:47:15,560
recently generated to your group

 the SuperWing data set. 

778
00:47:16,000 --> 00:47:17,520
I mean, where do you see this 
 
going? 

779
00:47:17,520 --> 00:47:21,720
And it's a little bit of a 
 
philosophical debate around data

780
00:47:21,720 --> 00:47:27,640
because on one hand you could 
 
argue that we the progress in 

781
00:47:27,640 --> 00:47:31,120
this field does seem limited by 
 data to a certain extent. 

782
00:47:31,720 --> 00:47:38,064
And, you know, does it mean that

 there should be some 

783
00:47:38,064 --> 00:47:41,652
coordinated effort to generate 
data? 
 

784
00:47:41,660 --> 00:47:47,178
And if so, who should be doing 
that? 
 

785
00:47:47,186 --> 00:47:50,812
And what should be the license 
attached to it, given that there

786
00:47:50,812 --> 00:47:54,400

 is clearly also a commercial 
benefit to be had? 
 

787
00:47:54,408 --> 00:47:57,060
That's a good question. 
I mean coming from from a 
 

788
00:47:57,068 --> 00:48:00,152
university, I think it's great 
if it's all public and as open 


789
00:48:00,160 --> 00:48:02,077
as possible. 
But given the commercial 
 

790
00:48:02,085 --> 00:48:05,942
interest and also by now the the
key outlook that this will be 
 

791
00:48:05,950 --> 00:48:10,144
useful, that would of course be 
great to have support from from 

792
00:48:10,144 --> 00:48:14,020
 commercial partners in this. 
I think it's getting better. 
 

793
00:48:14,028 --> 00:48:18,396
Also more, more companies see 
the need for AI and then also 
 

794
00:48:18,404 --> 00:48:22,026
the need for or benchmarks and 
data sets on that front. 
 

795
00:48:22,034 --> 00:48:25,480
So I think it's improving, but 
definitely an issue where things

796
00:48:25,480 --> 00:48:28,560

 could be done and could 
improve a lot. 
 

797
00:48:28,568 --> 00:48:32,362
And unfortunately there's so 
much to quickly wrap this up. 
 

798
00:48:32,370 --> 00:48:35,700
But I think there's so much 
proprietary data with IP rights 

799
00:48:35,700 --> 00:48:39,320
 and so on that are probably on 
some servers and companies that 

800
00:48:39,320 --> 00:48:43,020
 they that cannot be used. 
So it probably needs an effort 


801
00:48:43,028 --> 00:48:45,400
to generate data that's free and
and usable. 
 

802
00:48:45,408 --> 00:48:50,276
Yeah, that I keep jumping 
between that because on one hand

803
00:48:50,276 --> 00:48:54,640

 you could say that it is to 
the a bit like maybe open 

804
00:48:54,640 --> 00:48:58,149
science 
 experiments, you know,
with, with astronomy or, or, or 

805
00:48:58,149 --> 00:49:02,944
things 
 where there's a sort of
public good and a lot of it is 

806
00:49:02,944 --> 00:49:05,032
funded 
 through taxpayers. 
Ultimately, you know, through 
 

807
00:49:05,040 --> 00:49:08,896
sort of science funding and the 
data's made available. 
 

808
00:49:08,904 --> 00:49:15,614
Part of it feels that that would
help all companies, you know, to

809
00:49:15,614 --> 00:49:19,034

 be doing it. 
But at the same time, is it 
 

810
00:49:19,042 --> 00:49:23,774
really the responsibility of 
governments and science to do 
 

811
00:49:23,782 --> 00:49:27,852
this? 
It it's it's kind of. 
 

812
00:49:27,860 --> 00:49:31,686
Yeah, that's good to argue about
the companies. 
 

813
00:49:31,694 --> 00:49:34,501
So pay for it if they benefit 
from it, that's. 
 

814
00:49:34,509 --> 00:49:38,152
Yeah, that's it. 
I just look at, I don't know 
 

815
00:49:38,160 --> 00:49:40,868
about you, but I look at all the
supercomputers in Europe, just 


816
00:49:40,876 --> 00:49:46,088
as an example, you know, and you
think how much data could be 
 

817
00:49:46,096 --> 00:49:49,720
generated given that these 
systems now are quite large 
 

818
00:49:49,728 --> 00:49:55,056
because they're being scaled up 
for the task of also, you know, 

819
00:49:55,056 --> 00:49:57,143
 large language model training, 
etcetera. 
 

820
00:49:57,151 --> 00:50:00,224
You know, if you've got a 
cluster with 10,000 GPUs, you 
 

821
00:50:00,232 --> 00:50:05,160
think how much data could you 
generate, you know, from fluid 


822
00:50:05,168 --> 00:50:10,743
simulations to a 10,000 GPUs, 
even just, you know, for a few 


823
00:50:10,751 --> 00:50:16,258
weeks or, or you know, or a 
month, you know, that does feel 

824
00:50:16,258 --> 00:50:22,214
 like it could be a, a huge way 
of doing it. 
 

825
00:50:22,222 --> 00:50:26,480
And if, if it's only a 
commercial company that does it,

826
00:50:26,480 --> 00:50:30,656

 they probably have no 
incentive to release that data. 

827
00:50:30,656 --> 00:50:34,400
 
And it then becomes hard to 

828
00:50:34,400 --> 00:50:37,840
progress the field if only one 

company has that data where it 

829
00:50:37,840 --> 00:50:42,840
feels like with large language 

models, the data has actually 

830
00:50:42,840 --> 00:50:44,360
had quite an open movement, 
 
right? 

831
00:50:44,360 --> 00:50:47,240
There's quite a lot of open data

 obviously around this. 

832
00:50:47,720 --> 00:50:50,520
And so, yes, you still could say

 you can only make models if 

833
00:50:50,520 --> 00:50:54,334
you've got lots of compute, but 
 there was already a movement 

834
00:50:54,334 --> 00:50:59,580
now with open-source LLMs, you 
know, 
 going on, whereas I feel

835
00:50:59,580 --> 00:51:03,800
now how can there be a movement 
of open 
 source surrogate 

836
00:51:03,800 --> 00:51:09,352
models if the data is not there 
in an open 
 way, You know, it 

837
00:51:09,352 --> 00:51:13,176
it it's. 
It's a good point. 
 

838
00:51:13,184 --> 00:51:18,847
It would be a pity if that gets 
hidden away due to commercial 
 

839
00:51:18,855 --> 00:51:20,832
interests. 
So it's almost the. 
 

840
00:51:20,840 --> 00:51:24,285
Progression of science, I guess,
and this is, I don't know if 
 

841
00:51:24,293 --> 00:51:27,758
you've had this and certainly 
now my myself being a commercial

842
00:51:27,758 --> 00:51:31,622

 company or the last two 
companies that you know that I'm

843
00:51:31,622 --> 00:51:36,020

 at, there is always this 
debate of open source versus 
 

844
00:51:36,028 --> 00:51:40,080
commercial, and I still find 
that a tricky one. 
 

845
00:51:40,088 --> 00:51:43,122
I wondered how you've seen it. 
You know what, what at what 
 

846
00:51:43,130 --> 00:51:47,141
point does it help everybody to 
be open source? 
 

847
00:51:47,149 --> 00:51:51,270
Yeah, OK, right. 
I have AI have a good. 
 

848
00:51:51,278 --> 00:51:53,998
Argument for open source from 
from my career at least. 
 

849
00:51:54,006 --> 00:51:57,640
We used to work in computer 
graphics, basically, with movie 

850
00:51:57,640 --> 00:52:00,312
 studios. 
I remember, during my postdoc, 

851
00:52:00,312 --> 00:52:03,978
one of the first 
 papers was at
our conference, SIGGRAPH. 

852
00:52:03,978 --> 00:52:08,715
There, 
 we basically put out 
some open- source code, and 

853
00:52:08,715 --> 00:52:12,640
DreamWorks 
 at the time 
actually picked this up pretty 

854
00:52:12,640 --> 00:52:14,798
quickly 
 and directly used it 
for for one of the shots. 
 

855
00:52:14,806 --> 00:52:17,743
And I said, oh, can't we get, or
maybe they can put us in the 
 

856
00:52:17,751 --> 00:52:21,056
credit somewhere. 
They send us a poster in the end

857
00:52:21,056 --> 00:52:24,899

 and also quite a few of my 
friends and kind of laugh or 
 

858
00:52:24,907 --> 00:52:27,296
look all all you got for for 
this work. 
 

859
00:52:27,304 --> 00:52:29,868
You put up your code, you got a 
poster for it. 
 

860
00:52:29,876 --> 00:52:32,726
Why did you sell it or 
something? 
 

861
00:52:32,734 --> 00:52:36,808
A couple of years later, these, 
all these movies where it was 
 

862
00:52:36,816 --> 00:52:40,570
used were one of the key things 
to apply for one of these 
 tech

863
00:52:40,570 --> 00:52:42,430
Oscars. 
There's also application 
 

864
00:52:42,438 --> 00:52:44,284
procedure. 
But in the end, we could show, 


865
00:52:44,292 --> 00:52:46,440
look, it's been used in all 
these movies. 
 

866
00:52:46,448 --> 00:52:51,069
And this this Oscar actually did
help a lot also for applying for

867
00:52:51,069 --> 00:52:55,032

 faculty positions and so on. 
So this is largely due to to 
 

868
00:52:55,040 --> 00:52:58,252
open source availability. 
It just might impact in the 
 

869
00:52:58,260 --> 00:53:00,840
field. 
So I've had very good experience

870
00:53:00,840 --> 00:53:04,480

 with this and I'm actually 
very happy that now that, yeah, 

871
00:53:04,480 --> 00:53:09,114
I 
 feel this is so obvious. 15 
years ago it was very, it was 
 

872
00:53:09,122 --> 00:53:12,503
actually the exception that 
papers came with code and it's 


873
00:53:12,511 --> 00:53:14,120
great. 
It's changed so much. 
 

874
00:53:14,128 --> 00:53:17,544
Yeah, it, it does seem an 
interesting 1 though, because 
 

875
00:53:17,552 --> 00:53:21,935
it's, given the large amount of 
money that's required to develop

876
00:53:21,935 --> 00:53:27,196

 models and develop training 
data, you know, there has to be 

877
00:53:27,196 --> 00:53:30,555
 some route to that money coming
back. 
 

878
00:53:30,563 --> 00:53:35,024
Which is why for a commercial 
company, if they spend $100 
 

879
00:53:35,032 --> 00:53:38,916
million on compute and they, you
know, they have to think, well, 

880
00:53:38,916 --> 00:53:41,117
 how am I going to make a 
business out of this? 
 

881
00:53:41,125 --> 00:53:46,020
And so it's, it's kind of an 
interesting one that is the 
 

882
00:53:46,028 --> 00:53:51,021
business that they, that the 
value to the company is not the 

883
00:53:51,021 --> 00:53:54,980
 model, but how you sort of 
tweak the model, so to speak. 
 

884
00:53:54,988 --> 00:53:58,415
And therefore, just because 
there's a foundation model out 


885
00:53:58,423 --> 00:54:02,410
there which is open source, the 
real value is that every single 

886
00:54:02,410 --> 00:54:05,630
 company needs to customize it, 
which I think is the value at 
 

887
00:54:05,638 --> 00:54:08,680
the moment for open-source LLMs 
that ultimately companies still 

888
00:54:08,680 --> 00:54:14,464
 make money out of it because 
everybody realizes that it isn't

889
00:54:14,464 --> 00:54:19,060

 fully ready in its base form. 
And that actually you take it 
 

890
00:54:19,068 --> 00:54:23,038
and you tweak it and you 
optimize it and you know, you so

891
00:54:23,038 --> 00:54:25,623

 that so there's money still to
be made. 
 

892
00:54:25,631 --> 00:54:29,036
And that's what I kind of wonder
on the fluids surrogate 
 

893
00:54:29,044 --> 00:54:34,212
modelling side or even beyond 
that, it would it still be in 
 

894
00:54:34,220 --> 00:54:38,360
the interest of some company to 
make it because ultimately the 


895
00:54:38,368 --> 00:54:43,080
money will be made customizing 
it and therefore being open 
 

896
00:54:43,088 --> 00:54:48,688
helps the whole community to 
develop it faster and make it 
 

897
00:54:48,696 --> 00:54:50,790
more competitive against close 
models. 
 

898
00:54:50,798 --> 00:54:55,028
Right, that's definitely. 
But I hope that, that we can go 

899
00:54:55,028 --> 00:54:58,205
 towards some foundation models 
or some generalizing models in 


900
00:54:58,213 --> 00:55:03,000
the field that can then be 
fine-tuned and adaptive to 
 

901
00:55:03,008 --> 00:55:06,980
adapted to different, different 
applications quite easily. 
 

902
00:55:06,988 --> 00:55:12,234
I think that's, that's been 
super useful in the LLM field. 


903
00:55:12,242 --> 00:55:18,235
And I, I do see potential there 
on the PDE front and 
 fluids 

904
00:55:18,235 --> 00:55:22,334
front. 
So, and I think if, if at some 


905
00:55:22,342 --> 00:55:25,220
point it's clear that you can 
get a really good result, right,

906
00:55:25,220 --> 00:55:28,165

 if if you actually download 
this model and then right fine 

907
00:55:28,165 --> 00:55:32,296
tune 
 it on your couple of wing
data set cases or so for certain

908
00:55:32,296 --> 00:55:35,486

 regime you're interested in, 
then. 
 

909
00:55:35,494 --> 00:55:41,424
But there could be a clear 
commercial also benefit for a 
 

910
00:55:41,432 --> 00:55:45,240
company and the use case of 
supporting these for supporting 

911
00:55:45,240 --> 00:55:47,480
 these open models. 
Yeah. 
 

912
00:55:47,488 --> 00:55:51,620
So that could imagine in an 
infrastructure environment 
 

913
00:55:51,628 --> 00:55:55,852
working quite well also in this 
area, but that definitely needs 

914
00:55:55,852 --> 00:56:00,619
 a coordinated and a large scale
effort to build it up right now.

915
00:56:00,619 --> 00:56:02,400

 
Yeah, that that's seems that 

916
00:56:02,640 --> 00:56:05,840
there's lots of separate 
 
efforts, I guess going on, but 

917
00:56:05,880 --> 00:56:10,680
but not uncoordinated. 
 
I I mean, maybe moving to some 

918
00:56:10,680 --> 00:56:14,960
of the final questions for you 

is where well, look, I, I guess 

919
00:56:14,960 --> 00:56:17,840
a couple of things before we get

 to the future looking one. 

920
00:56:17,840 --> 00:56:24,080
I'm just interested, where's 
 
your side on the like startups 

921
00:56:24,080 --> 00:56:26,200
or academia industry? 
 
Because I've noticed there's a 

922
00:56:26,200 --> 00:56:33,600
lot of, in recent years, there's

 even more interest in startups

923
00:56:34,080 --> 00:56:36,120
and, and, and the value of 
 
startups. 

924
00:56:36,120 --> 00:56:40,720
But then that also in some ways 
 can conflict with the open 

925
00:56:40,720 --> 00:56:43,040
source academic, you know, 
 
mindset. 

926
00:56:43,520 --> 00:56:45,240
Where, where have you seen that 
 in the world? 

927
00:56:45,240 --> 00:56:48,320
Have you been tempted with 
 
startups of, of industry? 

928
00:56:48,320 --> 00:56:53,080
Where do you see that role of 
 
academia, start-ups and and 

929
00:56:53,080 --> 00:56:54,880
industry? 
 
Yeah, it's a good question. 

930
00:56:54,880 --> 00:56:57,040
Exactly. 
 
Especially in last one or two 

931
00:56:57,040 --> 00:57:00,200
years we've seen a very nice 
 
rise I think in terms of funding

932
00:57:00,200 --> 00:57:04,920
and and also just founding of 
 
all kinds of spin off companies 

933
00:57:05,040 --> 00:57:08,520
that are pivots of quite 
 
existing companies towards this 

934
00:57:08,520 --> 00:57:11,640
physics AI direction. 
 
So I think also in a way 

935
00:57:11,640 --> 00:57:15,360
confirmation that now it's 
 
really starting to work, right, 

936
00:57:15,720 --> 00:57:17,480
It's ready for practical 
 
applications. 

937
00:57:18,520 --> 00:57:21,720
I've definitely toyed with the 

idea so far or also on my side, 

938
00:57:21,720 --> 00:57:23,920
there are no immediate plans for

 this. 

939
00:57:24,240 --> 00:57:26,480
In a way. 
 
I, I had my experience with 

940
00:57:26,480 --> 00:57:31,355
industry for visual 
 effects, 
and therefore I'm quite happy 

941
00:57:31,355 --> 00:57:34,920
with the Open University 
 
research side. 

942
00:57:34,920 --> 00:57:38,160
So I've, I've always been a big 
 fan of open source and just 

943
00:57:38,160 --> 00:57:41,520
being able to put out things and

 for, for research. 

944
00:57:41,520 --> 00:57:45,040
I mean, it's also effectively a 
 market with this openness and 

945
00:57:45,040 --> 00:57:47,600
impact, right? 
 
You, if you can generate this 

946
00:57:47,600 --> 00:57:50,160
later on, you can say, look, 
 
give me more money for research 

947
00:57:50,160 --> 00:57:54,280
to do more of, of this. 
 
So it's not the commercial 

948
00:57:54,280 --> 00:57:58,800
market, but also their own 
 
market in terms of research 

949
00:57:58,800 --> 00:58:01,000
money. 
 
And but I like the openness of 

950
00:58:01,000 --> 00:58:04,840
that. 
 
So for now, I think for me 

951
00:58:04,840 --> 00:58:07,640
that's personally just because I

 like this openness the the 

952
00:58:07,640 --> 00:58:10,640
better direction. 
 
But we also, by now we're 

953
00:58:10,640 --> 00:58:14,880
working with all kinds of 
 
companies due to the commercial 

954
00:58:14,880 --> 00:58:16,320
interests. 
 
Yeah. 

955
00:58:16,480 --> 00:58:18,880
But we were basically trying to 
 do this on the more Open 

956
00:58:18,880 --> 00:58:21,520
University side and then work 
 
with individual companies for 

957
00:58:21,520 --> 00:58:25,720
maybe adopting or or adapting 
 
these these techniques. 

958
00:58:26,840 --> 00:58:29,560
Yeah, I was going to say because

 the decking jury if all top 

959
00:58:29,560 --> 00:58:32,512
academics create start-ups, then

 there will be no academics 

960
00:58:32,512 --> 00:58:36,160
left to do that. 
 
The sort of so that that I'm 

961
00:58:36,160 --> 00:58:38,800
sure that has been discussed a 

little bit in the broader LLM 

962
00:58:38,800 --> 00:58:41,600
space. 
 
If the only companies doing the 

963
00:58:41,600 --> 00:58:45,200
latest stuff because of the 
 
scale are the big tech 

964
00:58:45,200 --> 00:58:51,104
companies, then there is a risk 
 to open science, I guess 

965
00:58:51,104 --> 00:58:56,032
because everyone's tempted to go
to a 
 commercial company that 

966
00:58:56,032 --> 00:59:00,000
maybe aren't, as you said, 
 
incentivized to be open. 

967
00:59:00,080 --> 00:59:03,040
I guess as academia, you are 
 
incentivized to be open because 

968
00:59:03,880 --> 00:59:07,440
the reward structure is based 
 
around it publishing grant 

969
00:59:07,440 --> 00:59:09,680
money. 
 
Like if everything's closed, I 

970
00:59:09,680 --> 00:59:12,480
guess it may help the university

 a little bit because of grant 

971
00:59:12,480 --> 00:59:16,520
capture or something, but it's 

not really the the core aim, is 

972
00:59:16,520 --> 00:59:19,920
it? 
 
It's difficult actually for for 

973
00:59:19,920 --> 00:59:22,840
universities. 
 
I think in the past quite a few 

974
00:59:22,840 --> 00:59:25,920
labs had their in house solvers 
 and then basically earned money

975
00:59:25,920 --> 00:59:30,680
and and funding. 
 
With direct we basically support

976
00:59:30,680 --> 00:59:34,240
or consulting contracts with 
 
companies. 

977
00:59:35,320 --> 00:59:39,234
But right now I think especially

 in the fast moving AI field, 

978
00:59:39,234 --> 00:59:42,833
it was, it was closed solutions,

 you would see very little 

979
00:59:42,833 --> 00:59:48,118
impact at the at the moment, 
Stephanie. 
 

980
00:59:48,126 --> 00:59:52,560
Works nicely no I, I, I think 
that's why I'm always 
 

981
00:59:52,568 --> 00:59:56,624
championing the academic side 
because I think people look at 


982
00:59:56,632 --> 01:00:00,495
research coming out of tech 
companies or research coming out

983
01:00:00,495 --> 01:00:04,634

 of the but forget that 
actually most of the things 

984
01:00:04,634 --> 01:00:09,060
start in the 
 university most 
most fundamental ideas and. 
 

985
01:00:09,068 --> 01:00:15,358
They need to be kept being 
promoted and that's where I feel

986
01:00:15,358 --> 01:00:19,499

 the data's the issue, because 
the they if you don't have open 

987
01:00:19,499 --> 01:00:23,250
 data, then you can't actually 
universities can't progress and 

988
01:00:23,250 --> 01:00:25,160
 can't publish and can't 
advance. 

989
01:00:25,160 --> 01:00:27,760
The state-of-the-art. 
 
If everything's done closed 

990
01:00:27,760 --> 01:00:30,640
door, it actually hurts 
 
progress, doesn't it? 

991
01:00:31,360 --> 01:00:33,240
Yeah. 
 
So LLMs will be interesting to 

992
01:00:33,240 --> 01:00:34,960
some extent, right. 
 
That's there have been 

993
01:00:34,960 --> 01:00:39,480
initiatives of training open 
 
LLMs, but some of them I guess 

994
01:00:39,480 --> 01:00:42,560
did fairly well. 
 
But the top models are not open.

995
01:00:43,520 --> 01:00:47,520
And right now I think there are 
 very few people who really can 

996
01:00:47,520 --> 01:00:53,440
or try to train their own LLMs. 
 That's really become a very 

997
01:00:53,440 --> 01:00:54,840
tough field, at least to operate

 in. 

998
01:00:55,640 --> 01:00:59,960
Yes, yes. 
 
So where do you see if you have 

999
01:00:59,960 --> 01:01:03,800
your looking glass or your, you 
 know, crystal ball? 

1000
01:01:04,560 --> 01:01:08,400
If we were to talk again in five

 years time, where where do you

1001
01:01:08,400 --> 01:01:10,640
think the conversation will be 

going? 

1002
01:01:10,640 --> 01:01:14,160
What what do you think will be 

the breakthrough moments? 

1003
01:01:14,680 --> 01:01:17,040
Do you feel like we're going to 
 be just incrementally over the 

1004
01:01:17,040 --> 01:01:21,800
next five years or do you 
 
perceive some big breakthroughs 

1005
01:01:21,800 --> 01:01:23,400
or leap? 
 
Just hope that we're going to 

1006
01:01:23,400 --> 01:01:25,280
see this breakthrough in terms 

of adoption. 

1007
01:01:25,480 --> 01:01:29,960
So now also these all these 
 
initiatives on commercial side I

1008
01:01:29,960 --> 01:01:32,400
think point towards the field 
 
really starting to work. 

1009
01:01:32,400 --> 01:01:36,840
So I hope that five years we we 
 really see widespread adoption 

1010
01:01:36,840 --> 01:01:42,414
for for actual applications in a

 way sounds a bit boring, 

1011
01:01:42,414 --> 01:01:44,320
right, just being it used in in 
 practice. 

1012
01:01:44,320 --> 01:01:46,760
But I think right also working 

on this for almost 10 years. 

1013
01:01:46,760 --> 01:01:50,120
I think it's about time that we 
 that we do see that it's got 

1014
01:01:50,400 --> 01:01:53,840
it's really useful for for real 
 world world things. 

1015
01:01:54,920 --> 01:01:58,080
So I just hope that now we are 

at a stage where this really can

1016
01:01:58,080 --> 01:02:02,600
be can be pulled off, but in the

 field is really progressing 

1017
01:02:02,720 --> 01:02:05,920
extremely quickly. 
 
And so one of the, one of the 

1018
01:02:05,920 --> 01:02:08,760
outlooks that I've been 
 
intrigued about are these world 

1019
01:02:08,760 --> 01:02:11,360
models, right? 
 
Like vision by now, they're 

1020
01:02:11,440 --> 01:02:14,356
already going beyond foundation 
 models or specific language 

1021
01:02:14,356 --> 01:02:17,798
and, and just visual models 
towards 
 ones that just capture

1022
01:02:17,798 --> 01:02:20,440
the whole world. 
 
And right, also vision now 

1023
01:02:20,440 --> 01:02:22,600
notices actually physics is an 

important part. 

1024
01:02:22,600 --> 01:02:25,360
You only see so much. 
 
A lot of the complexity comes 

1025
01:02:25,360 --> 01:02:30,560
from things you don't directly 

see like air moving or so I 

1026
01:02:30,560 --> 01:02:34,320
think that's a really 
 
interesting outlook to to really

1027
01:02:34,320 --> 01:02:37,240
simulate basically the the whole

 world around us on a larger 

1028
01:02:37,240 --> 01:02:38,880
scale with the physics 
 
accurately. 

1029
01:02:39,120 --> 01:02:43,520
I think that could also be super

 useful for engineering and and

1030
01:02:43,720 --> 01:02:46,160
real applications in a bunch of 
 areas. 

1031
01:02:46,600 --> 01:02:51,000
But getting that right seems 
 
like a whole different level on 

1032
01:02:51,000 --> 01:02:54,007
top of what we currently dealing

 with just only doing the 

1033
01:02:54,007 --> 01:02:55,360
physics right with the 
foundation 
 models. 

1034
01:02:55,360 --> 01:02:56,840
But I think it's a really 
 
interesting outlook. 

1035
01:02:58,120 --> 01:02:59,520
Yeah, it's actually a good point

 there. 

1036
01:02:59,520 --> 01:03:00,840
Yeah, I forgot to ask you about 
 that. 

1037
01:03:00,840 --> 01:03:05,480
That I still don't fully get. 
 
I mean, I understand the world 

1038
01:03:05,480 --> 01:03:12,240
models are obviously one very 
 
big use case is robotics and 

1039
01:03:12,240 --> 01:03:15,680
driving drop, you know, 
 
autonomous vehicles and sort of 

1040
01:03:15,680 --> 01:03:19,360
giving that synthetic data to 
 
help them to operate in the 

1041
01:03:19,360 --> 01:03:26,200
world, The bit that I guess I'm 
 less sure on. 

1042
01:03:26,200 --> 01:03:29,800
And I wonder whether 
 
economically and scientifically 

1043
01:03:29,800 --> 01:03:35,120
is the physics side, you know, 

does it make a difference how 

1044
01:03:35,120 --> 01:03:39,400
good the physics is like, as 
 
long as it looks right? 

1045
01:03:39,560 --> 01:03:41,360
And this, This is why I'm 
 
interested, because your movie 

1046
01:03:41,360 --> 01:03:46,520
background, like does it, does 

it matter to the robots and to 

1047
01:03:46,520 --> 01:03:52,240
the autonomous vehicles? 
 
Actually, I do think that at 

1048
01:03:52,240 --> 01:03:56,280
least for robots, the physics 
 
play a big role because 

1049
01:03:56,680 --> 01:04:00,640
ultimately the, the action that 
 you're generating is, is the 

1050
01:04:00,640 --> 01:04:03,160
control right of your, of the 
 
actuators of the actual things 

1051
01:04:03,160 --> 01:04:07,920
that the robot should do. 
 
And you need to estimate surface

1052
01:04:07,920 --> 01:04:11,800
properties, the weight of an 
 
object, weight distribution 

1053
01:04:11,800 --> 01:04:16,680
stability and all these things 

to at least in a, in a variable 

1054
01:04:16,680 --> 01:04:19,440
environment to interact with 
 
the, with the world. 

1055
01:04:19,440 --> 01:04:24,720
So think for robots by now rigid

 body type or in some somewhat 

1056
01:04:24,720 --> 01:04:29,840
deformable and physics are 
 
playing quite a role. 

1057
01:04:29,960 --> 01:04:35,120
So I would I would guess that it

 does make a difference for 

1058
01:04:35,120 --> 01:04:37,840
autonomous driving. 
 
We could argue at the end you 

1059
01:04:37,880 --> 01:04:40,360
only need to take right? 
 
Is it a dangerous situation or 

1060
01:04:40,360 --> 01:04:41,435
would something have gone wrong?

 

1061
01:04:41,443 --> 01:04:46,488
You can probably do a lot 
without having to go through the

1062
01:04:46,488 --> 01:04:49,078

 whole crash or actually 
simulating how the car would 
 

1063
01:04:49,086 --> 01:04:51,210
tumble or crash into a building 
or so. 
 

1064
01:04:51,218 --> 01:04:53,840
That might not be so important, 
right, As long as you can detect

1065
01:04:53,840 --> 01:04:58,300

 this would have gone wrong. 
But for robots actually having 


1066
01:04:58,308 --> 01:05:00,820
to really interact with objects,
I can imagine that plays a 
 

1067
01:05:00,828 --> 01:05:03,730
larger role. 
And then again looking towards 


1068
01:05:03,738 --> 01:05:09,072
fluids, I, I think the challenge
also in, in the current 
 

1069
01:05:09,080 --> 01:05:12,290
development pipelines is this 
interaction with the real world.

1070
01:05:12,290 --> 01:05:14,480

 
I mean, I think in the end, the 

1071
01:05:14,480 --> 01:05:16,720
current validation and 
 
certification still goes through

1072
01:05:16,720 --> 01:05:19,580
a lot of testing status and then

 you have all kinds of real 

1073
01:05:19,580 --> 01:05:24,000
world flights and and experience
to 
 make sure that the systems 

1074
01:05:24,000 --> 01:05:25,840
really do what they should in 
 
the real world. 

1075
01:05:26,600 --> 01:05:32,800
But if you could shift some of 

that into an earlier stage, I 

1076
01:05:32,800 --> 01:05:35,840
think that that could also be 
 
definitely interesting. 

1077
01:05:35,840 --> 01:05:39,800
If right, you could have 
 
different environment 

1078
01:05:39,800 --> 01:05:42,760
conditions, different flow 
 
conditions, and you could have 

1079
01:05:42,760 --> 01:05:46,280
an have an object interact 
 
dynamically in in such an 

1080
01:05:46,280 --> 01:05:48,880
environment. 
 
Yeah, I guess maybe maybe you're

1081
01:05:48,880 --> 01:05:51,834
right to maybe the examples I've

 looked at, you know, if you 

1082
01:05:51,834 --> 01:05:58,960
look at a robot walking around a

 factory floor or it has no the

1083
01:05:59,000 --> 01:06:03,000
the drag, the lift, the thermal 
 properties are very second or 

1084
01:06:03,000 --> 01:06:05,680
third order effects. 
 
They don't really influence how 

1085
01:06:05,680 --> 01:06:08,160
it's performing. 
 
But I guess you're right. 

1086
01:06:08,160 --> 01:06:13,440
If it was like a drone in the 
 
sky or a boat in the water, then

1087
01:06:13,440 --> 01:06:17,640
the physics actually do make 
 
quite a bit of a difference in 

1088
01:06:18,720 --> 01:06:20,960
and then I guess it really is 
 
depending what's the use of 

1089
01:06:20,960 --> 01:06:26,600
these world models, if they 
 
really are like a true, true 

1090
01:06:26,600 --> 01:06:33,280
representation of the world for 
 that drone or or helicopter or,

1091
01:06:34,000 --> 01:06:36,760
you know, I guess the moment the

 autonomous vehicle, you're 

1092
01:06:36,760 --> 01:06:40,480
right, is it's more sensory. 
 
You know, like I see there's a 

1093
01:06:40,480 --> 01:06:43,560
person walking. 
 
It's not to design the car, 

1094
01:06:43,560 --> 01:06:46,680
right? 
 
It's not to say I see now that 

1095
01:06:46,680 --> 01:06:49,800
in this real world, by changing 
 the shape, the drag is lower 

1096
01:06:50,000 --> 01:06:53,120
because it's actually I guess 
 
maybe that's where you're 

1097
01:06:53,120 --> 01:06:57,760
thinking, right, that the world 
 model could be used as a true 

1098
01:06:57,760 --> 01:07:01,920
synthetic environment or cars, I

 guess if you could. 

1099
01:07:01,920 --> 01:07:05,320
Simulate a rainy and stormy 
 
situation and then really get 

1100
01:07:05,320 --> 01:07:09,000
feedback on how the rain would 

splash around a car, or how gust

1101
01:07:09,000 --> 01:07:13,587
of wind would influence driving 
 stability in extreme 

1102
01:07:13,587 --> 01:07:16,824
conditions. 
Or right, all kinds of of 
 

1103
01:07:16,832 --> 01:07:19,040
landing and starting scenarios 
for planes. 
 

1104
01:07:19,048 --> 01:07:22,922
If you could really get feedback
on how a deforming plane 
 with 

1105
01:07:22,922 --> 01:07:25,795
changing conditions would would 
behave. 
 

1106
01:07:25,803 --> 01:07:32,680
I think that's that's still 
beyond even regular simulation 


1107
01:07:32,688 --> 01:07:35,154
capabilities—that 
fluid-structure interactions 

1108
01:07:35,154 --> 01:07:39,126
with changing 
 conditions are 
really challenging. 
 

1109
01:07:39,134 --> 01:07:44,632
You could get some feedback on 
these and then ideally put an 
 

1110
01:07:44,640 --> 01:07:46,160
optimization loop around them, 
right? 
 

1111
01:07:46,168 --> 01:07:49,990
You you would want to actually 
optimize your, you know, your 
 

1112
01:07:49,998 --> 01:07:55,988
wing profile or the shape of of 
a vehicle to behave in the 
 

1113
01:07:55,996 --> 01:07:59,334
whole longer sequence. 
I feel like this is also the big

1114
01:07:59,334 --> 01:08:03,495

 debate, well not big debate, 
but a debate around the classic 

1115
01:08:03,495 --> 01:08:07,463
tool 
 calling thing. 
You know, like isn't LLM good at

1116
01:08:07,463 --> 01:08:11,095

 adding numbers? 
No, so just go and call a 
 

1117
01:08:11,103 --> 01:08:13,159
calculator. 
That's like ultimately, I guess 

1118
01:08:13,159 --> 01:08:16,376
 if you call ChatGPT to add 2 
numbers together, it's just 
 

1119
01:08:16,384 --> 01:08:20,131
calling a tool of a calculator, 
adding the 2 numbers together. 


1120
01:08:20,140 --> 01:08:24,910
I kind of wonder it's sometimes 
in this world of, you know, 
 

1121
01:08:24,917 --> 01:08:29,395
world models or optimizing, you 
know, is it, is it really that 


1122
01:08:29,403 --> 01:08:31,413
it's truly integrated into a 
world model? 
 

1123
01:08:31,421 --> 01:08:35,270
Or is it more likely that there 
is a sort of agent workflow 
 

1124
01:08:35,278 --> 01:08:39,493
where it's calling a surrogate 
model and maybe that's the where

1125
01:08:39,493 --> 01:08:42,509

 they're two separate models, 
if you know what I mean, rather 

1126
01:08:42,509 --> 01:08:45,560
 than just a world model. 
It's more of an agentic 
 way, 

1127
01:08:45,560 --> 01:08:48,911
if you know what I mean. 
Or not that that would already 


1128
01:08:48,919 --> 01:08:52,560
be a big step, right, If if you 
had any visual model that could 

1129
01:08:52,560 --> 01:08:56,840
 on demand call some physics sub
models surrogates to to get 
 

1130
01:08:56,848 --> 01:09:01,960
feedback on how these different 
pieces it recognises should 

1131
01:09:01,960 --> 01:09:04,268
interact. 
 
For rigid bodies, 
 actually 

1132
01:09:04,268 --> 01:09:06,800
these days you would just call 
it a rigid-body 
 simulator. 

1133
01:09:07,160 --> 01:09:09,359
You would probably not need a 
 
surrogate. 

1134
01:09:10,680 --> 01:09:14,680
Yeah, that's. 
 
That's kind of where I always 

1135
01:09:14,680 --> 01:09:16,640
debate it and that and that's 
 
the traditional one that you 

1136
01:09:16,640 --> 01:09:21,000
said if the solvers get so fast,

 like the spectral solver you 

1137
01:09:21,000 --> 01:09:24,439
mentioned or now with, you know,

 advancements in GPUs, some 

1138
01:09:24,439 --> 01:09:27,000
people are saying I can do a 
 
simulation in a minute. 

1139
01:09:27,880 --> 01:09:31,279
And then you think, well, do I 

need a surrogate? 

1140
01:09:31,359 --> 01:09:34,720
If if I can, if the solver can 

run so fast? 

1141
01:09:35,680 --> 01:09:38,045
Right, I think for world models,

 an interesting challenge will 

1142
01:09:38,045 --> 01:09:42,676
be to to quickly switch 
 from 
approximate fidelity for rigid 

1143
01:09:42,676 --> 01:09:46,573
bodies or 
 deformable objects, 
or winds, or different feedbacks

1144
01:09:46,573 --> 01:09:53,283
needed in 
 the environment. 
Sure, if I know I need perfect 


1145
01:09:53,292 --> 01:09:56,848
rigid bodies for a certain case 
that might be ideal, but but 
 

1146
01:09:56,856 --> 01:10:01,674
right I need some feedback on on
wind or how right my plastic cup

1147
01:10:01,674 --> 01:10:06,160

 should deform or should it 
break predictions across the all

1148
01:10:06,160 --> 01:10:08,360
these 
 different physical 
phenomena. 

1149
01:10:09,320 --> 01:10:13,080
I could imagine that a flexible 
 surrogate that at least gives a

1150
01:10:13,080 --> 01:10:16,440
gives the first prediction of 
 
estimates of of what what might 

1151
01:10:16,440 --> 01:10:18,440
happen there could be 
 
beneficial. 

1152
01:10:18,440 --> 01:10:21,880
I think that would be tricky to 
 do with classical simulators. 

1153
01:10:22,400 --> 01:10:26,680
Yes, yeah, yeah, yeah, that. 
 
I think this is where it's the 

1154
01:10:26,680 --> 01:10:29,280
classic what's the, what's the 

use case of real time? 

1155
01:10:29,280 --> 01:10:32,560
And sometimes if you go to an 
 
engineering team and I've had 

1156
01:10:32,560 --> 01:10:35,200
this where you tell them that a 
 surrogate model prediction in 

1157
01:10:35,200 --> 01:10:37,960
one second, sometimes they'll 
 
say, well, we don't need it to 

1158
01:10:37,960 --> 01:10:40,480
do in one second. 
 
You know, it's actually fine If 

1159
01:10:40,480 --> 01:10:44,160
it takes 5 minutes or even half 
 an hour, it's OK. 

1160
01:10:44,160 --> 01:10:47,480
It's not a bottleneck. 
 
Whereas I guess if it's 

1161
01:10:47,480 --> 01:10:50,840
integrated in a world model or 

something, you need it to be 

1162
01:10:51,240 --> 01:10:57,480
essentially real time or else 
 
the whole value breaks of some 

1163
01:10:57,480 --> 01:11:00,080
simulation. 
 
You know, like testing a car 

1164
01:11:00,080 --> 01:11:01,600
moving. 
 
If you have to sort of pause the

1165
01:11:01,600 --> 01:11:06,080
simulator for 30 minutes, then 

it's not that. 

1166
01:11:06,440 --> 01:11:09,560
Feel right? 
 
Yeah, I mean real time is is 

1167
01:11:09,560 --> 01:11:13,280
extremely challenging. 
 
Then you need I guess to to be a

1168
01:11:13,280 --> 01:11:16,360
realistic virtual environment 
 
for a person. 

1169
01:11:16,840 --> 01:11:20,560
Yes, you need milliseconds. 
 
Yeah, very real. 

1170
01:11:21,080 --> 01:11:26,440
Yeah. 
 
So maybe a final question to you

1171
01:11:26,880 --> 01:11:30,960
looking more if, if there's a 
 
student listening to this or 

1172
01:11:31,320 --> 01:11:34,440
someone doing HD or early in 
 
their career, given, given all 

1173
01:11:34,440 --> 01:11:39,200
these changes, what would you 
 
recommend someone was would 

1174
01:11:39,200 --> 01:11:41,480
study as a PhD topic? 
 
You know what was going to 

1175
01:11:41,480 --> 01:11:44,042
future proof them in this world.

 

1176
01:11:44,050 --> 01:11:48,434
Maybe I'm biased here, but I 
think these are really important

1177
01:11:48,434 --> 01:11:51,148

 techniques. 
I, I do see a lot of potential 


1178
01:11:51,156 --> 01:11:53,475
for AI based techniques. 
So I think it's important to 
 

1179
01:11:53,483 --> 01:11:56,367
have an understanding to know 
the classic, the classic physics

1180
01:11:56,367 --> 01:11:58,740

 numerics. 
But by now within these 
 

1181
01:11:58,748 --> 01:12:03,480
numerical tools, I think AI is, 
is a super important component. 

1182
01:12:03,480 --> 01:12:04,960
 
So I could, could highly 

1183
01:12:04,960 --> 01:12:09,920
recommend looking at at least 
 
combinations of classic and AI 

1184
01:12:09,920 --> 01:12:15,880
based methods. 
 
And my, my guess is also that in

1185
01:12:15,880 --> 01:12:21,560
the future or the next couple of

 of years, we were not going to

1186
01:12:21,560 --> 01:12:24,240
be completely replaced by I 
 
don't think ChatGPT is going to 

1187
01:12:24,240 --> 01:12:26,840
take over fluid simulation too 

soon. 

1188
01:12:26,840 --> 01:12:29,080
So having experts that we 
 
understand what's happening 

1189
01:12:29,080 --> 01:12:31,960
there, I think it's it's going 

to be needed for quite a while. 

1190
01:12:31,960 --> 01:12:34,720
So I think it's still a good 
 
field to to work on. 

1191
01:12:35,520 --> 01:12:39,760
Yeah, I guess that that's the 
 
the truth, I guess the, the 

1192
01:12:39,760 --> 01:12:43,000
short, medium, long term, I 
 
guess I see at least now the 

1193
01:12:43,000 --> 01:12:45,800
short term, there's even more 
 
need for specialists because 

1194
01:12:46,120 --> 01:12:51,866
frankly, it's almost become that

 simulation is more of a 

1195
01:12:51,866 --> 01:12:53,960
popular topic now. 
 
The fact that all these startups

1196
01:12:53,960 --> 01:12:56,440
are getting funded, the fact 
 
that, you know, there's more 

1197
01:12:56,440 --> 01:13:00,360
money in this space, there's 
 
more need to find out who's an 

1198
01:13:00,360 --> 01:13:03,760
expert in crash simulations, 
 
who's an expert in acoustics, 

1199
01:13:03,760 --> 01:13:07,760
who's an expert to help computer

 scientists, you know, to 

1200
01:13:07,760 --> 01:13:11,480
develop these models. 
 
I guess it's only in the very 

1201
01:13:11,480 --> 01:13:13,600
long term future that you could 
 imagine. 

1202
01:13:13,600 --> 01:13:18,120
Maybe you don't need as many 
 
specialists because so much of 

1203
01:13:18,120 --> 01:13:20,520
that knowledge is now ingrained 
 in these models. 

1204
01:13:21,080 --> 01:13:24,200
But there's a, in the short 
 
term, it's almost even more 

1205
01:13:24,360 --> 01:13:29,480
specialists to help. 
 
But deep, deep specialists, I 

1206
01:13:29,480 --> 01:13:34,360
guess you really understand it 

that, that at least that's the 

1207
01:13:34,360 --> 01:13:36,480
way I I'm seeing things at the 

moment. 

1208
01:13:36,480 --> 01:13:38,920
And I think that's usually a 
 
great topic for a PhD, really 

1209
01:13:38,920 --> 01:13:40,120
getting to the bottom of things.

 

1210
01:13:40,128 --> 01:13:43,259
So I just, it's a nice 
opportunity to work on one topic

1211
01:13:43,259 --> 01:13:46,934

 for a couple of years and 
yeah, get as much understanding 

1212
01:13:46,934 --> 01:13:48,616
as 
 possible. 
So on. 
 

1213
01:13:48,624 --> 01:13:50,014
Right now I think PhD is in this
area. 
 

1214
01:13:50,022 --> 01:13:52,428
Still a very good idea. 
Yeah, yeah. 
 

1215
01:13:52,436 --> 01:13:55,744
No, no, I agree. 
Well, thank you so much for 
 

1216
01:13:55,752 --> 01:13:58,970
taking the time to speak. 
I mean, I, I find all these 
 

1217
01:13:58,978 --> 01:14:02,733
topics fascinating and one of 
the things I'll do is put a list

1218
01:14:02,733 --> 01:14:06,370

 of the papers in the, in the 
show notes because I think 
 

1219
01:14:06,378 --> 01:14:09,384
actually there's a lot more to 
gather from looking in the 
 

1220
01:14:09,392 --> 01:14:12,286
details and, and looking at the 
codes that your, your, your 
 

1221
01:14:12,294 --> 01:14:14,232
group did. 
So I know we didn't have time to

1222
01:14:14,232 --> 01:14:16,746

 cover all in, you know, full 
detail, but I'll, I'll put a 
 

1223
01:14:16,754 --> 01:14:19,680
list and hopefully people can 
find the time to read through 
 

1224
01:14:19,688 --> 01:14:21,549
the papers. 
That's a good idea. 
 

1225
01:14:21,557 --> 01:14:24,480
Oh, great discussion. 
Thanks for the invite. 
 

1226
01:14:24,488 --> 01:14:26,702
Yeah, interesting. 
We'll speak again in two years 


1227
01:14:26,710 --> 01:14:29,620
time and we'll see if the 
predictions are. 
 

1228
01:14:29,628 --> 01:14:33,200
That's interesting. 
Thank you. 
 

1229
01:14:33,208 --> 01:14:33,720
Thanks.
