1
00:00:00,354 --> 00:00:02,895
Hi and welcome to the Neil 
Ashton Podcast. 

2
00:00:02,975 --> 00:00:06,586
In each episode, we explain some
of the fascinating ways that 

3
00:00:06,586 --> 00:00:09,538
science and engineering are 
 
changing the world around us. 

4
00:00:09,738 --> 00:00:13,388
We talk to leading engineers 
from elite level sports like 

5
00:00:13,388 --> 00:00:17,394
cycling and Formula One, to some

 of the world's top academics 

6
00:00:17,394 --> 00:00:20,307
to understand how fluid 
dynamics, machine learning, and 

7
00:00:20,307 --> 00:00:23,583
supercomputing are bringing in a
new era of discovery. 

8
00:00:23,904 --> 00:00:27,089
We also hear some of their life 
stories, their career advice, 

9
00:00:27,089 --> 00:00:29,112
and lessons they've learned 
 on
the way. 

10
00:00:29,112 --> 00:00:32,036
That I hope will be helpful to 
you too. 

11
00:00:32,036 --> 00:00:34,716
So sit back and enjoy this 
episode. 

12
00:00:39,822 --> 00:00:42,422
Welcome back to the Neil Ashton 
Podcast. 

13
00:00:42,562 --> 00:00:48,102
So this is, um, I guess a 
special episode, um, that I did 

14
00:00:48,102 --> 00:00:52,362
because I've just published a 
 
paper with Sid and Johannes. 

15
00:00:52,662 --> 00:00:54,202
They'll give their full intros 
later. 

16
00:00:54,202 --> 00:00:59,896
So I want to keep it a bit 
short, which is called Fluid 

17
00:00:59,896 --> 00:01:03,831
Intelligence: A Forward Look on 
AI Foundation Models in 

18
00:01:03,831 --> 00:01:06,482
Computational Fluid Dynamics. 
Bit of a long title. 

19
00:01:06,482 --> 00:01:09,251
I had to look exactly what we 
call the title now, so I didn't 

20
00:01:09,251 --> 00:01:11,810
say it wrong. 
But it's, it's something I'm 

21
00:01:11,810 --> 00:01:15,682
quite proud of because it's been
a real collaborative work to 
 

22
00:01:15,690 --> 00:01:20,245
try and give a vision of how you
could build a foundational model

23
00:01:20,245 --> 00:01:22,693
for CFD. 
And just to set the context, 

24
00:01:22,693 --> 00:01:26,185
what that means and the way I 
described it some people is 
 

25
00:01:26,193 --> 00:01:29,677
like a ChatGPT for fluids, you 
know, that, that dream of a 

26
00:01:29,677 --> 00:01:32,595
model that could predict any 
possible scenario of any plane, 

27
00:01:32,595 --> 00:01:36,986
of any car or of any data center
or of any pipe flow or anything.

28
00:01:38,452 --> 00:01:42,222
And, and not that our paper has 
the full answer to it, but we 

29
00:01:42,222 --> 00:01:44,105
tried to look at the scaling 
 
laws. 

30
00:01:44,105 --> 00:01:47,827
I mean, like what's the, what's 
the bottleneck to doing this 

31
00:01:47,827 --> 00:01:50,869
from a compute, from a data 
 
point of view. 

32
00:01:50,869 --> 00:01:54,661
Um, it is still a scientific 
piece of work. 

33
00:01:54,661 --> 00:01:58,693
It's a piece of research. 
It has hypotheses, but it's 

34
00:01:58,693 --> 00:02:01,451
something we deeply thought 
about, researched, discussed 

35
00:02:01,451 --> 00:02:05,778
with 
 people in the field, you 
know, to get their feedback. 

36
00:02:05,814 --> 00:02:10,699
And something that I think the 
three of us do stand behind as a

37
00:02:10,699 --> 00:02:13,138
combination of CFD, ML and 
applied maths. 

38
00:02:13,218 --> 00:02:16,000
And the paper's quite long. 
It's nearly 40 pages. 

39
00:02:16,000 --> 00:02:20,380
And so the point of this episode
was to have the three of us talk

40
00:02:20,380 --> 00:02:23,300
through it really, so 
 that 
hopefully if you're reading it, 

41
00:02:23,300 --> 00:02:26,837
you understand it um a bit more 
and you understand where we came

42
00:02:26,837 --> 00:02:29,712
from and some of the decisions 
we made, what we put in, what we

43
00:02:29,712 --> 00:02:33,018
didn't put 
 in. 
And so, yeah, I hope, I hope you

44
00:02:33,018 --> 00:02:35,296
enjoy this. 
This discussion does get into 

45
00:02:35,296 --> 00:02:38,224
the weeds, but it's a topic that
anyone who's listened to this 

46
00:02:38,224 --> 00:02:40,664
podcast for a 
 while has been a
key theme, right? 

47
00:02:40,664 --> 00:02:43,274
All the questions I've been 
asking guests have been about 

48
00:02:43,274 --> 00:02:46,181
foundational models. 
It's been about can AI, what can

49
00:02:46,181 --> 00:02:49,918
AI do for CFD? 
What can it do for things like 

50
00:02:49,918 --> 00:02:52,967
Formula One, the cycling, for 
aircraft design, for space. 

51
00:02:52,967 --> 00:02:56,870
So it's a combination of some of
that work, you know, over the 

52
00:02:56,870 --> 00:03:00,470
past couple of years, 
 thinking
in my head and Johannes and Sid 

53
00:03:00,470 --> 00:03:03,697
had similar sort of thought 
processes and we all came 

54
00:03:03,697 --> 00:03:06,550
together to. to write this 
paper. 

55
00:03:06,610 --> 00:03:10,850
I'll put a link in the notes so 
you can get this paper on arXiv.

56
00:03:10,850 --> 00:03:13,947
It's a pre-print or if you just 
Google something like Fluid 

57
00:03:13,947 --> 00:03:17,319
Intelligence and if you put in 

Neil Ashton and I mean, that'll 

58
00:03:17,319 --> 00:03:20,410
bring it up, but you could put 
any of other co-authors. 

59
00:03:20,410 --> 00:03:24,670
So I hope you enjoy this paper. 
I hope you found it useful. 

60
00:03:25,190 --> 00:03:29,348
It's been a fun piece of work 
and I hope this helps to explain

61
00:03:29,348 --> 00:03:33,467
some of the logic behind 
 it. 
So yeah, sit back and listen to 

62
00:03:33,467 --> 00:03:36,460
this special episode. 
Thanks, Sid, Johannes, for, for 

63
00:03:36,460 --> 00:03:39,828
joining. 
This is, I guess a special 

64
00:03:39,828 --> 00:03:45,768
episode because we put out a 
paper recently that I think has,

65
00:03:45,768 --> 00:03:48,738

 you know, resonated well with 
people. 

66
00:03:48,738 --> 00:03:51,841
And we said, why don't we 
actually talk about it? 

67
00:03:51,841 --> 00:03:56,379
Talk about the journey, how we 
got to this point, some of the 

68
00:03:56,379 --> 00:03:59,520
technical details, um, so 
 that
people, yeah, understand a 

69
00:03:59,520 --> 00:04:03,010
little bit more, um, than, just 
reading the paper alone. 

70
00:04:03,010 --> 00:04:05,989
But maybe before we begin, why 
don't we just set the background

71
00:04:05,989 --> 00:04:09,461
and for the two of you 
 just to
give a brief overview, you know,

72
00:04:09,461 --> 00:04:12,503
who you are, where you work, 
what you're doing and, and yeah,

73
00:04:12,503 --> 00:04:15,739
your background, I guess, and 
then we can get into the paper. 

74
00:04:15,739 --> 00:04:19,103
So, Sid, did you want to go 
first? 

75
00:04:19,125 --> 00:04:21,536
Yes, I can go first. 
I'm Sid Mishra. 

76
00:04:21,536 --> 00:04:24,236
I'm a professor for 
computational applied 

77
00:04:24,236 --> 00:04:26,930
mathematics at ETH Zurich in 
Switzerland. 

78
00:04:27,251 --> 00:04:30,293
And I run the computational 
applied mathematics laboratory 

79
00:04:30,293 --> 00:04:33,179
here. 
And what I do is my basic 

80
00:04:33,179 --> 00:04:35,325
background is in physics and 
math. 

81
00:04:35,325 --> 00:04:37,976
I did my undergraduate degrees 
in those subjects. 

82
00:04:37,977 --> 00:04:41,638
Then I shifted more or less to 
applied and computational math. 

83
00:04:41,799 --> 00:04:46,179
Then after my PhD, a lot of the 
research was on numerical 

84
00:04:46,179 --> 00:04:49,154
methods, numerical analysis, 
 
scientific computing. applied it

85
00:04:49,154 --> 00:04:53,188
to different areas, 
astrophysics, fluid dynamics, um

86
00:04:53,188 --> 00:04:55,492
material science, 
 different 
topics. 

87
00:04:55,492 --> 00:05:00,254
And for the last five, six 
years, I've been working on 

88
00:05:00,254 --> 00:05:04,574
using AI for solving physics 
 
problems and with some success. 

89
00:05:04,935 --> 00:05:08,031
First with physics-informed 
neural networks, and now for the

90
00:05:08,031 --> 00:05:10,783
last several years with 
 neural
operators, diffusion models, 

91
00:05:10,783 --> 00:05:14,068
foundation models, and so on. 
So it's been a lot of fun. 

92
00:05:14,306 --> 00:05:17,344
Yeah, nice. 
Johannes? 

93
00:05:17,944 --> 00:05:22,451
Hi, I'm Johannes Brandstetter. 
I am a professor at JKU in Linz 

94
00:05:22,451 --> 00:05:26,374
and I'm also a co-founder and 
chief scientist 
 of Emmi AI. 

95
00:05:26,542 --> 00:05:30,657
I've been in the field of 
surrogates for uh engineering, 

96
00:05:30,657 --> 00:05:35,178
if you want to call it like 
that, 
 for approximately three 

97
00:05:35,178 --> 00:05:37,960
years. 
I'm thinking about how to build 

98
00:05:37,960 --> 00:05:41,515
these surrogates that they 
really fit to industrial scale 


99
00:05:41,523 --> 00:05:43,095
problems, industrial size 
problems. 

100
00:05:43,136 --> 00:05:48,464
That is where I met uh you Neil,
think, approximately two years 

101
00:05:48,464 --> 00:05:52,859
ago, and also where the paths of
me and and Sid 
 were 

102
00:05:52,859 --> 00:05:54,345
overlapping. 
And it's a tremendous pleasure 

103
00:05:54,345 --> 00:05:57,481
for me to have a paper with both
of you because it was 
 very 

104
00:05:57,481 --> 00:06:00,608
high up on the bucket list to 
have a paper with each of you. 

105
00:06:00,608 --> 00:06:04,640
But to have it combined is like 
really true honor for me. 

106
00:06:05,516 --> 00:06:12,150
Yeah, I think, think we met 
because you gave it a finger and

107
00:06:12,150 --> 00:06:16,577
like an eye clear workshop. 
I think I was organizing, but I 

108
00:06:16,577 --> 00:06:21,689
remember where we really spoke 
was at a, um, what was the 
 

109
00:06:21,697 --> 00:06:26,375
restaurant was it pizza or 
Italian or something in, um, in 

110
00:06:26,375 --> 00:06:29,696
Silicon Valley, um, where we, 
where we met up. 

111
00:06:29,696 --> 00:06:32,515
And that was where I think we 
first were like properly getting

112
00:06:32,515 --> 00:06:34,577
into this topic. 
And that was actually. 

113
00:06:34,717 --> 00:06:37,550
Whilst I was still at AWS and I 
think I was just. 

114
00:06:37,550 --> 00:06:39,972
Was it like Christmas time or 
something or January time? 

115
00:06:39,972 --> 00:06:43,104
I can't remember what it was. 
So actually when I look back, 

116
00:06:43,104 --> 00:06:45,924
we'd sort of started having some
of these discussions almost 
 

117
00:06:45,932 --> 00:06:50,182
like a year ago, um, on this. 
And maybe what I thought we 

118
00:06:50,182 --> 00:06:55,122
could do is each of us maybe can
set the scene in what 
 

119
00:06:55,130 --> 00:06:57,022
prerequisite knowledge we 
brought, right? 

120
00:06:57,022 --> 00:07:03,288
What, before we really started 
to interact and, I guess, you 

121
00:07:03,288 --> 00:07:07,104
know, um, blend this knowledge 
together, which is, think 

122
00:07:07,104 --> 00:07:10,441
ultimately what the paper was. 
Maybe we can set the scene of 

123
00:07:10,441 --> 00:07:11,731
like what our thoughts were 
before. 

124
00:07:11,731 --> 00:07:16,553
So maybe if I can start it a 
little bit, I am not a machine 

125
00:07:16,553 --> 00:07:18,497
learning specialist. 
That's the first thing I'll put 

126
00:07:18,497 --> 00:07:22,969
my hand up and admit. am not 
anywhere like the two of you 

127
00:07:22,969 --> 00:07:27,633
really are like super deep in 
the knowledge, but I was 
 

128
00:07:27,641 --> 00:07:31,866
always intrigued on the, on how 
far AI could go. 

129
00:07:32,526 --> 00:07:37,615
And I sort of was thinking more 
from a very applied point of 

130
00:07:37,615 --> 00:07:39,870
view. 
You know, you, you start to 

131
00:07:39,870 --> 00:07:41,926
build these surrogates, you have
like 500 cases. 

132
00:07:41,926 --> 00:07:44,706
It has a, you know, a certain 
accuracy. 

133
00:07:45,286 --> 00:07:48,605
I like all of these ideas of 
scaling a little bit. 

134
00:07:48,605 --> 00:07:52,306
So like, if you take a, a 
solver, you know, it works well 

135
00:07:52,306 --> 00:07:55,714
for a simple problem, but will 

that same method work for a 

136
00:07:55,714 --> 00:07:58,870
really big problem? 
And we know if we do like a DNS 

137
00:07:58,870 --> 00:08:01,166
simulation, it works really well
for a small problem. 

138
00:08:01,208 --> 00:08:04,716
But the scaling means you can 
never really do it with a full 

139
00:08:04,716 --> 00:08:06,599
aircraft, regardless of how 
 
good it is. 

140
00:08:06,599 --> 00:08:10,951
And so with the AI, I was always
like tempted of, well, is this 

141
00:08:10,951 --> 00:08:14,532
just a matter of scale? 
You know, if you just throw more

142
00:08:14,532 --> 00:08:16,562
stuff at it, is that the 
solution? 

143
00:08:16,562 --> 00:08:20,163
Just more compute or is there 
something that ultimately stops 

144
00:08:20,163 --> 00:08:23,592
it? 
And I guess, um, the second bit 

145
00:08:23,592 --> 00:08:29,724
was when I did go to ICLR or 
NeurIPS, I was sort of 
 amazed 

146
00:08:29,724 --> 00:08:34,645
by the amazing talent. 
But at the same time, the sense 

147
00:08:34,645 --> 00:08:38,086
that they were two different 
worlds, that you were 
 speaking

148
00:08:38,086 --> 00:08:41,868
to these people who didn't have 
that practical sense of like 

149
00:08:41,868 --> 00:08:45,285
using CFD for engineering, you 
know, like actually designing 

150
00:08:45,285 --> 00:08:47,671
something, but we're looking 
more 
 theoretical. 

151
00:08:47,671 --> 00:08:51,251
And then you have the CFD side 
that were really good at that 

152
00:08:51,251 --> 00:08:54,551
stuff, but just had no clue 
 
about all this amazing work that

153
00:08:54,551 --> 00:08:57,461
was going on. 
And I think that's why I was, 

154
00:08:57,461 --> 00:09:00,965
I'm always excited when I speak 
to people like you two who 
 

155
00:09:00,973 --> 00:09:08,428
have much more of the So applied
math ML side and you can educate

156
00:09:08,428 --> 00:09:10,431
me. 
I feel like I'm absorbing off 

157
00:09:10,431 --> 00:09:12,620
you. 
Um, but how about you, Johannes?

158
00:09:12,620 --> 00:09:16,139
Where did you come into this? 
What was your, before we really 

159
00:09:16,139 --> 00:09:19,725
started to go in this sort of 
three way thing, what was 
 your

160
00:09:19,725 --> 00:09:21,528
thoughts? 
Yeah, approximately, think, no, 

161
00:09:21,528 --> 00:09:24,200
not approximately, pretty much 
exactly three years ago, 
 

162
00:09:24,208 --> 00:09:28,542
ChatGPT came out and at the very
same time I was at Microsoft 

163
00:09:28,542 --> 00:09:31,828
research and there I pushed 
really, really hard for these 

164
00:09:31,828 --> 00:09:34,082
foundation models. 
First, it was the ClimaX model 

165
00:09:34,082 --> 00:09:35,956
and then it was this Aurora 
model. 

166
00:09:36,197 --> 00:09:40,570
And the gist of it was basically
throwing as much data as 

167
00:09:40,570 --> 00:09:44,938
possible into the mix and trying

 to see what comes out, which 

168
00:09:44,938 --> 00:09:49,647
worked very well. um and which 
was a big motivation for me to 

169
00:09:49,647 --> 00:09:53,713
go into this engineering. 
However, there's a lot of uh 

170
00:09:53,713 --> 00:09:56,684
difficulties, uh differences. 
First of all, um Aurora worked 

171
00:09:56,684 --> 00:09:59,731
so well because it was all based
on vision transformers, 
 which 

172
00:09:59,731 --> 00:10:02,499
everyone at that point 
understood how they work and how

173
00:10:02,499 --> 00:10:06,550
to scale and how to operate. 
And there is not a lot of 

174
00:10:06,550 --> 00:10:09,465
control of the weather data 
because in the end, you download

175
00:10:09,465 --> 00:10:11,585

 some weather data from 
different uh offices at 

176
00:10:11,585 --> 00:10:14,578
different simulations and they 
kind of all model their similar 

177
00:10:14,578 --> 00:10:16,304
physics. 
So they have different 

178
00:10:16,304 --> 00:10:20,376
discretization schemes, but they
all try to have the same 

179
00:10:20,376 --> 00:10:22,411
turbulence modeling 
 schemes 
baked in. 

180
00:10:22,471 --> 00:10:25,765
However, when you go to 
engineering, you encounter two 

181
00:10:25,765 --> 00:10:29,770
problems on this axis. 
First, there is no scalable 

182
00:10:29,770 --> 00:10:34,351
architecture available for doing
a 200 million surface mesh, 
 

183
00:10:34,359 --> 00:10:37,403
volume mesh, Formula One car 
simulation. 

184
00:10:38,344 --> 00:10:41,214
So you have to build the 
scalable architectures and there

185
00:10:41,214 --> 00:10:44,646
is a big, big difference on the 
 data regime because data can 

186
00:10:44,646 --> 00:10:48,388
have so many aspects how you 
discretize what turbulence model

187
00:10:48,388 --> 00:10:52,920
you use, what machine you use, 
what 
 numerical scheme you use 

188
00:10:52,920 --> 00:10:58,266
and so on and so forth which 
makes this very very fragile and

189
00:10:58,266 --> 00:11:03,414
very hard to understand and on 
both of these fronts I tried to 

190
00:11:03,414 --> 00:11:08,166
do my research and I tried 
 to 
move forward and that's where 

191
00:11:08,166 --> 00:11:10,538
from my side the discussion 
started. 

192
00:11:11,694 --> 00:11:13,749
How about you, Sid? 
mean, obviously you've been 

193
00:11:13,749 --> 00:11:17,895
working on this for a while and 
there was like the Poseidon 
 

194
00:11:17,903 --> 00:11:21,000
paper and you've had these 
thoughts of foundational models,

195
00:11:21,000 --> 00:11:25,844
I guess. 
Yeah, so I come from it from 

196
00:11:25,844 --> 00:11:28,310
sort of similar viewpoint as 
Johannes. 

197
00:11:28,310 --> 00:11:31,411
That's why we discuss in general
very well. 

198
00:11:31,552 --> 00:11:35,332
So my perspective in the 
beginning was more I'm a 

199
00:11:35,332 --> 00:11:38,352
mathematician, so I wanted to 
 
essentially understand if 

200
00:11:38,352 --> 00:11:42,122
possible, rigorously prove how 
much of data is necessary for 

201
00:11:42,122 --> 00:11:46,438
and what is the model size that 
is necessary for an ML model or 

202
00:11:46,438 --> 00:11:50,365
an AI model for that matter to 

learn certain tasks, to excel at

203
00:11:50,365 --> 00:11:51,875
certain tasks, to generalize 
well. 

204
00:11:51,875 --> 00:11:55,993
So this was And I was not 
looking at vision and text like 

205
00:11:55,993 --> 00:11:59,833
what most people look at, but I 
was looking 
 at scientific data

206
00:11:59,833 --> 00:12:02,073
sets, physics, could be 
chemistry, and certainly 

207
00:12:02,073 --> 00:12:04,730
engineering, because this is 
what I've worked on for long 

208
00:12:04,730 --> 00:12:06,528
time. 
And this was sort of the 

209
00:12:06,528 --> 00:12:09,294
foundational question. 
And uh we already, some years 

210
00:12:09,294 --> 00:12:12,291
ago, had some theoretical 
results, which essentially said 

211
00:12:12,291 --> 00:12:15,657
 that, OK, things scale if you 
have enough data. 

212
00:12:15,657 --> 00:12:18,698
But as Johannes just expressed, 
you never have enough data. 

213
00:12:18,698 --> 00:12:22,536
mean, this is the... 
There's a challenge in some 

214
00:12:22,536 --> 00:12:24,916
sense, also the opportunity in 
this field. 

215
00:12:24,916 --> 00:12:27,767
And then this is where 
foundation models make sense. 

216
00:12:27,767 --> 00:12:30,946
The idea was that, you have a 
large corpus of pre-trained 

217
00:12:30,946 --> 00:12:33,540
data, and then maybe you can 
 
fine tune it. 

218
00:12:33,540 --> 00:12:38,010
And I have been working on this 
topic from different 

219
00:12:38,010 --> 00:12:42,477
perspectives, just interested in
the 
 fact that can AI systems 

220
00:12:42,477 --> 00:12:45,599
generalize to unseen 
circumstances, to unseen physics

221
00:12:45,599 --> 00:12:48,556
in particular, because that's my
interest. 

222
00:12:48,854 --> 00:12:51,867
And then of course, I discussed 
a little bit, Johannes 

223
00:12:51,867 --> 00:12:54,877
approximately a year back, I 
think 
 he was in Zurich. 

224
00:12:54,877 --> 00:12:58,462
And uh then to his credit, he's 
the one who somehow brought the 

225
00:12:58,462 --> 00:13:01,762
two of us together because 
 
Neil, I know your work very well

226
00:13:01,762 --> 00:13:05,338
from the past, but I never had 
the opportunity to sort of meet 

227
00:13:05,338 --> 00:13:07,544
you in person or interact with 
you before. 

228
00:13:07,544 --> 00:13:11,436
So I think it is Johannes who 
sort of deserves the credit. 

229
00:13:11,436 --> 00:13:16,234
Because you know, In a way, our 
knowledge base intersects pretty

230
00:13:16,234 --> 00:13:18,072
well. 
I know a little bit of CFD. 

231
00:13:18,072 --> 00:13:21,715
You are the real CFD expert. 
Johannes is a real ML expert. 

232
00:13:21,715 --> 00:13:25,333
I also know a little bit of ML. 
So we sort of complement each 

233
00:13:25,333 --> 00:13:28,489
other very well. 
But I think the questions that 

234
00:13:28,489 --> 00:13:32,502
sort of motivate us, drive us 
are very similar, I would 
 say,

235
00:13:32,502 --> 00:13:35,956
to some extent. 
So I thought maybe, um, we're 

236
00:13:35,956 --> 00:13:38,656
useful for people listening to 
this paper, Fluid Intelligence: 

237
00:13:38,656 --> 00:13:41,649
A Forward Look on AI Foundation 
Models in Computational Fluid 

238
00:13:41,649 --> 00:13:45,282
Dynamics, which by the way, I 
think it took us a while to 

239
00:13:45,282 --> 00:13:47,018
figure out what title we should 
have. 

240
00:13:47,018 --> 00:13:51,440
We had quite a few different 
titles. you know, maybe what we 

241
00:13:51,440 --> 00:13:54,877
should do is kind of go through 
it, section by section. 

242
00:13:55,018 --> 00:13:59,046
Obviously we're not going to be 
able to go in full depth, but 

243
00:13:59,046 --> 00:14:02,754
just describing maybe why, 
 you
know, we had some, some of the 

244
00:14:02,754 --> 00:14:04,954
things, and. 
And hopefully then by the end of

245
00:14:04,954 --> 00:14:07,439
it, people will get more like 
how we came to this 
 

246
00:14:07,447 --> 00:14:10,110
conclusion. 
Um, I think it's fair to say, 

247
00:14:10,110 --> 00:14:14,719
isn't it that we. 
Our, way the papers turned out 

248
00:14:14,719 --> 00:14:17,930
was not exactly how we went into
it. 

249
00:14:17,930 --> 00:14:25,212
It wasn't really the idea to do 
it this way. but I think it's 

250
00:14:25,212 --> 00:14:27,802
sort of organically, we kept 
having meetings and then we 

251
00:14:27,802 --> 00:14:31,426
would be like, 
 yes, for sure. 
This is how, you know, yes, 

252
00:14:31,426 --> 00:14:32,840
we've done it. 
Okay. 

253
00:14:32,840 --> 00:14:36,452
And then you're like, actually 
the data showing something 

254
00:14:36,452 --> 00:14:38,816
different. 
And, um, yeah, so it definitely 

255
00:14:38,816 --> 00:14:42,500
has been a bit of a journey, 
which is what all good 
 

256
00:14:42,508 --> 00:14:44,040
research should be. 
Right. 

257
00:14:44,040 --> 00:14:51,703
It was sort of live research 
over WhatsApp, essentially. but 

258
00:14:51,703 --> 00:14:57,420
yeah, maybe just, uh, I guess 
the first section, the whole CFD

259
00:14:57,420 --> 00:15:01,704
process, I think was 
 born out 
a little bit. 

260
00:15:02,578 --> 00:15:06,006
Um, or at least partially when, 
you know, Johannes and I were 

261
00:15:06,006 --> 00:15:08,856
talking originally, and I 
 
think there was this little bit 

262
00:15:08,856 --> 00:15:13,399
of sad, and I don't want to put 
words in your mouth hands, but I

263
00:15:13,399 --> 00:15:17,162
guess you don't come from a 
traditional CFD background. 

264
00:15:17,162 --> 00:15:21,805
So some of the stuff, you know, 
you, you kind of knew, but you 

265
00:15:21,805 --> 00:15:24,784
wasn't as obvious like the 
 
industrial side to it. 

266
00:15:24,784 --> 00:15:28,396
And I think that's when we 
realized that it might be useful

267
00:15:28,396 --> 00:15:32,298
for people to have almost that 

little bit of a, yeah, a go-to 

268
00:15:32,298 --> 00:15:36,058
reference. that highlighted this
idea that it isn't just uh a 

269
00:15:36,058 --> 00:15:39,163
single PDE that you just somehow
need 
 to model. 

270
00:15:39,484 --> 00:15:42,648
And so it is difficult and 
probably some people who are 

271
00:15:42,648 --> 00:15:45,518
reading this and they look at 
the 
 geometry, the physics 

272
00:15:45,518 --> 00:15:48,675
modeling, the meshing, people 
I'm sure will find holes or will

273
00:15:48,675 --> 00:15:50,633
find bits that couldn't be 
covered. 

274
00:15:50,633 --> 00:15:54,169
You would sort of need to write 
a textbook and people have wrote

275
00:15:54,169 --> 00:15:58,183
textbooks, right, on 
 this. 
But I think what we were trying 

276
00:15:58,183 --> 00:16:02,112
to get across is just how broad 
the input space is. 

277
00:16:02,306 --> 00:16:05,690
And I think that's where you, 
Johannes and Sid like this sort 

278
00:16:05,690 --> 00:16:08,503
of distributional way of 
 
looking at this, like trying to 

279
00:16:08,503 --> 00:16:13,826
look at it in a way that would 
set it up to have some link to 

280
00:16:13,826 --> 00:16:17,910
the LLMs, right? 
That was the way you wanted it, 

281
00:16:17,910 --> 00:16:21,423
was it right, Johannes? 
Yeah, fully 100%. 

282
00:16:21,423 --> 00:16:24,661
I think that for machine 
learning people, the most 

283
00:16:24,661 --> 00:16:28,610
important equation of the first 
half of the paper is equation 

284
00:16:28,610 --> 00:16:31,685
nine. 
So what equation nine is, we are

285
00:16:31,685 --> 00:16:35,177
basically saying you can 
deconstruct the CFD process into

286
00:16:35,177 --> 00:16:39,052

 an input vector, which uh 
well, which has the turbulence 

287
00:16:39,052 --> 00:16:42,539
model, the geometry, the meshing
and all these parts in it. 

288
00:16:42,539 --> 00:16:47,124
And this input vector you can 
use to have a look at the input 

289
00:16:47,124 --> 00:16:50,738
output relations. which makes it
very clear that if you want to 

290
00:16:50,738 --> 00:16:53,698
have a foundation model, it 
surely needs to 
 capture all 

291
00:16:53,698 --> 00:16:55,474
these variations in the input 
vector. 

292
00:16:55,474 --> 00:16:59,882
It also makes it clear that if 
you fix a few of these 

293
00:16:59,882 --> 00:17:03,272
conditions, for example, if you 
use 
 same inflow condition or 

294
00:17:03,272 --> 00:17:06,662
same boundary condition, that 
the space of variation gets just

295
00:17:06,662 --> 00:17:11,133
much smaller. 
And the first section, uh which 

296
00:17:11,133 --> 00:17:16,038
was mostly written by Neil, is 
actually really, um thought to 

297
00:17:16,038 --> 00:17:19,756
explain for machine learning 
people how to come to this input

298
00:17:19,756 --> 00:17:23,463
vector, which 
 makes a lot of 
sense because suddenly you don't

299
00:17:23,463 --> 00:17:25,484
think of complex CFD simulation 
anymore. 

300
00:17:25,484 --> 00:17:29,265
You think of input output 
relations and that makes it also

301
00:17:29,265 --> 00:17:32,695
easier to compare existing data 
 sets and to understand which 

302
00:17:32,695 --> 00:17:36,468
data sets can be actually mixed 
and which they're mixing is 

303
00:17:36,608 --> 00:17:39,928
resulting to very disjoint 
distributions. 

304
00:17:40,649 --> 00:17:44,507
And the distribution point of 
view is Sid's way of thinking, I

305
00:17:44,507 --> 00:17:47,843
got from him. 
Yeah, so maybe just to add to 

306
00:17:47,843 --> 00:17:50,363
the mix, actually engineers like
this thinking, right? 

307
00:17:50,363 --> 00:17:53,499
A systemic thinking. 
So you can imagine that there is

308
00:17:53,499 --> 00:17:57,332
a system and to a certain system
we feed some inputs and 
 we 

309
00:17:57,332 --> 00:18:00,566
have some outputs at the end of 
the system's work, right? 

310
00:18:00,566 --> 00:18:03,656
So if you think of machine 
learning or learning in 

311
00:18:03,656 --> 00:18:07,055
particular, it's sort of task 
specific 
 in that sense that a 

312
00:18:07,055 --> 00:18:10,753
learning system has an input or 
a set of inputs, input vectors, 

313
00:18:10,753 --> 00:18:14,110
and a set of outputs, the output
vector, output functions, 

314
00:18:14,110 --> 00:18:17,722
fields, whatever you call it. 
And equation nine is sort of 

315
00:18:17,722 --> 00:18:19,758
providing that, right? 
Where the distributional 

316
00:18:19,758 --> 00:18:23,740
perspective comes in is that you
cannot sort of have everything 


317
00:18:23,748 --> 00:18:27,508
under the sun, right? 
Means you have to sample from a 

318
00:18:27,508 --> 00:18:29,497
distribution. 
You can make this distribution 

319
00:18:29,497 --> 00:18:33,185
as broad as possible. 
But once you sample from an 

320
00:18:33,185 --> 00:18:36,336
input distribution, then in some
sense, your output 
 

321
00:18:36,344 --> 00:18:39,087
distribution is highly 
conditioned on it, right? 

322
00:18:39,087 --> 00:18:43,556
Because you might add some noise
at the time of measurement of 

323
00:18:43,556 --> 00:18:46,981
your system. 
So it's a very natural sort of 

324
00:18:46,981 --> 00:18:49,279
mathematical perhaps or 
algorithmic thinking about 

325
00:18:49,279 --> 00:18:52,719
machine 
 learning that you 
think of sampling from a 

326
00:18:52,719 --> 00:18:56,085
distribution, take the samples, 
feed the inputs into system, 

327
00:18:56,085 --> 00:19:00,583
observe its output, and then all
the AI system does is try to 

328
00:19:00,583 --> 00:19:03,696
learn how the 
 system sort of 
propagates or evolves. 

329
00:19:03,696 --> 00:19:06,977
And this was sort of the 
thinking that I always keep and 

330
00:19:06,977 --> 00:19:10,253
I also teach it to my students 
in 
 class that think of 

331
00:19:10,253 --> 00:19:12,368
everything in terms of systems. 
What's your input? 

332
00:19:12,368 --> 00:19:14,479
What's your output? 
What's your input distribution? 

333
00:19:14,479 --> 00:19:17,896
What's your output distribution?
And then there's a lot of 

334
00:19:17,896 --> 00:19:20,950
clarity because without this 
clarity, then we don't know what

335
00:19:20,950 --> 00:19:23,463

 you're talking about, right? 
It becomes vague. 

336
00:19:23,463 --> 00:19:26,484
And this is what we, I want that
mathematical precision. 

337
00:19:26,484 --> 00:19:30,916
And this is, think useful to 
have so that we know what our 

338
00:19:30,916 --> 00:19:34,038
target is. 
And probably, maybe I'll use 

339
00:19:34,038 --> 00:19:38,598
this point to jump forward a 
little bit, just cause it seems 

340
00:19:38,598 --> 00:19:42,771
 an opportune time around the 
appendix B and the whole notion.

341
00:19:42,771 --> 00:19:47,543
Cause as soon as, just to say 
appendix B is this idea of FLOPS

342
00:19:47,543 --> 00:19:51,117
per cell per step. 
And I felt that was important 

343
00:19:51,117 --> 00:19:54,736
because one of the objectives 
coming into this was to some 
 

344
00:19:54,744 --> 00:19:58,346
way quantify the, you know, if 
you're generating a data set. 

345
00:19:58,838 --> 00:20:00,759
Or you're trying to calculate 
the costs. 

346
00:20:00,759 --> 00:20:03,961
How do you do it? 
And it is obviously massively 

347
00:20:03,961 --> 00:20:07,081
dependent on those CFD inputs, 
you know, that equation nine. 

348
00:20:07,081 --> 00:20:12,563
But one of the things that it is
as well is the code itself. 

349
00:20:12,644 --> 00:20:16,125
And this is a, I still don't 
think we have it perfect. 

350
00:20:16,125 --> 00:20:23,307
Um, but it was an attempt to 
say, right, if you are running a

351
00:20:23,307 --> 00:20:28,940
RANS solver, you are going 
 to 
pick certain inputs to match. 

352
00:20:29,272 --> 00:20:31,183
Like they're not, they're never 
done in isolation. 

353
00:20:31,183 --> 00:20:34,243
If you pick RANS, you're 
probably going to pick an 

354
00:20:34,243 --> 00:20:36,997
unstructured grid because you're

 probably going to be doing 

355
00:20:36,997 --> 00:20:39,445
complex geometries that you need
to run fast. 

356
00:20:39,865 --> 00:20:43,539
And because it's a steady state,
you're probably going to go 

357
00:20:43,539 --> 00:20:45,206
implicit and because it's 
 
unstructured. 

358
00:20:45,206 --> 00:20:46,457
So there's these sort of 
secrets. 

359
00:20:46,457 --> 00:20:50,136
Yes, you could technically pick 
a different combination, but 

360
00:20:50,136 --> 00:20:53,623
then they're more extreme. 
So the idea was to say, well, if

361
00:20:53,623 --> 00:20:55,639
you do that, you're more likely 
to do that. 

362
00:20:55,639 --> 00:21:00,167
So let's clump that as one 
category. which was how we had 

363
00:21:00,167 --> 00:21:02,399
that sort of implicit RANS 
unstructured. 

364
00:21:02,399 --> 00:21:06,511
It's sort of, if you look at 
many of the ISV codes out today,

365
00:21:06,511 --> 00:21:08,855
they have sort of centered 
 
around that choice. 

366
00:21:09,055 --> 00:21:14,095
But then if you're going to be 
doing like a half a billion cell

367
00:21:14,095 --> 00:21:17,695
LES, whilst you could do 
 it 
within an implicit unstructured 

368
00:21:17,695 --> 00:21:20,926
code, it's not the optimum way 
of doing it. 

369
00:21:20,926 --> 00:21:23,828
And it would have massively 
misrepresented it. 

370
00:21:23,888 --> 00:21:28,494
So we thought, let's pick. 
You know, and there are examples

371
00:21:28,494 --> 00:21:32,894
of companies out there who do 
have, you know, an 
 explicit, 

372
00:21:32,894 --> 00:21:36,094
because now you are trying to 
time resolve. 

373
00:21:36,194 --> 00:21:38,134
So using implicit doesn't make 
as much sense. 

374
00:21:38,134 --> 00:21:40,751
You're probably going to do 
Cartesian because in this we 

375
00:21:40,751 --> 00:21:42,640
picked wall-modelled LES. 
So you don't need to resolve the

376
00:21:42,640 --> 00:21:45,159
boundary layer. 
So it's fine to use Cartesian 

377
00:21:45,159 --> 00:21:48,902
and it's a GPU solver. we, we, 
know, and again, they're 

378
00:21:48,902 --> 00:21:51,763
categorical choice in a way 
they're picking which inputs 
 

379
00:21:51,771 --> 00:21:54,652
and clumping them. 
And so there's a million other 

380
00:21:54,652 --> 00:21:57,239
combinations. 
But the idea was to say, and 

381
00:21:57,239 --> 00:21:59,658
maybe we'll get onto that in a 
minute. 

382
00:22:00,718 --> 00:22:03,872
A time step and a cell is not 
equivalent because in the 

383
00:22:03,872 --> 00:22:06,492
explicit, you're doing hundreds 
of 
 thousands of time steps and

384
00:22:06,492 --> 00:22:09,898
yet in the steady state, you may
be just doing hundreds or low 

385
00:22:09,898 --> 00:22:14,482
thousands, but that flops per 
step per cell, which is like how

386
00:22:14,482 --> 00:22:19,058
much it costs to do it is 
 the 
important multiplier in this. 

387
00:22:19,258 --> 00:22:24,952
And, um, it's one that I think 
there's lots of holes in it. 

388
00:22:24,952 --> 00:22:27,647
And you could calculate it in 
different ways, but I still 

389
00:22:27,647 --> 00:22:29,602
think it at least gives you a 
 
picture. 

390
00:22:29,602 --> 00:22:31,765
And I'm, I'm hoping that people 
listen to it. 

391
00:22:31,765 --> 00:22:37,152
I'd love to see people test that
theory, you know, like how close

392
00:22:37,152 --> 00:22:40,670
are we to it? 
And it'd be good to get 

393
00:22:40,670 --> 00:22:43,069
feedback, you know, if there's 
certain things people disagree 

394
00:22:43,069 --> 00:22:45,995
with 
 or, know, which is, 
guess, part of the point of 

395
00:22:45,995 --> 00:22:47,591
putting a pre-print out, isn't 
it? 

396
00:22:47,591 --> 00:22:52,954
It's to say, here's something 
give us feedback, but we got to 

397
00:22:52,954 --> 00:22:56,079
that bit. 
CFD and I was pushing for this 

398
00:22:56,079 --> 00:22:59,895
more and saying, I still think 
we need to explain AI, you 
 

399
00:22:59,903 --> 00:23:03,711
know, to the AI people reading 
this, they're like, I know that.

400
00:23:03,892 --> 00:23:08,138
So that fundamentals bit, how 
did you sort of think about 

401
00:23:08,138 --> 00:23:12,384
framing the whole, I think the, 
 um, maybe the sentence that 

402
00:23:12,384 --> 00:23:16,238
depicts it the most is CFD is 
not a language. 

403
00:23:16,779 --> 00:23:19,801
I can't remember which of you 
wrote that, but I that was 

404
00:23:19,801 --> 00:23:21,176
quite. 
I did, but... 

405
00:23:24,498 --> 00:23:26,840
Maybe let me add before we jump 
to that. 

406
00:23:26,840 --> 00:23:29,021
Let me add one thing for the ML 
community. 

407
00:23:29,021 --> 00:23:33,658
So important to understand and 
this is also when I started to 

408
00:23:33,658 --> 00:23:37,518
realize this, I don't know, 
 
one, two years ago in 

409
00:23:37,518 --> 00:23:41,378
engineering and CFD, it's not 
that you make choices on 

410
00:23:41,378 --> 00:23:44,128
fidelity and that basically 
results in everything what Neil 

411
00:23:44,128 --> 00:23:46,393
just said. 
It also problem specific. 

412
00:23:46,393 --> 00:23:50,653
So if you try to simulate the 
plane, which is in cruise 

413
00:23:50,653 --> 00:23:52,777
condition, certain CFD choices 
are 
 okay. 

414
00:23:52,777 --> 00:23:56,447
So you don't need higher 
resolution or better turbulence 

415
00:23:56,447 --> 00:24:00,146
model because certain turbulence
model are 
 capturing what you 

416
00:24:00,146 --> 00:24:05,071
actually want to have captured 
and then you can do all the 

417
00:24:05,071 --> 00:24:09,686
results in your parameter vector
is sort of limited. it's the, it

418
00:24:09,686 --> 00:24:13,774
really depends on the problem. 
It's very different than what we

419
00:24:13,774 --> 00:24:17,131
usually know from machine 
learning that you have, well, 
 

420
00:24:17,139 --> 00:24:21,597
my high quality data and my low 
quality data, there is quality 

421
00:24:21,597 --> 00:24:23,938
always comes with the problem 
set up. 

422
00:24:23,938 --> 00:24:27,554
And this is very important to 
convey to the machine learning 

423
00:24:27,554 --> 00:24:30,506
community that it really 
 
depends on the problem what 

424
00:24:30,506 --> 00:24:33,568
you're using. 
And that is, yeah, that's 

425
00:24:33,568 --> 00:24:40,354
already the going to the ML part
of things. um I think it was a 

426
00:24:40,354 --> 00:24:45,726
quite eye-opener discussing 
these LLMs. uh what token means 

427
00:24:45,726 --> 00:24:49,514
for LLMs? 
I in LLM, everything is so... 

428
00:24:50,196 --> 00:24:52,476
straightforward in a way you 
have your documents and your 

429
00:24:52,476 --> 00:24:54,974
documents you have your tokens 

you have a bunch of documents 

430
00:24:54,974 --> 00:24:57,471
you have many tokens this is how
you construct your learning 

431
00:24:57,471 --> 00:25:02,580
tasks and then we tried to map 
that to CFD which was much 

432
00:25:02,580 --> 00:25:06,500
harder because you have this 
 
data sets where you have 

433
00:25:06,500 --> 00:25:10,420
suddenly this huge geometries 
with half a billion meshes or 

434
00:25:10,420 --> 00:25:16,556
whatever so is this like one 
data point is this uh should you

435
00:25:16,556 --> 00:25:22,043
count the tokens but then 
 the 
Additionally, can subsample many

436
00:25:22,043 --> 00:25:25,046
different input combinations 
from this data point. 

437
00:25:25,046 --> 00:25:27,366
So how is this all coming 
together? 

438
00:25:27,366 --> 00:25:30,466
And at that point, we were 
really saying, okay, let's 

439
00:25:30,466 --> 00:25:34,186
really write down what people do
in 
 language and then make the 

440
00:25:34,186 --> 00:25:36,046
connection why CFD is not 
language. 

441
00:25:36,046 --> 00:25:39,306
And this is, I think, where Sid 
should comment. 

442
00:25:40,086 --> 00:25:42,448
Yeah, so why is CFD not 
language? 

443
00:25:42,448 --> 00:25:46,590
Well, for starters, I think uh 
there are so many obvious 

444
00:25:46,590 --> 00:25:49,492
differences, right? 
So in language modeling, you 

445
00:25:49,492 --> 00:25:53,792
have this entire notion of 
tokenization, which is a very 

446
00:25:53,792 --> 00:25:58,186
sort 
 of clear paradigm, right?
So what you have is you have 

447
00:25:58,186 --> 00:26:02,026
these words or pieces of words, 
and then you convert them 
 into

448
00:26:02,026 --> 00:26:03,946
vectors through the process of 
tokenization. 

449
00:26:03,946 --> 00:26:07,389
And in fact, you convert them 
into entries of a codebook 

450
00:26:07,389 --> 00:26:10,149
through what is called quantized
tokenization. 

451
00:26:10,149 --> 00:26:14,169
So your tokens are essentially 
living in some very large 

452
00:26:14,169 --> 00:26:18,585
codebook and they are sort of 
 
entries on that register, right?

453
00:26:18,585 --> 00:26:22,444
And then all that we do in 
language model is given a 

454
00:26:22,444 --> 00:26:25,333
distribution on tokens, you sort
of 
 sample from the 

455
00:26:25,333 --> 00:26:27,580
distribution, conditional 
distribution of the next token, 

456
00:26:27,580 --> 00:26:30,085
right? 
So it's very sort of uh 

457
00:26:30,085 --> 00:26:32,264
mathematically clear cut in some
sense. 

458
00:26:32,264 --> 00:26:33,986
In a way, it's very 
non-mathematical because 

459
00:26:33,986 --> 00:26:35,953
language, but thanks to 
tokenization, we have 
 been 

460
00:26:35,953 --> 00:26:39,385
able to push that into a very 
sort of occurred learning 

461
00:26:39,385 --> 00:26:42,490
objective. 
This is different in CFD or in 

462
00:26:42,490 --> 00:26:45,667
physics in general, right, 
because for us the learning 
 

463
00:26:45,675 --> 00:26:49,188
task, you everything is in 
equation nine in some sense, 

464
00:26:49,188 --> 00:26:51,832
that is our master equation and 
so on. 

465
00:26:51,832 --> 00:26:56,585
And what does it say that when 
you sample, let's say that we 

466
00:26:56,585 --> 00:26:59,505
fixed the categorical 
 
variables, this is always the 

467
00:26:59,505 --> 00:27:03,549
case you have RANS or LES or the
structured unstructured these 

468
00:27:03,549 --> 00:27:07,000
sort of categorical choices. 
Once you fix the categorical 

469
00:27:07,000 --> 00:27:11,016
choice, uh Then you condition 
that distribution, input 

470
00:27:11,016 --> 00:27:14,942
distribution upon this category.
And then when you sample from 

471
00:27:14,942 --> 00:27:17,398
this, you are essentially 
sampling functions, the shape, 

472
00:27:17,398 --> 00:27:20,468
for 
 instance, right, of your 
object, or the flow conditions, 

473
00:27:20,468 --> 00:27:23,343
which could be vectors or even 
single parameters, boundary 

474
00:27:23,343 --> 00:27:25,751
conditions and so on and so 
forth. 

475
00:27:25,751 --> 00:27:29,267
And your output could be a 
solution field, could be a sort 

476
00:27:29,267 --> 00:27:31,904
of quantity of interest, drag, 

lift, and so on. 

477
00:27:31,904 --> 00:27:34,385
So the learning objective is 
very different. 

478
00:27:34,385 --> 00:27:36,236
Just looking at equation number 
nine. 

479
00:27:36,236 --> 00:27:40,814
So we cannot Maybe someday we 
will be able to quantize 

480
00:27:40,814 --> 00:27:43,934
everything so that everything is
just written in 
 terms of 

481
00:27:43,934 --> 00:27:47,054
distributions of tokens and we 
can do some next token 

482
00:27:47,054 --> 00:27:49,115
prediction. 
But this is unclear to me at 

483
00:27:49,115 --> 00:27:52,365
least and I have worked quite a 
bit on this whether we have 
 

484
00:27:52,373 --> 00:27:55,355
the accuracy because you know 
it's not enough that the word is

485
00:27:55,355 --> 00:27:59,551
close enough to what because as 
humans we have a lot of slack in

486
00:27:59,551 --> 00:28:04,028
how we understand right, but if 
the drag is 20 
 % off we are 

487
00:28:04,028 --> 00:28:07,614
finished. 
So we have to be So to me, think

488
00:28:07,614 --> 00:28:11,334
it's very important to remember 
that the sort of input-output 

489
00:28:11,334 --> 00:28:15,420
setting, the 
 learning task is 
very different and that leads to

490
00:28:15,420 --> 00:28:18,848
a very different kind of 
formulation, which is what there

491
00:28:18,848 --> 00:28:21,764
are many similarities, right? 
Means, after all, means there 

492
00:28:21,764 --> 00:28:24,508
are more similarities and 
differences, but the differences

493
00:28:24,508 --> 00:28:28,048

 here are very, very important.
And I think the biggest 

494
00:28:28,048 --> 00:28:30,968
difference is that our 
distributions, our inputs, our 

495
00:28:30,968 --> 00:28:34,648
outputs are 
 very different. 
And just to keep that spirit is,

496
00:28:34,648 --> 00:28:38,219
Because as Johannes argued 
means, what does it mean? 

497
00:28:38,219 --> 00:28:44,661
See, if you take a huge corpus 
in language, you have a lot of 

498
00:28:44,661 --> 00:28:46,662
tokens. 
But if you look at the flow 

499
00:28:46,662 --> 00:28:49,082
field, a lot of it is going to 
be void, right? 

500
00:28:49,082 --> 00:28:51,162
So it's not very useful 
information. 

501
00:28:51,162 --> 00:28:55,453
So you need that sort of global 
coupling, which is uh probably 

502
00:28:55,453 --> 00:28:59,527
different from language. 
Maybe one day when you have very

503
00:28:59,527 --> 00:29:02,525
good tokenization for physics, 
this might change. 

504
00:29:02,525 --> 00:29:04,926
But at the moment, we are not 
yet there. 

505
00:29:05,162 --> 00:29:08,072
I don't know if we ever get 
there, we are not yet there. 

506
00:29:08,878 --> 00:29:12,068
And then I guess the next 
section was, again, it's 

507
00:29:12,068 --> 00:29:15,892
impossible for us to do a 
complete job 
 of it because our

508
00:29:15,892 --> 00:29:19,595
paper's dedicated to just this. 
But really that review of, I 

509
00:29:19,595 --> 00:29:24,171
mean, should say, I mean, we put
a, we had a much longer 
 

510
00:29:24,179 --> 00:29:28,735
version of this at one point, 
but then we cut it down, which 

511
00:29:28,735 --> 00:29:31,680
was the different use cases of 
AI. 

512
00:29:31,680 --> 00:29:37,140
And I think we were conscious 
that we are focusing on the, I 

513
00:29:37,140 --> 00:29:42,002
guess, surrogate use case. 
But it is fair to say that there

514
00:29:42,002 --> 00:29:46,778
are other use cases of AI for 
CFD around like, you know, 
 

515
00:29:46,786 --> 00:29:49,557
post-processing vision stuff, 
you know, trying to 

516
00:29:49,557 --> 00:29:52,277
automatically find patterns or 
initializing solutions, or that 

517
00:29:52,277 --> 00:29:55,956
I think the most famous one in 
the CFD world that I'm not 
 

518
00:29:55,964 --> 00:29:59,918
sure if both of you have tracked
much, but in some ways I feel 

519
00:29:59,918 --> 00:30:02,811
has hurt AI, which was the 
turbulence modelling. 

520
00:30:02,811 --> 00:30:06,534
I think there was this idea 
that, and I think, like Chris 

521
00:30:06,534 --> 00:30:08,704
Rumsey and Spalart and people 
 
from NASA. 

522
00:30:08,910 --> 00:30:11,150
I had some of these 
turbulence-modelling workshops 

523
00:30:11,150 --> 00:30:14,604
and, Karthik Duraisamy was one 
of the 
 first for this. think 

524
00:30:14,604 --> 00:30:17,898
Turbulence Modeling in the Age 
of Data or I think it was that 

525
00:30:17,898 --> 00:30:20,998
title, something like 
 that. 
And, um, really great piece of 

526
00:30:20,998 --> 00:30:25,392
work and it sort of made sense 
that a turbulence model being so

527
00:30:25,392 --> 00:30:28,768

 empirical in some ways tuned, 
hand tuned to coefficients that 

528
00:30:28,768 --> 00:30:33,380
ML could do a better job. 
And, um, you know, for years 

529
00:30:33,380 --> 00:30:37,395
people tried it and I think even
today it never really 
 

530
00:30:37,403 --> 00:30:40,721
generalized and. 
And when I speak to people who 

531
00:30:40,721 --> 00:30:45,173
are maybe not deep ML people, 
that sits in their mind as 
 

532
00:30:45,181 --> 00:30:48,882
like almost AI thought that it 
could build the ultimate 

533
00:30:48,882 --> 00:30:52,247
turbulence model. 
And so I think that's why, at 

534
00:30:52,247 --> 00:30:55,990
least personally, and I think 
you both agree, it's, steered 
 

535
00:30:55,998 --> 00:31:00,104
it towards a surrogate modeling 
side because it felt like a 

536
00:31:00,104 --> 00:31:04,872
well-defined problem that has 
got probably the most to gain. 

537
00:31:04,872 --> 00:31:08,244
know, like the, idea, like we 
say, if you build a surrogate, 

538
00:31:08,290 --> 00:31:11,451
The inference is in what seconds
or less than second. 

539
00:31:11,671 --> 00:31:15,893
Nowadays it can give you full 
volume, full surface prediction.

540
00:31:16,714 --> 00:31:20,167
and making an improved 
turbulence model is great, but 

541
00:31:20,167 --> 00:31:24,380
you're still going to be bound 
by the 
 same time constraints 

542
00:31:24,380 --> 00:31:28,567
that a CFD simulation gives you.
Um, so I think it's just, we, 

543
00:31:28,567 --> 00:31:30,719
it's probably, we want to make 
that clear. 

544
00:31:30,719 --> 00:31:35,785
This is not a, of all AI for 
CFD, this is in some way a 

545
00:31:35,785 --> 00:31:38,830
subset. 
Um, of the surrogate modeling, 

546
00:31:38,830 --> 00:31:41,131
which is, guess, the foundation 
model question. 

547
00:31:42,158 --> 00:31:45,515
Could I say something because 
since you raised this, maybe I 

548
00:31:45,515 --> 00:31:49,043
get this off my chest. 
I've always wondered or often 

549
00:31:49,043 --> 00:31:53,043
wondered why don't turbulence 
closure models with AI work so 


550
00:31:53,051 --> 00:31:55,859
well. 
And there is probably an obvious

551
00:31:55,859 --> 00:31:59,833
reason for that. 
And let me sort of state what I 

552
00:31:59,833 --> 00:32:01,908
believe could be a reason, 
right? 

553
00:32:01,908 --> 00:32:08,492
So the reason is that this 
mapping that takes data to a 

554
00:32:08,492 --> 00:32:12,636
turbulence closure model is a 
very, badly behaved operator if 

555
00:32:12,636 --> 00:32:15,948
it's a map, very, very 
 badly 
behaved map. 

556
00:32:15,948 --> 00:32:19,140
There are many, choices that you
can make, many knobs that you 

557
00:32:19,140 --> 00:32:22,328
can tune in the turbulence 
 
model that can give you the same

558
00:32:22,328 --> 00:32:25,518
flow outcome, right? 
So I think it's extremely 

559
00:32:25,518 --> 00:32:28,827
ill-behaved operator. 
And in a sense, is not making 

560
00:32:28,827 --> 00:32:31,851
use of, because eventually 
people are learning like three, 

561
00:32:31,851 --> 00:32:35,107
 four parameters, right? 
Or five, six parameters in these

562
00:32:35,107 --> 00:32:38,683
models. 
Machine learning is about big 

563
00:32:38,683 --> 00:32:41,120
things. 
If you use neural networks to 

564
00:32:41,120 --> 00:32:44,261
learn a one-dimensional 
function, it will always do very

565
00:32:44,261 --> 00:32:46,006

 poorly compared to its 
competition. 

566
00:32:46,006 --> 00:32:48,667
You can just take a polynomial 
and get a better fit. 

567
00:32:48,667 --> 00:32:53,033
So machine learning only works 
in very large dimensions, large 

568
00:32:53,033 --> 00:32:55,733
scale. 
And this is why I think the bad 

569
00:32:55,733 --> 00:32:58,563
behavior of the underlying map 
and the fact that you're 
 

570
00:32:58,571 --> 00:33:01,955
trying to solve a low or an 
artificially low, I think the 

571
00:33:01,955 --> 00:33:04,187
problem is really high 
dimensional, but you're trying 

572
00:33:04,187 --> 00:33:07,094
to solve it in this low 
dimensional setting. meant that 

573
00:33:07,094 --> 00:33:10,608
the chances of it generalizing 
beyond the specific flow regime 

574
00:33:10,608 --> 00:33:14,352
are very, very 
 low. perhaps 
that's one of the reasons. 

575
00:33:14,352 --> 00:33:18,180
Whereas in a surrogate model, on
the other hand, it's a really 

576
00:33:18,180 --> 00:33:21,599
high dimensional problem. 
Your input vector could be 

577
00:33:21,599 --> 00:33:24,837
millions of dimensions. 
Your output vector could be a 

578
00:33:24,837 --> 00:33:27,756
billion dimensions, Neil. 
In your DrivAerML data set, for 

579
00:33:27,756 --> 00:33:31,496
instance, if you learn the 
volume field, this is over a 
 

580
00:33:31,504 --> 00:33:35,630
billion points in principle. 
And this is where I think AI or 

581
00:33:35,630 --> 00:33:38,404
ML models can excel compared to 
turbulence modeling. 

582
00:33:38,404 --> 00:33:42,316
So our choice was perhaps based 
more on pragmatism, but I think 

583
00:33:42,316 --> 00:33:44,924
there is a deeper 
 underlying 
reason behind that. 

584
00:33:45,805 --> 00:33:48,462
Yeah, I am. 
And there's two more things 

585
00:33:48,462 --> 00:33:53,077
which we did not consider. 
So we were always in all our 

586
00:33:53,077 --> 00:33:56,922
learning tasks, we were always 
in the large data limit. 

587
00:33:56,922 --> 00:34:00,066
So basically, the model 
shouldn't learn any spurious 

588
00:34:00,066 --> 00:34:03,210
correlation, but it should 
really 
 generalize because it 

589
00:34:03,210 --> 00:34:07,138
has enough data, which is 
usually in this surrogate tasks 

590
00:34:07,138 --> 00:34:10,649
where you have 100 samples, 50 
samples, and then you should 

591
00:34:10,649 --> 00:34:12,161
learn to generalize across 
geometry. 

592
00:34:12,362 --> 00:34:15,748
That's quite hard for me. from a
machine learning point of view. 

593
00:34:15,748 --> 00:34:17,679
it's really in the large data 
limit. 

594
00:34:17,679 --> 00:34:23,577
It is fully um Transformer-based
in a sense that we train with a 

595
00:34:23,577 --> 00:34:28,107
simple MSE loss and no 
 physics
information is added, just 

596
00:34:28,107 --> 00:34:30,896
really standard training. 
And it's in distribution. 

597
00:34:30,896 --> 00:34:34,109
So no out of distribution 
questions are asked. 

598
00:34:34,109 --> 00:34:38,237
This is the setup. 
And in this setup, one can 

599
00:34:38,237 --> 00:34:41,333
formalize things, otherwise it 
gets very, very tricky. 

600
00:34:41,333 --> 00:34:44,353
And we got a lot of questions in
how How do you think out of 

601
00:34:44,353 --> 00:34:47,154
distribution works and so on and
so forth. think this is beyond 

602
00:34:47,154 --> 00:34:50,570
the scope and this is also where
experiments are needed. 

603
00:34:51,777 --> 00:34:55,828
Yeah. 
And I think, again, the point of

604
00:34:55,828 --> 00:35:01,600
the paper was really, this is on
the, every meeting I 
 have with

605
00:35:01,600 --> 00:35:05,596
every commercial company or 
research company is just trying 

606
00:35:05,596 --> 00:35:12,181
to get a sense of if is, is a 
foundational model possible. 

607
00:35:12,181 --> 00:35:14,631
And that's where you start to 
get into the numbers game. 

608
00:35:14,631 --> 00:35:18,766
Um, and that's where things did 
seem quite ill defined. 

609
00:35:18,766 --> 00:35:21,480
You know, people may read this 
and say, oh yeah, that's 

610
00:35:21,480 --> 00:35:23,910
obvious. 
Or I knew that, but I don't 

611
00:35:23,910 --> 00:35:26,406
think it was even obvious to us 
before. 

612
00:35:26,626 --> 00:35:30,520
And we kept sort of changing 
things around and we weren't 

613
00:35:30,520 --> 00:35:35,109
fully aware of the, yeah, until 
 you put pen to paper and start 

614
00:35:35,109 --> 00:35:38,286
to do some of the maths, it's 
not obvious. 

615
00:35:38,286 --> 00:35:42,336
And I think, um, you know, I 
think it's fair to say that Sid,

616
00:35:42,336 --> 00:35:45,226
this was the bit where I 
 
really learned from you. 

617
00:35:45,226 --> 00:35:48,938
And I think is the strength of 
the papers that Johannes and I, 

618
00:35:48,938 --> 00:35:51,926
when we were talking on the 
side, we were 
 like, wanted to 

619
00:35:51,926 --> 00:35:55,163
it, to bring some of the maths, 
you know, to actually bring some

620
00:35:55,163 --> 00:35:59,929
of the rigor, because you can 
do, you know, some back of the 

621
00:35:59,929 --> 00:36:03,802
envelope calculations, but, you 
need 
 to try to formalize it 

622
00:36:03,802 --> 00:36:07,005
more. 
So how did you go, you know, if 

623
00:36:07,005 --> 00:36:11,229
we go into like the section on 
the actual scaling laws, so 
 

624
00:36:11,237 --> 00:36:14,397
section five, so what, was your 
sort of process? 

625
00:36:14,397 --> 00:36:17,538
How did you come up? 
with these scales. 

626
00:36:17,538 --> 00:36:20,367
Yeah. 
Okay, this is a very, very good 

627
00:36:20,367 --> 00:36:22,474
question. 
Actually, this is a disclaimer, 

628
00:36:22,474 --> 00:36:26,553
we know this, but I think our 
listeners and viewers should 
 

629
00:36:26,561 --> 00:36:30,253
also know this, that till the 
night before the final 

630
00:36:30,253 --> 00:36:33,379
submission, we were still 
working out the final numbers 

631
00:36:33,379 --> 00:36:35,964
and so on. 
So it was a very iterative 

632
00:36:35,964 --> 00:36:37,515
process. 
Nothing was set in stone. 

633
00:36:37,515 --> 00:36:40,971
And I myself, I was surprised. 
And the whole point of research 

634
00:36:40,971 --> 00:36:45,332
is to be surprised, right? 
No, I have, as I said, I have 

635
00:36:45,332 --> 00:36:48,577
worked on a mathematical 
formulation of some of these 
 

636
00:36:48,585 --> 00:36:50,737
questions in different uh 
domains before. 

637
00:36:50,777 --> 00:36:56,252
So for me to understand, well, 
we should also say that these 

638
00:36:56,252 --> 00:37:00,391
are hypothesis, right? 
We don't have end to end proofs.

639
00:37:00,391 --> 00:37:04,263
have not yet built such a 
foundation model, but we believe

640
00:37:04,263 --> 00:37:07,783
that these are very reasonable 

hypothesis because they hold for

641
00:37:07,783 --> 00:37:11,280
small scale models, they hold 
for language models perhaps and 

642
00:37:11,280 --> 00:37:14,388
so on. 
So the basic formulation we had 

643
00:37:14,388 --> 00:37:18,328
very quickly, if you remember 
the first equations, how it 
 

644
00:37:18,336 --> 00:37:23,053
scales with model size and how 
it scales with data size, this 

645
00:37:23,053 --> 00:37:25,446
hypothesis we had pretty 
quickly. 

646
00:37:25,446 --> 00:37:30,044
I think what was fun was to sort
of try out different 

647
00:37:30,044 --> 00:37:33,874
implications of these ideas, of 
 these exponents and so on. 

648
00:37:33,874 --> 00:37:39,170
The fact that you get power laws
has been known for a while. m 

649
00:37:39,170 --> 00:37:41,978
You can show something with 
statistical learning theory, but

650
00:37:41,978 --> 00:37:45,722
in the LLMs you have the 
 
Kaplan et al., you know, the 

651
00:37:45,722 --> 00:37:48,218
famous scaling paper, have the 
Chinchilla scaling laws. 

652
00:37:48,218 --> 00:37:51,906
So there is quite a bit of work 
also in the domain that we work 

653
00:37:51,906 --> 00:37:54,728
on. 
It's much less, you know, um 

654
00:37:54,728 --> 00:37:57,684
this is surprising. 
Johannes and I are among the 

655
00:37:57,684 --> 00:37:59,908
very few people who even discuss
this question. 

656
00:37:59,908 --> 00:38:03,505
If you look at some of the 
foundational texts for 

657
00:38:03,505 --> 00:38:06,377
scientific machine learning, 
this 
 question of scaling is 

658
00:38:06,377 --> 00:38:10,171
not even there. 
If you look at some of original 

659
00:38:10,171 --> 00:38:14,347
papers, there's no scaling. 
You have a table where you have 

660
00:38:14,347 --> 00:38:18,681
at certain resolution or at a 
certain scale of data, these 
 

661
00:38:18,689 --> 00:38:22,618
are our errors and we compare 
with 10 different models. 

662
00:38:22,618 --> 00:38:25,643
Without any understanding that 
if you change the data or you 

663
00:38:25,643 --> 00:38:28,389
change the model size, the 
 
picture can be very, very 

664
00:38:28,389 --> 00:38:30,645
different. 
So I have been investigating 

665
00:38:30,645 --> 00:38:34,901
this both theoretically and 
empirically. to do it at this 

666
00:38:34,901 --> 00:38:38,740
scale with a proper foundation 
model to sort of have different 

667
00:38:38,740 --> 00:38:41,944
scenarios, 
 right? 
In the paper we have a low 

668
00:38:41,944 --> 00:38:45,162
fidelity, probably not the best 
way to express it. 

669
00:38:45,623 --> 00:38:49,766
But I think the key was to be 
able to formulate everything in 

670
00:38:49,766 --> 00:38:52,783
terms of a common quantity. 
So in the beginning we were 

671
00:38:52,783 --> 00:38:55,470
looking at the model error. 
If you remember, we had an 

672
00:38:55,470 --> 00:38:56,970
epsilon, which is the model 
error. 

673
00:38:56,970 --> 00:39:00,118
We formulated all the complexity
estimates in terms of the model 

674
00:39:00,118 --> 00:39:02,978
error so that the theory 
 can 
give us some predictions. 

675
00:39:02,978 --> 00:39:06,682
But at the last moment, I 
shifted it to the number of 

676
00:39:06,682 --> 00:39:08,222
samples so that everyone 
 
understands. 

677
00:39:08,382 --> 00:39:10,929
No one understands what's model 
error, but everyone understands 

678
00:39:10,929 --> 00:39:14,608
how much of data that you 
 have
so that you have a common 

679
00:39:14,608 --> 00:39:16,018
dictionary to compare different 
things. 

680
00:39:16,018 --> 00:39:18,250
And we had these three 
scenarios. 

681
00:39:18,250 --> 00:39:22,274
We had this, let's say, RANS, 
steady state RANS. 

682
00:39:22,274 --> 00:39:25,376
We have an LES, but we only look
at the time average. 

683
00:39:26,117 --> 00:39:30,569
And finally, we did this full 
transient uh LES, and we 

684
00:39:30,569 --> 00:39:33,002
compared the three. 
And we came up with. 

685
00:39:33,002 --> 00:39:36,647
I would say some surprising 
conclusions. uh Perhaps we are 

686
00:39:36,647 --> 00:39:39,918
going to talk more about it. 
But what really surprised me 

687
00:39:39,918 --> 00:39:43,794
later on, it was obvious, but in
the beginning, what is not 
 

688
00:39:43,802 --> 00:39:47,662
obvious is I maintain that data 
generation was going to be the 

689
00:39:47,662 --> 00:39:50,197
dominant cost. 
I think all three of us believe 

690
00:39:50,197 --> 00:39:52,364
that. 
And we wanted our theory to show

691
00:39:52,364 --> 00:39:54,258
that, and our numbers to show 
that. 

692
00:39:54,258 --> 00:39:58,658
But it's only by this iterative 
process of discussing that we 

693
00:39:58,658 --> 00:40:02,653
realized that, in the small 
 
data regime and small model 

694
00:40:02,653 --> 00:40:06,272
regime, You have this actually 
holds true, but there is this 

695
00:40:06,272 --> 00:40:09,492
crossover point, there is this 

critical limit, and that's why 

696
00:40:09,492 --> 00:40:11,422
the paper is really, really 
interesting. 

697
00:40:11,422 --> 00:40:13,586
Perhaps we should talk more 
about that. 

698
00:40:13,830 --> 00:40:17,650
Maybe let me add here one thing 
because Sid you mentioned that 

699
00:40:17,650 --> 00:40:21,148
people don't look at 
 scaling 
laws so they try everything at 

700
00:40:21,148 --> 00:40:23,056
low scale or with one 
resolution. 

701
00:40:23,157 --> 00:40:26,029
And then there's actually the 
other extreme where people say, 

702
00:40:26,029 --> 00:40:29,186
okay, we just burn millions 
 
and millions of dollars and we 

703
00:40:29,186 --> 00:40:32,568
get lots of different data, we 
bunch them together and what we 

704
00:40:32,568 --> 00:40:34,584
get out is a model which solves 
it all. 

705
00:40:34,804 --> 00:40:39,017
And that was my entry point into
this whole area. 

706
00:40:39,017 --> 00:40:43,058
That's why I pushed so hard to 
write everything as a composite 

707
00:40:43,058 --> 00:40:46,634
vector. because that makes it 
very clear that it just doesn't 

708
00:40:46,634 --> 00:40:52,329
work if you throw data together.
um And therefore, we had set up 

709
00:40:52,329 --> 00:40:58,225
this RANS versus LES type of 
things because we all agreed 
 

710
00:40:58,233 --> 00:41:03,585
that you probably cannot mix 
RANS and LES simulation so 

711
00:41:03,585 --> 00:41:08,912
easily and that the choice you 
make on this has a tremendous 

712
00:41:08,912 --> 00:41:12,951
impact in the data generation 
and in everything you do. and 

713
00:41:12,951 --> 00:41:16,774
all comes down to modeling error
you take into account. basically

714
00:41:16,774 --> 00:41:20,164
so before doing that, you 
already take a modeling error 

715
00:41:20,164 --> 00:41:23,554
into account, which has 
 the 
consequence of all your future 

716
00:41:23,554 --> 00:41:27,394
steps and all your scaling loss 
and everything you obtain. 

717
00:41:27,394 --> 00:41:30,266
And this is something which is 
not clear to the machine 

718
00:41:30,266 --> 00:41:32,615
learning community, which is 
very, 
 very different than how 

719
00:41:32,615 --> 00:41:35,928
machine learning people think. 
So that's why we have these 

720
00:41:35,928 --> 00:41:38,574
three different setups of this 
modeling paradigms. 

721
00:41:38,574 --> 00:41:41,050
Yeah. 
And I mean, we can come to that 

722
00:41:41,050 --> 00:41:43,502
a bit later as well. 
Some of the open questions and I

723
00:41:43,502 --> 00:41:45,518
think definitely the mixing of 
fidelities was something 
 we 

724
00:41:45,518 --> 00:41:48,422
kept going back and forth on, 
but maybe we should leave that a

725
00:41:48,422 --> 00:41:51,259
little bit towards the end. 
Cause I think that's definitely 

726
00:41:51,259 --> 00:41:55,494
the open question. 
One thing that I did want to 

727
00:41:55,494 --> 00:42:00,109
highlight is, that transient one
and maybe to explain a 
 little 

728
00:42:00,109 --> 00:42:05,137
bit more because, and again, 
it's true what Sid and you said 

729
00:42:05,137 --> 00:42:08,582
as well, Johannes, about our 
assumption going in. 

730
00:42:09,292 --> 00:42:13,962
I think I've even said it on 
this podcast in a few opinion 

731
00:42:13,962 --> 00:42:18,629
things where I said, I, it felt 
 to me that data generation was 

732
00:42:18,629 --> 00:42:22,219
the biggest bottleneck and the 
idea that you could use 

733
00:42:22,219 --> 00:42:26,191
transient data felt like 
intractable. 

734
00:42:26,932 --> 00:42:32,053
Um, but I changed my tune a 
little bit when I speak into the

735
00:42:32,053 --> 00:42:36,433
two of you, because I think 
 
one, we started to realize that 

736
00:42:36,433 --> 00:42:42,758
you run a RANS simulation. 
And you in some ways use all the

737
00:42:42,758 --> 00:42:46,781
information that the simulation 
gives you because you 
 can't 

738
00:42:46,781 --> 00:42:50,798
take an early mid simulation 
checkpoint because it doesn't 

739
00:42:50,798 --> 00:42:55,072
have any physical sense. 
The solution is converging to a 

740
00:42:55,072 --> 00:42:59,536
steady state with the transient.
On the other hand, we spend, you

741
00:42:59,536 --> 00:43:03,298
know, five times more, 10 times 
more, depending on the 
 solver 

742
00:43:03,298 --> 00:43:08,432
to generate this data. 
But we only use the time average

743
00:43:08,432 --> 00:43:11,112
checkpoint. 
All this data with, you know, 

744
00:43:11,112 --> 00:43:15,128
was sticking out. don't use. 
And one, some of the early 

745
00:43:15,128 --> 00:43:17,778
papers, I guess, uh, on like 
MeshGraphNets, you know, had 
 

746
00:43:17,786 --> 00:43:20,958
this idea that you would train 
on the time steps, you know, 

747
00:43:20,958 --> 00:43:25,019
cause it was like, know, a 
cylinder flow or something. 

748
00:43:25,019 --> 00:43:27,685
And that was always intractable.
Cause you said, I can't store 

749
00:43:27,685 --> 00:43:29,922
all that data. 
I can't use all that data. 

750
00:43:29,922 --> 00:43:34,467
And so I think it's to be fair 
today, all the state of the art 

751
00:43:34,467 --> 00:43:38,092
work in an industrial way 
 has 
always just been a steady state.

752
00:43:38,092 --> 00:43:41,672
or some time average. 
Um, and I thought, okay, you'd 

753
00:43:41,672 --> 00:43:44,318
need petabytes of storage, 
unbelievable amount. 

754
00:43:45,942 --> 00:43:50,891
And, but I think one of the 
findings, uh there's nuance is 

755
00:43:50,891 --> 00:43:55,835
the idea that if we do run a 
 
transient simulation, can we 

756
00:43:55,835 --> 00:44:01,191
take, let's say, you know, a 
50th or, you know, or some of 

757
00:44:01,191 --> 00:44:05,583
the intermediate data, which can
be extra samples. 

758
00:44:05,583 --> 00:44:09,519
And I think that was something 
you both mentioned, but Sid in 

759
00:44:09,519 --> 00:44:12,789
particular, I remember you 
 
mentioned that from some of your

760
00:44:12,789 --> 00:44:17,540
earlier work and that does seem 
to be key finding. 

761
00:44:18,903 --> 00:44:23,529
We're not sure how much you 
could use, but more than one. 

762
00:44:24,056 --> 00:44:26,838
Yeah, well, I can talk about it 
a little bit. 

763
00:44:26,838 --> 00:44:30,397
Again, I think a distributional 
perspective gives you some 

764
00:44:30,397 --> 00:44:33,347
insight here, right? 
Because what we're learning when

765
00:44:33,347 --> 00:44:36,876
we are learning a time-dependent
process is that we are 
 

766
00:44:36,884 --> 00:44:39,692
learning the evolution of the 
time-dependent operator, how 

767
00:44:39,692 --> 00:44:42,970
things vary in time. 
And in principle, we are 

768
00:44:42,970 --> 00:44:47,084
learning, given a snapshot of 
the current state, what would be

769
00:44:47,084 --> 00:44:49,698
a 
 future state given the lead 
time. 

770
00:44:49,698 --> 00:44:52,034
This can be done directly. 
This can be done 

771
00:44:52,034 --> 00:44:54,421
autoregressively. 
That's a detail, right? 

772
00:44:54,421 --> 00:44:58,384
But essentially what you do is 
that you sample the data 

773
00:44:58,384 --> 00:45:01,680
distribution in different ways. 
Of course, if you just have 

774
00:45:01,680 --> 00:45:05,046
very, very tiny time steps, if 
you learn every single time 
 

775
00:45:05,054 --> 00:45:07,794
step, this is useless because 
essentially you're learning the 

776
00:45:07,794 --> 00:45:10,596
identity then, right? 
But if you learn large time 

777
00:45:10,596 --> 00:45:14,050
steps, which is what surrogates 
allow you to do, you're really 


778
00:45:14,058 --> 00:45:15,934
sampling the distribution at 
different ways. 

779
00:45:15,934 --> 00:45:20,254
And I think that makes a lot of 
intuitive sense. my previous 

780
00:45:20,254 --> 00:45:23,728
work, we had demonstrated, think
others have also demonstrated, 

781
00:45:23,728 --> 00:45:27,588
you just 
 imagine that you're 
learning the steady state from 

782
00:45:27,588 --> 00:45:32,271
some inputs and just uh using 
this training where you learn 

783
00:45:32,271 --> 00:45:35,241
the transient information too, 
adds value. 

784
00:45:35,241 --> 00:45:38,059
You can reduce the number of 
samples, you can increase the 

785
00:45:38,059 --> 00:45:40,628
accuracy and so on. 
I think Johannes likes the 

786
00:45:40,628 --> 00:45:41,845
perspective of data 
augmentation. 

787
00:45:41,845 --> 00:45:44,608
This is one valid way to think 
about it. 

788
00:45:44,608 --> 00:45:48,837
I would just think that you are 
just getting more data from 

789
00:45:48,837 --> 00:45:51,301
sampling more data from the 
 
underlying distribution. 

790
00:45:51,301 --> 00:45:54,483
And because of that, then you 
have this behavior. 

791
00:45:54,483 --> 00:45:57,735
Of course, what happens is you 
cannot, means you lose 

792
00:45:57,735 --> 00:45:59,685
information if you sample too 
much. 

793
00:46:01,026 --> 00:46:04,710
But if you sample too little, 
down sample too much, if you 

794
00:46:04,710 --> 00:46:08,086
down sample too little, then 
 
it's too much in terms of 

795
00:46:08,086 --> 00:46:12,040
training and storage and so on. 
So the golden mean, which we 

796
00:46:12,040 --> 00:46:16,265
don't know, could be problem 
dependent, is to sample a few of

797
00:46:16,265 --> 00:46:20,137

 the things, 50th or whatever 
we said 50th, based on some 

798
00:46:20,137 --> 00:46:21,970
ballpark. 
And then you do have some gain. 

799
00:46:21,970 --> 00:46:25,641
And this is not 50 times better,
but it's perhaps square root of 

800
00:46:25,641 --> 00:46:27,333
50, which is seven times 
 
better. 

801
00:46:27,333 --> 00:46:31,132
This is what we put the number 
five, which is based on some 

802
00:46:31,132 --> 00:46:33,620
previous work. 
But I think it makes a lot of 

803
00:46:33,620 --> 00:46:37,798
sense. ah intuitively as well as
mathematically and I think it's 

804
00:46:37,798 --> 00:46:40,353
a very important, well one 
 
shouldn't throw away the 

805
00:46:40,353 --> 00:46:42,779
transient information. 
I'm pretty confident that it 

806
00:46:42,779 --> 00:46:49,620
will help you. uh For me, it was
very eye-opening because we 

807
00:46:49,620 --> 00:46:55,382
always had this discussion. um 
In LLM, you basically assume you

808
00:46:55,382 --> 00:46:59,454
see each token once. every token
is seen once. 

809
00:46:59,454 --> 00:47:02,499
And what does this mean in CFD? 
That means every data point 

810
00:47:02,499 --> 00:47:06,096
should be seen once if you go 
for large scale limits. 

811
00:47:06,096 --> 00:47:10,396
But if you have 20,000 
simulations, which is already a 

812
00:47:10,396 --> 00:47:14,264
pretty big data set, that's kind
of 
 not possible. 

813
00:47:14,264 --> 00:47:17,564
But then this traversing in 
time, what Sid mentioned that 

814
00:47:17,564 --> 00:47:21,192
you need this time dimension is 
 actually also what happens in 

815
00:47:21,192 --> 00:47:23,942
space. 
So just sub-sampling one uh data

816
00:47:23,942 --> 00:47:27,166
point is not giving you the 
whole information. 

817
00:47:27,607 --> 00:47:31,358
So therefore you can see the 
data points multiple times and 

818
00:47:31,358 --> 00:47:35,107
that's the same of traversing 
 
in space in time is the 

819
00:47:35,107 --> 00:47:38,507
traversing in space and that 
still corresponds to this LLM 

820
00:47:38,507 --> 00:47:42,161
paradigm. in the infinite data. 
this was quite enlightening for 

821
00:47:42,161 --> 00:47:46,506
me when I understood this by the
explanation of Sid over 
 the 

822
00:47:46,506 --> 00:47:50,128
temporal domain. 
I think the other one that maybe

823
00:47:50,128 --> 00:47:54,054
is a key thing is the, which 
comes later when we showed 
 the

824
00:47:54,054 --> 00:47:57,974
numbers is the down sampling. 
And I think this matters from 

825
00:47:57,974 --> 00:48:01,564
the training, but also the 
storage side that, you know, 

826
00:48:01,564 --> 00:48:05,154
when 
 I was generating the 
DrivAerML or Ahmed or whatever 

827
00:48:05,154 --> 00:48:09,770
data sets, you know, we, we 
provided the full data, the full

828
00:48:09,770 --> 00:48:12,816
volume, the full surface. 
And that was partially because 

829
00:48:12,816 --> 00:48:16,365
even at the time we were just 
unsure of like, it felt like 
 

830
00:48:16,373 --> 00:48:19,051
we were down sampling just 
because of memory limits. 

831
00:48:19,051 --> 00:48:21,434
And that was probably more 
because of the graph neural net 

832
00:48:21,434 --> 00:48:23,162
approaches that we were trying 

at the time. 

833
00:48:23,162 --> 00:48:27,777
Um, but it, it seems like the, 
and I think you have this, 

834
00:48:27,777 --> 00:48:31,674
Johannes, in a couple of your 
 
recent papers where there's like

835
00:48:31,674 --> 00:48:36,375
a saturation point of like how 
many tokens in this case do you 

836
00:48:36,375 --> 00:48:38,601
need? 
Where if you go beyond it or 

837
00:48:38,601 --> 00:48:41,001
lower, if you go beyond it, you 
don't gain anything. 

838
00:48:41,001 --> 00:48:42,741
And so I think, you know, what, 
what was that? 

839
00:48:42,741 --> 00:48:45,900
come up with a number now. 
What was it like on a hundred 

840
00:48:45,900 --> 00:48:49,352
million cell case with three 
million surface, you know, 

841
00:48:49,352 --> 00:48:52,799
you're only needing 
 hundreds 
of thousands of tokens rather 

842
00:48:52,799 --> 00:48:55,863
than the tokens equals the 
number of cells. 

843
00:48:55,863 --> 00:49:01,057
And that is a massive, um, 
difference, right? 

844
00:49:01,057 --> 00:49:07,885
It fundamentally alters the cost
of training versus if you had to

845
00:49:07,885 --> 00:49:11,278
take everyone. 
How much of that was, how 

846
00:49:11,278 --> 00:49:15,031
important do you think that is 
that, that down sampling 
 

847
00:49:15,039 --> 00:49:19,709
essentially? 
I see it again more from the 

848
00:49:19,709 --> 00:49:25,745
perspective of a person who does
not come from a CFD uh PhD. um 

849
00:49:25,745 --> 00:49:30,521
CFD needs this fine resolution 
because otherwise simulations 

850
00:49:30,521 --> 00:49:34,619
just don't converge. um We 
artificially have this fine mesh

851
00:49:34,619 --> 00:49:37,909
because we have to enforce 
physics onto the 
 computer and 

852
00:49:37,909 --> 00:49:42,186
we as humans cannot do that with
a mesh which a machine learning 

853
00:49:42,186 --> 00:49:45,310
model would understand. 
So we need this. fine resolution

854
00:49:45,310 --> 00:49:49,671
for the CFD to converge. 
I mean, this is a bit bluntly 

855
00:49:49,671 --> 00:49:53,413
put, but in a way. 
So the machine learning model 

856
00:49:53,413 --> 00:49:56,383
does not need this fine 
resolution to understand the 
 

857
00:49:56,391 --> 00:49:59,539
problem. in fact, the machine 
learning model is not solving 

858
00:49:59,539 --> 00:50:01,546
the CFD. 
It's not about convergence. 

859
00:50:01,546 --> 00:50:04,687
It's not about uh implicit time 
stepping or whatnot. 

860
00:50:04,687 --> 00:50:06,768
It's just about getting the 
information. 

861
00:50:06,768 --> 00:50:10,790
And information is not the same 
as a CFD convergence criteria. 

862
00:50:10,790 --> 00:50:16,392
um I have not understood this at
the first, but I think it boils 

863
00:50:16,392 --> 00:50:21,236
down to this that you can 
 have
the full information m in a down

864
00:50:21,236 --> 00:50:23,970
sampled case. 
And therefore I would assume 

865
00:50:23,970 --> 00:50:29,274
that if you do a RANS or an LES 
simulation, it would need 
 

866
00:50:29,282 --> 00:50:31,722
approximately the same number of
tokens. 

867
00:50:31,722 --> 00:50:34,578
However, you're modeling 
different physics and for the 

868
00:50:34,578 --> 00:50:38,145
CFD, you need much more, finer 

resolution for the LES. 

869
00:50:38,145 --> 00:50:41,464
So this is a very big difference
between uh simulation and 

870
00:50:41,464 --> 00:50:46,360
machine learning. 
Sid, is that how you say more or

871
00:50:46,360 --> 00:50:48,328
less. 
But from a purely mathematical 

872
00:50:48,328 --> 00:50:52,159
perspective, I just uh formalize
a little bit of what 
 Johannes 

873
00:50:52,159 --> 00:50:55,339
exactly said, right? 
So the reason why we have 

874
00:50:55,339 --> 00:50:59,601
extreme grids in space as well 
as very tiny time steps with 
 

875
00:50:59,609 --> 00:51:02,937
explicit methods is because of 
accuracy and stability 

876
00:51:02,937 --> 00:51:04,852
requirements. 
That's why we need extremely 

877
00:51:04,852 --> 00:51:07,918
fine grids so that we can 
resolve small vortices, we can 

878
00:51:07,918 --> 00:51:10,420
resolve and 
 in time you have 
the CFL conditions. 

879
00:51:10,420 --> 00:51:14,997
So if you have a very small 
spacing in space, mesh size in 

880
00:51:14,997 --> 00:51:17,813
space, you have consequently a 

small time step. 

881
00:51:17,813 --> 00:51:22,160
Now, once you have generated the
simulation and I urge everyone 

882
00:51:22,160 --> 00:51:26,110
to do this experiment, 
 they 
can themselves down sample and 

883
00:51:26,110 --> 00:51:30,060
then they can down sample, they 
can put it aside. 

884
00:51:30,060 --> 00:51:33,098
And then they can up sample 
using some simple interpolant, 

885
00:51:33,098 --> 00:51:35,249
right? 
It you don't even need a fancy 

886
00:51:35,249 --> 00:51:37,193
interpolant. 
And you'll see that the errors 

887
00:51:37,193 --> 00:51:40,273
that you make in this process 
are tiny compared to your 
 

888
00:51:40,281 --> 00:51:42,075
modeling error. 
Your modeling error is, of 

889
00:51:42,075 --> 00:51:45,299
course, if you just down sample 
at 10 points, when you have 10 


890
00:51:45,307 --> 00:51:47,035
million points, you make a huge 
error. 

891
00:51:47,035 --> 00:51:50,906
But as long as it is reasonable,
and that's what we sort of uh 

892
00:51:50,906 --> 00:51:54,218
argued in our exact 
 numbers, I
think we said that we down 

893
00:51:54,218 --> 00:51:57,738
sampled to like 8 million, and 
then we do something more. 

894
00:51:57,778 --> 00:51:59,609
And so we had some concrete 
numbers. 

895
00:51:59,609 --> 00:52:03,001
The interpolation errors are so 
tiny that no information is 

896
00:52:03,001 --> 00:52:05,518
lost. 
And this is the big difference 

897
00:52:05,518 --> 00:52:09,874
because in a CFD setting, you 
sort of simulate it at that 
 

898
00:52:09,882 --> 00:52:11,684
resolution, you get nothing, 
right? 

899
00:52:11,684 --> 00:52:15,080
But once you have simulated, you
can always down sample both in 

900
00:52:15,080 --> 00:52:18,677
space as well as in time. 
Sid, you should probably explain

901
00:52:18,677 --> 00:52:21,977
what you tell your PhD students 
what they should do with 
 a 

902
00:52:21,977 --> 00:52:24,449
data set because this is 
something I found very 

903
00:52:24,449 --> 00:52:26,813
interesting and which I learned 
from you. 

904
00:52:27,282 --> 00:52:32,378
My students, my students in my 
class, I always say when you get

905
00:52:32,378 --> 00:52:37,474
a sort of scientific data 
 set,
what you should really do is to 

906
00:52:37,474 --> 00:52:41,846
first see what is the essential 
spatial scale that you need and 

907
00:52:41,846 --> 00:52:44,677
what is the essential time scale
that you need, right? 

908
00:52:44,677 --> 00:52:49,507
And the best way you can do this
is that you take your data, you 

909
00:52:49,507 --> 00:52:53,752
keep on downsampling it 
 till 
you can arrive at 1 % error. 1 %

910
00:52:53,752 --> 00:52:55,296
is a number that I've plucked 
out here. 

911
00:52:55,296 --> 00:52:59,910
It could be 0.5 or 0.2 % error. 
And that is the information. 

912
00:52:59,910 --> 00:53:02,713
You don't need more than that. 
That's an extremis. 

913
00:53:02,713 --> 00:53:05,839
It's the same with time. 
You don't need every single time

914
00:53:05,839 --> 00:53:08,237
step. 
How much can you down sample? 

915
00:53:08,237 --> 00:53:11,552
And then you have an idea of the
spatial and temporal scales of 

916
00:53:11,552 --> 00:53:13,949
your problem. 
I also tell them to look at the 

917
00:53:13,949 --> 00:53:17,466
sort of variation of the data. 
So compute things like mean and 

918
00:53:17,466 --> 00:53:20,448
the variance. 
Say what is sort of the noise to

919
00:53:20,448 --> 00:53:23,298
signal ratio, what is standard 
deviation over mean so 
 that 

920
00:53:23,298 --> 00:53:27,508
They understand what the data 
has and it is very useful, but 

921
00:53:27,508 --> 00:53:30,308
uh as always, students don't 
 
always do it. 

922
00:53:31,631 --> 00:53:34,566
They want to take the data and 
they want to run a model and 

923
00:53:34,566 --> 00:53:43,072
yeah. m So if we go, you know, 
into maybe the numbers a little 

924
00:53:43,072 --> 00:53:48,969
bit, and if we, think what was, 
 it's true what Sid said and, 

925
00:53:48,969 --> 00:53:54,405
um, it's true, things were 
changing, but it was almost a uh

926
00:53:54,405 --> 00:53:58,332
good thing that we were doing 
that because we were trying, 

927
00:53:58,332 --> 00:54:02,259
there's always a temptation of, 
 um, like group think where you 

928
00:54:02,259 --> 00:54:05,029
all sort of don't want to 
challenge each other. 

929
00:54:05,029 --> 00:54:07,798
You've sort of got to a 
conclusion and nobody wants to 

930
00:54:07,798 --> 00:54:11,919
challenge it. 
Where I think we were all open 

931
00:54:11,919 --> 00:54:14,503
to challenging each other's 
perceptions. 

932
00:54:14,503 --> 00:54:18,364
And I know I, as you described, 
definitely came into this 

933
00:54:18,364 --> 00:54:22,217
thinking the dataset was the 
 
biggest and that seemed to be 

934
00:54:22,217 --> 00:54:26,950
true in the early numbers we 
had. but then, so if we go 

935
00:54:26,950 --> 00:54:29,533
through, so we have this low 
fidelity, high fidelity 

936
00:54:29,533 --> 00:54:32,403
transient, but 
 maybe the more 
interesting one is the table 

937
00:54:32,403 --> 00:54:36,168
for. which is this large, extra 
large, extra, extra large, and 

938
00:54:36,168 --> 00:54:39,862
then the graphs in figure three.
And, um, yeah, what we were 

939
00:54:39,862 --> 00:54:42,037
basically trying to show. 
So maybe for the data 

940
00:54:42,037 --> 00:54:45,109
generation, I can go over some 
of the numbers and maybe Sid, 

941
00:54:45,109 --> 00:54:47,923
you and Johannes can talk a 
little bit on the model 

942
00:54:47,923 --> 00:54:51,933
training. 
The, they all, they are by 

943
00:54:51,933 --> 00:54:55,214
definition. 
We did try to give some estimate

944
00:54:55,214 --> 00:54:59,218
based on the sort of error 
floor, I guess, but, but it is 


945
00:54:59,226 --> 00:55:01,676
still tricky to estimate the 
number of samples. 

946
00:55:01,676 --> 00:55:03,948
And that's why we did try a 
couple of approaches. 

947
00:55:03,948 --> 00:55:07,483
One was, I guess, the more 
statistical, more mathematical 

948
00:55:07,483 --> 00:55:10,893
way of doing it. 
And then the second way, which 

949
00:55:10,893 --> 00:55:14,884
is Appendix C was, I guess, the 
way that I think more 
 CFD 

950
00:55:14,884 --> 00:55:18,568
people have looked at it, which 
is just simply this sense of, 

951
00:55:18,568 --> 00:55:22,857
well, how can my model predict a
plane if it's only trained on 

952
00:55:22,857 --> 00:55:24,603
cars? 
You know, doesn't matter how 

953
00:55:24,603 --> 00:55:26,971
many cars you give it. 
It's never going to be able to 

954
00:55:26,971 --> 00:55:28,559
predict a plane, right? 
I mean, that's just common 

955
00:55:28,559 --> 00:55:30,846
sense. 
Um, so. 

956
00:55:31,778 --> 00:55:35,945
And one of the challenging 
theories of CFD and it's an 

957
00:55:35,945 --> 00:55:37,079
interesting intellectual 
exercise. 

958
00:55:37,100 --> 00:55:42,622
CFD is very broad, very broad. 
you know, the possible 

959
00:55:42,622 --> 00:55:46,822
geometries and flows and physics
that you can do with CFD is 

960
00:55:46,822 --> 00:55:49,538
quite 
 extreme. 
And so if you look at like 

961
00:55:49,538 --> 00:55:53,282
appendix C, I'm sure I missed 
out some, but it is amazing when

962
00:55:53,282 --> 00:55:56,738

 you start to go to thing and 
you think, okay, it's cars, 

963
00:55:56,738 --> 00:55:59,207
right. 
But then it's also lorries. and 

964
00:55:59,207 --> 00:56:04,368
trains and planes and space 
planes and rockets and engines, 

965
00:56:04,368 --> 00:56:06,432
data centers, buildings, 
 
combustion. 

966
00:56:06,432 --> 00:56:08,984
And then what's really 
interesting is some of that 

967
00:56:08,984 --> 00:56:12,230
chemical side. 
There's a huge industry all 

968
00:56:12,230 --> 00:56:15,975
doing these like multi-phase 
flows, multi-physics, cement 
 

969
00:56:15,983 --> 00:56:21,316
mixing, ice cream making. 
there's just so many and each of

970
00:56:21,316 --> 00:56:24,578
those. 
The input, you know, that we 

971
00:56:24,578 --> 00:56:27,448
described, which is essentially 
boundary conditions, initial 
 

972
00:56:27,456 --> 00:56:30,314
conditions. physics modeling, 
we, we purposely summarized it 

973
00:56:30,314 --> 00:56:33,594
in the sort of turbulence model,
because 
 turbulence model is 

974
00:56:33,594 --> 00:56:37,489
one of the biggest impacts. 
But once you get into these 

975
00:56:37,489 --> 00:56:40,414
chemical and process, the, the 
additional source terms you 
 

976
00:56:40,422 --> 00:56:43,989
need to model some of the 
physical processes, just get it 

977
00:56:43,989 --> 00:56:46,544
big. 
So, you know, with those 

978
00:56:46,544 --> 00:56:50,714
numbers, you can easily reach 
into the millions of samples. 

979
00:56:50,934 --> 00:56:56,691
Um, now what I think is not 
entirely clear to me is just how

980
00:56:56,691 --> 00:57:00,765
much cross learning there is 
between, you know, there is 

981
00:57:00,765 --> 00:57:05,649
obviously a, you could argue a 
low 
 speed plane shares quite a

982
00:57:05,649 --> 00:57:11,750
lot with a car in some ways, but
an ice cream is quite different.

983
00:57:11,750 --> 00:57:14,948
So there's obviously some, you 
know, learning, but anyway, 

984
00:57:14,948 --> 00:57:18,961
that's how we got into the 
 
millions. we said 200,000, a 

985
00:57:18,961 --> 00:57:23,177
million and 2 million. we 
purposely gave the Python code 

986
00:57:23,177 --> 00:57:25,336
to this. 
So you could change these 

987
00:57:25,336 --> 00:57:28,024
numbers yourself. 
Um, because they will change if 

988
00:57:28,024 --> 00:57:31,786
you went to now 20 million, it 
would change the numbers. 

989
00:57:32,126 --> 00:57:37,386
Um, we went for the 500 million 
cells. 

990
00:57:37,386 --> 00:57:43,416
This is based on some work that 
I've been doing, um, for, for an

991
00:57:43,416 --> 00:57:48,146
upcoming data set, which 
 was 
on an aircraft and the 500 

992
00:57:48,146 --> 00:57:53,306
million is sort of at the, it's 
a, it's a number that 

993
00:57:53,306 --> 00:57:59,292
encompasses quite a lot of cases
in terms of, you know, the 

994
00:57:59,292 --> 00:58:02,852
Reynolds numbers, the sort of 
structures 
 you get, know, 500 

995
00:58:02,852 --> 00:58:06,768
million is quite a good number 
to represent anything in the, 

996
00:58:06,768 --> 00:58:12,354
let's say, LES of buildings or 
planes of cars, or it, you could

997
00:58:12,354 --> 00:58:16,394
obviously change the number, but
it 
 is representative of that 

998
00:58:16,394 --> 00:58:19,498
sort of case. 
Steps is tricky because it 

999
00:58:19,498 --> 00:58:23,656
depends on the mesh size, the 
total convective time you'd want

1000
00:58:23,656 --> 00:58:26,874

 to go for. 
But we have to pick some numbers

1001
00:58:26,874 --> 00:58:29,470
to base it in. 
So that was a 200,000. 

1002
00:58:29,470 --> 00:58:32,010
The flops per cell per step is 
quite aggressive. 

1003
00:58:32,010 --> 00:58:36,585
If you do the maths on your own 
code, if anyone's listening to 

1004
00:58:36,585 --> 00:58:39,393
this, you'll probably find 
 
that that represents the 

1005
00:58:39,393 --> 00:58:41,850
ultimate today of efficiency of 
a code. 

1006
00:58:41,850 --> 00:58:44,850
You're probably fine if you run,
I don't know, OpenFOAM, you'll 

1007
00:58:44,850 --> 00:58:47,570
definitely not be at that 
 
flops per cell per step. 

1008
00:58:47,570 --> 00:58:51,327
It'll be quite a bit higher. 
But if you go through the maths,

1009
00:58:51,327 --> 00:58:53,810
it's not unreasonable, but it is
quite aggressive. 

1010
00:58:53,934 --> 00:58:57,938
Um, so I say that because the 
number of hours you may need 

1011
00:58:57,938 --> 00:59:02,731
would be more if your code 
 
wasn't that well optimized. then

1012
00:59:02,731 --> 00:59:07,675
in terms of just to finish off 
some of the logic of this, the, 

1013
00:59:07,675 --> 00:59:12,753
we got to the hours. 
Now this is tricky and does get 

1014
00:59:12,753 --> 00:59:17,940
into a bit of the HPC side, but 
we essentially took the 
 flops 

1015
00:59:17,940 --> 00:59:23,127
and then worked out if you were 
using FP 32, which is reasonable

1016
00:59:23,127 --> 00:59:27,627
for like modern day CFD. 
There are some cases that would 

1017
00:59:27,627 --> 00:59:30,797
need FP64, but quite a lot use 
FP32. 

1018
00:59:30,817 --> 00:59:35,003
So we took the flops that was 
available on the system, but we 

1019
00:59:35,003 --> 00:59:38,223
realized that because many 
 of 
these codes are memory bandwidth

1020
00:59:38,223 --> 00:59:42,209
bound, the actual flops they can
use on the system is still 

1021
00:59:42,209 --> 00:59:47,211
relatively low. 
So let's say 15%. em And that's 

1022
00:59:47,211 --> 00:59:52,640
ultimately then once you take a 
cost, And I took these costs. 

1023
00:59:52,640 --> 00:59:56,674
you go across all different 
cloud vendors, all the cloud 

1024
00:59:56,674 --> 00:59:59,092
computing vendors give their 
 
costs publicly. 

1025
00:59:59,212 --> 01:00:04,368
So I think $8 was like probably 
the best price with a reserved 

1026
01:00:04,368 --> 01:00:06,372
instance. 
You know, if you sign up and you

1027
01:00:06,372 --> 01:00:07,574
buy them upfront for three 
years. 

1028
01:00:07,574 --> 01:00:10,393
So, but my logic was if you're 
spending a hundred million 

1029
01:00:10,393 --> 01:00:12,953
dollars, you're probably going 

to do that sort of deal. 

1030
01:00:12,953 --> 01:00:16,017
You're not just going to pay 
like an off the shelf price. 

1031
01:00:16,017 --> 01:00:18,655
You are going to commit. 
And then the storage was the 

1032
01:00:18,655 --> 01:00:22,820
same. was like a one of 
cheapest. cloud pricing for 

1033
01:00:22,820 --> 01:00:25,986
storage. 
And so that's how you got to 

1034
01:00:25,986 --> 01:00:31,514
these 10, 50, a hundred million.
And just to finish off on the 

1035
01:00:31,514 --> 01:00:37,650
data generation side, what we 
did do is we put into Figure 3 

1036
01:00:37,650 --> 01:00:44,246
in, think the sub figures around
D, E and F that if you did 

1037
01:00:44,246 --> 01:00:50,821
change the amount of flops or 
the size of the mesh or the cell

1038
01:00:50,821 --> 01:00:56,108
size, how that would change. 
And so those numbers could go to

1039
01:00:56,108 --> 01:01:01,308
200 million to 500 million. 
Um, but they sort of set a 

1040
01:01:01,308 --> 01:01:05,930
ballpark of where we're at. then
for the model training, mean, 

1041
01:01:05,930 --> 01:01:10,554
one of the things that we were 
working on till the last 
 

1042
01:01:10,562 --> 01:01:13,634
moment was like the, the FLOPS 
question, right? 

1043
01:01:13,634 --> 01:01:17,631
That was a tricky one. 
Yeah, you want me to take it. 

1044
01:01:17,631 --> 01:01:21,562
But before that, maybe this is 
some people have commented on 

1045
01:01:21,562 --> 01:01:25,846
this to me privately is that 
 
our assumption is that they, so 

1046
01:01:25,846 --> 01:01:29,754
some of the practitioners, they 
claim that our numbers are too 

1047
01:01:29,754 --> 01:01:33,799
aggressive because most of the 
codes at the moment, or many of 

1048
01:01:33,799 --> 01:01:37,169
the codes don't have 
 the sort 
of GPU capacity yet. 

1049
01:01:37,169 --> 01:01:40,837
Right. 
And if you're doing this on a 

1050
01:01:40,837 --> 01:01:44,218
CPU, then this numbers change 
dramatically or they will 

1051
01:01:44,218 --> 01:01:46,636
change, right? 
So perhaps, maybe Neil, you are 

1052
01:01:46,636 --> 01:01:49,485
the expert in this. 
What's your thinking about this?

1053
01:01:49,485 --> 01:01:53,790
We have made a big assumption, a
big bet here that everything can

1054
01:01:53,790 --> 01:01:58,093
be generated on a GPU at 
 FP32 
with a very aggressive flops per

1055
01:01:58,093 --> 01:02:01,072
cell, which we don't vary much 
in figure three. 

1056
01:02:01,072 --> 01:02:05,100
So what happens if you are stuck
with a CPU code? 

1057
01:02:05,100 --> 01:02:07,071
Well, yeah, interesting 
question. 

1058
01:02:07,071 --> 01:02:11,567
And, know, obviously I have to, 
as you do in papers, put like, 

1059
01:02:11,567 --> 01:02:15,193
uh, what's your conflicting 
 
thing. work for NVIDIA, a GPU 

1060
01:02:15,193 --> 01:02:17,013
company, right? 
Obviously. 

1061
01:02:17,013 --> 01:02:21,167
But, you know, people hopefully 
will know that I'm not in any, 

1062
01:02:21,167 --> 01:02:23,935
this is not a marketing or 
 
sales exercise. 

1063
01:02:23,956 --> 01:02:28,168
I think it's pretty defensible 
to say that if you look across 

1064
01:02:28,168 --> 01:02:31,327
all of the supercomputing 
 
centers, all of the people 

1065
01:02:31,327 --> 01:02:34,446
developing codes, any new code. 
If you're going to start a code 

1066
01:02:34,446 --> 01:02:36,569
from today, are you going to 
write it on the CPU? 

1067
01:02:36,569 --> 01:02:38,532
No. 
I mean, that's, that's just the 

1068
01:02:38,532 --> 01:02:41,234
case. 
Um, so, but you are absolutely 

1069
01:02:41,234 --> 01:02:44,942
right, which is sort of the 
comment on OpenFOAM. 

1070
01:02:44,942 --> 01:02:48,733
If you were to run with like a 
CPU code, I'm pretty sure those 

1071
01:02:48,733 --> 01:02:50,353
numbers would be 10 times 
 
higher. 

1072
01:02:50,353 --> 01:02:54,524
Um, which would basically make 
it. 

1073
01:02:55,044 --> 01:02:58,317
You couldn't do it. 
That's sort of my logic, which 

1074
01:02:58,317 --> 01:03:03,773
is a little bit why the, It is 
irrelevant for the, for the GPU 

1075
01:03:03,773 --> 01:03:08,359
topic here. and, yeah, but 
you're right, probably in the 

1076
01:03:08,359 --> 01:03:12,341
revised version, we probably 
should give a 
 bit of a caveat 

1077
01:03:12,341 --> 01:03:14,535
of it. 
I just take it as an obvious 

1078
01:03:14,535 --> 01:03:17,265
thing, but maybe I'm a bit like 
surrounded by this data for 
 me

1079
01:03:17,265 --> 01:03:20,422
as a no-brainer that you do it. 
Um, but we could definitely add,

1080
01:03:20,422 --> 01:03:22,673
but it would make the numbers 
quite scary. 

1081
01:03:23,571 --> 01:03:28,491
This is the point that I wanted 
to make. maybe I can walk and 

1082
01:03:28,491 --> 01:03:32,125
Johannes can just add. 
I can walk us through the 

1083
01:03:32,125 --> 01:03:35,479
choices we made regarding the 
model training and what are the 

1084
01:03:35,479 --> 01:03:38,674
 different things that come in. 
So before that, regarding the 

1085
01:03:38,674 --> 01:03:42,694
number of samples, so it does 
turn out that it was a 
 

1086
01:03:42,702 --> 01:03:45,703
fortuitous coincidence that this
2 million is roughly, 2.5 

1087
01:03:45,703 --> 01:03:51,210
million is what Neil arrived at 
by just sampling this entire 

1088
01:03:51,210 --> 01:03:56,354
data set of, you know, chemical 
reactors and 
 rockets and cars 

1089
01:03:56,354 --> 01:03:58,689
and everything in between, 
right? 

1090
01:03:58,689 --> 01:04:03,848
But independently, because we 
have some or we have some clue 

1091
01:04:03,848 --> 01:04:07,996
about what is this. 
So there are two key numbers of 

1092
01:04:07,996 --> 01:04:11,571
exponents here, beta in table 
four, if someone has to look 
 

1093
01:04:11,579 --> 01:04:15,029
at the paper, and alpha. 
So beta is a coefficient by 

1094
01:04:15,029 --> 01:04:17,166
which things scale with respect 
to data. 

1095
01:04:17,166 --> 01:04:21,100
And more or less central limit 
theorem that we learn law of 

1096
01:04:21,100 --> 01:04:24,697
large numbers that we learn at 

an undergraduate level tells us 

1097
01:04:24,697 --> 01:04:28,826
that this is no better than 0.5.
There's usually a logarithmic 

1098
01:04:28,826 --> 01:04:31,441
correction. 
And so far, we looked at the 

1099
01:04:31,441 --> 01:04:35,075
literature, including my own 
work. we put the number 0.43 

1100
01:04:35,075 --> 01:04:39,285
because this was sort of a 
representative of several test 

1101
01:04:39,285 --> 01:04:41,705
cases. 
So it turns out that by using 

1102
01:04:41,705 --> 01:04:45,880
some nominal errors, if you want
to arrive at below 1 % error in 

1103
01:04:45,880 --> 01:04:49,422
the field, not in the integral 
quantities, then you would need 

1104
01:04:49,422 --> 01:04:52,964
2.7 
 million samples, which is 
very close to the number that 

1105
01:04:52,964 --> 01:04:56,625
you also arrived at, Neil. 
So this was not totally 

1106
01:04:56,625 --> 01:04:59,910
unscientific that these two 
things sort of coincide. 

1107
01:04:59,910 --> 01:05:04,404
So remember that we have this, 
the scale at which, or the 

1108
01:05:04,404 --> 01:05:08,144
exponent at which, power law at 
 which, model size contributes, 

1109
01:05:08,144 --> 01:05:11,510
sorry, the data set size 
contributes to the error. 

1110
01:05:11,510 --> 01:05:15,010
So that's an important thing. 
The other thing is, of course, 

1111
01:05:15,010 --> 01:05:17,495
the scaling exponent with 
respect to model size. 

1112
01:05:17,495 --> 01:05:23,447
Now, is a fact that you need to 
grow the models a lot more to be

1113
01:05:23,447 --> 01:05:27,158
able to ingest the data 
 that 
you provide to them. 

1114
01:05:27,158 --> 01:05:31,190
So because there is a slower 
decay with respect to model size

1115
01:05:31,190 --> 01:05:33,540
than with respect to data 
 set 
size. 

1116
01:05:33,540 --> 01:05:36,776
So that's why you see some of 
the big numbers when it comes to

1117
01:05:36,776 --> 01:05:39,563
model size in the paper. 
So these were the two key 

1118
01:05:39,563 --> 01:05:42,880
numbers that we fit in. 
We also added something like the

1119
01:05:42,880 --> 01:05:46,453
number of copies that you want 
to see in transient 
 training. 

1120
01:05:46,453 --> 01:05:48,815
What is the compression ratio 
that you are going to see? 

1121
01:05:48,815 --> 01:05:51,156
We put some reasonable numbers 
there. 

1122
01:05:51,236 --> 01:05:53,717
And the important thing was the 
flops. 

1123
01:05:53,717 --> 01:05:57,280
This is a similar question. 
This is a training flops per 

1124
01:05:57,280 --> 01:06:00,071
step. 
And this was finally we put, I 

1125
01:06:00,071 --> 01:06:03,597
think, very, very aggressive 
numbers based on some use case 


1126
01:06:03,605 --> 01:06:06,765
that we had in the lab on some 
GH200s. 

1127
01:06:06,765 --> 01:06:10,793
We have assumed GB200s. 
And we assumed a very, very sort

1128
01:06:10,793 --> 01:06:13,691
of aggressive scaling. 
In general, if someone is 

1129
01:06:13,691 --> 01:06:18,080
interested in figure three, I 
think we have provided where we,

1130
01:06:18,080 --> 01:06:22,459

 no, we didn't provide this, 
probably in the next revision we

1131
01:06:22,459 --> 01:06:25,204
can provide that. 
But this is also a caveat. 

1132
01:06:25,204 --> 01:06:29,006
We were very aggressive. 
So these numbers could grow. 

1133
01:06:29,006 --> 01:06:32,328
And also the values of alpha and
beta are not set in stone. 

1134
01:06:32,328 --> 01:06:38,214
Again, if you look at figures 
three, A, B, We sort of vary 

1135
01:06:38,214 --> 01:06:41,578
alpha, keeping beta fixed, vary 
beta, keeping alpha fixed. 

1136
01:06:41,578 --> 01:06:44,938
And then you can see how these 
different numbers change. 

1137
01:06:44,938 --> 01:06:48,027
And the important thing is 
whatever you do asymptotically, 

1138
01:06:48,027 --> 01:06:52,143
at some point for large enough 

amounts of data, the model size 

1139
01:06:52,143 --> 01:06:56,392
has to be very big. 
So training the model will be 

1140
01:06:56,392 --> 01:07:01,034
very expensive, and it will 
overtake or dominate the cost of

1141
01:07:01,034 --> 01:07:03,394

 data generation. 
And this is sort of the... 

1142
01:07:03,394 --> 01:07:06,922
We have a theoretical 
demonstration of this and we 

1143
01:07:06,922 --> 01:07:10,272
also see this empirically. 
So this was sort of the logic. 

1144
01:07:10,272 --> 01:07:12,476
Maybe, Johannes, you wanted to 
add something? 

1145
01:07:12,664 --> 01:07:16,581
mean, this was perfect. 
It's just one thing to 

1146
01:07:16,581 --> 01:07:20,218
underline. um All these 
coefficients can change. 

1147
01:07:20,618 --> 01:07:22,419
Change is not even the right 
word. 

1148
01:07:22,419 --> 01:07:27,402
They need to be um determined 
scientifically, experimentally. 

1149
01:07:27,542 --> 01:07:30,874
But what we are quite certain is
the slopes of these two curves. 

1150
01:07:30,874 --> 01:07:35,781
And that is, we were quite um 
surprised that data generation 

1151
01:07:35,781 --> 01:07:39,349
is a different slope than the 
 
model training. 

1152
01:07:39,349 --> 01:07:43,302
And the model training at some 
point, the slope overtakes the 

1153
01:07:43,302 --> 01:07:47,987
data-generation slope. 
And this has big impact in many,

1154
01:07:47,987 --> 01:07:50,514
many ways. 
And this is one of the big 

1155
01:07:50,514 --> 01:07:54,336
findings, I would say. 
Yes, so just to paraphrase, data

1156
01:07:54,336 --> 01:07:57,654
generation scales linearly with 
data set size. 

1157
01:07:57,654 --> 01:07:59,725
Model training scales 
superlinearly. 

1158
01:07:59,725 --> 01:08:03,652
How superlinear depends on this 
alpha and beta, but it will 

1159
01:08:03,652 --> 01:08:06,501
always go superlinearly 
 
because the model is always 

1160
01:08:06,501 --> 01:08:10,061
going to be slower ah in scaling
than the data. 

1161
01:08:10,061 --> 01:08:12,732
And then eventually you'll have 
a crossover point. 

1162
01:08:13,472 --> 01:08:16,609
Yeah. 
That's kind of the real teaser, 

1163
01:08:16,609 --> 01:08:21,225
I guess, because you're right. 
Nobody or, you know, publicly 

1164
01:08:21,225 --> 01:08:25,685
anyway, is doing the sort of 
huge data set generation. 

1165
01:08:25,685 --> 01:08:32,278
So it's kind of hard for us to, 
to put it out, but as soon as 

1166
01:08:32,278 --> 01:08:36,398
that does happen, it'd be 
 very
interesting to compare because 

1167
01:08:36,398 --> 01:08:41,580
as you say, if you take all 
you're doing is delaying it. 

1168
01:08:41,591 --> 01:08:45,041
So let's say, imagine the data 
generations on a CPU and all 

1169
01:08:45,041 --> 01:08:46,763
those numbers are 10 times 
 
higher. 

1170
01:08:46,944 --> 01:08:50,122
It just means from our 
prediction that the crossover 

1171
01:08:50,122 --> 01:08:54,006
point will come later, but it 
will 
 come at some point. 

1172
01:08:54,006 --> 01:08:57,357
And so it's almost then the 
question is, well, from a 

1173
01:08:57,357 --> 01:09:01,005
usefulness point of view, how 
big 
 does the model need to be 

1174
01:09:01,005 --> 01:09:04,349
to give a low enough error and a
broad enough generalization? 

1175
01:09:04,849 --> 01:09:11,098
And depending on that, will 
determine where the biggest cost

1176
01:09:11,098 --> 01:09:16,709
on investment is, but also, um, 
maybe this is a good segue 
 

1177
01:09:16,716 --> 01:09:22,212
into some of the open questions.
One of the topics that I went 

1178
01:09:22,212 --> 01:09:25,872
into writing this paper thinking
would be a bigger one, but 
 I 

1179
01:09:25,872 --> 01:09:28,921
think in the end, the scaling 
laws more dominated the 

1180
01:09:28,921 --> 01:09:30,921
discussion. 
And I think we agree that we 

1181
01:09:30,921 --> 01:09:33,335
just didn't have enough 
information to write on it was 

1182
01:09:33,335 --> 01:09:37,018
some 
 of the online training. 
You know, that sense that the 

1183
01:09:37,018 --> 01:09:41,519
data-generation cost is so big. 
that you want and the cost of 

1184
01:09:41,519 --> 01:09:45,166
storing the data is so high that
you need some way of doing 
 it 

1185
01:09:45,166 --> 01:09:49,447
online. 
And I still sort of get that 

1186
01:09:49,447 --> 01:09:51,444
feeling, but the time scales 
difference. 

1187
01:09:51,564 --> 01:09:55,395
We assume that the model 
training is super fast and 

1188
01:09:55,395 --> 01:09:59,225
generating the data takes ages, 
but at 
 some crossover point, 

1189
01:09:59,225 --> 01:10:03,812
when the model training takes 
forever or very long time in the

1190
01:10:03,812 --> 01:10:07,525
day, you you sort of wonder 
where these are going to come 

1191
01:10:07,525 --> 01:10:15,112
in. and it is interesting that 
the, I should also caveat. 

1192
01:10:15,320 --> 01:10:18,532
The storage numbers assume very 
heavy compression. 

1193
01:10:18,532 --> 01:10:25,108
Um, so we should put a sort of 
caveat out there that the 

1194
01:10:25,108 --> 01:10:31,465
storage may still be quite a 
 
large cost. also the storage I'm

1195
01:10:31,465 --> 01:10:37,834
picking here is, you know, quite
a cheap object, based storage 
 

1196
01:10:37,842 --> 01:10:40,229
system. 
This is not a like a Lustre file

1197
01:10:40,229 --> 01:10:43,405
system that you're dumping on. 
That would be like seven times 

1198
01:10:43,405 --> 01:10:44,946
more expensive. 
So. 

1199
01:10:45,434 --> 01:10:49,758
Um, I don't know what, what did,
what did you go, what did you 

1200
01:10:49,758 --> 01:10:53,146
feel as the key opening, 
 uh, 
the open questions, you know, 

1201
01:10:53,146 --> 01:10:54,994
maybe comment on the online 
training. 

1202
01:10:54,994 --> 01:10:59,974
Um, how much do you see that as 
being important, overblown? 

1203
01:11:00,364 --> 01:11:05,015
I just add one thing to the 
storage. um So there is this 

1204
01:11:05,015 --> 01:11:09,239
very interesting phenomenon that
we're training surrogates and 
 

1205
01:11:09,247 --> 01:11:11,879
surrogates always make errors, 
right? 

1206
01:11:11,879 --> 01:11:17,483
So usually in compression, if 
you do um text or images or 

1207
01:11:17,483 --> 01:11:21,683
videos, you basically want to go

 for lossless compression. 

1208
01:11:21,683 --> 01:11:25,588
But if your surrogate is, your 
model is making an error 

1209
01:11:25,588 --> 01:11:28,664
anyways, that means your 
 
compression can also have an 

1210
01:11:28,664 --> 01:11:32,252
error, this error just needs to 
be smaller than the error the 

1211
01:11:32,252 --> 01:11:35,233
surrogate is making, 
 which 
opens a huge field of research. 

1212
01:11:35,233 --> 01:11:38,737
I shouldn't probably say that 
loud, but I think that there's a

1213
01:11:38,737 --> 01:11:41,889
lot of potential in that. 
And the second point, which is 

1214
01:11:41,889 --> 01:11:44,140
super intriguing to me is the 
data mixing. 

1215
01:11:44,281 --> 01:11:49,588
I don't think what we did in 
Aurora, that bunching data 

1216
01:11:49,588 --> 01:11:54,890
together and just uh doing the 
mix, 
 is the way forward. 

1217
01:11:55,662 --> 01:12:00,216
I think that having a clever 
formulation of how you can 

1218
01:12:00,216 --> 01:12:03,939
leverage different fidelities in

 that sense, or different that 

1219
01:12:03,939 --> 01:12:09,013
you stay as cheap as possible in
the data generation side, while 

1220
01:12:09,013 --> 01:12:13,754
being as performant as possible 
with the highest fidelity on the

1221
01:12:13,754 --> 01:12:18,495
modeling side is 
 one of the 
key open research points to 

1222
01:12:18,495 --> 01:12:22,782
discuss, which is super, super 
hard because that means you have

1223
01:12:22,782 --> 01:12:26,938
to operate on scale. bring these
communities together. 

1224
01:12:26,938 --> 01:12:29,670
You cannot do that in a machine 
learning lab without CFD 

1225
01:12:29,670 --> 01:12:31,278
expertise. 
You probably cannot do that. 

1226
01:12:31,898 --> 01:12:34,598
But that's where the fun begins,
I would say. 

1227
01:12:35,402 --> 01:12:38,552
Yeah, and I don't know whether 
I'm allowed to say this aloud, 

1228
01:12:38,552 --> 01:12:40,386
but we discussed this quite 
 a 
bit. 

1229
01:12:40,386 --> 01:12:43,917
We even had different versions 
of the write-up where we had 

1230
01:12:43,917 --> 01:12:48,083
different models on how we can 

sort of combine data and so on. 

1231
01:12:48,083 --> 01:12:51,002
We different strategies. 
Maybe that should be a separate 

1232
01:12:51,002 --> 01:12:54,330
paper altogether at some point. 
But I agree with you, Johannes, 

1233
01:12:54,330 --> 01:12:57,510
that this is a super interesting
question and also this 
 

1234
01:12:57,518 --> 01:13:00,054
question of cross-learning that 
you raised, Neil, right? 

1235
01:13:00,054 --> 01:13:04,737
So how much of information can 
be In my day job as an academic,

1236
01:13:04,737 --> 01:13:06,204
I'm very, very interested in 
that. 

1237
01:13:06,204 --> 01:13:09,061
How much of physics can be 
learned from other physical 

1238
01:13:09,061 --> 01:13:11,020
effects? 
Can you learn diffusion from 

1239
01:13:11,020 --> 01:13:13,230
fluid flows? 
With Poseidon, we showed that 

1240
01:13:13,230 --> 01:13:16,310
this can be done because somehow
it is hidden inside. 

1241
01:13:16,310 --> 01:13:20,485
So these are the sort of 
scientific questions which will 

1242
01:13:20,485 --> 01:13:25,489
be hopefully answered in the 
year 
 or two, But uh we cannot 

1243
01:13:25,489 --> 01:13:29,230
wait for things to... 
So the best way to generalize is

1244
01:13:29,230 --> 01:13:32,874
to... make your distribution so 
large that everything is more or

1245
01:13:32,874 --> 01:13:35,758
less an in-distribution problem.
We have seen that with language 

1246
01:13:35,758 --> 01:13:38,304
modeling, right? 
So I think putting these numbers

1247
01:13:38,304 --> 01:13:41,684
out, telling people that, hey, 
look, this can be done 
 

1248
01:13:41,692 --> 01:13:45,402
provided that there are these, 
and people can make their own 

1249
01:13:45,402 --> 01:13:47,527
choices. 
Maybe I don't scale my model at 

1250
01:13:47,527 --> 01:13:50,568
all. 
And I simply say that my model 

1251
01:13:50,568 --> 01:13:53,347
has 50 billion parameters come 
what may. 

1252
01:13:53,347 --> 01:13:56,723
And then at some level, as you 
generate more and more and more 

1253
01:13:56,723 --> 01:13:59,858
data, All that it does is that 
the model doesn't improve 

1254
01:13:59,858 --> 01:14:02,154
because it's dominated by your 
modeling 
 error, right? 

1255
01:14:02,154 --> 01:14:04,485
So the model error, not the 
modeling error. 

1256
01:14:04,485 --> 01:14:08,361
So these are choices that people
can make and they will make 

1257
01:14:08,361 --> 01:14:11,588
pragmatic choices, but at 
 
least it gives a ballpark of 

1258
01:14:11,588 --> 01:14:16,070
what people should aim for. 
So I think that's what is very 

1259
01:14:16,070 --> 01:14:18,120
interesting outcome of this 
project. 

1260
01:14:18,540 --> 01:14:22,951
And I think, at least from my 
perspective, that that still 

1261
01:14:22,951 --> 01:14:24,955
feels like the unanswered 
 
question. 

1262
01:14:24,955 --> 01:14:30,331
We were hoping, or at least I 
was hoping that we would have a 

1263
01:14:30,331 --> 01:14:34,170
little bit more of a 
 
definitive answer on this topic,

1264
01:14:34,170 --> 01:14:40,289
which is, you know, if I train 
the model with the RANS with a 

1265
01:14:40,289 --> 01:14:44,503
certain level of error, and then
I also give it some LES with a 

1266
01:14:44,503 --> 01:14:48,108
similar error, how 
 can the ML 
model know that one was RANS? 

1267
01:14:48,108 --> 01:14:52,826
And one was LES and we touched 
this a little bit in the paper, 

1268
01:14:52,826 --> 01:14:56,868
but this, we have, I feel 
 set 
the question right with the 

1269
01:14:56,868 --> 01:15:00,228
equation nine, but we haven't 
still fully answered how you 

1270
01:15:00,228 --> 01:15:04,412
would use and mix all those 
different inputs. 

1271
01:15:04,412 --> 01:15:09,433
Um, that for me feels we, tried 
it and I think we felt it was 

1272
01:15:09,433 --> 01:15:12,773
just, we didn't have enough 
 
data to fully answer that 

1273
01:15:12,773 --> 01:15:15,111
question, but that feels like 
you're right. 

1274
01:15:15,111 --> 01:15:20,867
Cause the point being is that. 
Um, you may not have the luxury 

1275
01:15:20,867 --> 01:15:24,045
of generating all this data from
scratch. 

1276
01:15:24,045 --> 01:15:28,599
It may be pre-existing data 
where it has had a certain 

1277
01:15:28,599 --> 01:15:32,737
modeling error or a certain 
 
boundary conditions or a certain

1278
01:15:32,737 --> 01:15:36,030
fidelity. 
And you need a way of the model 

1279
01:15:36,030 --> 01:15:39,982
knowing that rather than saying,
everything will be an LES 
 and 

1280
01:15:39,982 --> 01:15:47,174
everything will be, that still 
feels like, uh Yeah, a key thing

1281
01:15:47,174 --> 01:15:53,120
which is not clear, right? 
My take is, so there's two 

1282
01:15:53,120 --> 01:15:58,229
opinions on that. take is, and I
think I agree here with Sid, 

1283
01:15:58,229 --> 01:16:01,991
that you had to get rid of 
categorical 
 distributions 

1284
01:16:01,991 --> 01:16:04,917
because categorical 
distributions is the arch enemy 

1285
01:16:04,917 --> 01:16:08,219
of generalization. 
On the other hand, people say, 

1286
01:16:08,219 --> 01:16:11,507
okay, it's basically a 
multi-task learning because the 

1287
01:16:11,507 --> 01:16:15,198
 model learns two different 
tasks and the weights share 

1288
01:16:15,198 --> 01:16:21,134
across these two tasks. um I 
cannot comment because it needs 

1289
01:16:21,134 --> 01:16:25,074
experiments, needs large scale 
experiments but I'm just 
 

1290
01:16:25,082 --> 01:16:29,010
convinced that categorical 
formalization is just very very 

1291
01:16:29,010 --> 01:16:34,662
hard because you never get then 
out of this distribution 

1292
01:16:34,662 --> 01:16:36,662
problem. 
Yeah. 

1293
01:16:36,662 --> 01:16:42,941
But the reason we did some of 
this scaling laws or example was

1294
01:16:42,941 --> 01:16:48,244
even if you could make the 
 
data generation half the cost, 

1295
01:16:48,244 --> 01:16:53,624
would still at some point have 
the training to be the biggest, 

1296
01:16:53,624 --> 01:16:55,734
right? 
So we're, there's no free lunch 

1297
01:16:55,734 --> 01:16:58,907
in that sense. 
You know, there's, if you do it 

1298
01:16:58,907 --> 01:17:02,460
all with RANS, yes, compared to 
the estimate we had, 
 maybe 

1299
01:17:02,460 --> 01:17:06,318
it's half the price of the data 
generation, but Um, if you still

1300
01:17:06,318 --> 01:17:08,698
believe that you need millions 
of samples, then you're just 

1301
01:17:08,698 --> 01:17:10,364
shifting it 
 more to the model 
training. 

1302
01:17:10,364 --> 01:17:14,862
Um, but then how do you fix the,
ultimate issue that the RANS 

1303
01:17:14,862 --> 01:17:19,014
error, you know, unless you 
 
come up, which is where the full

1304
01:17:19,014 --> 01:17:23,476
circle is interesting in my, and
we didn't talk about too much in

1305
01:17:23,476 --> 01:17:27,286
the paper. 
There is still a reason to come 

1306
01:17:27,286 --> 01:17:32,382
up with the ultimate RANS model.
If you could use AI to go with 

1307
01:17:32,382 --> 01:17:35,129
the ultimate RANS model. 
You would alternate the data 

1308
01:17:35,129 --> 01:17:39,001
generation costs quite a bit 
lower, but, the reality of that 

1309
01:17:39,001 --> 01:17:42,521
 happening, as you said, is a 
much harder problem, ironically.

1310
01:17:42,521 --> 01:17:46,083
Um, but there is, there is some 
motivation for it. 

1311
01:17:46,083 --> 01:17:49,708
Um, maybe one of the other final
topics that we should cover, 

1312
01:17:49,708 --> 01:17:53,030
maybe we should come back, 
 um,
and have another discussion on 

1313
01:17:53,030 --> 01:17:56,327
this when we've had more 
feedback from the community, um,

1314
01:17:56,327 --> 01:18:02,088
is the inductive bias one. 
I think we were quite deliberate

1315
01:18:02,088 --> 01:18:04,742
in saying. 
We'd already stretched ourselves

1316
01:18:04,742 --> 01:18:07,822
with making hypotheses and 
making assumptions that I 
 

1317
01:18:07,830 --> 01:18:12,819
think we didn't feel this was 
the right time to have too much 

1318
01:18:12,819 --> 01:18:15,525
of dedicated opinions of 
including physics. 

1319
01:18:15,525 --> 01:18:19,585
We wanted this to be a little 
bit more of a, like a data 

1320
01:18:19,585 --> 01:18:21,897
driven, like how much data do 
you 
 need? 

1321
01:18:21,897 --> 01:18:25,493
What's the scaling law? 
It's still an open question, 

1322
01:18:25,493 --> 01:18:28,325
right? 
Um, but it feels one way you 

1323
01:18:28,325 --> 01:18:32,652
would need quite a lot more 
paper space to, get into this. 

1324
01:18:32,652 --> 01:18:34,419
I don't know what you both 
think. 

1325
01:18:35,239 --> 01:18:38,878
I answer because I would burn if
I answer here. 

1326
01:18:39,787 --> 01:18:44,141
Okay, to put it as politely as 
possible, we don't know, right? 

1327
01:18:44,141 --> 01:18:48,894
This is a reality. 
I think the belief that one way 

1328
01:18:48,894 --> 01:18:52,568
to put physics is just to know 
what the governing 
 equations 

1329
01:18:52,568 --> 01:18:55,240
are and stick them into the loss
function. 

1330
01:18:55,240 --> 01:18:58,743
We know that there are some 
difficulties in training this. 

1331
01:18:58,743 --> 01:19:01,628
The training is ill-conditioned.
You have to precondition it 

1332
01:19:01,628 --> 01:19:04,288
somehow. 
There are some ways out there, 

1333
01:19:04,288 --> 01:19:09,072
but this is far from what we are
already able to do with the 

1334
01:19:09,072 --> 01:19:10,703
data-driven approach. 
Far from it, right? 

1335
01:19:10,703 --> 01:19:14,313
With the data-driven approach, 
we are able to predict the 

1336
01:19:14,313 --> 01:19:17,559
weather now, which is a 
 
physics-based problem, and it's 

1337
01:19:17,559 --> 01:19:19,719
completely done in a data-driven
approach. 

1338
01:19:19,719 --> 01:19:25,478
So I don't think the question of
how to add physics has been 

1339
01:19:25,478 --> 01:19:30,340
answered yet, even in an 
 
academic setting, let alone in a

1340
01:19:30,340 --> 01:19:33,926
large-scale setting, right? 
So you could argue that maybe we

1341
01:19:33,926 --> 01:19:35,854
can put in conservation laws, 
symmetries. 

1342
01:19:35,854 --> 01:19:39,225
I think there are some comments 
about that also on the post. 

1343
01:19:39,426 --> 01:19:42,952
Yes, that could be. 
We could use that as data 

1344
01:19:42,952 --> 01:19:46,489
augmentation and so on. 
But I think it would still be a 

1345
01:19:46,489 --> 01:19:49,229
low ball. 
It's not going to change things 

1346
01:19:49,229 --> 01:19:52,406
dramatically because maybe the 
physics has lot of hidden, 
 

1347
01:19:52,414 --> 01:19:55,230
explicit symmetries, but the 
boundary conditions break it, 

1348
01:19:55,230 --> 01:19:57,634
right? 
And then, where do you go? 

1349
01:19:57,634 --> 01:20:01,516
And so I... 
And sticking the sort of physics

1350
01:20:01,516 --> 01:20:07,249
equation based laws may so far 
has not been able to be 
 shown 

1351
01:20:07,249 --> 01:20:11,218
to scale at this limits, but 
maybe it's possible. 

1352
01:20:11,218 --> 01:20:15,887
I think the question of how to 
add physics into these models is

1353
01:20:15,887 --> 01:20:19,612
a very interesting one. 
In my opinion, it has not been 

1354
01:20:19,612 --> 01:20:25,492
answered yet. 
And I think one of the points of

1355
01:20:25,492 --> 01:20:30,075
this paper, which I was keen and
I think maybe sets it 
 apart 

1356
01:20:30,075 --> 01:20:32,891
from some other recent papers 
talking about foundational 

1357
01:20:32,891 --> 01:20:36,971
models is it is unashamedly a 
more industrial focused, like, 

1358
01:20:36,971 --> 01:20:43,458
you know, if you want to build a
model that could predict 
 a car

1359
01:20:43,458 --> 01:20:48,088
or a plane or a data center 
using the architectures 

1360
01:20:48,088 --> 01:20:50,362
available today. would you do 
that? 

1361
01:20:50,362 --> 01:20:54,896
And so that is why it doesn't 
focus on stuff that could be 

1362
01:20:54,896 --> 01:20:57,904
around. 
It is essentially saying there 

1363
01:20:57,904 --> 01:20:59,966
are architectures available 
today. 

1364
01:21:00,086 --> 01:21:04,569
How much compute and data would 
you need to throw at it to build

1365
01:21:04,569 --> 01:21:07,124
the model? 
That is essentially the like 

1366
01:21:07,124 --> 01:21:10,309
summary, isn't it? 
Um, and, and what is undeniable 

1367
01:21:10,309 --> 01:21:13,699
is there will be future 
architectures and future ways of

1368
01:21:13,699 --> 01:21:15,733

 doing things that could 
incorporate it. 

1369
01:21:15,733 --> 01:21:19,389
But I think what we're setting 
out is if you've got enough. you

1370
01:21:19,389 --> 01:21:23,726
know, compute budget and data 
budget. feels like this could be

1371
01:21:23,726 --> 01:21:27,564
done, right? 
I mean, that's the sort of 

1372
01:21:27,564 --> 01:21:33,996
takeaway I'm getting from this 
paper is that if you have a 
 

1373
01:21:34,004 --> 01:21:38,280
efficient enough data generation
code and a scalable 

1374
01:21:38,280 --> 01:21:42,756
architecture, the only 
limitation is compute. 

1375
01:21:43,996 --> 01:21:49,685
And, uh, I think because of 
that, I would make a prediction 

1376
01:21:49,685 --> 01:21:54,113
that we will see people try and 
build these models because those

1377
01:21:54,113 --> 01:21:58,172

 numbers are not so high that 
they are beyond, especially in 

1378
01:21:58,172 --> 01:22:03,060
the age of AI, you know, stuff. 
I hopefully that formalizes a 

1379
01:22:03,060 --> 01:22:07,548
little bit that at least from 
our numbers, this, and I say 
 

1380
01:22:07,556 --> 01:22:10,166
this, maybe this is the final 
point. 

1381
01:22:10,166 --> 01:22:13,672
It'd be good to get your final. 
There was an early version of 

1382
01:22:13,672 --> 01:22:17,082
the paper where I wrote the 
bottom of the abstract, 
 

1383
01:22:17,090 --> 01:22:20,484
something like. 
And we conclude that this is an 

1384
01:22:20,484 --> 01:22:23,588
intractable problem that you 
cannot actually build it. 

1385
01:22:23,728 --> 01:22:27,633
And, and I had to completely 
change it, you know, to 

1386
01:22:27,633 --> 01:22:31,889
basically the other way around 

because I believe now it is a 

1387
01:22:31,889 --> 01:22:34,142
tractable problem. 
We don't know how accurate it 

1388
01:22:34,142 --> 01:22:36,101
will be. 
Right. 

1389
01:22:36,101 --> 01:22:40,525
So I, so I think it can be done.
What's your final thoughts? 

1390
01:22:41,302 --> 01:22:44,484
Yeah, that's a beautiful ending 
and not much to add. 

1391
01:22:44,484 --> 01:22:50,367
Just saying um it can be done, 
but not for all CFD. 

1392
01:22:50,868 --> 01:22:56,664
So not for all the cases you 
listed in appendix C together, 

1393
01:22:56,664 --> 01:23:01,972
but for certain specified um 
 
set of cases either together or 

1394
01:23:01,972 --> 01:23:03,941
separately. 
It depends on what you're 

1395
01:23:03,941 --> 01:23:06,628
looking at. 
Industry doesn't need a model 

1396
01:23:06,628 --> 01:23:10,545
which works for all CFD. 
Industry needs a model which 

1397
01:23:10,545 --> 01:23:12,854
works for the type of problems 
they're looking at. 

1398
01:23:12,874 --> 01:23:17,580
This does not only apply for 
CFD, it applies for all sorts of

1399
01:23:17,580 --> 01:23:19,390
other problems, 
 
semiconductors, crash testing, 

1400
01:23:19,390 --> 01:23:24,826
blah, blah, blah. 
Yeah, being a sort of the more 

1401
01:23:24,826 --> 01:23:29,137
the scientist here, maybe I can 
say that, yeah, the 
 scientist 

1402
01:23:29,137 --> 01:23:32,265
dream is to have this large 
scale model. 

1403
01:23:32,265 --> 01:23:35,332
I think we are close. 
Maybe we, because our main 

1404
01:23:35,332 --> 01:23:39,039
assumption in terms of data 
generation, I think we have laid

1405
01:23:39,039 --> 01:23:42,007

 the case very well. 
In terms of model architecture, 

1406
01:23:42,007 --> 01:23:45,895
we have made the hypothesis very
clear that your model has 
 to 

1407
01:23:45,895 --> 01:23:49,782
scale in a certain manner with 
respect to data and with respect

1408
01:23:49,782 --> 01:23:52,435
to model size, right? 
There are architectures out 

1409
01:23:52,435 --> 01:23:56,388
there which can do that. 
And they can certainly do it in 

1410
01:23:56,388 --> 01:23:58,956
sort of specific domains, as 
Johannes rightly said. 

1411
01:23:58,956 --> 01:24:02,606
Whether they can do it 
cross-domain or not, this is 

1412
01:24:02,606 --> 01:24:05,173
essentially my research. 
I'm occupied with that question 

1413
01:24:05,173 --> 01:24:07,870
all the time. 
I think uh hopefully the answer 

1414
01:24:07,870 --> 01:24:11,186
is yes, but uh certainly in a 
sort of restricted setting 
 

1415
01:24:11,194 --> 01:24:14,497
that Johannes said, I think our 
numbers clearly show that this 

1416
01:24:14,497 --> 01:24:17,672
can certainly be done. 
But let's be optimistic and say 

1417
01:24:17,672 --> 01:24:21,656
that at least in some 
cross-domain settings, can still

1418
01:24:21,656 --> 01:24:24,571
do it. 
Soon, soon enough, rather than 

1419
01:24:24,571 --> 01:24:26,914
waiting for five years. 
It's been a pleasure. 

1420
01:24:26,914 --> 01:24:30,274
Um, we'll have to, do this 
again. 

1421
01:24:30,274 --> 01:24:32,819
So thanks guys. 
Thanks. 

1422
01:24:32,819 --> 01:24:36,046
It was an absolute pleasure not 
only writing this paper but also

1423
01:24:36,046 --> 01:24:38,293
having the podcast with 
 you. 
Yeah. 

1424
01:24:38,293 --> 01:24:40,042
Awesome. 
Hi, thanks.

