1
00:00:00,240 --> 00:00:02,360
Hi, and welcome to the Neil 
 
Ashton Podcast. 

2
00:00:03,000 --> 00:00:05,880
In each episode, we explained 
 
some of the fascinating ways 

3
00:00:05,880 --> 00:00:09,080
that science and engineering are

 changing the world around us. 

4
00:00:09,760 --> 00:00:12,720
We talked to leading engineers 

from elite level sports like 

5
00:00:12,800 --> 00:00:16,840
cycling and Formula One to some 
 of the world's top academics to

6
00:00:16,840 --> 00:00:20,960
understand how fluid dynamics, 

machine learning, supercomputing

7
00:00:21,360 --> 00:00:22,960
are bringing in a new era of 
 
discovery. 

8
00:00:23,920 --> 00:00:27,040
We also hear some of their life 
 stories, their career advice, 

9
00:00:27,640 --> 00:00:30,040
the lessons they've learned on 

the way that I hope will be 

10
00:00:30,040 --> 00:00:33,800
helpful to you too. 
 
So sit back and enjoy this 

11
00:00:33,800 --> 00:00:40,840
episode. 
 
Hi, and welcome back to the Neil

12
00:00:40,840 --> 00:00:44,040
Ashton Podcast. 
 
So today's guest is Professor 

13
00:00:44,040 --> 00:00:48,040
Ricardo Vinuesa. 
 
Ricardo is currently a Professor

14
00:00:48,040 --> 00:00:51,532
of Aerospace Engineering at the 
 University of Michigan, where 

15
00:00:51,532 --> 00:00:54,120
he leads research at the 
 
intersection of machine 

16
00:00:54,120 --> 00:00:57,440
learning, fluid mechanics, 
 
turbulence, and artificial 

17
00:00:57,440 --> 00:01:00,480
intelligence. 
 
Before joining Michigan, he was 

18
00:01:00,480 --> 00:01:03,720
a professor at KTH Royal 
 
Institute of Technology in 

19
00:01:03,720 --> 00:01:06,400
Stockholm, one of Europe's 
 
leading engineering 

20
00:01:06,400 --> 00:01:10,240
universities, and he built an 
 
internationally recognized 

21
00:01:10,240 --> 00:01:14,400
research program focused on 
 
turbulent flow control and 

22
00:01:14,400 --> 00:01:17,360
data-driven methods for fluid 
 
mechanics. 

23
00:01:18,600 --> 00:01:21,040
He's definitely established 
 
himself as one of the leading 

24
00:01:21,040 --> 00:01:24,640
researchers looking into this. 

And one of the topics that we 

25
00:01:24,640 --> 00:01:27,800
bring on today, of course, is 
 
around like explainable AI, 

26
00:01:27,800 --> 00:01:30,600
causality, reinforcement 
 
learning, reduced order 

27
00:01:30,600 --> 00:01:35,440
modelling, and of course the 
 
topic of the moment foundation 

28
00:01:35,440 --> 00:01:40,240
models for fluid mechanics. 
 
He, along with people like Steve

29
00:01:40,240 --> 00:01:43,313
Brunton, has done a lot to 
really educate the community, 

30
00:01:43,313 --> 00:01:47,320
also the 
 public in terms of 
these fluid mechanics and how 

31
00:01:47,320 --> 00:01:48,760
machine 
 learning can be used 
for. 

32
00:01:48,760 --> 00:01:51,892
It's really, he's got some great

 YouTube videos and 

33
00:01:51,892 --> 00:01:54,064
explanations. 
And what I particularly enjoy 
 

34
00:01:54,072 --> 00:01:58,030
about his work is he doesn't 
just ask whether AI can make 
 

35
00:01:58,038 --> 00:02:00,729
predictions faster. 
He's looking at the question, 
 

36
00:02:00,737 --> 00:02:05,420
can AI help us to understand the
mechanics of turbulence and to 


37
00:02:05,428 --> 00:02:07,618
ultimate accelerate scientific 
discovery itself. 
 

38
00:02:07,626 --> 00:02:10,842
So, you know, in this 
discussion, which as with any of

39
00:02:10,842 --> 00:02:13,484

 them, is never long enough to 
fully cover everything, you 
 

40
00:02:13,492 --> 00:02:17,312
know, we dived into some of the 
bigger questions facing CFD 
 

41
00:02:17,320 --> 00:02:19,700
today. 
So you know whether fluid 
 

42
00:02:19,708 --> 00:02:23,488
mechanics will have its ChatGPT 
moment, how close are we to 
 

43
00:02:23,496 --> 00:02:26,055
foundation models can be 
generalized across different 
 

44
00:02:26,063 --> 00:02:29,880
flow problems, and why he's 
actually more optimistic about 


45
00:02:29,888 --> 00:02:32,407
that possibility than than ever 
before. 
 

46
00:02:32,415 --> 00:02:36,909
As I mentioned before, the 
explainable AI causality, 
 

47
00:02:36,917 --> 00:02:39,772
understanding where they come 
from. 
 

48
00:02:39,780 --> 00:02:42,876
Can you explain how AI is 
getting to it? 
 

49
00:02:42,884 --> 00:02:47,080
And one of the things that we 
also look about is focusing on 


50
00:02:47,088 --> 00:02:50,388
the role of reinforcement 
learning in terms of 
 

51
00:02:50,396 --> 00:02:54,949
optimization and control and not
just using AI as a pure 
 

52
00:02:54,957 --> 00:02:56,923
predictor. 
And then towards the end, we, we

53
00:02:56,923 --> 00:03:00,000

 sort of zoom out and talk 
about the future of the field, 

54
00:03:00,000 --> 00:03:02,660
agentic 
 AI, autonomous 
scientific discovery, and also 

55
00:03:02,660 --> 00:03:06,880
how 
 universities should be 
adapting education given that AI

56
00:03:06,880 --> 00:03:09,280
is 
 changing the whole 
landscape. 

57
00:03:09,280 --> 00:03:12,840
So I, I think this is one of 
 
these conversations that, you 

58
00:03:12,840 --> 00:03:17,040
know, I learned a lot from it 
 
and, and I hope you do too. 

59
00:03:17,040 --> 00:03:20,920
So sit back and enjoy this 
 
episode with Professor Vinuesa. 

60
00:03:21,280 --> 00:03:23,080
Thanks very much for agreeing to

 do this. 

61
00:03:23,640 --> 00:03:25,902
You are definitely on the list. 
 

62
00:03:25,910 --> 00:03:29,379
Of people who everyone says, oh,
you should be speaking to him. 


63
00:03:29,387 --> 00:03:31,380
You know, I read the papers, I 
see stuff. 
 

64
00:03:31,388 --> 00:03:35,448
And you're the sort of person 
where when I read the paper, I 


65
00:03:35,456 --> 00:03:38,810
was like, yeah, I really should 
have thought of that. 
 

66
00:03:38,818 --> 00:03:43,228
That's a really good idea. 
So, yeah, thank you. 
 

67
00:03:43,236 --> 00:03:46,072
No thanks for having me. 
It's a real pleasure. 
 

68
00:03:46,080 --> 00:03:48,115
And yeah, looking forward to the
conversation. 
 

69
00:03:48,123 --> 00:03:54,976
So maybe we could get straight 
into it, which is the, I guess 


70
00:03:54,984 --> 00:04:00,840
the hot conversation now. 
Which has only got more intense.

71
00:04:00,840 --> 00:04:03,240

 
Is this foundation models? 

72
00:04:03,280 --> 00:04:05,680
You know, can we somehow make a 
 model that could? 

73
00:04:06,720 --> 00:04:10,400
You know, compute any fluid 
 
flow, whether it's, you know, a 

74
00:04:10,400 --> 00:04:12,520
geometry variation, A boundary 

condition. 

75
00:04:14,040 --> 00:04:20,079
Where do you think we are along 
 that journey and how close? 

76
00:04:20,480 --> 00:04:22,520
How close do you think we are? 

To to get in there. 

77
00:04:23,880 --> 00:04:28,440
Yeah, I think few years ago I 
 
would have probably given a 

78
00:04:28,440 --> 00:04:33,024
different answer, but things are

 changing quickly in a in a 

79
00:04:33,024 --> 00:04:36,320
good way. 
 
And I think we're getting closer

80
00:04:36,320 --> 00:04:39,480
to having and it all depends on 
 the type of question that you 

81
00:04:39,480 --> 00:04:41,480
want to answer, right. 
 
What type of predictions and 

82
00:04:41,480 --> 00:04:45,000
what level of accuracy do you 
 
want to use the systems for 

83
00:04:45,000 --> 00:04:48,160
design, Do you want to use the 

systems for, for scientific 

84
00:04:48,280 --> 00:04:52,040
insight? 
 
And of course what quantities do

85
00:04:52,040 --> 00:04:54,120
you want to get with what level 
 of accuracy? 

86
00:04:54,320 --> 00:04:57,720
So there's many possible 
 
questions and directions to go 

87
00:04:57,720 --> 00:05:00,080
into. 
 
But in general, I think that 

88
00:05:00,080 --> 00:05:03,640
we're getting pretty close to 
 
having systems that can perform 

89
00:05:03,640 --> 00:05:08,200
very well even for quite 
 
complicated quantities and quite

90
00:05:08,200 --> 00:05:11,640
some nuanced phenomena on on our

 system. 

91
00:05:13,480 --> 00:05:19,280
Especially because depending on 
 how you define your model, and 

92
00:05:19,280 --> 00:05:22,160
depending on the type of 
 
question that you may be 

93
00:05:22,160 --> 00:05:27,080
interested in answering, you 
 
don't really need to mimic all 

94
00:05:27,080 --> 00:05:29,259
the physics of the real system. 
 

95
00:05:29,267 --> 00:05:33,624
You might just focus on a subset
of questions or a subset of 
 

96
00:05:33,632 --> 00:05:37,388
mechanisms that could be 
represented encapsulated in a 
 

97
00:05:37,396 --> 00:05:40,640
smart way. 
And that's basically what your 


98
00:05:40,648 --> 00:05:44,295
model does, and not necessarily 
all the intricate interscale 
 

99
00:05:44,303 --> 00:05:48,388
mechanisms that give rise to 
those particular phenomena. 
 

100
00:05:48,396 --> 00:05:53,205
So yeah, I think we're getting 
closer to having systems that 
 

101
00:05:53,213 --> 00:05:56,640
truly can give us the right 
accuracy, the right performance 

102
00:05:56,640 --> 00:05:58,740
 in in quite challenging 
problems. 
 

103
00:05:58,748 --> 00:06:03,627
Yeah. 
So I'm intrigued to know what's 

104
00:06:03,627 --> 00:06:07,605
 changed your mind. 
We said a few years ago you 
 

105
00:06:07,613 --> 00:06:09,294
would have given a different 
answer. 
 

106
00:06:09,302 --> 00:06:13,452
Can you pinpoint anything that 
has, yeah, changed over the past

107
00:06:13,452 --> 00:06:18,184

 few years that's changed that?
There's a couple of things. 
 

108
00:06:18,192 --> 00:06:24,358
The first one is we realised 
that we could identify those key

109
00:06:24,358 --> 00:06:29,696

 mechanisms pretty well in 
quite complex systems using 

110
00:06:29,696 --> 00:06:34,635
causality, 
 explainability. 
So not just in idealised 
 

111
00:06:34,643 --> 00:06:39,320
versions of the whole flow or in
reduced-order representations, 


112
00:06:39,328 --> 00:06:44,624
but in really full fidelity high
order versions of the system, we

113
00:06:44,624 --> 00:06:47,120

 could really interrogate the 
data, identify what are the 
 

114
00:06:47,128 --> 00:06:51,122
mechanisms that are the most 
important and focus on those for

115
00:06:51,122 --> 00:06:56,068

 control, for modelling. 
So I think those almost 
 

116
00:06:56,076 --> 00:07:00,060
surprising capabilities of this 
explainability and causal 
 

117
00:07:00,068 --> 00:07:03,180
frameworks to really 
characterize the systems has 
 

118
00:07:03,188 --> 00:07:07,260
been one big pillar. 
And the other one, the 
 

119
00:07:07,268 --> 00:07:10,615
performance of some of the 
generative deep learning 
 

120
00:07:10,623 --> 00:07:14,892
methods, especially in the last 
few years, diffusion, flow 
 

121
00:07:14,900 --> 00:07:18,680
matching, conditional latent 
diffusion, these systems are 
 

122
00:07:18,688 --> 00:07:23,464
really capable of generalizing 
to an extent that I mean not so 

123
00:07:23,464 --> 00:07:26,907
 long ago I would say that it 
was not really possible. 
 

124
00:07:26,915 --> 00:07:31,010
There's still quite some more to
do and there needs to be some 
 

125
00:07:31,018 --> 00:07:34,700
smart way of formulating these 
problems in order to achieve 
 

126
00:07:34,708 --> 00:07:39,725
real generalization, but I would
say that that we are really in a

127
00:07:39,725 --> 00:07:41,900

 stage where things are pretty 
impressive. 
 

128
00:07:41,908 --> 00:07:46,890
Maybe you could explain, pardon 
the pun, what you mean by 
 

129
00:07:46,898 --> 00:07:50,578
explainable AI. 
Yeah, that's a, that's a good a 

130
00:07:50,578 --> 00:07:54,908
 good question. 
So what we mean with the systems

131
00:07:54,908 --> 00:07:59,386

 and what constitutes an 
explanation and even the, the, 


132
00:07:59,394 --> 00:08:03,715
you know, the, the grammar and 
the definition changes a little 

133
00:08:03,715 --> 00:08:08,010
 bit depending on the, on the 
field you have explainable AI, 


134
00:08:08,018 --> 00:08:10,107
explainable deep learning 
interpretability. 
 

135
00:08:10,115 --> 00:08:14,071
So there's different kind of 
flavours to it. 
 

136
00:08:14,079 --> 00:08:19,928
But the way that I use this term
explainable deep learning would 

137
00:08:19,928 --> 00:08:24,951
 be kind of assessing for a 
particular model what features 


138
00:08:24,959 --> 00:08:30,820
of your input matter the most to
the prediction of that model. 
 

139
00:08:30,828 --> 00:08:35,744
So really identifying kind of 
like in a feature attribution 
 

140
00:08:35,751 --> 00:08:40,360
sense, what, what, which ones of
those components of the input 
 

141
00:08:40,368 --> 00:08:44,620
really matter to your output. 
And why is that interest in the 

142
00:08:44,620 --> 00:08:47,321
 context of high dimensional 
chaotic engineering systems? 
 

143
00:08:47,329 --> 00:08:51,620
Because if you create a model, 
an input output model, where I'm

144
00:08:51,620 --> 00:08:53,860

 just going to give you an 
example. 
 

145
00:08:53,868 --> 00:08:59,580
If I have the wing of an 
aircraft and I really want to 
 

146
00:08:59,588 --> 00:09:03,785
see what flow motions are really
affecting the most the drag, 
 

147
00:09:03,793 --> 00:09:08,322
well, then I can create a model 
which given the flow, it just 
 

148
00:09:08,330 --> 00:09:12,715
predicts the drag on the wing. 
And that's a predictive model 
 

149
00:09:12,723 --> 00:09:14,710
that hopefully performs 
reasonably well. 
 

150
00:09:14,718 --> 00:09:19,815
But if I use this explainability
methods on that model, then I 
 

151
00:09:19,823 --> 00:09:24,055
can go point by point and really
identify which points are 
 

152
00:09:24,063 --> 00:09:27,624
contributing the most to the 
drag on that wing. 
 

153
00:09:27,632 --> 00:09:33,236
And that idea can really allow 
us to identify volumes of 
 

154
00:09:33,244 --> 00:09:36,560
importance of importance for 
that particular question, for 
 

155
00:09:36,568 --> 00:09:40,943
the drag in this case. 
And what we have found is 2 
 

156
00:09:40,951 --> 00:09:44,570
interesting things. 
One is by using these 
 

157
00:09:44,578 --> 00:09:49,640
explainability tools, we realize
that the classical approaches to

158
00:09:49,640 --> 00:09:54,900

 study turbulence and people 
have been looking at vortices, 


159
00:09:54,908 --> 00:09:58,561
streaks, Reynolds stresses, many
other quantities, right? 
 

160
00:09:58,569 --> 00:10:04,560
In fact, when people want to 
show off a bit and they do a 
 

161
00:10:04,568 --> 00:10:07,965
very nice CFD visualisation, 
very colourful visualisation, 
 

162
00:10:07,973 --> 00:10:10,925
they will show you lambda-2, 
right, or something similar. 
 

163
00:10:10,933 --> 00:10:13,510
They show you vortices because 
it's very spectacular and it's 


164
00:10:13,518 --> 00:10:16,555
very turbulent. 
And what we realise with our 
 

165
00:10:16,563 --> 00:10:20,045
methods is that the vortices 
don't matter so much actually, 


166
00:10:20,053 --> 00:10:22,625
at least for the friction and 
for the drag. 
 

167
00:10:22,633 --> 00:10:25,697
They matter for other things, 
but not for the stuff that you 


168
00:10:25,705 --> 00:10:29,916
are trying to show them for. 
And the reason is the fact that 

169
00:10:29,916 --> 00:10:33,024
 you can see something does not 
mean that it's important. 
 

170
00:10:33,032 --> 00:10:37,515
You simply can't see it. 
And a lot of the turbulence, a 


171
00:10:37,523 --> 00:10:41,769
lot of the turbulence research 
has been a bit misled by what 
 

172
00:10:41,777 --> 00:10:44,844
you could see in experiments, 
which of course it has been very

173
00:10:44,844 --> 00:10:46,909

 helpful and it has been very 
illustrative. 
 

174
00:10:46,917 --> 00:10:52,504
But sometimes you need to step 
away a bit from the physical 
 

175
00:10:52,512 --> 00:10:55,360
intuition and just simply 
interrogate the data in a more 


176
00:10:55,368 --> 00:10:58,420
diagnostic way and let the data 
tell you what matters and what 


177
00:10:58,428 --> 00:11:02,175
doesn't matter. 
So what we've found is the 
 

178
00:11:02,183 --> 00:11:04,983
classical approaches to 
turbulence research. 
 

179
00:11:04,991 --> 00:11:08,618
They were only telling a part of
the story. 
 

180
00:11:08,626 --> 00:11:13,854
So in different regions, close 
to the wall, the the Reynolds 
 

181
00:11:13,862 --> 00:11:16,660
stresses were important. 
Farther away, the streaks were 


182
00:11:16,668 --> 00:11:18,376
important. 
Farther away there were all the 

183
00:11:18,376 --> 00:11:19,755
 Reynolds stresses that were 
important. 
 

184
00:11:19,763 --> 00:11:24,148
So the classical views were not 
wrong because obviously, you 
 

185
00:11:24,156 --> 00:11:26,247
know, there's a lot of time and 
effort. 
 

186
00:11:26,255 --> 00:11:28,560
They brought it to it. 
They were just telling part of 


187
00:11:28,568 --> 00:11:31,665
the story, right? 
Like the three blind men and the

188
00:11:31,665 --> 00:11:34,040

 elephant. 
They were, you know, giving 
 

189
00:11:34,048 --> 00:11:37,560
different partial views of what 
the elephant look like. 
 

190
00:11:37,568 --> 00:11:40,520
But all the classical 
perspectives together is what 
 

191
00:11:40,528 --> 00:11:44,888
actually comes close to what we 
identify with our explainability

192
00:11:44,888 --> 00:11:48,310

 methods. 
And why are we so sure that 
 

193
00:11:48,318 --> 00:11:51,120
these methods are giving us 
something that is interesting 
 

194
00:11:51,128 --> 00:11:54,015
and physical? 
Because when we devise control 


195
00:11:54,023 --> 00:11:58,095
mechanisms, when we manipulate 
the flow to diminish the 
 

196
00:11:58,103 --> 00:12:01,692
presence of these mechanisms 
that we identify, that's when we

197
00:12:01,692 --> 00:12:03,000

 get the highest drag 
reduction. 

198
00:12:03,720 --> 00:12:09,880
So it's not just that purely we 
 can find nice colourful volumes

199
00:12:09,880 --> 00:12:13,840
of structures is that those 
 
structures in a causal way are 

200
00:12:13,880 --> 00:12:16,960
actually affecting the drag in 

this case the most. 

201
00:12:17,280 --> 00:12:23,120
And they are tackling those 
 
structures is really looking at 

202
00:12:23,120 --> 00:12:28,480
the root cause of the drag, so 

the disease and not the symptom.

203
00:12:29,000 --> 00:12:32,760
And that's why we think that 
 
this is really a good way to 

204
00:12:32,760 --> 00:12:38,120
look at these mechanisms and not

 only from kind of controlled 

205
00:12:38,120 --> 00:12:42,320
perspective or from a 
 
satisfactory knowledge point of 

206
00:12:42,320 --> 00:12:45,600
view, also from a modelling 
 
point of view, because those 

207
00:12:45,600 --> 00:12:49,680
mechanisms are the ones that 
 
contain the crucial information 

208
00:12:49,680 --> 00:12:52,240
if you want to be a good model 

for that phenomenon. 

209
00:12:53,800 --> 00:13:02,120
So how do you see the the 
 
challenge that I still be not 

210
00:13:02,120 --> 00:13:04,280
got my head around which is the 
 following. 

211
00:13:05,160 --> 00:13:09,280
If we assume that for a 
 
foundational model or you know, 

212
00:13:09,360 --> 00:13:14,120
surrogate model, we're assuming 
 that we have lots of data of 

213
00:13:14,200 --> 00:13:17,040
wings of cars, of buildings, 
 
etcetera. 

214
00:13:18,160 --> 00:13:23,440
At least today, you know, most 

of those have, you know, a 

215
00:13:23,440 --> 00:13:27,680
decent error to what is the real

 turbulence, what is the real 

216
00:13:27,680 --> 00:13:30,880
flow? 
 
And you're asking the model to 

217
00:13:30,880 --> 00:13:33,040
learn it. 
 
And sometimes the data comes 

218
00:13:33,040 --> 00:13:36,960
from different sources. 
 
You know, the wing may have been

219
00:13:36,960 --> 00:13:39,344
done with this turbulence model,

 The car may have done with 

220
00:13:39,344 --> 00:13:45,160
this. 
So how do you see the foundation

221
00:13:45,160 --> 00:13:49,188

 model? 
Is it ultimately yeah. 
 

222
00:13:49,196 --> 00:13:54,375
Do we have to train on DNS? 
But then we can't train on DNS 


223
00:13:54,383 --> 00:13:55,749
because it will just be too 
expensive. 
 

224
00:13:55,757 --> 00:14:00,518
Do you need some DNS? 
I still feel like the foundation

225
00:14:00,518 --> 00:14:05,802

 model is ultimately going to 
not be as good as people think. 

226
00:14:05,802 --> 00:14:07,680
 
Yeah, because no, I totally see 

227
00:14:07,680 --> 00:14:11,920
your your point. 
 
And I think here the key is 

228
00:14:11,920 --> 00:14:15,720
going back a bit to the one of 

the first questions, what do we 

229
00:14:15,720 --> 00:14:21,040
want this model for, right? 
 
And I mean, I don't think that 

230
00:14:21,040 --> 00:14:24,760
we are trying to build these 
 
models to replace DNS, right? 

231
00:14:24,760 --> 00:14:28,280
I mean we're not going to 
 
because that computationally 

232
00:14:28,280 --> 00:14:30,240
would not really make sense, 
 
right? 

233
00:14:30,560 --> 00:14:34,240
So I think that we're trying to 
 build these models to either 

234
00:14:34,240 --> 00:14:39,840
accelerate design and 
 
optimization and or achieve some

235
00:14:39,840 --> 00:14:42,520
sort of scientific insight in 
 
our system. 

236
00:14:43,000 --> 00:14:49,320
So how can we handle data from 

different fidelities and 

237
00:14:49,320 --> 00:14:53,120
different modalities? 
 
That's something that in what 

238
00:14:53,120 --> 00:14:56,120
we're building in my group, 
 
we're trying to embed the 

239
00:14:56,280 --> 00:14:59,409
uncertainty quantification with 
 active learning loops in a 

240
00:14:59,409 --> 00:15:03,080
sense that we know that not all 
all 
 our data is of the same 

241
00:15:03,080 --> 00:15:08,000
fidelity. 
 
When we are trying to generalize

242
00:15:08,000 --> 00:15:12,560
and we're trying to produce data

 for new cases, the system will

243
00:15:12,560 --> 00:15:14,720
tell us, look for this 
 
particular case that you're 

244
00:15:14,720 --> 00:15:19,760
trying to produce data in some 

latent representation, your 

245
00:15:19,760 --> 00:15:23,000
uncertainty is very high. 
 
So you should probably go back 

246
00:15:23,000 --> 00:15:26,640
and if you know, if you're 
 
trying to do a helicopters, 

247
00:15:26,640 --> 00:15:29,000
well, your helicopter data is 
 
actually pretty bad. 

248
00:15:29,000 --> 00:15:32,680
So you should be there, have 
 
higher fidelity or, you know, 

249
00:15:32,680 --> 00:15:35,440
run more experiments or really 

improve your model there. 

250
00:15:35,680 --> 00:15:38,520
And I think that's one strength 
 actually that we can 

251
00:15:38,680 --> 00:15:42,240
progressively keep improving our

 model as new data becomes 

252
00:15:42,240 --> 00:15:47,280
available. 
 
Now your data for helicopters is

253
00:15:47,280 --> 00:15:50,080
pretty bad for what, right, 
 
Because it depends on what 

254
00:15:50,080 --> 00:15:53,360
you're trying to do. 
 
And here is when we need to be a

255
00:15:53,360 --> 00:15:58,360
bit a bit realistic about the 
 
metrics and about the targets 

256
00:15:58,360 --> 00:16:02,440
that we want to hit. 
 
So if we are trying to get lift 

257
00:16:02,440 --> 00:16:06,760
and drag, right, that's one type

 of accuracy and one type of 

258
00:16:06,760 --> 00:16:09,200
model. 
 
If we're trying to get the 

259
00:16:09,200 --> 00:16:12,800
spectrum right and all the 
 
interscale mechanisms, right, 

260
00:16:12,800 --> 00:16:14,360
then that's a different type of 
 model, right? 

261
00:16:14,720 --> 00:16:17,640
So I think in that sense, 
 
depending on the type of 

262
00:16:17,920 --> 00:16:19,840
approach and the type of 
 
figurative that we want to 

263
00:16:19,840 --> 00:16:24,000
achieve, we may have to resort 

to different data sets and 

264
00:16:24,000 --> 00:16:26,328
different ways of assessing the 
 error and the uncertainty of 

265
00:16:26,328 --> 00:16:30,560
our systems. 
 
Yeah, no, I, I agree. 

266
00:16:30,560 --> 00:16:34,760
I think the there'll be probably

 one side of the community that

267
00:16:35,520 --> 00:16:40,920
happily uses, you know, Reynolds

 averaged or even panel methods

268
00:16:40,920 --> 00:16:44,440
or or or lower and are, you 
 
know, quite comfortable with 

269
00:16:44,440 --> 00:16:47,960
models that are not of DNS 
 
accuracy. 

270
00:16:48,400 --> 00:16:53,280
But yeah, yeah. 
 
On the other hand, I, I, I can 

271
00:16:53,280 --> 00:16:57,160
imagine just as companies spend 
 a fortune constantly trying to 

272
00:16:57,160 --> 00:17:00,160
get to higher and higher 
 
fidelity CFD that maybe the 

273
00:17:00,160 --> 00:17:02,440
models will, you know, go with 

it. 

274
00:17:03,480 --> 00:17:06,480
One question, because I think 
 
some of your background, you 

275
00:17:06,480 --> 00:17:11,760
know, was on, I guess, would it 
 be fair to say like reduced 

276
00:17:11,760 --> 00:17:18,440
order modelling and and some of 
 that side 2, maybe a more like 

277
00:17:19,240 --> 00:17:20,680
lay audience? 
 
I'm not saying everyone 

278
00:17:20,680 --> 00:17:21,920
listening to this is a lay 
 
audience. 

279
00:17:21,920 --> 00:17:23,839
That's quite the opposite. 
 
But there are some people who 

280
00:17:23,839 --> 00:17:29,194
are maybe not as grounded in all

 the latest common question I 

281
00:17:29,194 --> 00:17:32,520
get is, well, machine learning 
is 
 just reduced order 

282
00:17:32,520 --> 00:17:36,540
modelling. 
How would you chart the 
 

283
00:17:36,548 --> 00:17:42,116
evolution, you know, from like 
what people would call ROMs 
 

284
00:17:42,124 --> 00:17:46,152
before to machine learning? 
And what is it about modern 
 

285
00:17:46,160 --> 00:17:49,848
machine learning that separates 
it from reduced-order modelling?

286
00:17:49,848 --> 00:17:52,920

 
Unless you classify them as to 

287
00:17:52,920 --> 00:17:53,600
say. 
 
Yeah. 

288
00:17:53,840 --> 00:17:58,640
No, that's a good question. 
 
Depends a little bit on what 

289
00:17:58,640 --> 00:18:01,520
you're trying to do with your 
 
machine learning, right, Because

290
00:18:01,520 --> 00:18:04,400
also machine learning is a quite

 broad area. 

291
00:18:05,080 --> 00:18:11,120
So, I mean, and we have some 
 
work that we've done before on 

292
00:18:11,120 --> 00:18:14,880
machine learning for CFD. 
 
So how should you or what are 

293
00:18:14,880 --> 00:18:17,680
the areas where machine learning

 can help CFD? 

294
00:18:18,080 --> 00:18:21,000
And this is work that I did with

 my friend Steve Brunton some 

295
00:18:21,000 --> 00:18:23,880
years ago where we found three 

areas. 

296
00:18:23,880 --> 00:18:28,520
We found accelerate DNS, 
 
improve models, both RANS and 

297
00:18:28,520 --> 00:18:32,960
LES and improve reduced-order 
 
models or surrogate models. 

298
00:18:33,120 --> 00:18:36,090
And, and, and, and that's if we 
 are thinking of machine 

299
00:18:36,090 --> 00:18:39,591
learning for CFD, something that
we made 
 quite clear at the 

300
00:18:39,591 --> 00:18:43,936
time, this is 4 years ago was 
that machine 
 learning in 

301
00:18:43,936 --> 00:18:45,868
principle would not replace CFD.

 

302
00:18:45,876 --> 00:18:49,600
It's not about, and I think that
many people when they think 
 

303
00:18:49,608 --> 00:18:52,080
machine learning for CFD, 
they're thinking automatically 


304
00:18:52,088 --> 00:18:55,170
replacing CFD, right. 
I, I don't think that it's 
 

305
00:18:55,178 --> 00:18:57,850
really about replacing CFD, but 
rather complementing and 
 

306
00:18:57,858 --> 00:19:00,528
helping. 
So reduced-order modelling would

307
00:19:00,528 --> 00:19:05,134

 be just one area within all 
the spectrum of possibilities 

308
00:19:05,134 --> 00:19:09,584
for 
 machine learning can help 
just within CFD, but we can also

309
00:19:09,584 --> 00:19:12,721

 think about control and 
optimization where this 
 

310
00:19:12,729 --> 00:19:16,545
fantastic reinforcement learning
work, for example, on really 
 

311
00:19:16,553 --> 00:19:18,170
finding new controller 
strategies. 
 

312
00:19:18,178 --> 00:19:25,100
So it's very broad, but I would 
say that if we focus on on 
 

313
00:19:25,108 --> 00:19:31,064
surrogate models, traditional 
models, one key has been the 
 

314
00:19:31,072 --> 00:19:35,460
possibility of handling well, 
nonlinearities in your in your 


315
00:19:35,468 --> 00:19:37,374
model development, right in your
surrogate. 
 

316
00:19:37,382 --> 00:19:43,241
And of course POD, DMD, I mean, 
they're mostly linear, although 

317
00:19:43,241 --> 00:19:45,080
you can 
 have nonlinearities in
there. 

318
00:19:45,560 --> 00:19:50,960
But having neural network based 
 compression at scale, which is 

319
00:19:50,960 --> 00:19:54,138
what you can have now with 
 
modern autoencoder architectures

320
00:19:54,138 --> 00:19:58,880
that really allows 
 you to have
very impressive compression 

321
00:19:58,880 --> 00:20:02,400
rates, right. 
 
So quite some capability to to 

322
00:20:02,640 --> 00:20:05,520
distill the essential physics of

 your system. 

323
00:20:06,160 --> 00:20:10,640
And something that we found also

 in some of our work is in, I 

324
00:20:10,640 --> 00:20:12,680
mean, autoencoders are good at 

compressing, right? 

325
00:20:12,760 --> 00:20:16,070
Compressing anything, images of 
 cats and dogs and images of 

326
00:20:16,070 --> 00:20:20,160
turbulent flows if you wish. 
 
But turbulent flows, they're 

327
00:20:20,160 --> 00:20:23,480
images, their flow fields, they 
 contain much more information. 

328
00:20:23,480 --> 00:20:25,960
They're much richer than cats 
 
and dogs, right? 

329
00:20:26,080 --> 00:20:27,880
There's a spectrum, there's a 
 
range of scales. 

330
00:20:28,240 --> 00:20:32,840
So you should somehow, when you 
 compress, you should somehow 

331
00:20:32,840 --> 00:20:36,360
disentangle the latent 
 
representation, because if you 

332
00:20:36,360 --> 00:20:39,640
do that and you can use beta- 
 
VAEs, you can use hierarchical 

333
00:20:39,640 --> 00:20:41,720
priors. 
 
There's different ways to do 

334
00:20:42,080 --> 00:20:46,840
that disentanglement, but that 

really allows you to encapsulate

335
00:20:46,840 --> 00:20:49,160
in different latent 
 
representations and different 

336
00:20:49,160 --> 00:20:51,800
latent vectors, different 
 
physical phenomena of your 

337
00:20:51,800 --> 00:20:54,960
problem. 
 
And that's going to be more 

338
00:20:54,960 --> 00:20:57,600
effective from the 
 
interpretability point of view 

339
00:20:58,120 --> 00:21:01,640
of your surrogate system, but 
 
also from a predictive point of 

340
00:21:01,640 --> 00:21:03,320
view. 
 
If you want to make predictions 

341
00:21:03,320 --> 00:21:06,680
in in the latent space, that 
 
disentanglement is actually 

342
00:21:06,680 --> 00:21:08,560
going to be quite, quite 
 
helpful. 

343
00:21:08,960 --> 00:21:13,200
So, yeah, I mean, I guess it was

 a bit of a long answer to your

344
00:21:13,200 --> 00:21:15,800
question, but machine learning 

is not just reduced-order 

345
00:21:15,800 --> 00:21:17,200
modelling. 
 
Let's say that there's many 

346
00:21:17,240 --> 00:21:21,080
other things that one can do and

 within reduced-order modelling

347
00:21:21,080 --> 00:21:25,920
that, you know, capability of 
 
really having autoencoders at 

348
00:21:25,920 --> 00:21:29,240
scale and disentangling the 
 
latent representations has been 

349
00:21:29,240 --> 00:21:32,480
quite critical. 
 
And also Transformers, that's 

350
00:21:32,480 --> 00:21:37,400
another, well dimension let's 
 
say, or another topic, which 

351
00:21:37,400 --> 00:21:40,019
also has helped in building the 
 temporal dynamics of these 

352
00:21:40,019 --> 00:21:43,440
reduced-order systems. 
 
I mean, where do you see? 

353
00:21:43,840 --> 00:21:50,760
At least my impression has been 
 that even though the turbulent 

354
00:21:50,760 --> 00:21:57,040
modelling started off as being 

one of the the big focuses of 

355
00:21:57,040 --> 00:22:02,480
the fluids community, it does 
 
seem to have slightly petered 

356
00:22:02,480 --> 00:22:07,400
out a little bit and the focus 

has shifted a little bit more to

357
00:22:07,400 --> 00:22:09,640
the surrogate modelling, to the 
 sort of reduced-order 

358
00:22:09,640 --> 00:22:11,966
modelling. 
Is that something that you've 
 

359
00:22:11,974 --> 00:22:14,582
also noticed? 
And do you think that's just 
 

360
00:22:14,590 --> 00:22:19,456
because of a lack of progress? 
Do you think it's because the, 


361
00:22:19,464 --> 00:22:23,890
you know, the improvements? 
I'm just wondering if how you've

362
00:22:23,890 --> 00:22:28,427

 seen that? 
Yeah, no, that's another good 
 

363
00:22:28,435 --> 00:22:33,320
point. 
I think that so turbulence 
 

364
00:22:33,328 --> 00:22:37,200
modelling can be done in 
different ways. 
 

365
00:22:37,208 --> 00:22:43,998
And One Direction that has been 
adopted by many people has been 

366
00:22:43,998 --> 00:22:48,840
 to start where more classical 
approaches to turbulence model 


367
00:22:48,848 --> 00:22:52,936
modelling start. 
Basically to which in a way is 


368
00:22:52,944 --> 00:22:56,318
just making an empirical 
assumption of how the, you know,

369
00:22:56,318 --> 00:22:58,180

 the Reynolds stresses need to 
behave. 
 

370
00:22:58,188 --> 00:23:00,700
To some extent. 
There is some empirical 
 

371
00:23:00,708 --> 00:23:03,980
assumption, more or less 
sophisticated and then try to 
 

372
00:23:03,988 --> 00:23:07,428
use machine learning systems to 
continue from there. 
 

373
00:23:07,436 --> 00:23:11,616
So in, in other words, to fit 
coefficients on classical 
 

374
00:23:11,624 --> 00:23:16,016
models, right? 
And well, that's not maybe the 


375
00:23:16,024 --> 00:23:20,530
most revolutionary thing, right?
Because if you do that, and 
 

376
00:23:20,538 --> 00:23:22,845
that's what many people have 
done. 
 

377
00:23:22,853 --> 00:23:26,662
And I think that's why a bit of 
the ML community could have got 

378
00:23:26,662 --> 00:23:31,132
 a bit disappointed with ML for 
fluids, because they were using 

379
00:23:31,132 --> 00:23:35,090
 ML to adjust coefficients in 
two turbulence models that have 

380
00:23:35,090 --> 00:23:38,346
 been around for decades and 
which have based on assumptions 

381
00:23:38,346 --> 00:23:40,698
 that are a bit crude, perhaps, 
right. 
 

382
00:23:40,706 --> 00:23:44,508
But for a reason, they were 
crude because there were not 
 

383
00:23:44,516 --> 00:23:48,040
other approaches to, to 
modelling that were, you know, 


384
00:23:48,048 --> 00:23:49,760
computationally possible at that
time. 

385
00:23:51,240 --> 00:23:54,720
So I think that could be why 
 
there has been a bit of slow 

386
00:23:54,720 --> 00:23:57,280
down. 
 
But at the same time, I've seen 

387
00:23:57,280 --> 00:24:01,800
some promising approaches. 
 
I mean, I'm, I'm a bit of a fan 

388
00:24:01,800 --> 00:24:04,440
of reinforcement learning. 
 
I've, I've been developing many 

389
00:24:04,440 --> 00:24:07,760
methods within reinforcement 
 
learning for control, for 

390
00:24:07,760 --> 00:24:10,800
optimization and also for 
 
modelling. 

391
00:24:10,800 --> 00:24:13,800
We have some, some projects on 

reinforcement learning for 

392
00:24:13,800 --> 00:24:16,880
modelling and there's other 
 
groups who are doing great 

393
00:24:16,880 --> 00:24:24,040
things in this space. 
 
And I think that the idea is why

394
00:24:25,400 --> 00:24:30,600
start with a model that has been

 around for, you know, 20305060

395
00:24:30,600 --> 00:24:35,960
years and try to adjust those 
 
coefficients when we can maybe 

396
00:24:35,960 --> 00:24:38,360
take one step back. 
 
So instead of asking my machine 

397
00:24:38,360 --> 00:24:42,360
learning model to say for the 
 
Smagorinsky model, what should 

398
00:24:42,360 --> 00:24:48,400
be my coefficient, Can you find 
it 
 to, to, or even say to my 

399
00:24:48,400 --> 00:24:54,080
machine learning model or just 

find the subgrid-scale tensor to

400
00:24:54,080 --> 00:24:58,150
be able to fit the statistics of

 this this particular channel 

401
00:24:58,150 --> 00:25:01,800
or particular flow? 
 
Maybe we can be a bit more 

402
00:25:02,720 --> 00:25:07,320
abstract and say, look, I don't 
 know what the mean flow or the 

403
00:25:07,320 --> 00:25:10,000
fluctuations will look like in 

this particular case. 

404
00:25:10,080 --> 00:25:13,280
I, I don't know what happened. 

I've never run this wing or this

405
00:25:13,280 --> 00:25:18,000
aircraft, but what I know is 
 
that turbulence is characterized

406
00:25:18,000 --> 00:25:21,240
by certain energy transfer 
 
mechanisms across scales, and 

407
00:25:21,240 --> 00:25:23,750
that energy goes in this 
 
direction and there is an 

408
00:25:23,750 --> 00:25:26,280
inverse cascade. 
 
And there's certain phenomena 

409
00:25:26,680 --> 00:25:29,440
that need to be true from a 
 
spectral point of view. 

410
00:25:29,440 --> 00:25:33,767
And from a mechanistic point of 
 view, whatever your model is 

411
00:25:33,767 --> 00:25:38,197
and whatever the forcing term 
that 
 you get each step from 

412
00:25:38,197 --> 00:25:40,720
your reinforcement learning or 

whatever optimizer that you're 

413
00:25:40,720 --> 00:25:43,680
using, just respect those 
 
physical constraints. 

414
00:25:44,000 --> 00:25:46,560
Because my flow needs to be 
 
physically correct. 

415
00:25:47,120 --> 00:25:49,880
And physically correct does not 
 mean this is your mean flow. 

416
00:25:49,880 --> 00:25:52,600
You need to adjust to this mean 
 flow because that's a circular 

417
00:25:52,600 --> 00:25:53,960
argument, then you need the mean

 flow, right? 

418
00:25:54,480 --> 00:25:59,720
But rather not only a circular 

argument, it's some more, it's a

419
00:25:59,720 --> 00:26:04,040
much more constraining goal to 

fit the statistics right. 

420
00:26:04,360 --> 00:26:09,680
But something a bit more 
 
flexible, such as making sure 

421
00:26:09,680 --> 00:26:13,480
that the energy fluxes are 
 
correct step by step is 

422
00:26:13,480 --> 00:26:15,720
something that is less impose 
 
imposing. 

423
00:26:16,400 --> 00:26:21,480
And it's also something that is 
 a bit more manageable for a for

424
00:26:21,480 --> 00:26:25,040
a system for an optimization 
 
system to to be able to adapt 

425
00:26:25,040 --> 00:26:27,440
step by step such that those 
 
fluxes are OK. 

426
00:26:27,760 --> 00:26:31,680
So I think the second direction 
 of making statements that are a

427
00:26:31,680 --> 00:26:36,520
bit more general and still 
 
correct from a physical point of

428
00:26:36,520 --> 00:26:39,600
view, that could be more 
 
promising to develop more. 

429
00:26:41,880 --> 00:26:46,400
I mean, one of the things that 

I've started contemplating, and 

430
00:26:46,400 --> 00:26:50,746
I'm interested to see if you've 
 been the same is for for a 

431
00:26:50,746 --> 00:26:56,084
while I was only thinking about,
you 
 know, the machine 

432
00:26:56,084 --> 00:27:01,560
learning, either in the context 
of 
 learning what the turbulent

433
00:27:01,560 --> 00:27:04,651
viscosity should be or, or more 
 broadly, what a surrogate 

434
00:27:04,651 --> 00:27:08,780
model, you know, should be in 
terms of 
 a transformer based 

435
00:27:08,780 --> 00:27:11,278
or some neural operator or, or 
whatever. 
 

436
00:27:11,286 --> 00:27:17,540
And the criticism has been that 
it's, how can it match the, the 

437
00:27:17,540 --> 00:27:22,074
 brains, you know, the thought 
process that, that, that we've 


438
00:27:22,082 --> 00:27:25,334
had. 
And I saw at an event, I was at 

439
00:27:25,334 --> 00:27:30,200
 the British Library, there was 
an AI, somebody organized UKTC, 

440
00:27:30,200 --> 00:27:33,252
 the Turbulence Consortium in 
the UK. 
 

441
00:27:33,260 --> 00:27:37,880
And it was Luca Magri from 
Imperial was organizing with 
 

442
00:27:37,888 --> 00:27:41,140
some people. 
And there was a Professor, 
 

443
00:27:41,148 --> 00:27:44,911
Michael Leschziner, who, you 
know, is a sort of pioneer of 
 

444
00:27:44,919 --> 00:27:47,220
turbulence modelling developed 
with Brian Launder, you know, 

445
00:27:47,220 --> 00:27:48,280
all 
 these Reynolds-stress 
models. 

446
00:27:48,680 --> 00:27:52,840
And, and he rightly said to me 

that he was, you know, gently 

447
00:27:52,840 --> 00:27:58,560
wondering how could these AI 
 
models, you know, know, all 

448
00:27:58,560 --> 00:28:00,720
these things that, that, that 
 
they've done. 

449
00:28:01,440 --> 00:28:05,340
And then I started to think, is 
 this really, and I don't know 

450
00:28:05,340 --> 00:28:09,312
if the capabilities today where 
the 
 agentic thing comes in 

451
00:28:09,312 --> 00:28:16,920
because in some ways is it more 
 realistic to ask an LLM to 

452
00:28:16,920 --> 00:28:21,502
reason through the process of 
 
how you would develop a 

453
00:28:21,502 --> 00:28:27,540
turbulence model and have links 
in to 
 software tools to 

454
00:28:27,540 --> 00:28:31,526
explore it and go through that 
reasoning 
 process and 

455
00:28:31,526 --> 00:28:34,520
reinforcement learning. 
 
Then just say, hey, try and find

456
00:28:35,240 --> 00:28:38,120
me the the latent representation

 of this. 

457
00:28:38,200 --> 00:28:39,880
Do you know what I mean? 
 
Do you think almost the 

458
00:28:39,880 --> 00:28:45,760
scientific discovery that piece 
 is the better route rather than

459
00:28:45,760 --> 00:28:47,160
just expecting machine that is? 
 

460
00:28:47,168 --> 00:28:49,749
So I don't know if I've 
explained that, but. 
 

461
00:28:49,757 --> 00:28:53,480
That makes perfect sense. 
I think it's, it's a very fair 


462
00:28:53,488 --> 00:28:58,088
question and we are trying to 
well to really first of all 
 

463
00:28:58,096 --> 00:29:01,975
understand the potential of 
agentic systems and 2nd, we 
 

464
00:29:01,983 --> 00:29:07,330
are learning to ask the right 
questions to these systems. 
 

465
00:29:07,338 --> 00:29:12,478
I think before getting into 
that, I think that a preliminary

466
00:29:12,478 --> 00:29:19,100

 point is what is the right way
to represent the data that we as

467
00:29:19,100 --> 00:29:23,240

 engineers and scientists in 
you know, higher order chaotic 


468
00:29:23,248 --> 00:29:25,714
complex systems. 
What is the right, the right way

469
00:29:25,714 --> 00:29:29,140

 to represent the data and such
an exploration from these agents

470
00:29:29,140 --> 00:29:32,960

 or whatever or, or any 
scientist, what's the right 
 

471
00:29:32,968 --> 00:29:35,627
platform to represent the data, 
right? 
 

472
00:29:35,635 --> 00:29:40,840
Because there's quite some work 
on LLMs to do this and to try to

473
00:29:40,840 --> 00:29:43,496

 achieve discovery, whatever we
define by discovery. 
 

474
00:29:43,504 --> 00:29:49,984
But you know, if we think of the
of turbulent flows, of course, 


475
00:29:49,992 --> 00:29:53,760
this is these are systems that 
are having the broadband in 
 

476
00:29:53,768 --> 00:29:57,328
terms of the spectrum, they have
multiple scales, multiple energy

477
00:29:57,328 --> 00:29:59,485

 fluxes. 
Perhaps text is not the right 
 

478
00:29:59,493 --> 00:30:02,590
way to represent this very 
complex data, right? 
 

479
00:30:02,598 --> 00:30:09,040
And perhaps using LLMs directly 
and text as a platform to try to

480
00:30:09,040 --> 00:30:13,748

 explore complex questions in 
this space might not be the the 

481
00:30:13,748 --> 00:30:17,598
 most suitable approach. 
And what we have been thinking 


482
00:30:17,606 --> 00:30:21,980
and exploring is what if we can 
create latent representations 
 

483
00:30:21,988 --> 00:30:27,344
that are particularly attuned to
represent important properties 


484
00:30:27,352 --> 00:30:30,962
of these turbulent-flow systems.
And, and that's an approach that

485
00:30:30,962 --> 00:30:34,838

 we have been developing in our
group where we are trying to 
 

486
00:30:34,846 --> 00:30:39,928
build a foundation models, kind 
of latent representations of 
 

487
00:30:39,936 --> 00:30:46,360
very broad ranges of cases such 
that we can use those latent 
 

488
00:30:46,368 --> 00:30:50,500
spaces to explore the, the, you 
know, the, the complexity of 
 

489
00:30:50,508 --> 00:30:56,145
these, of these turbulent flows.
One guiding principle has been 


490
00:30:56,153 --> 00:30:59,180
information, information fluxes,
causality. 
 

491
00:30:59,188 --> 00:31:05,370
So what if we can create latent 
spaces where causal relations 
 

492
00:31:05,378 --> 00:31:09,004
are maximized or where the 
disentanglement of the variables

493
00:31:09,004 --> 00:31:12,452

 is such that I can very easily
explore physical mechanisms 
 

494
00:31:12,460 --> 00:31:16,680
through that latent space. 
And I believe and again, this is

495
00:31:16,680 --> 00:31:20,712

 stuff that we are very excited
about because we are we're 
 

496
00:31:20,720 --> 00:31:24,640
developing it with pretty 
promising results that using a 


497
00:31:24,648 --> 00:31:29,615
agentic systems in that latent 
representation might be a good 


498
00:31:29,623 --> 00:31:37,430
way to achieve discovery. 
So in, in other words, I have a 

499
00:31:37,430 --> 00:31:42,360
 data from wings aircraft flows 
in cities a compressors and 
 

500
00:31:42,368 --> 00:31:44,908
these are very different fluid 
mechanics problems. 
 

501
00:31:44,916 --> 00:31:49,826
But if we find smart way to 
compress all that data and 
 

502
00:31:49,834 --> 00:31:53,176
express it in the same latent 
representation, so now suddenly 

503
00:31:53,176 --> 00:31:57,355
 I don't have a wing and a city,
but everything is expressed in 


504
00:31:57,363 --> 00:31:59,549
the same language on the same 
table. 
 

505
00:31:59,557 --> 00:32:04,740
And then I can allow this 
agentic systems to explore this 

506
00:32:04,740 --> 00:32:08,460
 latent space very efficiently. 
Because of course this AI 
 

507
00:32:08,468 --> 00:32:12,416
systems can't really find one 
side compress everything so 
 

508
00:32:12,424 --> 00:32:15,412
much. 
I can find patterns across cases

509
00:32:15,412 --> 00:32:20,075

 that might not be obvious in 
in the physical space for us 

510
00:32:20,075 --> 00:32:22,800
humans 
 and neither for agents,
right. 

511
00:32:22,800 --> 00:32:26,000
Because finding a connection 
 
between a wing and a city, I 

512
00:32:26,000 --> 00:32:28,823
mean, yeah, maybe if I look like

 that, I see a vortex that 

513
00:32:28,823 --> 00:32:30,600
looks similar. 
 
But you know, it's, it's a bit 

514
00:32:30,600 --> 00:32:32,360
more anecdotal, anecdotal than 

anything. 

515
00:32:32,600 --> 00:32:36,920
But when I express everything in

 that space and then I can do 

516
00:32:36,920 --> 00:32:40,600
that automatic and autonomous 
 
exploration with these agents, 

517
00:32:41,200 --> 00:32:44,640
then I can try to get insight. 

Then I can try to find causal 

518
00:32:44,640 --> 00:32:46,880
relations that are key in that 

latent space. 

519
00:32:47,480 --> 00:32:52,480
And then our job with 
 
explainability tools is to 

520
00:32:52,720 --> 00:32:55,640
express that insight in the 
 
latent space, which is 

521
00:32:55,640 --> 00:32:59,840
completely non understandable 
 
for us back in the physical 

522
00:32:59,840 --> 00:33:02,880
space and say, look, this is 
 
what insight looks like in the 

523
00:33:02,880 --> 00:33:05,800
latent space. 
 
That means that this vortex from

524
00:33:05,800 --> 00:33:08,120
the shear layer of the wing is 

actually very similar to this 

525
00:33:08,120 --> 00:33:10,800
vortex from the shear layer in 

the separation of that building 

526
00:33:10,800 --> 00:33:13,160
in the city. 
 
So that's actually a common 

527
00:33:13,160 --> 00:33:14,920
mechanism. 
 
That's interesting. 

528
00:33:15,280 --> 00:33:19,800
And I can find that through 
 
compressing, expressing in the 

529
00:33:19,800 --> 00:33:24,080
common space, A autonomously 
 
exploring and then going back to

530
00:33:24,080 --> 00:33:27,800
the physical representation. 
 
Yeah, that's, I really like the 

531
00:33:28,600 --> 00:33:31,600
I've you've hit the really good 
 explanation there because this 

532
00:33:31,600 --> 00:33:37,080
is something I've been thinking 
 in my head is and I I'm not, I 

533
00:33:37,080 --> 00:33:39,080
need to look into it more to be 
 completely honest. 

534
00:33:39,080 --> 00:33:42,720
But I almost see 2 slightly 
 
different approaches 

535
00:33:42,720 --> 00:33:48,259
traditionally, which is 1, which

 is kind of what we wrote in 

536
00:33:48,259 --> 00:33:50,840
this paper. 
 
We put out recently this idea 

537
00:33:50,840 --> 00:33:53,880
that, oh, you, you know, we're 

going to run lots of simulations

538
00:33:53,880 --> 00:33:57,680
of planes and cars and cities 
 
and turbo machineries and 

539
00:33:57,680 --> 00:34:01,560
basically every single possible 
 geometry and boundary condition

540
00:34:01,560 --> 00:34:07,240
that we see in real life, which 
 feels in some ways the more 

541
00:34:07,920 --> 00:34:12,080
logical and achievable way. 
 
It's just like you just have to 

542
00:34:12,080 --> 00:34:15,159
have a lot of them. 
 
But then the other side is that 

543
00:34:16,040 --> 00:34:21,840
actually it doesn't matter 
 
whether it's a plane or a city 

544
00:34:21,840 --> 00:34:25,239
or whatever, It's just you're 
 
just trying to learn the flow 

545
00:34:25,239 --> 00:34:28,440
physics. 
 
And actually you may massively 

546
00:34:28,440 --> 00:34:32,547
have way too much information by

 trying to do all possible 

547
00:34:32,547 --> 00:34:35,719
ones. 
Which then I thought, couldn't 


548
00:34:35,728 --> 00:34:38,344
you achieve this? 
Do you really need wings? 
 

549
00:34:38,351 --> 00:34:41,760
Couldn't you just create lots of
random shapes with lots of 
 

550
00:34:41,768 --> 00:34:45,543
random boundary conditions? 
And as long as you take every 
 

551
00:34:45,552 --> 00:34:48,219
possible consequence of shear 
flow and pressure gradients, it 

552
00:34:48,219 --> 00:34:51,106
 doesn't actually matter if it 
looks like a plane because 
 

553
00:34:51,114 --> 00:34:53,406
you're just trying to learn the 
physics. 
 

554
00:34:53,414 --> 00:34:56,610
That's a very good question, I 
would say. 
 

555
00:34:56,618 --> 00:35:01,080
I mean, that can be helpful 
because what you want is to be 


556
00:35:01,088 --> 00:35:05,048
able to explore such a latent 
representation very broadly, you

557
00:35:05,048 --> 00:35:08,810

 know, and that you can really 
investigate many configurations 

558
00:35:08,810 --> 00:35:13,398
 and calculations and so on. 
At the end, well, we want to 
 

559
00:35:13,406 --> 00:35:16,720
revert back to something that we
can understand, right? 
 

560
00:35:16,728 --> 00:35:21,188
And then probably having some 
random shapes with some random 


561
00:35:21,196 --> 00:35:24,015
pressure gradient and curvature 
distributions would be helpful 


562
00:35:24,023 --> 00:35:27,330
for the exploration because that
would allow us to identify 
 

563
00:35:27,338 --> 00:35:29,580
different mechanisms that 
probably we haven't seen just 
 

564
00:35:29,588 --> 00:35:32,272
because they were not ideal 
dynamic, but maybe they were 
 

565
00:35:32,280 --> 00:35:35,500
helpful for other things. 
But probably in our data sets, 


566
00:35:35,508 --> 00:35:37,980
we want to have shapes that 
we're familiar with and 
 

567
00:35:37,988 --> 00:35:41,252
applications that we're familiar
with just for the kind of 
 

568
00:35:41,260 --> 00:35:45,052
interpretation point of view, so
such that we can connect it. 
 

569
00:35:45,060 --> 00:35:49,215
But from the latent space 
generation, yes, the most the 
 

570
00:35:49,223 --> 00:35:52,854
more diverse and the more crazy,
the better. 
 

571
00:35:52,862 --> 00:35:57,554
In fact, what what we are doing,
what we are building is systems 

572
00:35:57,554 --> 00:36:01,445
 that the agent not only does 
the exploration but also 

573
00:36:01,445 --> 00:36:03,520
generates 
 new geometries by 
itself. 

574
00:36:04,520 --> 00:36:09,120
So we have this problem that we 
 like very much and we're having

575
00:36:09,120 --> 00:36:11,920
a pre print coming out very soon

 on this. 

576
00:36:12,160 --> 00:36:16,333
So it's basically two cylinders,
to 
 the keep it simple for now,

577
00:36:16,333 --> 00:36:19,418
but it's 2 cylinders with 
different 
 radii and different 

578
00:36:19,418 --> 00:36:23,579
separations. 
And we just want the our system 

579
00:36:23,579 --> 00:36:29,605
 agentic AI with a foundation 
model system to to explore the 


580
00:36:29,613 --> 00:36:34,728
wake of this tandem cylinder 
arrangement and just look at the

581
00:36:34,728 --> 00:36:38,598

 recovery of the wake. 
And then what this does is, 
 

582
00:36:38,606 --> 00:36:42,286
well, in order to do that, and 
this is trained on few cases, 
 

583
00:36:42,294 --> 00:36:43,924
right? 
It's not trained on many cases, 

584
00:36:43,924 --> 00:36:45,970
 but I need to explore the whole
space, right? 
 

585
00:36:45,978 --> 00:36:49,468
So these agents, what they do 
is, oh, but I need to look at 
 

586
00:36:49,476 --> 00:36:52,120
this configuration and calculate
my wake characteristics and 
 

587
00:36:52,128 --> 00:36:54,720
internal quantities and you 
know, displacement thickness, 
 

588
00:36:54,728 --> 00:36:56,738
momentum thickness, okay, I know
this. 
 

589
00:36:56,746 --> 00:36:59,520
But then I need to look at a 
different case and a different 


590
00:36:59,528 --> 00:37:03,232
case and a different case. 
And then it starts to 
 

591
00:37:03,240 --> 00:37:06,199
sequentially understand what the
scaling of the wake recovery 
 

592
00:37:06,207 --> 00:37:09,168
looks like as a function of the 
parameters of your problem. 
 

593
00:37:09,176 --> 00:37:12,388
And you haven't not had to 
create hundreds of cases. 
 

594
00:37:12,396 --> 00:37:14,754
You only have to create a few 
cases. 
 

595
00:37:14,762 --> 00:37:18,864
But it's the foundation model 
combined with the agent that is 

596
00:37:18,864 --> 00:37:22,304
 doing that autonomously is 
doing the discovery for you by 


597
00:37:22,312 --> 00:37:25,184
identifying what regions are 
important and creating and 
 

598
00:37:25,192 --> 00:37:27,410
generating new data in those 
regions. 
 

599
00:37:27,418 --> 00:37:31,734
In those regions. 
I think that's where we are 
 

600
00:37:31,742 --> 00:37:37,226
going so that we can even find a
discovery of cases that we don't

601
00:37:37,226 --> 00:37:41,570

 even think of because, because
we have been looking at patterns

602
00:37:41,570 --> 00:37:43,855

 influenced by the past 
history, right? 
 

603
00:37:43,863 --> 00:37:47,235
And kind of like the need of 
applications, but autonomously 


604
00:37:47,243 --> 00:37:52,014
exploring the whole space as 
agents can do. 
 

605
00:37:52,022 --> 00:37:55,510
They can really multiply the the
the insights and new solutions 


606
00:37:55,518 --> 00:37:58,680
and new mechanisms that we can 
actually discover. 
 

607
00:37:58,688 --> 00:38:04,782
And and that's why I'd feel and 
I'm guilty of this myself, that 

608
00:38:04,782 --> 00:38:09,325
 we are still probably too, too 
much splitting the fluids task 


609
00:38:09,333 --> 00:38:13,632
or the data generation task from
the machine learning task. 
 

610
00:38:13,640 --> 00:38:17,552
And so, you know, I'll be 
transparent for the data sets 
 

611
00:38:17,560 --> 00:38:21,306
that I've typically generated, 
they have been kind of segmented

612
00:38:21,306 --> 00:38:24,184

 a little bit for practical 
reasons, for people reasons. 
 

613
00:38:24,192 --> 00:38:27,280
So you're like, right, Well, you
know, you have a bunch of 
 

614
00:38:27,288 --> 00:38:28,734
meetings, you decide what you're
going to do. 
 

615
00:38:28,742 --> 00:38:31,135
You then kick off the data 
generation exercise, you 
 

616
00:38:31,143 --> 00:38:34,484
generate lots of data, and then 
you train the model. 
 

617
00:38:34,492 --> 00:38:38,162
And where, as you say, what you 
really want to be doing is 
 

618
00:38:38,170 --> 00:38:41,980
constantly them being in a loop 
together and only generating the

619
00:38:41,980 --> 00:38:46,382

 data it needs to generate. 
But I don't know if you found 
 

620
00:38:46,390 --> 00:38:50,216
this in your work, but one of 
the problems is that even the 
 

621
00:38:50,224 --> 00:38:53,485
machine learning frameworks are 
typically not even written in 
 

622
00:38:53,493 --> 00:38:56,383
the same language that the data 
generation is. 
 

623
00:38:56,391 --> 00:39:00,972
And doing that on a big HPC 
machine and training and 
 

624
00:39:00,980 --> 00:39:04,615
translating information, it's, 
it's quite challenging and maybe

625
00:39:04,615 --> 00:39:08,136

 brings the topic that, you 
know, we briefly discussed 

626
00:39:08,136 --> 00:39:12,760
before we 
 started recording, 
which is this fluid community, 

627
00:39:12,760 --> 00:39:16,040
computer 
 science, ML 
community. 

628
00:39:16,840 --> 00:39:21,720
How have you seen those two 
 
communities come closer? 

629
00:39:22,480 --> 00:39:25,640
Still not close enough, You 
 
know, how how have you assessed 

630
00:39:25,640 --> 00:39:27,480
this? 
 
And, and, and to be fair, you 

631
00:39:27,480 --> 00:39:31,400
have done a great deal. 
 
And to be fair with people like 

632
00:39:31,520 --> 00:39:34,720
who you collaborate, Steve 
 
Brunton of of trying to, you 

633
00:39:34,720 --> 00:39:37,200
know, bring a little bit 
 
together. 

634
00:39:38,000 --> 00:39:40,520
But yeah, how how do you see 
 
those two communities at the 

635
00:39:40,520 --> 00:39:42,480
moment? 
 
Well, thank you. 

636
00:39:43,280 --> 00:39:48,520
Thank you so much. 
 
First, I think the problem that 

637
00:39:48,520 --> 00:39:53,080
has happened for a while and 
 
still present, although probably

638
00:39:53,080 --> 00:39:56,480
getting a bit better, is that 
 
they have been a bit 

639
00:39:57,200 --> 00:40:02,120
disconnected in really 
 
understanding the the depths of 

640
00:40:02,120 --> 00:40:04,080
the problems in fluid mechanics.

 

641
00:40:04,088 --> 00:40:07,840
And perhaps the fluid mechanics 
community, when trying to adopt 

642
00:40:07,840 --> 00:40:12,540
 ML methods, has been sometimes 
a bit naive in just. 
 

643
00:40:12,548 --> 00:40:15,624
Adopting very vanilla 
architectures and maybe giving 


644
00:40:15,632 --> 00:40:18,220
up a bit too quickly without 
fully understanding. 
 

645
00:40:18,228 --> 00:40:21,880
So, you know, if the mechanics 
community is guilty in part, and

646
00:40:21,880 --> 00:40:26,530

 I think what has happened also
from the CS community is that, 


647
00:40:26,538 --> 00:40:30,748
you know, I mean aerodynamics 
from I am at the aerospace 
 

648
00:40:30,756 --> 00:40:33,070
engineering department right 
here in Michigan. 
 

649
00:40:33,078 --> 00:40:36,894
We train engineers for many 
years to understand the 
 

650
00:40:36,902 --> 00:40:39,368
mechanics, aerodynamics, 
structures and many other 
 

651
00:40:39,376 --> 00:40:40,600
things. 
But of course fluid mechanics is

652
00:40:40,600 --> 00:40:45,096

 a big part of it. 
So you can't understand all the 

653
00:40:45,096 --> 00:40:48,925
 complexities of fluid 
mechanics, you know, in the 

654
00:40:48,925 --> 00:40:52,670
course of 2-3 
 months, right? 
Because there is a lot to look 


655
00:40:52,678 --> 00:40:56,880
at. 
These are very entangled 
 

656
00:40:56,888 --> 00:41:02,212
problems and sometimes the 
questions and the problems that 

657
00:41:02,212 --> 00:41:05,996
 we have in fluid mechanics are 
not what typically the CS 
 

658
00:41:06,004 --> 00:41:08,349
community has looked at 
necessarily. 
 

659
00:41:08,357 --> 00:41:13,740
Because you know, the type of 
data that has been used in the 


660
00:41:13,748 --> 00:41:16,840
CS community to develop methods 
is not necessarily of the same 


661
00:41:16,848 --> 00:41:18,246
characteristics as the fluid 
mechanics data. 
 

662
00:41:18,254 --> 00:41:22,240
And this is very rich data with 
a very broadband spectrum, 

663
00:41:22,240 --> 00:41:24,720
multi- 
 scale. 
This is really, really complex 


664
00:41:24,728 --> 00:41:28,940
physics, right? 
So maybe looking at the MSE is 


665
00:41:28,948 --> 00:41:32,682
not the right metric, right? 
Because this only tells you a 
 

666
00:41:32,690 --> 00:41:34,235
very, very partial view of the 
story. 
 

667
00:41:34,243 --> 00:41:37,612
And maybe even having a model 
that predicts the MSE, well, 
 

668
00:41:37,620 --> 00:41:39,960
that might be useless in some 
cases, right? 
 

669
00:41:39,968 --> 00:41:43,415
We maybe want other things. 
Maybe we just want to look at, 


670
00:41:43,423 --> 00:41:46,072
you know, systems that predict 
separation, right or that 
 

671
00:41:46,080 --> 00:41:50,360
predict certain aspects of the 
mixing or maybe, you know, 
 

672
00:41:50,368 --> 00:41:53,959
certain structures, certain 
mechanisms as I was mentioning 


673
00:41:53,967 --> 00:41:57,743
before for, you know, for a drag
producing mechanisms. 
 

674
00:41:57,751 --> 00:42:02,702
So it could be that for my 
purpose, if I want to reduce 
 

675
00:42:02,710 --> 00:42:05,380
drag, having a good MSE is 
actually counterproductive 
 

676
00:42:05,388 --> 00:42:09,400
because the MSE is going to 
drive me to certain events, 
 

677
00:42:09,408 --> 00:42:13,426
which may be very energetic and 
they may be contributing a lot 


678
00:42:13,434 --> 00:42:17,056
to the MSE, but not so relevant 
towards the drag generation, 
 

679
00:42:17,064 --> 00:42:21,770
which is the stuff that I want 
to, you know, focus on in my 
 

680
00:42:21,778 --> 00:42:24,942
application. 
Therefore, a model with a much 


681
00:42:24,950 --> 00:42:29,762
worse MSE but getting the error 
right in the right mechanisms is

682
00:42:29,762 --> 00:42:34,040

 way better, right? 
So I think the this type of this

683
00:42:34,040 --> 00:42:37,860

 type of interaction is what is
really going to help us build 
 

684
00:42:37,868 --> 00:42:40,232
more useful models, at least 
from the fluid mechanics 
 

685
00:42:40,240 --> 00:42:41,996
community. 
And of course, this is not just 

686
00:42:41,996 --> 00:42:44,168
 for fluid mechanics. 
This is applicable to any area 


687
00:42:44,176 --> 00:42:47,848
of, of science, right? 
But of course, I can speak, you 

688
00:42:47,848 --> 00:42:51,888
 know, with more depth on the 
fluid mechanics problems because

689
00:42:51,888 --> 00:42:56,040

 what you want, and this is 
another question, you know, 
 

690
00:42:56,048 --> 00:42:59,988
let's build a foundation model. 
Let's build a big ML system for 

691
00:42:59,988 --> 00:43:03,442
 fluids, OK. 
But for fluids for what, right. 

692
00:43:03,442 --> 00:43:06,280
 
I mean, within fluids and within

693
00:43:07,000 --> 00:43:10,880
aerodynamics and within 
 
combustion, I mean, there are so

694
00:43:10,880 --> 00:43:14,400
many questions that we can look 
 at and probably the model that 

695
00:43:14,400 --> 00:43:17,360
you build is going to be quite 

different for all those 

696
00:43:17,360 --> 00:43:20,440
questions, right? 
 
So I think that's really 

697
00:43:20,440 --> 00:43:24,800
understanding in more depth what

 the fluid mechanics problems 

698
00:43:24,800 --> 00:43:27,240
require. 
 
What are the questions that we 

699
00:43:27,240 --> 00:43:28,520
have? 
 
What are the questions that we 

700
00:43:28,520 --> 00:43:32,200
don't have? 
 
How to build models that can 

701
00:43:32,200 --> 00:43:36,560
really effectively tackle those 
 questions, that can be a much 

702
00:43:36,560 --> 00:43:38,760
more effective use of 
 
everybody's time, basically. 

703
00:43:40,160 --> 00:43:47,080
Where do you sit on the debate 

around a data-driven versus 

704
00:43:47,080 --> 00:43:51,360
physics driven models? 
 
You know, again, the usual 

705
00:43:51,480 --> 00:43:58,400
criticism from people who are 
 
maybe from a fluid background 

706
00:43:58,400 --> 00:44:02,680
is, you know, any of these 
 
data-driven approaches cannot 

707
00:44:02,680 --> 00:44:05,640
guarantee mass conservation, 
 
energy conservation. 

708
00:44:05,640 --> 00:44:09,160
They can't guarantee, you know, 
 we should be implicitly or 

709
00:44:09,160 --> 00:44:13,015
explicitly imposing, which I 
 
guess is where some of the 

710
00:44:13,015 --> 00:44:16,664
PINNs, you know, things started.

 

711
00:44:16,672 --> 00:44:24,620
But then the successes seem to 
have been not as strong in the 


712
00:44:24,628 --> 00:44:27,840
sort of physics informed, 
inspired, conditioned. 
 

713
00:44:27,848 --> 00:44:34,530
So yeah, where where do you sit 
on that side of the fence? 
 

714
00:44:34,538 --> 00:44:38,232
That's, well, that's, that's a 
big question, right? 
 

715
00:44:38,240 --> 00:44:43,495
Of course you need to, you need 
to use any physical information 

716
00:44:43,495 --> 00:44:46,624
 that you have about your system
when building your models, 
 

717
00:44:46,632 --> 00:44:50,535
right. 
But if you just use the 
 

718
00:44:50,543 --> 00:44:54,560
equations and nothing else, then
you're building a numerical 
 

719
00:44:54,568 --> 00:44:57,768
solver, right, with a 
data-driven numerical solver. 
 

720
00:44:57,776 --> 00:45:04,856
And if you just use data, then 
you are a problem, you know, be 

721
00:45:04,856 --> 00:45:07,820
 being very wasteful because 
there's a lot of information 
 

722
00:45:07,828 --> 00:45:10,784
that you have about your system 
that you need to relearn. 
 

723
00:45:10,792 --> 00:45:14,094
Like, you know that your system 
has certain symmetries, certain 

724
00:45:14,094 --> 00:45:17,710
 conservation properties. 
If you don't embed that in your 

725
00:45:17,710 --> 00:45:20,880
 system, well, first of all, 
your model will be wrong because

726
00:45:20,880 --> 00:45:25,460
you 
 will never be, you know, 
conservative exactly right or as

727
00:45:25,460 --> 00:45:28,942

 exactly as possible. 
But second, even if it's a very,

728
00:45:28,942 --> 00:45:32,172

 it's a system that is very 
conservative, you have to learn 

729
00:45:32,172 --> 00:45:34,560
 that right. 
So you need to use data and 
 

730
00:45:34,568 --> 00:45:38,040
compute to get to that point. 
And of course, that's something 

731
00:45:38,040 --> 00:45:39,675
 that you knew from the 
beginning. 
 

732
00:45:39,683 --> 00:45:43,201
So it's not very smart to to 
waste that information. 
 

733
00:45:43,209 --> 00:45:47,625
So like with everything in life,
right, no extreme is going to be

734
00:45:47,625 --> 00:45:49,480

 optimal. 
And I think that both extremes 


735
00:45:49,488 --> 00:45:52,746
have been explored and we have 
learned a lot from it. 
 

736
00:45:52,754 --> 00:45:57,188
To me, being able to use 
symmetries, conservation laws in

737
00:45:57,188 --> 00:46:00,416

 your problems. 
If you're trying to solve for 
 

738
00:46:00,424 --> 00:46:03,449
several velocity components, I 
mean, do you know that there is 

739
00:46:03,449 --> 00:46:05,848
 the flow is incompressible, Do 
you know that you can use 
 

740
00:46:05,856 --> 00:46:08,584
incompressibility to get the 
velocity component or use that 


741
00:46:08,592 --> 00:46:11,343
in the way that you're building 
your system? 
 

742
00:46:11,351 --> 00:46:15,974
But also going back to an idea 
that I like very much, this 
 

743
00:46:15,982 --> 00:46:17,772
whole explainability, causality 
mechanisms. 
 

744
00:46:17,780 --> 00:46:22,862
If we know these things, why 
don't we try to build systems 
 

745
00:46:22,870 --> 00:46:26,890
that are really focused on this?
And one example is, let's 
 

746
00:46:26,898 --> 00:46:31,292
imagine that I want to create a 
system that predicts very well 


747
00:46:31,300 --> 00:46:34,672
the acoustic field of an airfoil
or a drone, right? 
 

748
00:46:34,680 --> 00:46:35,965
Because drones can be very 
noisy. 
 

749
00:46:35,973 --> 00:46:39,100
I want to be able to control the
acoustic field, the noise 
 

750
00:46:39,108 --> 00:46:44,804
produced by this drone. 
Why don't I use some of these 
 

751
00:46:44,812 --> 00:46:48,636
causality explainability methods
to identify and pinpoint the 
 

752
00:46:48,644 --> 00:46:52,076
mechanisms producing the noise? 
And by mechanisms, I'm talking 


753
00:46:52,084 --> 00:46:54,987
about flow regions interacting 
with each other in physical 
 

754
00:46:54,995 --> 00:46:58,015
space. 
And then when I build a 
 

755
00:46:58,023 --> 00:47:02,896
predictive model, I don't look 
at the, you know, trying to get 

756
00:47:02,896 --> 00:47:06,495
 neither just the governing 
equations like in a PINN way or 

757
00:47:06,495 --> 00:47:08,596
 just use data to predict the 
flow fields. 
 

758
00:47:08,604 --> 00:47:12,020
But rather, why don't I try to 
predict just those structures, 


759
00:47:12,028 --> 00:47:16,736
those mechanisms that produce 
the acoustic field in that way, 

760
00:47:16,736 --> 00:47:21,200
 you're really encapsulating in 
those structures a lot of 
 

761
00:47:21,208 --> 00:47:23,488
information. 
And that's where I think that 
 

762
00:47:23,496 --> 00:47:25,864
one can have computational 
savings with respect to, you 
 

763
00:47:25,872 --> 00:47:28,170
know, doing a much more resolved
simulation. 
 

764
00:47:28,178 --> 00:47:31,639
Because those structures 
producing the acoustic field are

765
00:47:31,639 --> 00:47:34,690

 the result of a lot of, you 
know, interscale energy 
 

766
00:47:34,698 --> 00:47:39,306
interactions that you need to 
resolve when you do a DNS. 
 

767
00:47:39,314 --> 00:47:44,820
But the model, I mean, once that
DNS is run, those structures 
 

768
00:47:44,828 --> 00:47:48,594
that are producing the acoustic 
field and so on, that the result

769
00:47:48,594 --> 00:47:51,474

 of those interactions, right? 
So you don't need to re simulate

770
00:47:51,474 --> 00:47:54,160

 them all the time. 
Once you know what the 
 

771
00:47:54,168 --> 00:47:58,415
mechanisms are, you can try to 
develop models that target those

772
00:47:58,415 --> 00:48:02,694

 physical mechanisms precisely 
and not everything else. 
 

773
00:48:02,702 --> 00:48:06,030
Of course. 
That's why I repeat, sometimes I

774
00:48:06,030 --> 00:48:09,240

 want a model for what? 
If your model is for the 
 

775
00:48:09,248 --> 00:48:11,620
acoustic field, then my advice 
would be, well, identify what 
 

776
00:48:11,628 --> 00:48:14,456
are the mechanisms producing the
acoustic field and try to get 
 

777
00:48:14,464 --> 00:48:17,015
those very well, right? 
Because probably you will not 
 

778
00:48:17,023 --> 00:48:20,295
get a perfect model, but you 
will get something that is 
 

779
00:48:20,303 --> 00:48:22,910
pretty accurate and pretty 
efficient because you don't need

780
00:48:22,910 --> 00:48:26,100

 to solve everything. 
But if your goal is to have 
 

781
00:48:26,108 --> 00:48:28,942
something that gets the acoustic
field very well and the gusts 
 

782
00:48:28,950 --> 00:48:32,952
and the turbulence and the heat 
transfer well, then you need a 


783
00:48:32,960 --> 00:48:35,903
multi physics DNS, right? 
Then, then you need to solve 
 

784
00:48:35,911 --> 00:48:38,360
everything. 
So it depends a little bit on 
 

785
00:48:38,368 --> 00:48:42,456
what you want and of course, 
embedding that physical insight 

786
00:48:42,456 --> 00:48:46,616
 in your model, not as 
equations, but as mechanisms 

787
00:48:46,616 --> 00:48:50,890
that you can 
 then, you know, 
recreate with your system. 
 

788
00:48:50,898 --> 00:48:54,864
I think that that's a pretty 
promising way to go. 
 

789
00:48:54,872 --> 00:48:59,476
And how from a, you know, 
you're, you're obviously a 
 

790
00:48:59,484 --> 00:49:02,632
professor at university. 
What do you think? 
 

791
00:49:02,640 --> 00:49:08,015
What's the role of academia in 
this, focusing more on like the 

792
00:49:08,015 --> 00:49:13,752
 students in the education side,
you know, do you try and team up

793
00:49:13,752 --> 00:49:15,960

 more with the computer science
departments? 
 

794
00:49:15,968 --> 00:49:21,502
Is it a matter of swapping 
people over like, but but then 


795
00:49:21,510 --> 00:49:25,690
equally, I guess you don't want 
to not teach some of your 
 

796
00:49:25,698 --> 00:49:28,290
existing content. 
So how do you balance getting 
 

797
00:49:28,298 --> 00:49:31,740
rid of some content, bringing 
new ones in? 
 

798
00:49:31,748 --> 00:49:36,480
That's that's a key question. 
And we're having many 
 

799
00:49:36,488 --> 00:49:39,822
discussions around this. 
For example, here in Michigan, I

800
00:49:39,822 --> 00:49:43,160

 mean, we're having a lot of 
conversations around our new 
 

801
00:49:43,168 --> 00:49:49,360
curriculum and how to well adapt
to AI era, LLM-based era. 
 

802
00:49:49,368 --> 00:49:52,991
So we're having many 
conversations around this. 
 

803
00:49:52,999 --> 00:49:58,355
So first of all, we have a very 
good interaction with Los Alamos

804
00:49:58,355 --> 00:50:02,675

 National Lab. 
We actually have a bunch of Los 

805
00:50:02,675 --> 00:50:06,660
 Alamos scientists sitting on 
our campus and working with them

806
00:50:06,660 --> 00:50:08,795

 daily and interacting with 
them constantly. 
 

807
00:50:08,803 --> 00:50:13,500
So we really get access to that 
synergy and that cooperation 
 

808
00:50:13,508 --> 00:50:18,032
with computer scientists, which 
are world class, right? 
 

809
00:50:18,040 --> 00:50:21,606
So with with really incredible 
computational facilities. 
 

810
00:50:21,614 --> 00:50:26,229
So I think such an interface and
such a close connection is, is 


811
00:50:26,237 --> 00:50:29,610
essential. 
But the other thing is from the 

812
00:50:29,610 --> 00:50:33,849
 educational point of view, we 
need to be of course aware of 
 

813
00:50:33,857 --> 00:50:36,688
the fact that students have 
access to LLMs, right? 
 

814
00:50:36,696 --> 00:50:40,760
And and that, you know, the way 
that we teach, the way that we 


815
00:50:40,768 --> 00:50:43,484
assess is different from what it
used to be. 
 

816
00:50:43,492 --> 00:50:47,440
But in a way, and this is 
conversations that we're having 

817
00:50:47,440 --> 00:50:50,509
 with my colleagues and I think 
that there's quite some 
 

818
00:50:50,517 --> 00:50:53,242
interesting, I'm promising 
directions in here. 
 

819
00:50:53,250 --> 00:50:58,360
We, we don't need to necessarily
see this as a disadvantage, but 

820
00:50:58,360 --> 00:51:02,800
 rather as an opportunity to use
these systems to encourage 
 

821
00:51:02,808 --> 00:51:09,032
students for more exploration, 
for deeper, for, for deeper 
 

822
00:51:09,040 --> 00:51:14,160
questions and deeper assessments
of the content almost in a 
 

823
00:51:14,168 --> 00:51:17,599
customized way, right? 
So, so I think that rather than 

824
00:51:17,599 --> 00:51:20,822
 thinking that students are 
going to be lazy and just look 

825
00:51:20,822 --> 00:51:23,856
up 
 their answers and not 
think, rather use these these 

826
00:51:23,856 --> 00:51:26,120
systems 
 to help students think
in a way. 

827
00:51:28,480 --> 00:51:32,440
But do you feel that? 
 
I guess flicking around the 

828
00:51:32,440 --> 00:51:36,560
other way, do do computer 
 
science students need to learn 

829
00:51:36,560 --> 00:51:39,280
more about the sciences? 
 
You know, there's a lot of like 

830
00:51:39,440 --> 00:51:42,240
AI for science. 
 
If the engineers are all 

831
00:51:42,240 --> 00:51:44,880
learning some of the computer 
 
science being, then what are the

832
00:51:44,880 --> 00:51:48,360
computer scientists doing? 
 
Is there, is there an equal 

833
00:51:48,640 --> 00:51:50,760
thing on the computer science 
 
say hey you all need to be 

834
00:51:50,760 --> 00:51:53,360
learning some science domain 
 
because AI is going to do the 

835
00:51:53,360 --> 00:51:55,800
some of the core CS stuff, you 

know? 

836
00:51:56,680 --> 00:52:02,728
I was, well, I'm, I can tell you

 that there is, I mean, there 

837
00:52:02,728 --> 00:52:07,600
is an increase of enrollment of 
 students in, in areas like 

838
00:52:07,600 --> 00:52:12,200
aerospace engineering and, and 

part of it comes from computer 

839
00:52:12,200 --> 00:52:14,200
science, although of course the 
 computer science degrees, 

840
00:52:14,480 --> 00:52:18,600
they've been excellent at 
 
adapting to, to new technologies

841
00:52:18,600 --> 00:52:22,400
and new well, developments 
 
within AI for science. 

842
00:52:23,400 --> 00:52:27,880
But there is more and more 
 
overlap and more and more, you 

843
00:52:27,880 --> 00:52:32,400
know, joint educational programs

 across systems and across 

844
00:52:32,720 --> 00:52:36,160
engineering disciplines. 
 
And yes, I would, I would agree 

845
00:52:36,160 --> 00:52:40,080
completely that if computer 
 
scientists are going to be the 

846
00:52:40,520 --> 00:52:45,920
developing AI for science or AI 
 for engineering methods, then 

847
00:52:45,920 --> 00:52:48,760
they're going to need more 
 
background in these, in these 

848
00:52:48,760 --> 00:52:50,560
topics. 
 
And that's something that is 

849
00:52:50,560 --> 00:52:57,720
really that's multidisciplinary 
 wealth and and richness is 

850
00:52:57,720 --> 00:53:02,440
what's going to bring really 
 
stellar research in the next 

851
00:53:02,440 --> 00:53:04,560
years, right? 
 
That 1 is really capable of 

852
00:53:04,560 --> 00:53:07,800
finding completely new 
 
directions like the foundation 

853
00:53:07,800 --> 00:53:11,120
model and the agent are finding 
 patterns across data sets that 

854
00:53:11,120 --> 00:53:14,720
you would not imagine, right. 
 
When we start to interact in 

855
00:53:14,720 --> 00:53:17,800
disciplines at a different 
 
level, we're going to find such 

856
00:53:18,320 --> 00:53:20,520
connections that maybe we did 
 
not expect at the beginning. 

857
00:53:22,360 --> 00:53:25,640
So one of the things that, you 

know, you mentioned a few times 

858
00:53:25,640 --> 00:53:29,720
is on the like control and 
 
optimization side of things. 

859
00:53:30,120 --> 00:53:33,760
And I feel at least myself a 
 
little bit guilty that I, I'm 

860
00:53:33,760 --> 00:53:37,080
always thinking of, you know, 
 
the, the impact of machine 

861
00:53:37,080 --> 00:53:39,880
learning or reduced order models

 in terms of a predicted 

862
00:53:39,880 --> 00:53:41,920
capability. 
 
You know, we can predict this 

863
00:53:41,920 --> 00:53:47,160
flow much faster. 
 
But I guess for industry or even

864
00:53:47,160 --> 00:53:49,880
for scientific discovery, but 
 
particularly for industry, it's 

865
00:53:49,880 --> 00:53:53,680
almost always a kind of 
 
optimization problem in a way 

866
00:53:53,680 --> 00:53:57,720
isn't it's like I need to have 

the best design or the lowest 

867
00:53:57,720 --> 00:54:01,160
drag or the it's never just a 
 
prediction. 

868
00:54:02,360 --> 00:54:06,640
So, So what progress have you 
 
seen with the sort of automatic 

869
00:54:06,640 --> 00:54:10,280
differentiation, the the ability

 of machine learning to maybe 

870
00:54:10,280 --> 00:54:14,200
help advance that optimization 

process? 

871
00:54:15,000 --> 00:54:18,760
Yeah, I I think that prediction 
 is important. 

872
00:54:18,760 --> 00:54:20,960
It's interesting. 
 
It's not the most impactful 

873
00:54:20,960 --> 00:54:22,660
application of machine learning.

 

874
00:54:22,668 --> 00:54:24,623
It's in optimization and 
control. 
 

875
00:54:24,631 --> 00:54:28,060
That's where we can really have 
the, the biggest, the biggest 
 

876
00:54:28,068 --> 00:54:30,189
impact. 
So as I mentioned before, I, I 


877
00:54:30,197 --> 00:54:31,910
worked a lot on reinforcement 
learning. 
 

878
00:54:31,918 --> 00:54:38,722
So that's an area where we have 
been able to control cases that 

879
00:54:38,722 --> 00:54:44,151
 were not possible three or four
years ago or at least with this 

880
00:54:44,151 --> 00:54:48,659
 degree of control authority. 
And not only. 
 

881
00:54:48,667 --> 00:54:53,912
So I would say not just from a 
purely design point of view of 


882
00:54:53,920 --> 00:54:56,487
reducing drag or enhancing 
mixing. 
 

883
00:54:56,495 --> 00:55:01,178
I would argue that for 
discovery, for really 
 

884
00:55:01,186 --> 00:55:05,400
understanding the mechanisms and
the building blocks of these 
 

885
00:55:05,408 --> 00:55:09,148
physical systems, you can use 
optimization and, and, and 
 

886
00:55:09,156 --> 00:55:13,255
reinforcement learning. 
For example, I mean, imagine a 


887
00:55:13,263 --> 00:55:18,900
turbulent flow where a made-up 
turbulent flow where the energy 

888
00:55:18,900 --> 00:55:24,152
 fluxes are restricted to 
certain scales and to certain 

889
00:55:24,152 --> 00:55:27,986
wave 
 numbers dynamically, 
right That you do it as the flow

890
00:55:27,986 --> 00:55:31,296
is running 
 and, and, and you 
want to do that in physical 

891
00:55:31,296 --> 00:55:32,920
space. 
 
So you want to really affect 

892
00:55:32,920 --> 00:55:35,040
certain structures with a 
 
certain forcing. 

893
00:55:35,400 --> 00:55:38,120
Well, you are probably going to 
 have to do it with some 

894
00:55:38,120 --> 00:55:40,280
optimization technique. 
 
And reinforcement learning is 

895
00:55:40,280 --> 00:55:42,840
good for optimizing systems that

 are changing dynamically. 

896
00:55:43,200 --> 00:55:47,160
So you you're going to end up 
 
with a with a modified system or

897
00:55:47,160 --> 00:55:50,800
a system where production is 
 
minimized and dissipation is 

898
00:55:50,800 --> 00:55:53,360
maximized. 
 
How does a flow like that look 

899
00:55:53,360 --> 00:55:56,000
like? 
 
And let me study, let me study 

900
00:55:56,000 --> 00:55:58,040
that new flow, right? 
 
I mean, I have created a 

901
00:55:58,040 --> 00:56:00,880
completely made-up flow thanks 

to optimization. 

902
00:56:01,640 --> 00:56:03,720
And now I can look at the 
 
Spectra, I can look at the 

903
00:56:03,720 --> 00:56:06,840
structures and learn something 

about my problem, right. 

904
00:56:07,160 --> 00:56:12,720
So yes, these capabilities with 
 optimization are key also for 

905
00:56:12,720 --> 00:56:14,480
discovery and so for getting 
 
insight. 

906
00:56:14,800 --> 00:56:18,040
And going back to your previous 
 point about how new systems and

907
00:56:18,040 --> 00:56:22,760
differentiable solvers, I mean, 
 we, we are really having the 

908
00:56:22,760 --> 00:56:30,120
chance of solving problems at 
 
scale of this optimization type 

909
00:56:30,120 --> 00:56:32,680
of system, which was not 
 
possible before. 

910
00:56:33,040 --> 00:56:37,920
And inverse problems where we 
 
can, you know, optimal sensing, 

911
00:56:37,920 --> 00:56:41,200
we can really look at optimal 
 
initial conditions and optimal 

912
00:56:41,200 --> 00:56:45,280
perturbations really for very, 

very challenging configurations.

913
00:56:45,920 --> 00:56:49,600
All of that is possible now 
 
thanks to the scale that we can 

914
00:56:49,600 --> 00:56:54,200
achieve with these systems. 
 
So I think that's, yeah, I mean,

915
00:56:54,200 --> 00:56:56,280
prediction is, is good. 
 
It's interesting. 

916
00:56:56,440 --> 00:57:02,600
It's only part of the problem. 

It's really in the control and 

917
00:57:02,600 --> 00:57:06,080
optimization and in the 
 
explainability of those 

918
00:57:06,080 --> 00:57:09,360
predictions. 
 
Like, OK, I have predicted very 

919
00:57:09,360 --> 00:57:12,080
well in my system, but why are 

those predictions happening? 

920
00:57:12,080 --> 00:57:15,520
And what are the really 
 
important physics and effects 

921
00:57:15,720 --> 00:57:16,880
that I can observe with my 
 
system? 

922
00:57:17,160 --> 00:57:19,320
That's that's the stuff that I 

think is is is actually 

923
00:57:19,320 --> 00:57:22,760
valuable. 
 
So maybe a sort of final 

924
00:57:22,760 --> 00:57:27,040
question or, or topic looking 
 
towards the future a little bit.

925
00:57:27,040 --> 00:57:30,000
You, you said right at the 
 
beginning that if, if I had 

926
00:57:30,000 --> 00:57:33,720
asked you a few years ago, you'd

 have said that we're, we're 

927
00:57:33,720 --> 00:57:35,560
actually, you know, further 
 
away. 

928
00:57:35,960 --> 00:57:44,360
So if you look forward in time 

and let's say to 2030, so 4 

929
00:57:44,360 --> 00:57:49,493
years out, yeah, four years out,

 where do you think we'll be 

930
00:57:49,493 --> 00:57:52,572
at? 
Do you perceive there being a a 

931
00:57:52,572 --> 00:57:57,864
 sort of ChatGPT moment in fluid
or do you think it'll be just 
 

932
00:57:57,872 --> 00:58:02,480
more incremental improvements? 
Yeah, that's a tough question. 


933
00:58:02,488 --> 00:58:09,342
I think one key to go towards 
ChatGPT moment in fluids was to 

934
00:58:09,342 --> 00:58:14,912
 realise that we need a good 
latent representations for our 


935
00:58:14,920 --> 00:58:16,984
problem. 
And I think that now there's 
 

936
00:58:16,992 --> 00:58:19,004
several groups of several people
working in that direction. 
 

937
00:58:19,012 --> 00:58:21,896
So I think we're probably in the
right path. 
 

938
00:58:21,904 --> 00:58:28,088
I would not. 
So I don't know if we want to 
 

939
00:58:28,096 --> 00:58:32,460
aim at being having a ChatGPT 
situation in fluids, maybe 
 

940
00:58:32,468 --> 00:58:37,400
something beyond maybe something
more than what ChatGPT can can 


941
00:58:37,408 --> 00:58:41,383
do. 
What I think that we are aiming 

942
00:58:41,383 --> 00:58:45,245
 at, and I don't know if that 
would be in 2030, but I think 
 

943
00:58:45,253 --> 00:58:48,595
that it's actually reasonable to
think that it should happen in 


944
00:58:48,603 --> 00:58:51,774
the next years, is more 
autonomous discovery, a more 
 

945
00:58:51,782 --> 00:58:57,341
automatic discovery. 
So one, I like the, the one 
 

946
00:58:57,349 --> 00:59:02,605
example of how big changes in in
the paradigm of physics have 
 

947
00:59:02,613 --> 00:59:06,548
taken place and of course, well,
relativity departing from more 


948
00:59:06,556 --> 00:59:10,293
classical mechanics. 
I mean, it takes, if you think 


949
00:59:10,301 --> 00:59:15,624
about it, a lot of a lot of 
coincidences to line up in order

950
00:59:15,624 --> 00:59:19,958

 for that to happen. 
And you need someone who is 
 

951
00:59:19,966 --> 00:59:23,720
smart enough and also crazy 
enough to propose something very

952
00:59:23,720 --> 00:59:25,915

 different. 
And of course, Einstein lacked 


953
00:59:25,923 --> 00:59:30,176
all the math background to be 
able to, you know, develop the 


954
00:59:30,184 --> 00:59:33,576
theory of relativity properly. 
So he had to learn the math and 

955
00:59:33,576 --> 00:59:36,342
 interact with the right people 
to be able to do this properly. 

956
00:59:36,342 --> 00:59:38,120
 
And then, you know, to have the 

957
00:59:38,120 --> 00:59:42,320
right experiments of the of the 
 solar eclipse to really look at

958
00:59:42,320 --> 00:59:46,240
the deviation of the sunbeams. 

And there's some fun stories 

959
00:59:46,240 --> 00:59:49,320
about how those measurements 
 
took place because there were 

960
00:59:49,480 --> 00:59:51,880
several groups doing the 
 
experiments at the same time, 

961
00:59:52,200 --> 00:59:55,200
and they had different types of 
 problems to really get the 

962
00:59:55,200 --> 00:59:59,120
measurements that eventually 
 
corroborated that light was, you

963
00:59:59,120 --> 01:00:03,360
know, displaced by the sun in a 
 way consistent with relativity.

964
01:00:03,720 --> 01:00:08,739
There's a lot of serendipity in 
 discovery, a lot of 

965
01:00:08,739 --> 01:00:13,574
coincidences lined up, and a lot
of human trial 
 and error to 

966
01:00:13,574 --> 01:00:16,320
achieve truly transformative 
discovery in 
 science. 

967
01:00:16,840 --> 01:00:20,960
And I think that the sort of 
 
systems where we can line up on 

968
01:00:20,960 --> 01:00:24,640
the same space, the data from 
 
different disciplines and having

969
01:00:24,640 --> 01:00:30,360
agents systematically probing 
 
and analysing these data sets, 

970
01:00:30,840 --> 01:00:34,440
that can really be the way in 
 
which we can accelerate that. 

971
01:00:34,440 --> 01:00:37,640
But it's this acceleration of 
 
discovery, which is something 

972
01:00:37,640 --> 01:00:40,760
that, you know, many people are 
 talking about with different 

973
01:00:40,760 --> 01:00:42,680
meanings. 
 
And I think it's a little bit 

974
01:00:42,680 --> 01:00:45,400
shallow the way that it's been 

portrayed sometimes. 

975
01:00:45,960 --> 01:00:52,600
To me, what it means is being 
 
able to probe the data and find 

976
01:00:52,600 --> 01:00:58,600
patterns in a more systematic 
 
way than what it takes us as 

977
01:00:58,600 --> 01:01:01,720
humans to take those leaps. 
 
Because we're it's really 

978
01:01:01,720 --> 01:01:06,280
decades and decades and 
 
centuries to line up all those 

979
01:01:06,280 --> 01:01:09,600
coincidences necessary for that 
 leap to happen. 

980
01:01:10,280 --> 01:01:14,200
And if we can do this in a way 

that is much more autonomous and

981
01:01:14,200 --> 01:01:18,800
much more systematic in that 
 
sense, that can be accelerated, 

982
01:01:18,800 --> 01:01:23,040
right? 
 
And this sort of systems I, you 

983
01:01:23,040 --> 01:01:26,920
know, I think that we are really

 getting there actually. 

984
01:01:27,080 --> 01:01:31,994
And, and, and this is not really

 about replacing human insight 

985
01:01:31,994 --> 01:01:34,360
or human input. 
 
This is really about 

986
01:01:34,360 --> 01:01:37,348
accelerating the exploration on 
 the capability of coming up 

987
01:01:37,348 --> 01:01:41,680
with mechanisms with hypothesis 
with 
 a really patterns that 

988
01:01:41,680 --> 01:01:45,760
are non trivial across across 
 
disciplines and across cases. 

989
01:01:46,320 --> 01:01:50,040
So yeah, I think that we're 
 
heading in that direction of 

990
01:01:50,040 --> 01:01:55,160
really being able to make 
 
autonomous discovery thanks to 

991
01:01:55,720 --> 01:01:58,840
the possibility of compressing 

data and interrogating it 

992
01:01:58,840 --> 01:02:06,400
systematically in a way. 
 
Yeah, it does seem that there is

993
01:02:06,400 --> 01:02:10,600
this, I would say inflection 
 
point now where I would largely 

994
01:02:10,600 --> 01:02:12,040
agree with you, you know, a few 
 years. 

995
01:02:12,040 --> 01:02:14,120
Ago when I. 
 
Started to think about some of 

996
01:02:14,120 --> 01:02:17,760
like a foundational model, you 

know, it it did for various 

997
01:02:17,760 --> 01:02:22,920
reasons just felt quite Yeah, 
 
just like impossible almost. 

998
01:02:22,920 --> 01:02:27,640
It would, you know, may still 
 
be, but I I also get the feeling

999
01:02:27,640 --> 01:02:34,200
that there is a, a general 
 
movement of replicating the sort

1000
01:02:34,200 --> 01:02:37,000
of ChatGPT success, you 
 know, 
the the set. 

1001
01:02:37,000 --> 01:02:39,760
And this is broader than just 
 
fluids across science, across 

1002
01:02:39,760 --> 01:02:45,120
physics, across all, you know, 

possible disciplines that, you 

1003
01:02:45,120 --> 01:02:48,960
know, it worked for that which 

and, and I'm not I haven't 

1004
01:02:48,960 --> 01:02:51,680
studied a a great deal, but I, 

you know, I know enough to know 

1005
01:02:51,680 --> 01:02:55,240
that many people thought that it

 couldn't be possible, right. 

1006
01:02:55,600 --> 01:02:57,960
You know, people weren't saying,

 oh, ChatGPT Oh yeah, we knew 

1007
01:02:57,960 --> 01:03:00,280
that was going to come. 
 
You know, it was quite a big 

1008
01:03:00,800 --> 01:03:05,280
impact and, and it, it was done 
 at a scale that was larger than

1009
01:03:05,280 --> 01:03:09,640
anyone thought possible. 
 
And I, I do have a feel that 

1010
01:03:10,280 --> 01:03:13,560
made people more ambitious, you 
 know, and, and think, well, if,

1011
01:03:13,800 --> 01:03:16,000
if it could work for that, it, 

it could work. 

1012
01:03:16,440 --> 01:03:21,280
But I do like your argument on 

that. 

1013
01:03:21,280 --> 01:03:24,400
Sometimes you need to look at 
 
things a bit differently because

1014
01:03:24,400 --> 01:03:29,560
I have a funny feeling that the 
 approach that ultimately makes 

1015
01:03:29,560 --> 01:03:34,440
it may not be the one that is 
 
mainstream today. 

1016
01:03:34,640 --> 01:03:39,560
You know that that like in all 

in science, there's always, it's

1017
01:03:39,560 --> 01:03:42,400
always the person who takes a 
 
little bit of a, an odd 

1018
01:03:42,400 --> 01:03:46,600
direction at the time which 
 
turns out to actually be the 

1019
01:03:46,600 --> 01:03:51,600
one, you know? 
 
An agent to come up with that 

1020
01:03:51,600 --> 01:03:53,360
direction, maybe. 
 
But that's what I meant. 

1021
01:03:53,360 --> 01:03:58,932
Like the, I can totally see how 
 with humans sort of in the 

1022
01:03:58,932 --> 01:04:04,310
loop, you know, in, in, in terms
of 
 guiding things, that the 

1023
01:04:04,310 --> 01:04:09,720
maybe it is ultimately the power
of 
 ChatGPT to find the next 

1024
01:04:09,720 --> 01:04:12,920
ChatGPT. 
 
You know, that those LLM 

1025
01:04:12,920 --> 01:04:15,760
technology and the agents and 
 
the way they work together with,

1026
01:04:15,760 --> 01:04:18,840
of course, the way that you 
 
structure the data, you know, 

1027
01:04:18,960 --> 01:04:20,720
may help. 
 
And, and then maybe that's 

1028
01:04:20,880 --> 01:04:21,564
artificial general intelligence.

 

1029
01:04:21,572 --> 01:04:25,507
And then we can all just, you 
know, go and retire and, and, 
 

1030
01:04:25,515 --> 01:04:30,932
and chill out. 
But yeah, I really appreciate 
 

1031
01:04:30,940 --> 01:04:34,628
having the chance to, to, to 
chat about this. 
 

1032
01:04:34,636 --> 01:04:39,710
What I'm going to do for people 
listening or watching is to put 

1033
01:04:39,710 --> 01:04:44,400
 a bunch of links into the, the 
comments because there's a few 


1034
01:04:44,408 --> 01:04:48,462
papers that you alluded to that 
are really good, you know, great

1035
01:04:48,462 --> 01:04:51,875

 reads that talk in more detail
about some of that explainable 


1036
01:04:51,883 --> 01:04:54,740
AI and some of the work that, 
you know, you've done. 
 

1037
01:04:54,748 --> 01:04:58,460
So I'll, we'll put them in the 
link and highly recommend that 


1038
01:04:58,468 --> 01:05:03,620
people read through them to go 
even deeper than we discussed 
 

1039
01:05:03,628 --> 01:05:06,286
now. 
But yeah, I, I really appreciate

1040
01:05:06,286 --> 01:05:08,582

 it. 
And I think your work is going 


1041
01:05:08,590 --> 01:05:12,280
to be seen as a major 
contributing factor to hopefully

1042
01:05:12,280 --> 01:05:15,405

 achieving that ChatGPT moment.
Well, thank you very much. 
 

1043
01:05:15,413 --> 01:05:18,090
I appreciate it. 
Maybe that's a bit optimistic, 


1044
01:05:18,098 --> 01:05:21,424
the money work, but I appreciate
that much. 
 

1045
01:05:21,432 --> 01:05:25,157
And at the end we, we have fun 
with what we do. 
 

1046
01:05:25,165 --> 01:05:28,020
I think that's kind of like the 
key, the key idea and we keep 
 

1047
01:05:28,028 --> 01:05:31,030
learning. 
So that's that's pretty much it.

1048
01:05:31,030 --> 01:05:31,920

 
Exactly, exactly. 

1049
01:05:32,480 --> 01:05:33,760
Great. 
 
Thank you so much. 

1050
01:05:34,120 --> 01:05:34,640
Thank you so much.
