1
00:00:00,280 --> 00:00:02,400
Hi, and welcome to the Neil 
 
Ashton Podcast. 

2
00:00:03,080 --> 00:00:05,920
In each episode, we explained 
 
some of the fascinating ways 

3
00:00:05,920 --> 00:00:09,080
that science and engineering are

 changing the world around us. 

4
00:00:09,800 --> 00:00:12,760
We talked to leading engineers 

from elite level sports like 

5
00:00:12,840 --> 00:00:16,920
cycling and Formula One to some 
 of the world's top academics to

6
00:00:16,920 --> 00:00:20,490
understand how fluid dynamics, 

machine learning and 

7
00:00:20,490 --> 00:00:23,000
supercomputing are bringing in a
new era of 
 discovery. 

8
00:00:23,960 --> 00:00:27,080
We also hear some of their life 
 stories, their career advice, 

9
00:00:27,640 --> 00:00:30,080
the lessons they've learned on 

the way that I hope will be 

10
00:00:30,080 --> 00:00:33,840
helpful to you too. 
 
So sit back and enjoy this 

11
00:00:33,840 --> 00:00:41,560
episode. 
 
Hi, and welcome back to the Neil

12
00:00:41,560 --> 00:00:45,320
Ashton Podcast. 
 
So on today's episode, we have 

13
00:00:45,320 --> 00:00:48,760
Professor Michael Mahoney, who 

is one of the world's leading 

14
00:00:48,760 --> 00:00:54,640
experts on machine learning, 
 
mathematics, and computer 

15
00:00:54,640 --> 00:00:57,600
science. 
 
He's also somebody who I've got 

16
00:00:57,600 --> 00:01:00,480
to know over the past year and 

has been a great help in 

17
00:01:00,480 --> 00:01:04,040
understanding this space. 
 
He's among other things, an 

18
00:01:04,040 --> 00:01:05,800
Amazon scholar. 
 
So we've managed to work 

19
00:01:05,800 --> 00:01:09,160
together a little bit. 
 
And hopefully, as you can tell 

20
00:01:09,160 --> 00:01:12,440
from this episode, he's a 
 
really nice guy and has a great 

21
00:01:12,440 --> 00:01:15,360
sense of humor and such an 
 
intelligent person. 

22
00:01:15,800 --> 00:01:18,880
Where do I start to talk about 

what he's done? 

23
00:01:18,880 --> 00:01:24,600
Well, he's a professor at UC 
 
Berkeley in the Department of 

24
00:01:24,600 --> 00:01:28,440
Statistics, but he's also at 
 
Lawrence Berkeley National Lab. 

25
00:01:28,680 --> 00:01:32,120
And as I mentioned, he's also an

 Amazon scholar and, and has a 

26
00:01:32,120 --> 00:01:38,320
few more hats as well. 
 
He, if you look at his Google 

27
00:01:38,320 --> 00:01:41,240
Scholar, which for a lot of 
 
academics is a way of, you know,

28
00:01:41,320 --> 00:01:43,960
getting a look at what they've 

done, he has some pretty 

29
00:01:43,960 --> 00:01:47,160
impressive statistics. 
 
And one that sort of comes to 

30
00:01:47,160 --> 00:01:52,520
mind, not only does he have more

 than 36,000 citations, which, 

31
00:01:52,560 --> 00:01:57,600
which is a lot, his h-index is 

84, which is very high. 

32
00:01:58,240 --> 00:02:02,720
But what's more impressive is, 

and again, this is probably 

33
00:02:02,800 --> 00:02:04,600
going into the details, but it 

makes a difference. 

34
00:02:04,600 --> 00:02:06,200
If you go on Google Scholar and 
 you look at the number of 

35
00:02:06,200 --> 00:02:10,479
citations, his is exponentially 
 growing. 

36
00:02:10,479 --> 00:02:14,040
So not only does he have more 
 
than, you know, 30,000 

37
00:02:14,480 --> 00:02:18,660
citations, h-index of 84, which 
is 
 already in the very high, 

38
00:02:18,660 --> 00:02:22,920
sort of top 1% or probably even 
 higher, it is increasing like 

39
00:02:22,920 --> 00:02:27,600
this, which basically shows that

 his research is becoming ever 

40
00:02:27,600 --> 00:02:29,580
more relevant every single year.

 

41
00:02:29,588 --> 00:02:31,920
It's not plateauing or going 
down. 
 

42
00:02:31,928 --> 00:02:35,485
That's actually very impressive 
and that shows why he is so 
 

43
00:02:35,493 --> 00:02:38,744
highly regarded because his work
is really cutting edge. 
 

44
00:02:38,752 --> 00:02:43,628
There are many, I think it's 
fair to say that because he has 

45
00:02:43,628 --> 00:02:45,708
 that mathematics background and
that computer science 
 

46
00:02:45,716 --> 00:02:48,255
background, he's put his hand to
many different things. 
 

47
00:02:48,263 --> 00:02:51,956
But one thing maybe relevant for
the listeners or viewers of this

48
00:02:51,956 --> 00:02:55,320

 podcast is a paper he did with
some colleagues on 
 

49
00:02:55,328 --> 00:02:58,272
characterizing possible failure 
modes in physics-informed neural

50
00:02:58,272 --> 00:03:00,690

 networks. 
And this is, you hear many 
 

51
00:03:00,698 --> 00:03:02,739
people talk about PINNs: 
physics- informed neural 

52
00:03:02,739 --> 00:03:04,440
networks. 
 
And so they did a great paper a 

53
00:03:04,440 --> 00:03:07,800
couple of years ago looking at 

some places where it may not do 

54
00:03:07,800 --> 00:03:11,080
so well. 
 
We talk actually about that in 

55
00:03:11,080 --> 00:03:12,840
the podcast. 
 
We go through some of the themes

56
00:03:12,840 --> 00:03:16,120
I've discussed with other 
 
leading ML experts. 

57
00:03:16,920 --> 00:03:19,468
And we talk about, for example, 
 his opinion on, you know, do 

58
00:03:19,468 --> 00:03:23,120
you really need to include the 

physics in this AI for science 

59
00:03:23,120 --> 00:03:25,560
regime or is it good enough just

 to use data? 

60
00:03:26,440 --> 00:03:29,200
We discussed this at length. 
 
And we also get into the topic 

61
00:03:29,200 --> 00:03:31,960
of foundational models for 
 
science, something that he's 

62
00:03:31,960 --> 00:03:34,400
actually really being a leading 
 voice on. 

63
00:03:34,400 --> 00:03:37,200
He's given numerous keynotes, 
 
important seminars, and 

64
00:03:37,200 --> 00:03:39,880
published papers in this area. 

So we have a good debate about 

65
00:03:39,880 --> 00:03:41,840
that. 
 
But we actually start off the 

66
00:03:41,840 --> 00:03:44,680
conversation talking about 
 
something he's also very well 

67
00:03:44,680 --> 00:03:48,920
known for, which is randomized 

numerical linear algebra, which 

68
00:03:48,920 --> 00:03:52,120
is quite a complex topic, to be,

 to be completely honest. 

69
00:03:52,320 --> 00:03:55,083
And I don't think we got to the 
 bottom of it in this short 

70
00:03:55,083 --> 00:03:58,784
talk, but it really shows that 
some of 
 his work from the pure

71
00:03:58,784 --> 00:04:02,222
math side or pure math side is 
coming 
 through and being 

72
00:04:02,222 --> 00:04:05,245
relevant as we increasingly look
towards lower 
 precision 

73
00:04:05,245 --> 00:04:09,415
methods and ways of doing 
mathematical tricks to try 
 and

74
00:04:09,415 --> 00:04:13,872
improve the speed of many linear
solvers, which are 
 common, of 

75
00:04:13,872 --> 00:04:17,720
course in many, many fields of 
science, including 
 CFD. 

76
00:04:18,920 --> 00:04:24,560
We also talk about his general 

opinions or advice for people 

77
00:04:24,560 --> 00:04:26,440
looking to move between academia

 and industry. 

78
00:04:26,440 --> 00:04:29,160
Just general career advice, 
 
something he's well placed given

79
00:04:29,160 --> 00:04:31,720
that he has an incredible 
 
academic record, but he's also 

80
00:04:31,720 --> 00:04:36,040
very close to industry and sort 
 of industrial applications as 

81
00:04:36,040 --> 00:04:39,480
well. 
 
An hour or an hour and a bit is 

82
00:04:39,840 --> 00:04:44,680
not really enough to go through 
 all of this, but I, I certainly

83
00:04:44,680 --> 00:04:47,600
learnt a lot and I really 
 
enjoyed talking to him like I do

84
00:04:47,600 --> 00:04:50,520
every time I speak to him. 
 
And I hope you'll, you'll find 

85
00:04:50,520 --> 00:04:53,240
the same thing. 
 
I will add some links if you're 

86
00:04:53,240 --> 00:04:58,000
watching this on YouTube because

 again, he's published so many 

87
00:04:58,000 --> 00:05:01,160
papers and done so many great 
 
talks that I really want you 

88
00:05:01,720 --> 00:05:04,067
after watching this, listening 
to 
 this to go to his website 

89
00:05:04,067 --> 00:05:06,812
and and look through all of 
those. 
 

90
00:05:06,820 --> 00:05:10,395
And obviously if you are 
watching this on YouTube, you 
 

91
00:05:10,403 --> 00:05:14,440
may prefer to listen to it. 
Some people don't realise that 


92
00:05:14,448 --> 00:05:17,642
these podcasts are in video form
and YouTube, but they're also on

93
00:05:17,642 --> 00:05:19,300

 Spotify and Apple for audio 
and vice versa. 
 

94
00:05:19,308 --> 00:05:21,842
If you normally listen to this 
and you didn't realise there's a

95
00:05:21,842 --> 00:05:24,720

 full video version, you can go
to YouTube. 
 

96
00:05:24,728 --> 00:05:28,880
So yeah, I hope you enjoy this 
conversation with Professor 
 

97
00:05:28,888 --> 00:05:32,756
Michael Mahoney. 
One topic that you mentioned to 

98
00:05:32,756 --> 00:05:36,875
 me when we were chatting in the
past that I was fascinated by, 


99
00:05:36,883 --> 00:05:40,008
but if I'm being completely 
honest, I didn't fully 
 

100
00:05:40,016 --> 00:05:42,735
understand. 
So it's good for my purpose and 

101
00:05:42,735 --> 00:05:47,774
 for everybody else was around 
the randomized linear algebra. 


102
00:05:47,782 --> 00:05:51,306
What exactly? 
For people who don't know what, 

103
00:05:51,306 --> 00:05:55,470
 what is it and why is this 
becoming even mentioned on a 
 

104
00:05:55,478 --> 00:05:59,540
Netflix show? 
So that yeah. 
 

105
00:05:59,548 --> 00:06:03,864
So that's, this was a sign of 
success when it was finally 
 

106
00:06:03,872 --> 00:06:07,118
mentioned on Netflix—no, the 
Lincoln Lawyer a couple of years

107
00:06:07,118 --> 00:06:10,744

 ago, someone sent it to me and
it was mentioned in in one of 
 

108
00:06:10,752 --> 00:06:15,116
the courtroom scenes. 
So randomized linear algebra, 
 

109
00:06:15,124 --> 00:06:19,420
randomized numerical linear 
algebra is basically an area 
 

110
00:06:19,428 --> 00:06:23,600
that uses randomness as an 
algorithmic resource to solve 
 

111
00:06:23,608 --> 00:06:28,099
linear algebra problems. 
A lot, a lot of linear algebra 


112
00:06:28,107 --> 00:06:30,356
problems are under the hood, 
whether you're doing machine 
 

113
00:06:30,364 --> 00:06:32,704
learning or scientific 
computing, they manifest 
 

114
00:06:32,712 --> 00:06:35,944
themselves in different ways. 
So the exact questions and 
 

115
00:06:35,952 --> 00:06:39,106
numerical issues and so on are, 
are different in those areas. 
 

116
00:06:39,114 --> 00:06:42,940
And that's the source of maybe 
tension in the area, but also 
 

117
00:06:42,948 --> 00:06:45,975
synergy and, and, and, and, you 
know, part of the reason a lot 


118
00:06:45,983 --> 00:06:47,845
of people are interested in it 
because it, it holds the 
 

119
00:06:47,853 --> 00:06:50,480
potential to solve a lot of 
problems people look at, But 
 

120
00:06:50,488 --> 00:06:53,369
it's, it's solving core linear 
algebra problems and, and linear

121
00:06:53,369 --> 00:06:55,256

 algebra ones appear in a lot 
of places. 
 

122
00:06:55,264 --> 00:06:57,555
If you're solving partial 
differential equations, it's, 
 

123
00:06:57,563 --> 00:07:00,720
you know, linear operators or 
iterated forms of linear 

124
00:07:00,720 --> 00:07:02,716
operators 
 appear. 
If you're solving machine 
 

125
00:07:02,724 --> 00:07:06,120
learning, you may be interested 
in support vector machines or 
 

126
00:07:06,128 --> 00:07:09,618
ensembling methods or these days
deep neural networks. 
 

127
00:07:09,626 --> 00:07:12,924
And so matrix multiplications 
are at the core of a lot of that

128
00:07:12,924 --> 00:07:14,640

 stuff. 
So very core linear algebra 
 

129
00:07:14,648 --> 00:07:16,952
problems. 
I think historically the way 
 

130
00:07:16,960 --> 00:07:19,340
people thought about the 
relationship between linear 
 

131
00:07:19,348 --> 00:07:22,800
algebra and and randomness or 
noise was that there's 
 

132
00:07:22,808 --> 00:07:25,680
randomness and noise in the 
world, meaning the data you 
 

133
00:07:25,688 --> 00:07:29,450
measure, think of least squares 
and it's your job to clean it 
 

134
00:07:29,458 --> 00:07:32,132
up. 
And then you call a linear 
 

135
00:07:32,140 --> 00:07:34,760
algebra problem least squares or
low rank approximation. 
 

136
00:07:34,768 --> 00:07:37,986
And you get more or less a 
deterministic answer and you get

137
00:07:37,986 --> 00:07:41,964

 more or less the exact answer.
I mean, and I say exact in 
 

138
00:07:41,972 --> 00:07:44,492
scare quotes because there's 
numerical issues and you 
 can't

139
00:07:44,492 --> 00:07:47,546
represent square root of two on 
a computer, but you know, in so 

140
00:07:47,546 --> 00:07:50,660
 far as machine precision is 
exact and you get an exact 
 

141
00:07:50,668 --> 00:07:53,476
deterministic answer. 
And for a lot of things that's 


142
00:07:53,484 --> 00:07:56,680
just overkill. 
And so you can use randomness as

143
00:07:56,680 --> 00:07:59,789

 an algorithmic resource, 
meaning inside the algorithm to 

144
00:07:59,789 --> 00:08:04,344
speed up 
 computation. 
And this may be most notable 
 

145
00:08:04,352 --> 00:08:08,020
historically, like in Monte 
Carlo or Markov Chain Monte 
 

146
00:08:08,028 --> 00:08:10,800
Carlo, where you run simulations
of of fluid dynamics. 
 

147
00:08:10,808 --> 00:08:14,020
I mean, the Metropolis algorithm
was developed in that context, 


148
00:08:14,028 --> 00:08:17,442
but this is for core numerical 
linear algebra problems. 
 

149
00:08:17,450 --> 00:08:21,386
And so most people, if they're 
running computations, sit on top

150
00:08:21,386 --> 00:08:25,265

 of BLAS and LAPACK and, and, 
and related software. 
 

151
00:08:25,273 --> 00:08:28,660
And if you're calling Python, 
you're calling something else, 


152
00:08:28,668 --> 00:08:30,628
you're calling something that's 
calling something that's calling

153
00:08:30,628 --> 00:08:33,990

 them typically. 
And so these are core libraries.

154
00:08:33,990 --> 00:08:36,679

 
And so the question is, can you,

155
00:08:36,880 --> 00:08:41,000
you know, use good theory from 

randomness and measure 

156
00:08:41,000 --> 00:08:44,080
concentration and high 
 
dimensional probability to 

157
00:08:44,080 --> 00:08:46,880
improve those algorithms or 
 
improve, you know, variance of 

158
00:08:46,880 --> 00:08:48,280
those algorithms? 
 
The short answer is that you 

159
00:08:48,280 --> 00:08:49,840
can. 
 
So how? 

160
00:08:49,840 --> 00:08:53,160
But so how does it work in 
 
practice then? 

161
00:08:53,200 --> 00:08:57,480
So if I'm solving a PDE and I'm 
 normally using some sort of 

162
00:08:57,480 --> 00:09:02,120
linear algebra library and it 
 
takes me this time or this 

163
00:09:02,120 --> 00:09:06,280
amount of flops or this compute,

 how is it that the randomness 

164
00:09:06,600 --> 00:09:12,280
reduces that? 
 
Yeah, I mean, when you're 

165
00:09:12,280 --> 00:09:16,080
solving a PDE, you're typically 
 solving it for a particular 

166
00:09:16,080 --> 00:09:21,480
application and you're calling 

certain core linear algebraic 

167
00:09:21,680 --> 00:09:23,800
primitives in in one way or 
 
another. 

168
00:09:23,800 --> 00:09:26,240
So for example, if you're 
 
solving it with a splitting 

169
00:09:26,240 --> 00:09:29,240
method or predictor corrector 
 
method, or you're solving a 

170
00:09:29,240 --> 00:09:32,259
finite element or finite volume 
 or you're solving a second 

171
00:09:32,259 --> 00:09:35,640
order optimisation, too, as a 
piece of 
 that PDE solver. 

172
00:09:37,000 --> 00:09:40,840
These have least squares, you 
 
know, linear solves things like 

173
00:09:40,840 --> 00:09:43,080
this under the hood. 
 
So typically you're improving 

174
00:09:43,080 --> 00:09:44,400
those. 
 
You could, you could improve the

175
00:09:44,400 --> 00:09:46,062
modeling setup for the PDE also.

 

176
00:09:46,070 --> 00:09:49,464
But but you know, you could, you
could try and solve the, you 
 

177
00:09:49,472 --> 00:09:51,080
know, the improve the core 
primitive. 
 

178
00:09:51,088 --> 00:09:54,860
So take least squares as an 
example, a low-rank 
 

179
00:09:54,868 --> 00:09:58,448
approximation. 
A common motif in in a lot of 
 

180
00:09:58,456 --> 00:10:01,755
these algorithms is you have the
data and you want to solve the 


181
00:10:01,763 --> 00:10:03,252
least squares of the low-rank 
problem. 
 

182
00:10:03,260 --> 00:10:08,228
And you could solve it exactly 
machine precision or whatever. 


183
00:10:08,236 --> 00:10:12,442
But you might want to on the 
other hand, get what they call a

184
00:10:12,442 --> 00:10:15,005

 sketch of the of the data, 
which is roughly a small number 

185
00:10:15,005 --> 00:10:17,470
of 
 data points. 
It could be a small number of 
 

186
00:10:17,478 --> 00:10:20,612
actual data points, or it could 
be what they call a random 
 

187
00:10:20,620 --> 00:10:22,302
projection. 
And if you're familiar with like

188
00:10:22,302 --> 00:10:24,544

 if you're familiar with signal
processing and electrical 

189
00:10:24,544 --> 00:10:28,528
engineering and, 
 and physics, 
think of this as like a 

190
00:10:28,528 --> 00:10:31,120
randomized version of a, 
 of a 
Fourier transform. 

191
00:10:31,120 --> 00:10:34,840
So it takes some signal that 
 
that might be localized in space

192
00:10:34,840 --> 00:10:36,280
and, and spreads it out 
 
everywhere. 

193
00:10:36,920 --> 00:10:38,800
But you don't need this to be a 
 physical space. 

194
00:10:38,800 --> 00:10:41,120
This is just a linear algebraic 
 problem. 

195
00:10:41,120 --> 00:10:46,280
And so you have, you know, you 

have columns and rows, so you 

196
00:10:46,280 --> 00:10:49,480
could select the important 
 
columns or rows, or you could do

197
00:10:49,480 --> 00:10:52,560
this random projection, which is

 essentially a, a basically a 

198
00:10:52,560 --> 00:10:55,440
random type rotation that 
 
spreads the information out and 

199
00:10:55,440 --> 00:10:57,600
then sample uniformly in that 
 
rotated space. 

200
00:10:57,880 --> 00:11:00,840
And so you take this sketch and 
 you can do one of a couple 

201
00:11:00,840 --> 00:11:03,009
things that so some communities,

 the more theoretically 

202
00:11:03,009 --> 00:11:05,512
inclined communities want to 
take that 
 sketch and solve the

203
00:11:05,512 --> 00:11:07,920
subproblem exactly. 
 
And there's a range of 

204
00:11:07,920 --> 00:11:10,920
theoretical work that says the 

solution to the exact, the exact

205
00:11:10,920 --> 00:11:13,720
solution to that subproblem 
 
computed any which way. 

206
00:11:13,720 --> 00:11:16,840
A traditional solver or 
 
something else is epsilon-close 

207
00:11:16,840 --> 00:11:19,080
to the exact solution to the 
 
original problem if you set 

208
00:11:19,080 --> 00:11:22,320
things up right. 
 
Now, in a lot of cases that is 

209
00:11:22,320 --> 00:11:24,800
sort of coarse because you 
 
can't. 

210
00:11:24,800 --> 00:11:27,120
It's hard to get machine 
 
precision just by drawing a 

211
00:11:27,120 --> 00:11:30,120
sample and solving the sub 
 
problem unless you take a huge 

212
00:11:30,120 --> 00:11:33,600
number of samples, just because 
 Monte Carlo methods tend to 

213
00:11:33,600 --> 00:11:36,280
converge slowly as a function of

 the error parameter. 

214
00:11:36,840 --> 00:11:40,440
So you could take that, you 
 
know, low quality, low 

215
00:11:40,440 --> 00:11:44,560
precision, but not trivially bad

 solution and ask yourself what

216
00:11:44,560 --> 00:11:46,640
is a preconditioner? 
 
And by preconditioner I just 

217
00:11:46,640 --> 00:11:50,880
mean a preconditioner for a PDE 
 solver or or least squares or a

218
00:11:50,880 --> 00:11:52,640
linear solver. 
 
And a preconditioner is 

219
00:11:52,640 --> 00:11:55,000
basically a low quality solution

 that you refine. 

220
00:11:55,360 --> 00:11:58,920
So you can actually take this 
 
preconditioner that the that 

221
00:11:58,920 --> 00:12:00,760
this is a sketch. 
 
The theoretically inclined 

222
00:12:00,760 --> 00:12:03,640
people just say good, done; I'm 
 epsilon-good. 

223
00:12:04,120 --> 00:12:06,240
And you can say, now I'm going 

to take that epsilon-good where 

224
00:12:06,240 --> 00:12:09,320
epsilon is 0.1 and just use any 
 of a range of traditional 

225
00:12:09,320 --> 00:12:12,760
iterative solvers and drive that

 epsilon down to 10⁻⁸ or 10⁻¹⁶.

226
00:12:12,760 --> 00:12:15,960

 
And, and just like most 

227
00:12:15,960 --> 00:12:20,030
preconditioners, if, if the cost

 to construct it is less than 

228
00:12:20,030 --> 00:12:23,590
the cost to, you know, the cost 
to 
 construct it plus iterate 

229
00:12:23,590 --> 00:12:26,580
is less than you know, the, the 
 algorithm you're competing 

230
00:12:26,580 --> 00:12:28,360
with, you win. 
 
And it's a good preconditioner. 

231
00:12:29,120 --> 00:12:31,520
So there's a range of ways 
 
people solve it that way. 

232
00:12:31,520 --> 00:12:33,720
So those are called sketch-and- 
 solve and sketch-and- 

233
00:12:33,720 --> 00:12:37,000
precondition, respectively. 
 
Increasingly for machine 

234
00:12:37,000 --> 00:12:41,360
learning applications that want 
 medium precision, also relevant

235
00:12:41,360 --> 00:12:45,200
for scientific computing when 
 
you're interested in low 

236
00:12:45,200 --> 00:12:47,840
precision data representations, 
 you know, going to half- or 

237
00:12:47,840 --> 00:12:50,840
quarter precision, which is, is 
 increasingly seen in hardware. 

238
00:12:51,360 --> 00:12:54,080
There's a more subtle interplay 
 where you can get the solution 

239
00:12:54,080 --> 00:12:56,360
and then you can iterate it and 
 get a second sketch and toggle 

240
00:12:56,360 --> 00:12:57,960
back and forth. 
 
And if you set the parameters 

241
00:12:57,960 --> 00:13:00,960
differently, you can get 
 
intermediate solutions 10 to the

242
00:13:00,960 --> 00:13:04,579
10⁻², 10⁻⁴ and 10⁻⁶ 
 quality. 
So there's a range of ways you 


243
00:13:04,587 --> 00:13:06,998
can use the sketches, but 
roughly the idea is you get most

244
00:13:06,998 --> 00:13:09,494

 of the information in the 
sketch and and solve it or do 

245
00:13:09,494 --> 00:13:13,440
something 
 with it. 
And is there any, you know, 
 

246
00:13:13,448 --> 00:13:17,640
orders of magnitude of the 
savings that you could get? 
 

247
00:13:17,648 --> 00:13:23,344
You know, if I'm doing a Navier–
Stokes solver or solving some 
 

248
00:13:23,352 --> 00:13:27,032
neural net using these 
randomized, you know, numerical 

249
00:13:27,032 --> 00:13:31,240
 linear algebra versus BLAS or 
LINPACK, are we talking, 
 you 

250
00:13:31,240 --> 00:13:36,589
know, a 10% saving? 
Is it an order of magnitude or 


251
00:13:36,597 --> 00:13:40,490
is it still under research? 
And so it's not delivering the 


252
00:13:40,498 --> 00:13:42,190
full yet. 
Yeah. 
 

253
00:13:42,198 --> 00:13:46,440
I mean, I think the question of 
of how well you'd improve upon 


254
00:13:46,448 --> 00:13:49,228
something depends strongly on 
what your baseline is. 
 

255
00:13:49,236 --> 00:13:52,664
And if you feed this into a big 
PDE solver, there's a lot of 
 

256
00:13:52,672 --> 00:13:55,080
moving parts and you're 
competing with very mature code 

257
00:13:55,080 --> 00:13:58,472
 and they're doing a little bit 
better or even a lot better on a

258
00:13:58,472 --> 00:14:01,080

 solver may or may not matter 
for the downstream use case that

259
00:14:01,080 --> 00:14:03,982

 that the scientist that is 
looking at the PDE solver is 
 

260
00:14:03,990 --> 00:14:06,048
interested in. 
Or go to the other extreme. 
 

261
00:14:06,056 --> 00:14:10,200
And you want to say, I want to 
compete with BLAS or 
 LAPACK. 

262
00:14:10,920 --> 00:14:14,800
These are extremely optimized 
 
pieces of code and it's, you 

263
00:14:14,800 --> 00:14:17,360
know, it's hard to beat them. 
 
So one of the big successes in 

264
00:14:17,360 --> 00:14:21,198
the area was with Blendenpik. 
And then soon after that was 

265
00:14:21,198 --> 00:14:24,760
LSRN 
 and Blendenpik, which 
said we wanted to ask whether 

266
00:14:24,760 --> 00:14:29,132
these 
 randomized sketches, you
know, not in theory, not Big O 


267
00:14:29,140 --> 00:14:32,598
notation, not whatever can they 
beat LAPACK Because this is not 

268
00:14:32,598 --> 00:14:36,104
 boutique code you have or I 
have or a solver that you don't 

269
00:14:36,104 --> 00:14:38,476
share 
 with the community. 
So well, this is something 
 

270
00:14:38,484 --> 00:14:40,080
that's been stress-tested for 
decades. 
 

271
00:14:40,088 --> 00:14:43,680
And the answer is yes. 
I mean, basically on any tall 
 

272
00:14:43,688 --> 00:14:46,281
dense matrix, you can beat 
LAPACK with these techniques. 
 

273
00:14:46,289 --> 00:14:48,742
You got to set parameters right 
and be a little bit careful. 
 

274
00:14:48,750 --> 00:14:51,217
But, but the short answer is 
yes, how much you improve it by 

275
00:14:51,217 --> 00:14:54,589
 depends on the aspect ratio and
depends on properties of of the 

276
00:14:54,589 --> 00:14:56,912
 matrix. 
But think of it as ballpark 2 to

277
00:14:56,912 --> 00:15:00,550

 10. 
Now this was back in 2010 and so

278
00:15:00,550 --> 00:15:03,940

 a lot's changed since then, in
particular in the hardware 
 

279
00:15:03,948 --> 00:15:06,756
landscape with respect to GPUs 
and hardware heterogeneity and 


280
00:15:06,764 --> 00:15:09,475
the sunsetting of Moore's law 
and Dennard scaling and these 
 

281
00:15:09,483 --> 00:15:12,920
sorts of things. 
And so it is the case that you 


282
00:15:12,928 --> 00:15:15,510
know, you can be a factor of 10 
better. 
 

283
00:15:15,518 --> 00:15:19,269
Similarly in in in low rank type
approximations. 
 

284
00:15:19,277 --> 00:15:23,589
I think the the the way 
scientists versus machine 
 

285
00:15:23,597 --> 00:15:25,920
learners use low rank 
approximations is very 
 

286
00:15:25,928 --> 00:15:28,880
different. 
Scientists tend to say low rank 

287
00:15:28,880 --> 00:15:33,558
 means you know 99% of the 
Frobenius norm, meaning 99.9% of

288
00:15:33,558 --> 00:15:37,680

 the Frobenius norm and 99.99 I
mean basically the whole matrix.

289
00:15:37,680 --> 00:15:39,320

 
When machine learners say low 

290
00:15:39,320 --> 00:15:41,440
rank, they mean sorry, sorry, 
 
the spectral norm. 

291
00:15:41,760 --> 00:15:44,920
When machine learners say low 
 
rank, they mean, you know, 50% 

292
00:15:44,920 --> 00:15:48,640
of the Frobenius norm, in which 
 case, you know, you lose a lot 

293
00:15:48,640 --> 00:15:51,440
of information, but you might 
 
iterate and and do something 

294
00:15:51,440 --> 00:15:53,000
better. 
 
So think of sort of scaling laws

295
00:15:53,000 --> 00:15:56,760
and the neural scaling context 

with, with neural networks. 

296
00:15:57,160 --> 00:16:00,520
And so the way you'd use those 

low rank algorithms and feed 

297
00:16:00,520 --> 00:16:02,800
them into other solvers 
 would 
be very different. 

298
00:16:03,120 --> 00:16:05,120
And then that would lead to very

 different levels of 

299
00:16:05,120 --> 00:16:06,978
improvement. 
And so I think that's largely 
 

300
00:16:06,986 --> 00:16:08,920
open. 
And then with the memory wall 
 

301
00:16:08,928 --> 00:16:12,120
that you're increasingly seeing 
in, in general, but most 
 

302
00:16:12,128 --> 00:16:14,936
egregiously in the machine 
learning applications, one of 
 

303
00:16:14,944 --> 00:16:18,655
the big wins and it's just 
starting to be explored for the 

304
00:16:18,655 --> 00:16:21,760
 randomized techniques is 
basically improving the memory 


305
00:16:21,768 --> 00:16:24,200
properties. 
And so this could be just 
 

306
00:16:24,208 --> 00:16:25,920
reordering algorithms in the 
sense of communication-avoiding 

307
00:16:25,920 --> 00:16:28,832
 linear algebra could be much 
broader because the randomness 


308
00:16:28,840 --> 00:16:32,195
sort of decouples things, makes 
it easier to parallelise certain

309
00:16:32,195 --> 00:16:35,059

 sorts of computations. 
And so you see much bigger than 

310
00:16:35,059 --> 00:16:36,734
 factors of 2 or 10 improvement 
there. 
 

311
00:16:36,742 --> 00:16:40,352
But then again the baseline's 
changing year to year as as as 


312
00:16:40,360 --> 00:16:43,452
people explore different 
representations and and low 
 

313
00:16:43,460 --> 00:16:46,325
precision and intermediate 
precision and stuff. 
 

314
00:16:46,333 --> 00:16:50,135
Interesting, so how was it 
mentioned in Netflix then going 

315
00:16:50,135 --> 00:16:53,960
 back to the former? 
Well, the particular I, I think 

316
00:16:53,960 --> 00:16:58,036
 so, I'm not a movie star or a 
movie producer, but I, I have a 

317
00:16:58,036 --> 00:17:00,802
 sense that they, they want to, 
I, I mentioned this to someone 


318
00:17:00,810 --> 00:17:02,040
and they said, you know what 
that means? 
 

319
00:17:02,048 --> 00:17:04,730
They tried to take the most 
exotic thing that wouldn't make 

320
00:17:04,730 --> 00:17:07,576
 sense to anyone and, and, and, 
and cited it. 
 

321
00:17:07,584 --> 00:17:11,390
So the context was I, I think it
was early on and the, the, the 


322
00:17:11,397 --> 00:17:14,369
main character, who was this, 
the attorney needed to represent

323
00:17:14,369 --> 00:17:19,048

 someone and the person had to 
get out of, of, of whatever they

324
00:17:19,048 --> 00:17:22,528

 were charged with, which was 
not a, a, you know, relatively 

325
00:17:22,528 --> 00:17:25,320
minor 
 thing, but it was the 
10th time they did it. 
 

326
00:17:25,329 --> 00:17:27,736
And so they, the, the, the 
Lincoln lawyer said, why do you 

327
00:17:27,736 --> 00:17:29,400
 need to get out? 
And they said, I have a defense 

328
00:17:29,400 --> 00:17:30,500
 on Thursday. 
I need to get out. 
 

329
00:17:30,508 --> 00:17:32,345
And they asked what's the, 
what's the topic? 
 

330
00:17:32,353 --> 00:17:35,188
And they said randomization and 
numerical linear algebra. 
 

331
00:17:35,196 --> 00:17:38,992
So that was the context of how 
it was mentioned there. 
 

332
00:17:39,000 --> 00:17:41,120
I wonder how they found that 
out. 
 

333
00:17:41,128 --> 00:17:42,240
I wouldn't know how they 
actually. 
 

334
00:17:42,248 --> 00:17:44,550
Yeah, I don't know. 
They they, they, they didn't 
 

335
00:17:44,558 --> 00:17:46,940
call me up. 
So they must have had someone in

336
00:17:46,940 --> 00:17:49,909

 movie producer land or 
something that did did a Google 

337
00:17:49,909 --> 00:17:52,680
search on 
 on something exotic 
and came across it. 
 

338
00:17:52,688 --> 00:17:53,487
That's fine. 
Yeah. 
 

339
00:17:53,495 --> 00:17:57,076
That reminds me of the Did you 
ever like Star Trek? 
 

340
00:17:57,084 --> 00:18:00,160
Yeah, I used to see if there's a
bifurcation which maybe would 
 

341
00:18:00,168 --> 00:18:04,000
get us in into a separate 
discussion whether you like the 

342
00:18:04,000 --> 00:18:06,960
 original version or the the 
second version. 
 

343
00:18:06,968 --> 00:18:08,636
But then there was this 
explosion of different, 
 

344
00:18:08,644 --> 00:18:13,692
different versions of it. 
Yeah, I was more the next 
 

345
00:18:13,700 --> 00:18:18,240
generation person, but I the 
reason I mention it is it was 
 

346
00:18:18,248 --> 00:18:21,660
often whenever they wanted to 
explain how they were breaking 


347
00:18:21,668 --> 00:18:26,376
the laws of physics, some 
stupidly complicated phrase was 

348
00:18:26,376 --> 00:18:28,735
 used. 
And I'm sure it exists somewhere

349
00:18:28,735 --> 00:18:32,280

 and it just reminds me of 
that, you know, explanation of 

350
00:18:32,280 --> 00:18:33,834
how 
 they're exceeding the warp
drive. 
 

351
00:18:33,842 --> 00:18:36,320
Oh, it's because the something 
something. 
 

352
00:18:36,328 --> 00:18:38,000
Yeah. 
So I think it was a little bit 


353
00:18:38,008 --> 00:18:39,378
like that. 
I mean, I, I think it's a sign 


354
00:18:39,386 --> 00:18:41,934
of the times that what they're 
citing is not warp drive and 
 

355
00:18:41,942 --> 00:18:44,256
quantum gravity, but randomized 
numerical linear algebra. 
 

356
00:18:44,264 --> 00:18:48,189
This is a core statement about 
how areas are progressing and so

357
00:18:48,189 --> 00:18:50,020

 on. 
Exactly. 
 

358
00:18:50,028 --> 00:18:52,040
Exactly. 
Yeah, yeah. 
 

359
00:18:52,048 --> 00:18:59,019
So maybe one of the topics would
be good to chat with you about, 

360
00:18:59,019 --> 00:19:04,880
 and it's definitely a hot topic
at the moment, is the the 
 

361
00:19:04,888 --> 00:19:10,656
foundational models for science.
I've seen quite different 
 

362
00:19:10,664 --> 00:19:14,180
viewpoints. 
There's one argument that is, 
 

363
00:19:14,188 --> 00:19:18,528
well, large language models take
all this data from the Internet.

364
00:19:18,528 --> 00:19:19,600

 
They train and they're 

365
00:19:19,600 --> 00:19:22,000
remarkably good at doing what 
 
they do. 

366
00:19:24,160 --> 00:19:27,560
What about if we could do it for

 scientific disciplines? 

367
00:19:27,680 --> 00:19:31,560
Wouldn't we be able to just, you

 know, ask a prompt to go and 

368
00:19:31,640 --> 00:19:33,960
simulate me a plane or, or, or 

do something? 

369
00:19:34,680 --> 00:19:39,222
And then there's the the other 

avenue, which seems to be more 

370
00:19:39,222 --> 00:19:44,595
the replacing simulation tools 
or 
 augmenting simulation tools

371
00:19:44,595 --> 00:19:48,104
like GraphCast or FourCastNet 
or, 
 you know, these sort of 

372
00:19:48,104 --> 00:19:50,970
seminal pieces of work that have
come 
 out, but they were for 

373
00:19:50,970 --> 00:19:54,120
very, they had to be trained, 
you 
 know, for something. 

374
00:19:56,200 --> 00:20:01,400
Where do you see this? 
 
Where do you see this going? 

375
00:20:01,760 --> 00:20:05,680
Are foundational models 
 
possible? 

376
00:20:05,960 --> 00:20:08,960
Is it we just don't have enough 
 data? 

377
00:20:09,080 --> 00:20:11,880
This is a big topic. 
 
So yeah, I'm not sure how we 

378
00:20:11,880 --> 00:20:14,680
want to start it. 
 
I mean, we're in the middle of 

379
00:20:14,680 --> 00:20:18,120
this fray and I guess that's why

 we were asking about this. 

380
00:20:19,080 --> 00:20:20,960
And I should start off with 
 
saying I don't know what a 

381
00:20:20,960 --> 00:20:23,280
foundation model is. 
 
Actually. 

382
00:20:23,280 --> 00:20:26,360
I know what different people say

 a foundation model is and, and

383
00:20:26,360 --> 00:20:29,080
it means very different things 

to to different communities. 

384
00:20:29,880 --> 00:20:33,880
And so I think articulating that

 a little bit helps articulate 

385
00:20:36,680 --> 00:20:38,608
some of the possible directions 
 and some of the questions 

386
00:20:38,608 --> 00:20:41,520
you're asking. 
 
I mean, I think one way to think

387
00:20:41,520 --> 00:20:44,920
about a lot of these machine 
 
learning methods and it's 

388
00:20:45,120 --> 00:20:47,360
highlighted particularly cleanly

 with these foundation models 

389
00:20:47,360 --> 00:20:51,760
is machine learning methods. 
 
And in particular, these 

390
00:20:51,760 --> 00:20:53,920
foundation models are, they're 

an infrastructure, right? 

391
00:20:53,920 --> 00:20:58,280
The, the, the, the Stanford 
 
report said, call it foundation 

392
00:20:58,280 --> 00:21:01,320
models, not foundational models,

 because it's a foundation in 

393
00:21:01,320 --> 00:21:05,200
which you build things and as 
 
and as opposed to a half dozen 

394
00:21:05,200 --> 00:21:06,400
other terms they could have 
 
used. 

395
00:21:06,400 --> 00:21:08,600
Now, whether that's right or 
 
not, I mean, but that, that's 

396
00:21:08,600 --> 00:21:11,320
what they said. 
 
And so since then people have 

397
00:21:11,560 --> 00:21:13,960
said, well, it has to be this 
 
big and that big and, and 

398
00:21:13,960 --> 00:21:17,440
whatever. 
 
And, and then people want to get

399
00:21:17,440 --> 00:21:20,872
a foundation model, not for all 
 of science, but for areas A, B 

400
00:21:20,872 --> 00:21:24,224
and C and subareas D, E and F 
and get finer 
 and finer 

401
00:21:24,224 --> 00:21:26,216
granularity. 
So in a sense it's, it's a it's 

402
00:21:26,216 --> 00:21:28,156
 a term. 
And, and if you ask people, most

403
00:21:28,156 --> 00:21:30,580

 people using it what it means,
they sort of will acknowledge 
 

404
00:21:30,588 --> 00:21:32,740
that, that they don't know. 
And it's not a standard 
 

405
00:21:32,748 --> 00:21:34,102
definition. 
I think probably the best way to

406
00:21:34,102 --> 00:21:35,940

 think about it is it's, it's, 
it's infrastructure, it's a 
 

407
00:21:35,948 --> 00:21:37,800
foundation on which you build 
things. 
 

408
00:21:37,808 --> 00:21:42,520
And so the computer's 
infrastructure, I mean, you can 

409
00:21:42,520 --> 00:21:45,374
 do different things with the 
computer than you can do with 
 

410
00:21:45,382 --> 00:21:49,130
pencil and paper. 
And so it wasn't obvious the way

411
00:21:49,130 --> 00:21:51,732

 computer science, even before 
it existed, would evolve. 
 

412
00:21:51,740 --> 00:21:54,902
And at one point, post World War
2, the US thought they'd need 

413
00:21:54,902 --> 00:21:56,280
five 
 computers and that would 
be it. 

414
00:21:56,400 --> 00:21:58,280
And then it would solve all the 
 nation's needs, right? 

415
00:21:58,280 --> 00:22:00,040
So it evolved in ways different 
 than expected. 

416
00:22:00,520 --> 00:22:03,480
And it evolved in particular 
 
because it, it could complement 

417
00:22:03,480 --> 00:22:06,360
what people did. 
 
I mean, you could ask different 

418
00:22:06,360 --> 00:22:09,280
scientific questions than you 
 
could with pencil and paper and 

419
00:22:09,280 --> 00:22:11,000
you could ask different 
 
engineering questions than you 

420
00:22:11,000 --> 00:22:12,320
could. 
 
I mean, now you can simulate 

421
00:22:12,440 --> 00:22:15,200
whole things and, and the very, 
 very last step is building it, 

422
00:22:15,200 --> 00:22:17,240
right. 
 
But in addition, you could do 

423
00:22:17,240 --> 00:22:19,160
lots of other things that were 

driven by industry. 

424
00:22:19,160 --> 00:22:21,320
And, and so there's a 
 
bifurcation in the area whether 

425
00:22:21,320 --> 00:22:23,560
you're doing more numerical 
 
things that are continuous or 

426
00:22:23,560 --> 00:22:25,760
whether you're doing discrete 
 
things that, that, that are not 

427
00:22:25,760 --> 00:22:28,320
continuous that were driven 
 
largely by business needs. 

428
00:22:29,080 --> 00:22:31,800
So I think the, the right way to

 think about this is it's that 

429
00:22:31,800 --> 00:22:36,000
sort of infrastructure and it's 
 tied to the data more closely 

430
00:22:36,000 --> 00:22:39,440
than than, you know, database 
 
systems because the foundation 

431
00:22:39,440 --> 00:22:41,280
models have information cooked 

into them. 

432
00:22:42,000 --> 00:22:45,240
And so it wasn't obvious some 
 
number of years ago that just 

433
00:22:45,240 --> 00:22:47,560
with language you, you could do 
 what you know, has caught the 

434
00:22:47,560 --> 00:22:49,480
popular attention in the last 
 
couple of years. 

435
00:22:50,160 --> 00:22:53,680
So I think so to anyone who says

 that confidently, you, you can

436
00:22:53,680 --> 00:22:58,320
or you can't with science, I 
 
mean, probably you can't say 

437
00:22:58,320 --> 00:23:01,680
that they're sort of reliably 
 
saying what what will hold for 

438
00:23:01,680 --> 00:23:02,840
the future 'cause I think it's 

not at all. 

439
00:23:02,960 --> 00:23:04,280
I think it's not at all clear 
for 
 two reasons. 

440
00:23:04,280 --> 00:23:07,360
One, there's certain technical 

things that really do matter. 

441
00:23:07,400 --> 00:23:11,320
I mean that that are different 

about scientific data than, you 

442
00:23:11,320 --> 00:23:12,880
know, language and, and vision 

data. 

443
00:23:12,880 --> 00:23:15,200
And then there's just very 
 
different cultural, there's very

444
00:23:15,200 --> 00:23:17,800
different cultural things. 
 
I think every scientist thinks 

445
00:23:17,800 --> 00:23:21,120
that their data is special and 

unique just like they think that

446
00:23:21,120 --> 00:23:24,640
they're special and unique. 
 
If ML's taught you anything. 

447
00:23:24,640 --> 00:23:26,800
These algorithms can predict 
 
what movie you'll watch and what

448
00:23:26,800 --> 00:23:28,320
stuff you'll buy better than you

 can. 

449
00:23:28,320 --> 00:23:31,240
And so you're maybe slightly 
 
less unique, you know, than than

450
00:23:31,240 --> 00:23:33,960
you thought in some sense. 
 
And so I think there's a 

451
00:23:33,960 --> 00:23:37,040
question, you know, what does 
 
foundation model mean and how 

452
00:23:37,040 --> 00:23:40,040
can it be used so narrowly? 
 
You know, I have a big model I 

453
00:23:40,040 --> 00:23:43,600
can train it in in domain A 
 
domain A could be weather and 

454
00:23:43,600 --> 00:23:46,240
climate have gotten attention, 

but it could be fluid dynamics. 

455
00:23:46,240 --> 00:23:48,640
It could be something else about

 stellar formation. 

456
00:23:48,640 --> 00:23:51,920
It could be properties of 
 
materials and doing density 

457
00:23:51,920 --> 00:23:54,400
functional theory and so on. 
 
And then there's a question 

458
00:23:54,400 --> 00:23:57,480
about whether you could maybe 
 
learn broad based models that 

459
00:23:57,480 --> 00:23:59,280
cut across domains, which would 
 of course be the more 

460
00:23:59,280 --> 00:24:03,440
interesting thing. 
 
So there's something, Chronos 

461
00:24:03,440 --> 00:24:06,360
said, I was involved with the 
 
AWS people that sort of says, I,

462
00:24:06,360 --> 00:24:10,400
I don't want to be the best at 

at the most extreme things. 

463
00:24:10,400 --> 00:24:12,800
I don't want to be the 1% that's

 predicting the most extreme 

464
00:24:12,800 --> 00:24:16,120
things I want to be. 
 
I want to do as well as as the 

465
00:24:16,120 --> 00:24:20,000
majority of users for for time 

series analysis and oversimplify

466
00:24:20,000 --> 00:24:21,240
the story. 
 
I mean, time series is a 

467
00:24:21,240 --> 00:24:24,680
complicated area clearly of 
 
interest in a lot of cases, but 

468
00:24:24,680 --> 00:24:26,320
it's it's a little bit, you 
 
know, you got to be careful 

469
00:24:26,320 --> 00:24:28,880
about off-by-one errors and a 
 
range of other technical things.

470
00:24:29,440 --> 00:24:33,880
And what that says basically is 
 take a language model, language

471
00:24:33,880 --> 00:24:37,262
model structure with a bunch of 
 time series and just 

472
00:24:37,262 --> 00:24:38,816
mean-centre it and 
variance-normalise it. 
 

473
00:24:38,824 --> 00:24:42,950
So just do the simplest possible
things you could and you got to 

474
00:24:42,950 --> 00:24:46,070
 do some data augmentation stuff
and boom, you know, you do sort 

475
00:24:46,070 --> 00:24:48,940
 of comparably and, and better 
than a a wide range of public 
 

476
00:24:48,948 --> 00:24:51,290
benchmarks. 
Now, the public benchmarks are 


477
00:24:51,298 --> 00:24:54,339
not extraordinarily high because
a lot of time series data is 
 

478
00:24:54,347 --> 00:24:56,416
valuable and so companies tend 
not to release it. 
 

479
00:24:56,424 --> 00:25:00,592
But but the fact that you can do
that just by sort of variance- 


480
00:25:00,600 --> 00:25:03,558
normalising and and mean- 
centring is is, you know, not 
 

481
00:25:03,566 --> 00:25:08,215
obvious and sort of interesting.
It turns out sort of on a 
 

482
00:25:08,223 --> 00:25:11,524
separate thread that you could 
say, could I do well compared to

483
00:25:11,524 --> 00:25:14,620

 the 1%, you know, the most 
extreme things and, and the the 

484
00:25:14,620 --> 00:25:17,460
 answers that you can there too.
You got to use different 
 

485
00:25:17,468 --> 00:25:21,012
techniques than you do there. 
And so that leads to questions 


486
00:25:21,020 --> 00:25:25,602
as to whether you could have a 
foundation model for sort of a 


487
00:25:25,610 --> 00:25:28,728
broad swath of, of time series 
and forecasting analysis. 
 

488
00:25:28,736 --> 00:25:32,238
And so, as you know, I said, I'm
an Amazon Scholar working with 


489
00:25:32,246 --> 00:25:34,545
the Scott team and we're looking
at that there. 
 

490
00:25:34,553 --> 00:25:37,484
And one of the interesting 
things I think about there, if 


491
00:25:37,492 --> 00:25:40,630
you look at the details of the 
model, not the Chronos, but some

492
00:25:40,630 --> 00:25:44,112

 other things, why would you 
expect that articles scraped 

493
00:25:44,112 --> 00:25:48,500
from 
 Wikipedia, you know, 
public, publicly available 

494
00:25:48,500 --> 00:25:51,260
language 
 models? 
What would, why would that allow

495
00:25:51,260 --> 00:25:53,441

 you to do better job 
predicting dog food demand? 
 

496
00:25:53,449 --> 00:25:55,940
You know, I mean, it's not 
obvious they have anything to do

497
00:25:55,940 --> 00:25:58,575

 with anything. 
You know, it turns out, I 
 

498
00:25:58,583 --> 00:26:01,980
suspect the hypothesis is that 
the text, the linguistic 
 

499
00:26:01,988 --> 00:26:05,940
structure that you're learning 
from the Wikipedia articles is 


500
00:26:05,948 --> 00:26:08,928
strongly related to sequence to 
sequence modeling, which, you 
 

501
00:26:08,936 --> 00:26:11,904
know, if you think of 1 
dimension as a metric space, 
 

502
00:26:11,912 --> 00:26:15,012
it's a very, very special metric
space, a very specially 
 

503
00:26:15,020 --> 00:26:18,420
structured metric space. 
And so you can learn not just 
 

504
00:26:18,428 --> 00:26:21,125
recent information and not just 
low frequency information over 


505
00:26:21,133 --> 00:26:25,000
the past, but maybe in 
information at different scales.

506
00:26:25,000 --> 00:26:26,480

 
Because, you know, sometimes 

507
00:26:26,560 --> 00:26:30,760
articles in text refer back 10 

words or 100 words or 1000 words

508
00:26:30,760 --> 00:26:32,240
or 10,000 words. 
 
So you can learn these sort of 

509
00:26:32,240 --> 00:26:35,720
heavy-tailed sort of structures 
 and that gives you a better set

510
00:26:35,720 --> 00:26:38,800
of basis functions to learn dog 
 food demand or whatever else. 

511
00:26:38,800 --> 00:26:42,000
So in a sense, these, the 
 
language models give you better 

512
00:26:42,000 --> 00:26:46,880
embeddings, you know, Fourier 
 
analysis and Laplace transforms 

513
00:26:46,880 --> 00:26:49,080
and, and these sort of things 
 
are great for pencil and paper. 

514
00:26:49,080 --> 00:26:50,800
They're all developed in the 
 
1800s, right? 

515
00:26:52,080 --> 00:26:54,720
These are data-driven embeddings

 that are good if you have 

516
00:26:54,720 --> 00:26:57,440
computers and and data not 
 
necessary for pencil and paper, 

517
00:26:57,440 --> 00:26:59,840
but they're good for that. 
 
And so that leads you to the 

518
00:26:59,840 --> 00:27:03,880
question about if, if I'm really

 careful about doing foundation

519
00:27:03,880 --> 00:27:05,880
model work in science, what do I

 need to worry about? 

520
00:27:06,160 --> 00:27:08,520
It's not the details of this PDE

 or that PDE. 

521
00:27:08,800 --> 00:27:12,480
It might be that I want to 
 
reproduce what happens in the 

522
00:27:12,480 --> 00:27:16,160
natural language processing 
(NLP) models, which is roughly 

523
00:27:16,160 --> 00:27:18,120
scale model and data and 
 
compute. 

524
00:27:18,120 --> 00:27:20,480
So none of them saturate. 
 
That's very different than the 

525
00:27:20,480 --> 00:27:23,000
usual strong scaling and weak 
 
scaling and high performance 

526
00:27:23,000 --> 00:27:26,080
computing. 
 
I want to change the amount of 

527
00:27:26,080 --> 00:27:28,840
data I'm putting in. 
 
I want to change the size of the

528
00:27:28,840 --> 00:27:30,160
model. 
 
I want to change the size as it 

529
00:27:30,160 --> 00:27:31,800
computes. 
 
So none of them saturate and, 

530
00:27:31,800 --> 00:27:34,680
and the conjecture would be that

 if if any of them saturate, 

531
00:27:34,680 --> 00:27:36,680
it's going to be harder to do 
 
transfer learning. 

532
00:27:36,680 --> 00:27:40,600
So you really need the non 
 
saturation and you do that one 

533
00:27:40,600 --> 00:27:44,240
way with NLP (natural language 

processing) and CV (computer 

534
00:27:44,240 --> 00:27:48,280
vision). 
 
But in NLP and CV, there's much 

535
00:27:48,280 --> 00:27:52,400
weaker control you have and that

 you need on the spatiotemporal

536
00:27:52,400 --> 00:27:55,800
geometry than PDEs, right? 
 
Is, is, you know, PDEs, if 

537
00:27:55,800 --> 00:27:58,458
you're below a Mach 
 transition
or some other physical 

538
00:27:58,458 --> 00:28:00,000
transition, things are 
 sort of
smooth. 

539
00:28:00,000 --> 00:28:01,640
If you're above it, things are 

very messy. 

540
00:28:02,040 --> 00:28:04,200
Dealing with that transition is 
 hard and people spend their 

541
00:28:04,200 --> 00:28:05,300
whole careers dealing with that.

 

542
00:28:05,308 --> 00:28:08,070
So can you come up with data 
generation or tokenisation 
 

543
00:28:08,078 --> 00:28:10,752
mechanisms that that respect the
spatiotemporal properties? 
 

544
00:28:10,760 --> 00:28:14,560
So I think if you're going to 
have a foundation model that 
 

545
00:28:14,568 --> 00:28:17,920
applies across a broad range of 
sciences that's trained on 
 

546
00:28:17,928 --> 00:28:21,420
weather and, and, and climate 
and, and simulations of 
 

547
00:28:21,428 --> 00:28:23,796
different grid structures from 
the machine learning 
 

548
00:28:23,804 --> 00:28:26,660
perspective, it's not so 
different than satellite data 
 

549
00:28:26,668 --> 00:28:30,096
from satellites looking down, 
you know, it's a different grid 

550
00:28:30,096 --> 00:28:32,732
 and it's an image. 
And then you say, how does that 

551
00:28:32,732 --> 00:28:33,800
 couple with the spatiotemporal 
properties? 
 

552
00:28:33,808 --> 00:28:37,304
So I think if, if you want to 
deliver on the promise in the 
 

553
00:28:37,312 --> 00:28:40,040
same way as computer science had
to work out a range 
 of 

554
00:28:40,040 --> 00:28:42,088
numerical methods to really 
match and beat state-of-the-art,

555
00:28:42,088 --> 00:28:44,055

 you're going to have to do a 
similar thing here. 
 

556
00:28:44,063 --> 00:28:47,300
And so partly depends on these 
technical issues, but partly 
 

557
00:28:47,308 --> 00:28:50,322
depends on cultural issues. 
But how much do you think? 
 

558
00:28:50,330 --> 00:28:53,384
I mean you raised the 1% versus 
the 50%? 
 

559
00:28:53,392 --> 00:29:00,314
I guess that is one argument as 
well, which is how close does it

560
00:29:00,314 --> 00:29:03,628

 need to be to machine 
precision to be useful. 
 

561
00:29:03,636 --> 00:29:07,340
You could argue that any of 
these large language models, the

562
00:29:07,340 --> 00:29:11,228

 way that most people use them,
there is still a correction you 

563
00:29:11,228 --> 00:29:14,328
 need to add. 
You don't typically ask it to 

564
00:29:14,328 --> 00:29:16,840
write 
 a document and literally
take it word for word. 
 

565
00:29:16,848 --> 00:29:19,364
You usually go in and go. 
That's, that's remarkably close,

566
00:29:19,364 --> 00:29:22,320

 but I'm still going to go and 
fix it. 
 

567
00:29:22,328 --> 00:29:25,640
Or it writes you Python code. 
It's unusual, isn't it, that 
 

568
00:29:25,648 --> 00:29:27,864
it's perfect. 
There's usually something you 
 

569
00:29:27,872 --> 00:29:31,742
have to correct. 
So I guess with the science 
 

570
00:29:31,750 --> 00:29:38,080
side, maybe the argument is how 
close does it need to be to be 


571
00:29:38,088 --> 00:29:42,072
useful And the cost of getting 
the incremental increase in 
 

572
00:29:42,080 --> 00:29:45,905
accuracy, you know, is it, is 
there a like a trade off where 


573
00:29:45,913 --> 00:29:49,497
you need so much more? 
Data, I mean, again, I think the

574
00:29:49,497 --> 00:29:52,344

 best analogy is look at the 
history of computer science and 

575
00:29:52,344 --> 00:29:54,520
 how that evolved, I think, I 
think framing the question to 
 

576
00:29:54,528 --> 00:29:57,347
say how close does it need to be
to be useful? 
 

577
00:29:57,355 --> 00:30:03,280
You're already framing it in a 
way that that makes certain one 

578
00:30:03,280 --> 00:30:06,480
 group comfortable, the 
numerical analysts and the PDE 

579
00:30:06,480 --> 00:30:08,240
people that 
 frame a question a
certain way. 

580
00:30:08,640 --> 00:30:11,280
That's not how people would have

 framed the question before. 

581
00:30:11,320 --> 00:30:14,920
I mean that, you know, before 
 
you had, you know, represented 

582
00:30:15,000 --> 00:30:19,080
continuous numbers on, on a 
 
computer discretely take a step 

583
00:30:19,080 --> 00:30:21,120
back and say, I want to solve a 
 certain problem. 

584
00:30:21,120 --> 00:30:24,530
And, and the question is now 
 
that I have very different 

585
00:30:24,530 --> 00:30:28,030
trade- off in terms of compute 
versus 
 data and I have this 

586
00:30:28,030 --> 00:30:29,960
new infrastructure that's 
language 
 models. 

587
00:30:30,280 --> 00:30:34,920
Can I ask a different question 

and, and, and push the science 

588
00:30:34,920 --> 00:30:37,440
for it? 
 
I mean, so in, for example, in, 

589
00:30:39,160 --> 00:30:42,240
in chemistry historically, but 

also nuclear physics and, and in

590
00:30:42,240 --> 00:30:44,480
fluid mechanics, there's this 
 
notion of a semi-empirical 

591
00:30:44,480 --> 00:30:48,520
theory, which is a theory sort 

of derived, you know, it's not 

592
00:30:48,520 --> 00:30:51,560
just curve fitting, it's derived

 from an underlying more 

593
00:30:51,560 --> 00:30:54,200
fundamental theory, maybe 
 
phenomenologically with 

594
00:30:54,200 --> 00:30:56,880
parameters that are then fit 
 
empirically or semi-empirically.

595
00:30:57,160 --> 00:31:01,920
And So what you need there is 
 
not the theory to be right, but 

596
00:31:01,920 --> 00:31:04,000
sort of right enough at the 
 
level of in the, in the 

597
00:31:04,000 --> 00:31:06,320
chemistry is chemical accuracy, 
 which is, is however many 

598
00:31:06,320 --> 00:31:08,760
kilovolts or whatever depends on

 the reaction you're interested

599
00:31:08,760 --> 00:31:11,320
in so on. 
 
And so certain methods that 

600
00:31:11,320 --> 00:31:14,000
maybe were more principled were 
 a bit too coarse for that. 

601
00:31:14,000 --> 00:31:15,880
And other methods, when you 
 
combined it with other 

602
00:31:15,880 --> 00:31:18,040
techniques achieved chemical 
 
accuracy. 

603
00:31:18,320 --> 00:31:20,520
And so none of them were low 
 
enough in the stack that you 

604
00:31:20,520 --> 00:31:23,720
they were better QR codes for QR

 from, from linear algebra. 

605
00:31:24,200 --> 00:31:26,440
But they solved the downstream 

problem at that level of 

606
00:31:26,440 --> 00:31:28,200
actually, probably what you'd 
 
see here is that right? 

607
00:31:28,200 --> 00:31:31,800
So it's, it's probably not going

 to be so successful to say I 

608
00:31:31,800 --> 00:31:36,300
want to get 10⁻¹⁶ or 
 10⁻² of 
my large language model, but but

609
00:31:36,300 --> 00:31:39,187
a foundation 
 model for science
need not be an LLM. 
 

610
00:31:39,195 --> 00:31:41,175
It, it, it, you could be 
literally learning the 
 

611
00:31:41,183 --> 00:31:43,075
embeddings from a range of PDEs.

 

612
00:31:43,083 --> 00:31:45,608
You could, you could train to 
PDEs of different types. 
 

613
00:31:45,616 --> 00:31:49,508
And you could say, I want to 
port if I, if I know the basics 

614
00:31:49,508 --> 00:31:52,168
 of hyperbolic and parabolic. 
And I know that in the transport

615
00:31:52,168 --> 00:31:53,910

 equation, this can enter in a 
certain way. 
 

616
00:31:53,918 --> 00:31:56,184
And you have nonlinearities and 
certain types of forcing 
 

617
00:31:56,192 --> 00:31:59,800
functions, you know, model each 
of those and and solve each of 


618
00:31:59,808 --> 00:32:01,546
the components separately. 
So I think the question is, 
 

619
00:32:01,554 --> 00:32:03,384
what's the right level of 
abstraction to do that? 
 

620
00:32:03,392 --> 00:32:06,324
And, and one of the modalities 
you could use is language, but 


621
00:32:06,332 --> 00:32:09,201
of course you could use image or
PDEs or simulations, any of a 
 

622
00:32:09,209 --> 00:32:12,696
range of things. 
So I think in the context of, of

623
00:32:12,696 --> 00:32:17,170

 scientific problems that the 
foundation model need not be an 

624
00:32:17,170 --> 00:32:19,180
 LLM based. 
I mean, there's some people are 

625
00:32:19,180 --> 00:32:21,658
 pushing that, I think, but but 
it certainly need not be LLM 
 

626
00:32:21,666 --> 00:32:22,900
based. 
We just query the model and say,

627
00:32:22,900 --> 00:32:26,622

 you know, tell me what to look
for, the top quark or whatever. 

628
00:32:26,622 --> 00:32:29,360
 
Well that that's the bit maybe 

629
00:32:29,360 --> 00:32:31,400
you brought up would be 
 
interesting to explore. 

630
00:32:31,400 --> 00:32:36,280
There is, I would say, a 
 
relatively fierce or contested 

631
00:32:36,400 --> 00:32:42,280
argument around physics-informed

 or physics in it versus 

632
00:32:42,280 --> 00:32:48,600
data-driven. 
 
Where do you sit on that, on the

633
00:32:48,600 --> 00:32:49,440
argument? 
 
Do you, do you? 

634
00:32:49,880 --> 00:32:53,240
Does the model need to 
 
implicitly have some of the 

635
00:32:53,240 --> 00:32:56,080
boundary conditions, some 
 
awareness of continuity of 

636
00:32:56,080 --> 00:32:57,880
energy? 
 
Or is it good enough just to 

637
00:32:57,880 --> 00:33:02,880
give it so much data that it 
 
essentially learns those laws by

638
00:33:02,880 --> 00:33:06,840
itself? 
 
Yeah, I mean, I, I think you 

639
00:33:06,840 --> 00:33:09,640
could ask the same question 
 
about in natural language 

640
00:33:09,640 --> 00:33:15,280
processing, computer vision did 
 you just, I mean, there's a, 

641
00:33:15,320 --> 00:33:17,640
there's a common story people 
 
tell which is just get more data

642
00:33:17,640 --> 00:33:20,800
and everything works. 
 
But that ignores the fact as you

643
00:33:20,800 --> 00:33:23,200
know, that a lot of companies 
 
and universities, a lot of 

644
00:33:23,200 --> 00:33:27,800
people have put a lot of 
 
resources into NLP and CV to try

645
00:33:27,800 --> 00:33:29,000
and figure out how to make it 
 
work. 

646
00:33:29,440 --> 00:33:33,400
And you know, convolutions 
 
convolve things, they spread 

647
00:33:33,400 --> 00:33:35,440
stuff out. 
 
You know, if you have an image 

648
00:33:35,440 --> 00:33:38,320
in 2D that may make sense, you 

might want to average over 

649
00:33:38,320 --> 00:33:40,120
nearby things. 
 
But if there's, if there's a 

650
00:33:40,120 --> 00:33:44,080
sharp corner like, you know, the

 shirt against the background 

651
00:33:44,080 --> 00:33:47,640
wall in this picture of 
 me, 
you may smudge stuff out. 

652
00:33:47,640 --> 00:33:49,840
And so there's a range of ways 

to try and sort of deal with 

653
00:33:49,840 --> 00:33:51,680
that. 
 
That's core structure. 

654
00:33:51,680 --> 00:33:54,080
The 2D structure that you're 
 
trying to average over is very 

655
00:33:54,080 --> 00:33:56,080
different than sequence to 
 
sequence learning. 

656
00:33:56,560 --> 00:33:59,200
That's also much more discrete 

in, in natural language 

657
00:33:59,200 --> 00:34:00,800
processing. 
 
So it's not like computer vision

658
00:34:00,800 --> 00:34:02,720
or natural language processing, 
 just learn stuff. 

659
00:34:03,160 --> 00:34:05,440
You gave it very particular 
 
architectures that had very 

660
00:34:05,440 --> 00:34:08,520
particular inductive biases and 
 then you're careful about the 

661
00:34:08,520 --> 00:34:12,000
data and, and trying to get gobs

 of it and, and then then it 

662
00:34:12,000 --> 00:34:13,719
worked. 
 
So I think if you, you know, 

663
00:34:13,719 --> 00:34:15,080
you'd have to follow the same 
 
path here. 

664
00:34:15,080 --> 00:34:18,040
You can't just take an existing 
 architecture and press a button

665
00:34:18,040 --> 00:34:20,520
and hope it works. 
 
I mean, time series, forecasting

666
00:34:20,520 --> 00:34:24,760
space—temporal modelling, 
 
these sorts of things will have 

667
00:34:24,760 --> 00:34:27,440
to be, you have to figure that 

out, how to do that. 

668
00:34:27,719 --> 00:34:31,600
So I think the question, I mean,

 some people, you know, I think

669
00:34:31,600 --> 00:34:35,040
are, and, and maybe sometimes 
 
people that have a vested 

670
00:34:35,040 --> 00:34:37,239
interest in, in moving forward 

or not. 

671
00:34:37,520 --> 00:34:39,360
I mean, some people say, you 
 
know, just ignore everything, 

672
00:34:39,360 --> 00:34:41,000
let the data do it all. 
 
Some people say, Oh, you'll 

673
00:34:41,000 --> 00:34:43,960
never match what I can do with 

my careful PDE solver. 

674
00:34:45,159 --> 00:34:47,600
I mean, I think that's neither 

of those questions is the right 

675
00:34:47,600 --> 00:34:48,880
one. 
 
I mean, the question is, given 

676
00:34:48,880 --> 00:34:51,239
the fact that there's an 
 
infrastructure of code and 

677
00:34:51,239 --> 00:34:53,760
experience numerically, but but 
 you're also at a very different

678
00:34:53,760 --> 00:34:55,520
place. 
 
You've been sitting on top of 

679
00:34:55,520 --> 00:35:00,600
certain essentially hardware 
 
and, and, and linear algebraic 

680
00:35:00,600 --> 00:35:02,600
advances. 
 
I mean, for years, for decades. 

681
00:35:02,960 --> 00:35:07,280
People working on PDEs call 
 
certain linear algebra libraries

682
00:35:07,280 --> 00:35:10,080
and have essentially had not to 
 pay the technical debt 

683
00:35:10,080 --> 00:35:13,080
associated with the fairly high 
 level of of complexity with 

684
00:35:13,120 --> 00:35:15,360
developing new linear algebra 
 
because just wait a year or two 

685
00:35:15,360 --> 00:35:19,160
and the machines are faster. 
 
And so if, if, if that's not the

686
00:35:19,160 --> 00:35:20,840
case, you're going to have to 
 
start thinking a little bit more

687
00:35:20,840 --> 00:35:23,880
carefully about the underlying 

linear algebraic computations. 

688
00:35:24,120 --> 00:35:26,640
Maybe there's different 
 
trade-off points in the space 

689
00:35:27,040 --> 00:35:28,600
and maybe bring something else 

to bear. 

690
00:35:28,600 --> 00:35:33,400
You know, a model that has 
 
learned certain coarse-type 

691
00:35:33,400 --> 00:35:36,760
functions of one dimension, say 
 from the language model that I 

692
00:35:36,760 --> 00:35:39,920
alluded to before that aren't 
 
Fourier modes or aren't Laplace 

693
00:35:39,920 --> 00:35:41,920
modes or aren't something else 

like that that they're more 

694
00:35:41,920 --> 00:35:45,360
familiar with, learn a different

 grid type discretization. 

695
00:35:45,800 --> 00:35:50,268
So I think the real challenge in

 delivering on the promise of 

696
00:35:50,268 --> 00:35:53,207
all this stuff is figuring out 
what 
 the right way to combine 

697
00:35:53,207 --> 00:35:54,680
those two. 
 
So you can imagine rather than 

698
00:35:54,680 --> 00:35:57,520
just writing down a physics- 
 
informed loss and pressing a 

699
00:35:57,520 --> 00:36:00,240
button saying, well, the way 
 
people actually would solve this

700
00:36:00,240 --> 00:36:02,160
is to have some sort of 
 
splitting method and deal with 

701
00:36:02,160 --> 00:36:04,379
the two types of physical things

 and the diffusion and 

702
00:36:04,379 --> 00:36:06,800
advection in two slightly 
different ways. 
 

703
00:36:06,808 --> 00:36:09,800
And machine-learn one and take 
the other as numerical. 
 

704
00:36:09,808 --> 00:36:12,658
And then you couple them 
together, like maybe in a 
 

705
00:36:12,666 --> 00:36:13,630
differentiable end-to-end. 
Yeah. 
 

706
00:36:13,638 --> 00:36:17,005
The, the reason I ask this is 
because it does feel certainly 


707
00:36:17,013 --> 00:36:22,090
at least in the CFD community, 
that there's a, there's a real 


708
00:36:22,098 --> 00:36:24,903
sort of split within the 
community. 
 

709
00:36:24,911 --> 00:36:29,462
And in a sense that on one hand 
you have people saying, well, we

710
00:36:29,462 --> 00:36:34,208

 have all these PDE solvers, 
you know, developed for decades.

711
00:36:34,208 --> 00:36:38,600

 
And you know what convince me 

712
00:36:39,320 --> 00:36:45,080
how your machine learning or AI 
 method is going to be as good 

713
00:36:46,200 --> 00:36:52,840
and faster and cheaper versus on

 the other hand, you have quite

714
00:36:52,840 --> 00:36:58,280
sensational claims of, you know,

 10,000 times faster, you know,

715
00:36:58,280 --> 00:37:01,000
sort of real-time. 
 
Well, of course, when you dig 

716
00:37:01,000 --> 00:37:03,120
in, you know, you start asking 

the questions, well, how much 

717
00:37:03,120 --> 00:37:06,040
data did you have? 
 
What was the time to create the 

718
00:37:06,040 --> 00:37:08,920
data? 
 
And but there is a bigger 

719
00:37:08,920 --> 00:37:14,127
question, I think more when you 
 look to the future, which is, 

720
00:37:14,127 --> 00:37:19,800
is it just a matter of time? 
 
Is it a bit like in the 1980s, 

721
00:37:20,440 --> 00:37:25,280
they could only simulate a plane

 to a reasonably low accuracy 

722
00:37:25,280 --> 00:37:27,360
because the computers just 
 
weren't powerful enough. 

723
00:37:27,680 --> 00:37:30,920
And as you said, you just almost

 keep the same effort almost. 

724
00:37:31,320 --> 00:37:35,280
And you just literally add 20 
 
years of compute on top and now 

725
00:37:35,280 --> 00:37:41,640
you can simulate. 
 
Or is it a we'll never get there

726
00:37:41,640 --> 00:37:46,000
because we'll it's like an it's 
 like impossible to have that 

727
00:37:46,000 --> 00:37:48,440
much data? 
 
This is the bit I think a lot of

728
00:37:48,440 --> 00:37:52,560
VCs and start-ups are also 
 
asking that question if we pump 

729
00:37:52,560 --> 00:37:55,280
100 million into this or a 
 
billion, will we fix it? 

730
00:37:55,720 --> 00:37:58,240
Or is it a trillion dollar 
 
problem and it's not worth 

731
00:37:58,240 --> 00:37:59,960
doing? 
 
Yeah, Yeah. 

732
00:37:59,960 --> 00:38:01,600
I mean, OK. 
 
So I think there's a there's a 

733
00:38:01,600 --> 00:38:02,960
bunch of things in the question 
 there. 

734
00:38:03,600 --> 00:38:06,600
Sorry. 
 
We'll we'll CFD people to BC and

735
00:38:06,600 --> 00:38:10,280
and you know that they have two 
 very different utility. 

736
00:38:10,680 --> 00:38:13,720
You. 
 
Know, I think, OK, so it when 

737
00:38:13,720 --> 00:38:16,640
someone says to me do this, 
 
convince me of that I have these

738
00:38:16,640 --> 00:38:19,480
metrics that have been I've 
 
worked on for decades, I just 

739
00:38:19,480 --> 00:38:20,880
say thank you, good to talk to 

you. 

740
00:38:20,880 --> 00:38:22,934
I mean, because, because there's

 no way you're going to win 

741
00:38:22,934 --> 00:38:24,920
that battle because they're very

 fine-tuned metrics. 

742
00:38:24,920 --> 00:38:26,800
So they're very one little use 

case. 

743
00:38:27,280 --> 00:38:29,760
Even if you do win that battle, 
 they're not going to admit it 

744
00:38:30,040 --> 00:38:32,320
because they'll, they'll tweak 

their method and, and you know, 

745
00:38:32,320 --> 00:38:35,280
be slightly better than you. 
 
And there's greener pastures few

746
00:38:35,280 --> 00:38:37,720
everywhere. 
 
And this is not a hypothetical 

747
00:38:37,720 --> 00:38:39,120
statement. 
 
This is a very practical thing. 

748
00:38:39,120 --> 00:38:41,040
You're asking about this in the 
 context of PDEs. 

749
00:38:41,280 --> 00:38:43,440
We saw this before in randomized

 linear algebra. 

750
00:38:43,440 --> 00:38:46,520
I mean, so literally, I remember

 it was very clear that the 

751
00:38:46,520 --> 00:38:49,960
techniques could potentially be 
 useful, not just in theory, but

752
00:38:49,960 --> 00:38:53,120
in practice, and that the reason

 the numerical people were 

753
00:38:53,120 --> 00:38:54,680
saying, oh, they won't work, You

 know, really. 

754
00:38:54,680 --> 00:38:57,920
I mean, those are those number 

of methods that clearly wouldn't

755
00:38:57,920 --> 00:38:59,560
work. 
 
But but then there's the next 

756
00:38:59,560 --> 00:39:01,080
generation of methods. 
 
And that's sort of when I 

757
00:39:01,080 --> 00:39:02,840
entered the the area and it was 
 clear that they could 

758
00:39:02,840 --> 00:39:04,800
potentially work and the 
 
objections people had before 

759
00:39:04,800 --> 00:39:07,000
wouldn't work. 
 
And so this was right around the

760
00:39:07,000 --> 00:39:09,200
time of the Blendenpik paper 
 
that I mentioned that that 

761
00:39:09,200 --> 00:39:10,680
basically said, can you beat 
LAPACK? 

762
00:39:10,680 --> 00:39:13,920
And the short answer was yes. 
 
So now, for example, if you look

763
00:39:13,920 --> 00:39:16,120
at the SIAM Linear Algebra 
 
meeting, half the talks are on 

764
00:39:16,120 --> 00:39:18,480
this topic, but there's been 
 
this process of creative 

765
00:39:18,480 --> 00:39:21,720
destruction where this they're 

still asking the same questions,

766
00:39:21,720 --> 00:39:24,120
but putting randomness there. 
 
And so that's one group of 

767
00:39:24,200 --> 00:39:25,640
community. 
 
I mean, a very different 

768
00:39:25,640 --> 00:39:28,360
community is the is the people 

that say, I'm going to not do 

769
00:39:28,360 --> 00:39:29,440
that. 
 
I'm going to go and apply it to 

770
00:39:29,440 --> 00:39:31,040
machine learning and 10 other 
 
problems. 

771
00:39:31,480 --> 00:39:35,120
And, and, and so there's that 
 
tension I was alluding to 

772
00:39:35,120 --> 00:39:38,320
before. 
 
So I think if you, if you try 

773
00:39:38,320 --> 00:39:42,560
and, and satisfy an old school 

metric, it's, it's good to have 

774
00:39:42,560 --> 00:39:44,520
a few examples of that as a 
 
proof of principle. 

775
00:39:44,520 --> 00:39:47,240
And, and this is 15 years later,

 and I mentioned it twice here,

776
00:39:47,240 --> 00:39:48,560
right? 
 
So this was a clear proof of 

777
00:39:48,560 --> 00:39:52,800
principle of the area. 
 
The area sort of accelerated 

778
00:39:52,800 --> 00:39:54,520
after that. 
 
And there's been a lot of theory

779
00:39:54,520 --> 00:39:57,360
and empirical development. 
 
So I think there's been a few 

780
00:39:57,360 --> 00:39:59,720
things that, if not that are 
 
starting to look almost like 

781
00:39:59,720 --> 00:40:02,040
that in this general scientific 
 ML area. 

782
00:40:04,160 --> 00:40:06,440
I I think. 
 
History is moving in that 

783
00:40:06,440 --> 00:40:08,640
direction. 
 
So it's clear then that, you 

784
00:40:08,640 --> 00:40:10,440
know, randomness would be 
 
important for linear algebra. 

785
00:40:10,440 --> 00:40:12,680
Now it's even more so for all 
 
the historical and hardware 

786
00:40:12,680 --> 00:40:15,440
trends you were talking about. 

I think likely the same thing's 

787
00:40:15,560 --> 00:40:18,920
going to happen here, right? 
 
And so a question you could ask 

788
00:40:18,920 --> 00:40:22,280
is, are the people who know the 
 linear algebra, are the people 

789
00:40:22,280 --> 00:40:24,400
who know the computational fluid

 dynamics? 

790
00:40:24,760 --> 00:40:26,400
Are they at the table developing

 the methods? 

791
00:40:26,400 --> 00:40:28,360
Are they just going to pooh- 
 
pooh it and say, well, they're 

792
00:40:28,360 --> 00:40:30,640
not going to work because it's 

not going to satisfy my 

793
00:40:30,640 --> 00:40:34,960
particular measure as opposed to

 here's a broad class of 

794
00:40:34,960 --> 00:40:38,000
techniques and it's hard to 
 
imagine that there's nothing in 

795
00:40:38,000 --> 00:40:39,640
my area that will be improved by

 them. 

796
00:40:40,320 --> 00:40:42,760
And so I think that's the that 

that latter question is the 

797
00:40:42,760 --> 00:40:44,280
question to ask. 
 
And that latter is the question 

798
00:40:44,280 --> 00:40:47,040
the questions VCs and and other 
 people will ask. 

799
00:40:47,640 --> 00:40:53,200
And I think the, an important 
 
maybe determinant of how this 

800
00:40:53,200 --> 00:40:56,880
will evolve is, is how the 
 
different players interact with 

801
00:40:56,880 --> 00:41:01,280
this in, in the sense that a lot

 of the machine learning, more 

802
00:41:01,280 --> 00:41:03,720
broadly, scientific machine 
 
learning, not just for the 

803
00:41:03,720 --> 00:41:06,680
foundation model, but more 
 
broadly machine learning really 

804
00:41:06,680 --> 00:41:11,840
is, really is a horizontal. 
 
I mean, it's, it's designed to 

805
00:41:11,840 --> 00:41:15,480
say, what can I do with lots of 
 data and be relatively ignorant

806
00:41:15,480 --> 00:41:18,000
about your application area? 
 
Because who knows why people 

807
00:41:18,000 --> 00:41:20,800
click on ads or click on links 

on their social media account, 

808
00:41:20,800 --> 00:41:21,840
right? 
 
I mean, you can tell the story, 

809
00:41:21,840 --> 00:41:28,320
but really who knows why? 
 
And so a lot of machine learners

810
00:41:28,320 --> 00:41:30,560
implicitly or explicitly will 
 
say, I'm not interested in 

811
00:41:30,560 --> 00:41:33,280
solving your scientific problem 
 if it's just a one off problem.

812
00:41:33,280 --> 00:41:34,960
I mean, if you're a scientist 
 
using machine learning, you 

813
00:41:34,960 --> 00:41:36,508
might be interested in that, but

 that's because you're 

814
00:41:36,508 --> 00:41:38,120
interested in the domain. 
 
But if you're a machine learner,

815
00:41:38,520 --> 00:41:41,320
you might say, I mean, I see 
 
some commonalities between your 

816
00:41:41,320 --> 00:41:44,560
fluid dynamics and this other 
 
solid-mechanics problem. 

817
00:41:44,960 --> 00:41:48,240
And I see some similarities 
 
there between that and something

818
00:41:48,240 --> 00:41:51,880
in in 3D in the atmosphere as 
 
opposed to, you know, something 

819
00:41:51,880 --> 00:41:54,640
very different. 
 
And so identifying those 

820
00:41:54,640 --> 00:41:56,080
commonalities I think is 
 
important. 

821
00:41:56,240 --> 00:41:58,880
What I see in, in in some cases,

 and this is hampered by 

822
00:41:58,880 --> 00:42:01,630
universities and it's hampered 
by 
 industry, and it's hampered

823
00:42:01,630 --> 00:42:05,240
in a couple government labs in, 
in a 
 couple different ways, is

824
00:42:05,240 --> 00:42:09,560
a lot of people want to solve a 
 scientific problem. 

825
00:42:09,560 --> 00:42:12,640
And so they're going to say, I'm

 only going to invest, invest 

826
00:42:12,640 --> 00:42:15,200
literally or figuratively in 
 
machine learning in that one 

827
00:42:15,200 --> 00:42:17,240
area. 
 
And that's just not how machine 

828
00:42:17,240 --> 00:42:19,720
learning works. 
 
You don't see the, I mean, think

829
00:42:19,720 --> 00:42:22,080
of it as horizontal business 
 
with horizontals and verticals. 

830
00:42:22,360 --> 00:42:24,160
If machine learning is a 
 
horizontal, like high 

831
00:42:24,160 --> 00:42:26,760
performance computing that will 
 solve a wide range of problems,

832
00:42:27,120 --> 00:42:28,960
you got to apply it to a 
 
vertical to justify it. 

833
00:42:28,960 --> 00:42:32,120
That's that's where that that's 
 where you know, the rubber hits

834
00:42:32,120 --> 00:42:35,280
the road typically. 
 
And so you don't see the 

835
00:42:35,280 --> 00:42:39,280
successes in machine learning, 

scientific machine learning at 

836
00:42:39,920 --> 00:42:43,120
vertical companies, you know, 
 
companies working in genetics or

837
00:42:43,120 --> 00:42:47,640
or oil exploration or any of a 

range of other particular domain

838
00:42:47,640 --> 00:42:49,760
verticals. 
 
You see it in the horizontal 

839
00:42:49,760 --> 00:42:51,320
companies. 
 
The companies have built out the

840
00:42:51,320 --> 00:42:53,960
technology infrastructure and 
 
they get a team of people that 

841
00:42:53,960 --> 00:42:56,120
know this is a vertical and 
 
applies it there. 

842
00:42:56,320 --> 00:42:59,240
And then, you know, a team that 
 knows a different vertical and 

843
00:42:59,240 --> 00:43:01,680
applies it there. 
 
And so I think the question is, 

844
00:43:01,680 --> 00:43:03,360
how does that play out? 
 
And that's going to be a non- 

845
00:43:03,360 --> 00:43:06,640
trivial evolution in terms of. 

And your. 

846
00:43:06,800 --> 00:43:09,120
That day. 
 
Your comment was interesting as 

847
00:43:09,120 --> 00:43:13,840
well about the difference 
 
between sort of Transformers 

848
00:43:13,840 --> 00:43:17,880
versus CNNs and you know, if you

 just throw the whole world's 

849
00:43:17,880 --> 00:43:22,480
Internet at CNN, that isn't you 
 needed to have the right 

850
00:43:22,480 --> 00:43:26,440
architecture to be able to throw

 a lot of data at it. 

851
00:43:26,960 --> 00:43:29,120
I guess the question is more for

 the science. 

852
00:43:29,160 --> 00:43:33,760
Is there a common architecture 

across the sciences that then 

853
00:43:33,760 --> 00:43:37,720
would allow you to sort of build

 out that big which could be 

854
00:43:37,720 --> 00:43:41,240
applied to verticals? 
 
Or is it the case that the needs

855
00:43:41,240 --> 00:43:47,160
for weather or genetics or 
 
chemistry are so different that 

856
00:43:47,160 --> 00:43:50,400
you would need a different 
 
architecture for for each 

857
00:43:50,400 --> 00:43:52,800
different science? 
 
And therefore the, the, the 

858
00:43:52,920 --> 00:43:54,960
dream of a true foundational 
 
thing. 

859
00:43:54,960 --> 00:43:58,040
Besides, it will never happen 
 
because you know. 

860
00:43:58,960 --> 00:44:00,200
Yeah. 
 
So that's the big question. 

861
00:44:00,200 --> 00:44:04,080
And I, I think there's a 
 
technical component to that and 

862
00:44:04,080 --> 00:44:06,080
then non-technical component in 
 terms of the areas, right? 

863
00:44:06,080 --> 00:44:09,160
And so technically it's 
 
open-ended. 

864
00:44:09,160 --> 00:44:12,720
I don't know, right? 
 
I mean, we're, we're in the 

865
00:44:12,720 --> 00:44:14,920
middle of the fray and we're, 
 
we're voting with our feet and 

866
00:44:14,920 --> 00:44:17,920
placing a bet there that the, 
 
that the answer is or could be 

867
00:44:17,920 --> 00:44:20,640
yes. 
 
To try and address that, you're 

868
00:44:20,640 --> 00:44:22,320
going to have to deal with the 

issues we were talking about 

869
00:44:22,320 --> 00:44:24,920
before. 
 
You know, one dimension is very,

870
00:44:24,920 --> 00:44:27,400
very special. two dimensions, 
the 
 surface area to volume 

871
00:44:27,400 --> 00:44:29,605
properties are very different. 
Three 
 dimensions is very 

872
00:44:29,605 --> 00:44:32,272
different. 
By the time you get four in 
 

873
00:44:32,280 --> 00:44:34,460
scientific computing, that's 
sort of infinite, right? 
 

874
00:44:34,468 --> 00:44:37,238
Machine learners, by the time 
you get to a million, you 
 

875
00:44:37,246 --> 00:44:38,669
know, it still seems pretty 
small. 
 

876
00:44:38,677 --> 00:44:42,744
So mapping the methods to one 
versus two versus three 

877
00:44:42,744 --> 00:44:45,078
dimensions, 
 open question, 
maybe there'll be no trade-off 

878
00:44:45,078 --> 00:44:47,320
point where the 
 surface area 
to volume and the isoperimetry 

879
00:44:47,320 --> 00:44:49,000
and other things 
 like that 
wins. 

880
00:44:49,000 --> 00:44:51,360
I suspect not. 
 
I suspect it could, but I mean, 

881
00:44:51,760 --> 00:44:55,240
open question. 
 
And then non-technically, if you

882
00:44:55,240 --> 00:44:57,800
look at how computer science 
 
evolved, numerical analysis and 

883
00:44:57,800 --> 00:45:00,480
scientific computing was core to

 computer science. 

884
00:45:00,480 --> 00:45:02,560
Look at how numerical analysis 

and computer science evolved. 

885
00:45:03,200 --> 00:45:06,520
Yeah, these areas gave birth to 
 computer science and then 

886
00:45:06,520 --> 00:45:08,520
computer science banished them 

because it was easier to build 

887
00:45:08,520 --> 00:45:11,400
the structure, the theory 
 
around, around Turing machines 

888
00:45:11,400 --> 00:45:13,920
and, and, and the Lambda 
 
calculus and, and discrete 

889
00:45:13,920 --> 00:45:16,720
notions that talked about 
 
complexity independent of the 

890
00:45:16,720 --> 00:45:17,560
data. 
 
And if you're talking about 

891
00:45:17,560 --> 00:45:19,240
complexity independent of the 
 
data, you don't need 

892
00:45:19,240 --> 00:45:21,320
preconditioners because it's 
 
independent of the data. 

893
00:45:22,080 --> 00:45:25,200
And so there's a cultural aspect

 that that says, you know, does

894
00:45:25,200 --> 00:45:28,120
the vertical or the horizontal 

culture win here? 

895
00:45:28,120 --> 00:45:30,920
And so even if you stipulate 
 
that the answer is technically 

896
00:45:30,920 --> 00:45:32,760
yes, I think there's an open 
 
question on how it'll look, 

897
00:45:32,800 --> 00:45:36,040
it'll evolve and whether that'll

 win the day and solve the 

898
00:45:36,040 --> 00:45:38,880
problems that you're asking. 
 
But I think it's not, it's not 

899
00:45:38,880 --> 00:45:42,480
so much that I solved PDE type A

 and applied it to PDE type B. 

900
00:45:42,480 --> 00:45:43,720
It would be certain coarse 
 
things. 

901
00:45:43,720 --> 00:45:47,520
I mean, as you know, there's a 

class of things for which finite

902
00:45:47,520 --> 00:45:49,240
element methods work. 
 
There's a different class of 

903
00:45:49,240 --> 00:45:51,240
things for which finite volume 

methods work. 

904
00:45:51,760 --> 00:45:55,160
And you know, you, you take a 
 
course, you learn a 101 method 

905
00:45:55,160 --> 00:45:57,160
and then boom, they bifurcate 
 
after that. 

906
00:45:57,160 --> 00:45:58,920
So what's the analogous taxonomy

 here? 

907
00:45:58,920 --> 00:45:59,960
And it's probably neither of 
 
those. 

908
00:45:59,960 --> 00:46:02,720
It's probably something that's 

more data-driven, but what's the

909
00:46:02,720 --> 00:46:05,640
right way to taxonomize that? 
 
And even if there's a 

910
00:46:05,640 --> 00:46:08,640
technically a correct way, I 
 
mean, you know, computer science

911
00:46:08,640 --> 00:46:11,520
is not computational science. 
 
Computer science is a coherent 

912
00:46:11,520 --> 00:46:14,040
area. 
 
Computational science is spread 

913
00:46:14,040 --> 00:46:15,920
over many departments and many 

divisions. 

914
00:46:15,920 --> 00:46:17,960
And so I think it's an open 
 
question about how this area 

915
00:46:17,960 --> 00:46:20,360
evolves. 
 
Yeah, it is certainly. 

916
00:46:21,520 --> 00:46:24,360
It is certainly interesting. 
 
I think one of the sort of key 

917
00:46:25,840 --> 00:46:33,680
topics to all of this though is 
 the data availability and 

918
00:46:34,760 --> 00:46:37,228
licensing and financial reward. 
 

919
00:46:37,236 --> 00:46:44,365
So one thing that I've seen in I
guess the early days is nobody 


920
00:46:44,373 --> 00:46:48,152
knew what was going on. 
So I guess the companies were 
 

921
00:46:48,160 --> 00:46:51,514
scraping half the Internet. 
There were a few CAPTCHAs or 
 

922
00:46:51,522 --> 00:46:54,165
there was a few sort of, you 
know, firewalls to stop things. 

923
00:46:54,165 --> 00:46:55,960
 
But in general, it was a sort of

924
00:46:55,960 --> 00:46:59,880
free-for-all it seemed, Whereas 
 now people have, you know, 

925
00:46:59,880 --> 00:47:03,979
cottoned on to that and some of 
 the companies having to do 

926
00:47:03,979 --> 00:47:06,788
deals with publishing companies 
or do 
 deals with, you know, to

927
00:47:06,788 --> 00:47:10,640
try and legally get access or or

 permissively get access to 

928
00:47:10,640 --> 00:47:13,920
things. 
 
I saw the example on the weather

929
00:47:14,760 --> 00:47:19,760
side that, you know, some of 
 
those databases were released 

930
00:47:19,840 --> 00:47:24,240
openly because they were 
 
released openly enabled, you 

931
00:47:24,240 --> 00:47:26,920
know, these various teams to go 
 and do things like, you know, 

932
00:47:26,920 --> 00:47:32,320
FourCastNet and GraphCast. 
 
But those data may not be open 

933
00:47:32,320 --> 00:47:38,280
in the future. 
 
So how much you know, and you 

934
00:47:38,280 --> 00:47:42,600
work in academia, you see people

 want to have their own data. 

935
00:47:42,600 --> 00:47:43,720
Some people want to release 
 
later. 

936
00:47:45,560 --> 00:47:48,400
Do you see that that will 
 
ultimately be the blocker? 

937
00:47:48,400 --> 00:47:51,840
I mean, how do you incentivize 

people to want to make their 

938
00:47:51,840 --> 00:47:56,000
data available and should it 
 
even be done that way, you know,

939
00:47:56,000 --> 00:47:58,680
that the whole open source data 
 versus closed? 

940
00:47:58,880 --> 00:48:01,920
You know, this is this, I think,

 is the elephant in the room in

941
00:48:01,920 --> 00:48:06,360
terms of this area because a lot

 of government grants require 

942
00:48:06,360 --> 00:48:09,040
that you make the data publicly 
 available in some sense of the 

943
00:48:09,040 --> 00:48:10,720
word. 
 
Now, what does in some sense of 

944
00:48:10,720 --> 00:48:12,680
the word mean for a lot of 
 
scientists? 

945
00:48:12,680 --> 00:48:14,760
If the data is more than a year 
 or two old, it's worthless 

946
00:48:14,760 --> 00:48:16,880
because they're trying to push 

the cutting edge of the science 

947
00:48:17,480 --> 00:48:19,720
and then they don't maintain the

 data for more than a couple 

948
00:48:19,720 --> 00:48:21,480
years ago. 
 
And so no one maintains and it's

949
00:48:21,480 --> 00:48:23,160
hard to access. 
 
So there's a lot of public data 

950
00:48:23,160 --> 00:48:27,680
that's not de facto public or 
 
accessible, you know, in a lot 

951
00:48:27,720 --> 00:48:31,000
in those cases and maybe in, in,

 but, but certainly in, in 

952
00:48:31,000 --> 00:48:33,640
machine learning sort of 
 
applications more generally in, 

953
00:48:33,640 --> 00:48:35,640
in Internet and social media and

 so on. 

954
00:48:36,880 --> 00:48:40,080
Having older data is good, but 

you can't get more than a couple

955
00:48:40,080 --> 00:48:42,320
decades older, right? 
 
And so oftentimes the the most 

956
00:48:42,320 --> 00:48:44,440
recent data is the most valuable

 and you just design a method 

957
00:48:44,440 --> 00:48:48,440
around that. 
 
And so it's not obvious to me 

958
00:48:48,440 --> 00:48:50,880
that the way the data are 
 
generated and, and collected, 

959
00:48:50,880 --> 00:48:53,834
distributed now will be the one 
 that is going to be developed 

960
00:48:53,834 --> 00:48:55,880
going forward. 
 
I mean, if you get better 

961
00:48:56,040 --> 00:49:01,120
telescopes, if you get better, 

you know, molecular sort of 

962
00:49:01,120 --> 00:49:06,520
robotic systems to, to to make 

molecules, I can easily imagine 

963
00:49:06,520 --> 00:49:10,280
a situation where data that 
 
exists now or prior to a few 

964
00:49:10,280 --> 00:49:12,200
years ago is basically obsolete 
 and irrelevant. 

965
00:49:12,200 --> 00:49:15,080
And so all of that data sits at 
 whoever generated it. 

966
00:49:15,080 --> 00:49:17,520
And so then the question is, is 
 does a government lab generate 

967
00:49:17,520 --> 00:49:20,040
it? 
 
And if so, do they make it easy 

968
00:49:20,040 --> 00:49:22,840
de facto to access? 
 
Is it done at universities 

969
00:49:23,600 --> 00:49:25,840
where, you know, you have a big 
 infrastructure capital 

970
00:49:25,840 --> 00:49:29,280
investment in in particular, 
 
essentially in particular 

971
00:49:29,280 --> 00:49:30,640
science and verticals to 
 
generate it? 

972
00:49:31,080 --> 00:49:36,680
Is it done at companies that are

 in different verticals or is 

973
00:49:36,680 --> 00:49:40,320
there some sort of collaboration

 or interaction with, you know,

974
00:49:40,360 --> 00:49:44,334
horizontal infrastructure 
 
companies, hyperscalers and then

975
00:49:44,334 --> 00:49:46,240
they get their fingers in 
 that
and have access to it? 

976
00:49:46,920 --> 00:49:49,080
So I think these are all open 
 
questions in in terms of how 

977
00:49:49,080 --> 00:49:51,960
this evolves. 
 
Yeah, it it seems to be very 

978
00:49:51,960 --> 00:49:56,840
much, you know, traditionally 
 
and I guess if this is for me, 

979
00:49:56,840 --> 00:50:00,585
the difference between the 
pretrained, quote- 
 unquote 

980
00:50:00,585 --> 00:50:03,752
foundation models versus just 
letting 
 someone go and train 

981
00:50:03,752 --> 00:50:06,095
them. 
Although I mean, I guess most 
 

982
00:50:06,103 --> 00:50:09,840
people you and I would not go 
and train our own LLM. 
 

983
00:50:09,848 --> 00:50:12,820
We're just going to go and use 
someone who's audit and we're 
 

984
00:50:12,828 --> 00:50:19,492
just doing an inference on it. 
That's and I could see if I'm 
 

985
00:50:19,500 --> 00:50:23,740
BMW and I have data from 
previous car designs or 
 

986
00:50:23,748 --> 00:50:26,710
simulations or whatever, how 
they could train a model using 


987
00:50:26,718 --> 00:50:32,678
their data. 
But I'm sort of wondering in my 

988
00:50:32,678 --> 00:50:37,642
 head, how could it ever become 
the point where some person has 

989
00:50:37,642 --> 00:50:43,862
 access to not only BMW's data, 
but Fords and Audi's and Volvos 

990
00:50:43,862 --> 00:50:48,372
 and Boeings and Airbuses and 
you know, ExxonMobil and Shell 

991
00:50:48,372 --> 00:50:52,210
and 
 everyone else. 
Like how how could they collect 

992
00:50:52,210 --> 00:50:55,217
 the data? 
What would be the incentive for 

993
00:50:55,217 --> 00:50:58,414
 that to happen to or that's 
what I mean. 
 

994
00:50:58,422 --> 00:51:01,722
Or is it the case that that's 
why there will never be a 
 

995
00:51:01,730 --> 00:51:04,196
foundation and it will just be 
that individual people train 
 

996
00:51:04,204 --> 00:51:07,435
their own models because the 
data is just so sensitive and 
 

997
00:51:07,443 --> 00:51:09,738
difficult to collect in one? 
Yeah. 
 

998
00:51:09,746 --> 00:51:12,252
I mean, this is, I think this is
the big elephant in the room. 
 

999
00:51:12,260 --> 00:51:14,107
I mean, it's not clear to me 
that one. 
 

1000
00:51:14,115 --> 00:51:16,776
OK. 
So I think if you were to try 
 

1001
00:51:16,784 --> 00:51:19,361
and generate a meaningful 
scientific foundation model, 
 

1002
00:51:19,369 --> 00:51:23,744
you'd need data from a pretty 
broad range of use cases. 
 

1003
00:51:23,752 --> 00:51:28,320
And so you gave, you know, 
several that in particular are 


1004
00:51:28,328 --> 00:51:31,538
with a strong sort of financial 
backing. 
 

1005
00:51:31,546 --> 00:51:33,480
I'm always reminded of the 
story. 
 

1006
00:51:33,488 --> 00:51:39,464
Jim Gray was, I guess, the most 
prominent person in the 
 

1007
00:51:39,472 --> 00:51:42,282
Microsoft Research division 
years ago, such that, you know, 

1008
00:51:42,282 --> 00:51:44,916
 he would have to have regular 
meetings with Bill Gates, the 
 

1009
00:51:44,924 --> 00:51:47,026
CEO at the time. 
And Bill Gates would say, why is

1010
00:51:47,026 --> 00:51:49,600

 the most prominent person in 
my research division applying 
 

1011
00:51:49,608 --> 00:51:51,515
database techniques to 
astronomy? 
 

1012
00:51:51,523 --> 00:51:55,060
This is when he started working 
with Alex Szalay and some others

1013
00:51:55,060 --> 00:51:57,600

 on, on databases for, for 
astronomy and, and cosmology. 
 

1014
00:51:57,608 --> 00:52:01,295
Why don't you work on something 
more useful like, you know, any 

1015
00:52:01,295 --> 00:52:03,420
 of our products? 
And, and Gray's answer was 
 

1016
00:52:03,428 --> 00:52:06,144
basically, I'm working on it 
because it's worthless, because 

1017
00:52:06,144 --> 00:52:08,364
 there's no financial sort of 
incentive behind it. 
 

1018
00:52:08,372 --> 00:52:10,852
I don't have to talk to the 
lawyers and we can work out IP 


1019
00:52:10,860 --> 00:52:13,630
and all these things. 
It's public data and, and I can 

1020
00:52:13,630 --> 00:52:16,170
 work on database techniques and
figure out latency and bandwidth

1021
00:52:16,170 --> 00:52:19,848

 and all these trade-offs. 
So I don't know, but it would be

1022
00:52:19,848 --> 00:52:22,532

 interesting to see. 
Will you have an analogous thing

1023
00:52:22,532 --> 00:52:24,850

 here? 
So there's there, there is 
 

1024
00:52:24,858 --> 00:52:28,087
public, there's, there's various
places where public data is 
 

1025
00:52:28,095 --> 00:52:30,122
available. 
And I think if you can 
 

1026
00:52:30,130 --> 00:52:33,020
illustrate that you would have a
proof-of-principle foundation- 


1027
00:52:33,028 --> 00:52:36,455
like model. 
Then the question is, will 
 

1028
00:52:36,463 --> 00:52:39,600
people come to it? 
Now clearly the incentives for 


1029
00:52:39,608 --> 00:52:42,395
scientists will be different 
than than for companies there. 


1030
00:52:42,403 --> 00:52:46,260
But I could imagine a situation 
where you build that out and 
 

1031
00:52:46,268 --> 00:52:49,294
then companies could take it and
fine-tune to their particular 
 

1032
00:52:49,302 --> 00:52:51,335
data, right. 
It's not obvious to me that any 

1033
00:52:51,335 --> 00:52:54,092
 of the companies that you 
mentioned have data that that's,

1034
00:52:54,092 --> 00:52:56,990

 that is that special in a 
scientific sense? 
 

1035
00:52:56,998 --> 00:53:00,336
I mean, clearly the particular 
data they have is particular to 

1036
00:53:00,336 --> 00:53:03,045
 their particular situation and 
then the particular, you know, 


1037
00:53:03,053 --> 00:53:07,196
product they have. 
But I, I can imagine that, you 


1038
00:53:07,204 --> 00:53:12,020
know, the PDEs from early solar 
system formation and, and 
 

1039
00:53:12,028 --> 00:53:16,335
hydrogen bomb explosions and 
hydrodynamics inside stars, you 

1040
00:53:16,335 --> 00:53:19,544
 know, is, has some similarities
and some differences with 
 

1041
00:53:19,552 --> 00:53:21,601
weather and climate and some 
similarities and some 
 

1042
00:53:21,609 --> 00:53:23,820
differences with crack 
propagation in the Earth. 
 

1043
00:53:23,828 --> 00:53:26,155
And that'll have some 
similarities and differences 
 

1044
00:53:26,163 --> 00:53:29,268
with, with various sorts of 
sensing modalities that you 
 

1045
00:53:29,276 --> 00:53:31,320
would, that have a geospatial 
component. 
 

1046
00:53:31,328 --> 00:53:34,964
And so can you stress-test that 
and, and get even even a proof 


1047
00:53:34,972 --> 00:53:37,659
of principle that, that if you 
would attack those, then, then 

1048
00:53:37,659 --> 00:53:40,670
the 
 answer would be yes, that,
that you could have a path to 

1049
00:53:40,670 --> 00:53:42,817
have a 
 more foundation model. 
If the answer is yes, then the 


1050
00:53:42,825 --> 00:53:44,000
question is you talk to various 
stakeholders. 
 

1051
00:53:44,008 --> 00:53:46,472
Some of them are government 
entities or government sponsors,

1052
00:53:46,472 --> 00:53:49,335

 some of them are corporate and
try and figure that out. 
 

1053
00:53:49,343 --> 00:53:52,482
And I could, I could imagine 
that that basically it's too big

1054
00:53:52,482 --> 00:53:55,676

 a lift in the area stalls. 
So I think this is something to 

1055
00:53:55,676 --> 00:53:58,199
 be determined. 
But I think having a nucleus of 

1056
00:53:58,199 --> 00:54:01,775
 something is going to be 
important because right now I at

1057
00:54:01,775 --> 00:54:05,020

 least what I've seen, I mean, 
in most of companies like that, 

1058
00:54:05,020 --> 00:54:07,700
 there's, there's, there may be 
some people a little bit 
 

1059
00:54:07,708 --> 00:54:09,960
intrigued, but there's not a 
sufficiently heavy lift that 
 

1060
00:54:09,968 --> 00:54:13,456
they go screaming at their CEO 
to say we need to do this, 
 

1061
00:54:13,464 --> 00:54:15,887
probably because it's not 
extremely strong baselines that 

1062
00:54:15,887 --> 00:54:18,645
 it can happen. 
And their competitors by talking

1063
00:54:18,645 --> 00:54:21,300

 to these public models and 
fine- tuning on their 

1064
00:54:21,300 --> 00:54:23,000
proprietary data 
 will do 
better. 

1065
00:54:23,000 --> 00:54:25,600
I think at that point they'll be

 you can see a change. 

1066
00:54:26,440 --> 00:54:29,080
And I think this is, I mean, I 

guess this is affecting the 

1067
00:54:29,080 --> 00:54:34,168
current generative 
 AI, which 
is how do you make money out of 

1068
00:54:34,168 --> 00:54:36,160
this? 
 
You know, like you can do it for

1069
00:54:36,160 --> 00:54:37,800
the sake of science. 
 
That's nice. 

1070
00:54:38,080 --> 00:54:41,920
I developed this thing. 
 
But I'm always interested to see

1071
00:54:41,920 --> 00:54:46,640
like the money that you put in 

to create the data to train the 

1072
00:54:46,640 --> 00:54:52,727
model, can you ultimately make a

 profitable thing out of it 

1073
00:54:52,727 --> 00:54:56,400
that justifies its cost? 
 
Now, with standard, I guess, 

1074
00:54:57,280 --> 00:55:00,840
simulation software, there's an 
 argument, well, if I simulate 

1075
00:55:00,840 --> 00:55:06,880
it, I don't have to build it. 
 
If I can predict the weather, I 

1076
00:55:06,880 --> 00:55:10,440
can help people figure out if 
 
it's there's going to be a 

1077
00:55:10,440 --> 00:55:11,880
hurricane. 
 
And I can say, you know, there's

1078
00:55:11,880 --> 00:55:15,800
a, there's a clear thing. 
 
But because we can already 

1079
00:55:15,800 --> 00:55:18,960
simulate things, we can already 
 predict the weather. 

1080
00:55:18,960 --> 00:55:21,560
We can always do it. 
 
For me, the machine learning bit

1081
00:55:21,560 --> 00:55:25,300
is just the argument, well, it's

 faster or, or, or it's 

1082
00:55:25,300 --> 00:55:28,172
cheaper. 
That's where the amount of data 

1083
00:55:28,172 --> 00:55:29,840
 do. 
Do you know what I mean? 
 

1084
00:55:29,848 --> 00:55:33,422
Is there a commercial? 
But I, I think, I think there's 

1085
00:55:33,422 --> 00:55:36,450
 a faster question and I think 
it's tempting to say, can I be 


1086
00:55:36,458 --> 00:55:37,775
faster? 
Because faster means there's a 


1087
00:55:37,783 --> 00:55:39,544
number you're comparing to. 
So am I bigger? 
 

1088
00:55:39,552 --> 00:55:44,601
Am I more on some axis? 
And back to the, the, the 
 

1089
00:55:44,609 --> 00:55:47,260
randomized linear algebra 
example, I mean, I, I was 
 

1090
00:55:47,268 --> 00:55:49,440
talking to various people and 
saying, don't worry about the 
 

1091
00:55:49,448 --> 00:55:51,513
randomness. 
You know, if, if you do this and

1092
00:55:51,513 --> 00:55:54,182

 that, that'll work. 
And the people would, you know, 

1093
00:55:54,182 --> 00:55:56,490
 basically when you have a new 
idea, they'll beat you up. 
 

1094
00:55:56,498 --> 00:55:58,155
And so they beat us up and so 
on. 
 

1095
00:55:58,163 --> 00:56:02,020
And a few years later, one of 
those people in particular had 


1096
00:56:02,028 --> 00:56:04,696
some methods and showed it 
worked, you know, not the blend 

1097
00:56:04,696 --> 00:56:06,880
 people, you know, something 
around the same time and their 


1098
00:56:06,888 --> 00:56:09,100
students came back to me and 
didn't know who I was. 
 

1099
00:56:09,108 --> 00:56:10,900
was lecturing me and saying, oh,
don't worry about the 
 

1100
00:56:10,908 --> 00:56:12,445
randomness, it'll all be fine. 
And so on. 
 

1101
00:56:12,453 --> 00:56:16,372
There's just a cultural shift. 
And so there was a cultural 
 

1102
00:56:16,380 --> 00:56:20,594
shift that really was not, is as
faster as this is as better. 
 

1103
00:56:20,602 --> 00:56:23,980
And so I, I, I don't think 
you're going to see the win here

1104
00:56:23,980 --> 00:56:25,272

 just because something's 
faster and better. 
 

1105
00:56:25,280 --> 00:56:28,142
You might need that to point to,
but the win will happen because 

1106
00:56:28,142 --> 00:56:30,520
 you can solve other things that
you couldn't solve before. 
 

1107
00:56:30,528 --> 00:56:33,520
So just as an example, say 
you're, you know, you want a 
 

1108
00:56:33,528 --> 00:56:36,005
better battery and you want to 
understand the way crack 
 

1109
00:56:36,013 --> 00:56:39,356
propagates better, you know, so 
that I mean, you want, you want,

1110
00:56:39,356 --> 00:56:41,280

 you want a longer lifetime, 
you know, warranty. 
 

1111
00:56:41,288 --> 00:56:43,250
So you, you're working on 
developing batteries. 
 

1112
00:56:43,258 --> 00:56:46,696
One of the ways this fits in, 
not in a scientific, but in an 


1113
00:56:46,704 --> 00:56:49,906
engineering or commercial 
context is you want to write a 


1114
00:56:49,914 --> 00:56:53,108
warranty to say that that the 
batteries are going to last so 


1115
00:56:53,116 --> 00:56:54,740
and so long. 
And, and otherwise you'll, you 


1116
00:56:54,748 --> 00:56:58,132
know, you'll eat the cost. 
And so can you predict battery 


1117
00:56:58,140 --> 00:57:00,428
lifetimes. 
And maybe if you have a certain 

1118
00:57:00,428 --> 00:57:02,720
 crack propagation heterogeneity
structure, you know that it 
 

1119
00:57:02,728 --> 00:57:04,855
correlates better with how long 
the batteries will last. 
 

1120
00:57:04,863 --> 00:57:07,100
So now there's a fair 
variability and a particular run

1121
00:57:07,100 --> 00:57:10,972

 on the extreme tails. 
So that sounds sort of very 
 

1122
00:57:10,980 --> 00:57:14,868
different than I want to do oil 
exploration and better fracking 

1123
00:57:14,868 --> 00:57:16,642
 or something. 
I send sound waves down into the

1124
00:57:16,642 --> 00:57:19,284

 earth and I look at how they 
bounce back and I try and learn 

1125
00:57:19,284 --> 00:57:21,680
 something about it. 
But I could imagine that the 
 

1126
00:57:21,688 --> 00:57:25,440
embeddings that you get from a 
core model that cuts across a 
 

1127
00:57:25,448 --> 00:57:28,320
couple different domains learn 
spatiotemporal properties in 
 

1128
00:57:28,328 --> 00:57:32,355
particular of complex systems 
that have sort of long-range, 
 

1129
00:57:32,363 --> 00:57:35,112
non-trivial, heavy-tailed 
correlations, which is what both

1130
00:57:35,112 --> 00:57:39,200

 of those examples have. 
And so say that you can come up 

1131
00:57:39,200 --> 00:57:40,870
 with a model that captures 
those embeddings. 
 

1132
00:57:40,878 --> 00:57:45,645
Now, a company here may tweak it
to their data, do some transfer 

1133
00:57:45,645 --> 00:57:49,142
 learning in the battery space. 
A company over here can do a 
 

1134
00:57:49,150 --> 00:57:52,755
better job adjusting it to their
data, which is, you know, 
 

1135
00:57:52,763 --> 00:57:55,372
spatiotemporally very different 
in terms of, you know, better 
 

1136
00:57:55,380 --> 00:57:57,230
fracking. 
And, you know, for all I know, 


1137
00:57:57,238 --> 00:57:59,410
the same embeddings could be 
useful for, you know, you had a 

1138
00:57:59,410 --> 00:58:02,050
 very different use case, right?
And certainly in the United 
 

1139
00:58:02,058 --> 00:58:05,196
States, there's various places 
where insurance markets don't 
 

1140
00:58:05,204 --> 00:58:08,774
exist for floods or for 
earthquakes or for a range of 
 

1141
00:58:08,782 --> 00:58:11,215
things like that. 
I could imagine that you get 
 

1142
00:58:11,223 --> 00:58:13,344
these coarse embeddings. 
These are coarse functions. 
 

1143
00:58:13,352 --> 00:58:15,360
They're not sinusoids and 
Fourier modes. 
 

1144
00:58:15,368 --> 00:58:16,632
They're something that's 
data-driven embeddings. 
 

1145
00:58:16,640 --> 00:58:20,760
And you say on top of that, I 
want to do transfer learning to 

1146
00:58:20,760 --> 00:58:23,320
 the relatively small amount of 
data that I have. 
 

1147
00:58:23,328 --> 00:58:25,920
That's very high touch. 
You know, you have 10,000 
 

1148
00:58:25,928 --> 00:58:28,632
stations across the country that
measure something in chemicals, 

1149
00:58:28,632 --> 00:58:31,740
 whatever I want to do transfer 
learning to that and I model, 
 

1150
00:58:31,748 --> 00:58:34,855
you know, model the system, the 
way information flows in the 
 

1151
00:58:34,863 --> 00:58:39,320
Mississippi Delta or something. 
And, and if I do that, I can do 

1152
00:58:39,320 --> 00:58:42,560
 a better job, you know, 
predicting the tails of flood 
 

1153
00:58:42,568 --> 00:58:43,195
events. 
Why? 
 

1154
00:58:43,203 --> 00:58:46,450
Because as a model, I don't know
how things propagate, but you 
 

1155
00:58:46,458 --> 00:58:50,238
know, subtle effects in terms of
the density of the soil or, or 


1156
00:58:50,246 --> 00:58:53,700
the moisture of the soil coupled
with, you know, 10,000 other 
 

1157
00:58:53,708 --> 00:58:56,360
things that the neural network 
learns in a complicated way. 
 

1158
00:58:56,368 --> 00:58:59,108
What's hard to reason about 
scientifically means that now I 

1159
00:58:59,108 --> 00:59:01,384
 can create that market. 
And that's, that's something 
 

1160
00:59:01,392 --> 00:59:04,462
that any of those companies I 
could imagine try and you know, 

1161
00:59:04,462 --> 00:59:06,512
 buy out a startup company that 
does that. 
 

1162
00:59:06,520 --> 00:59:07,947
So I think you're going to see a
win. 
 

1163
00:59:07,955 --> 00:59:10,718
Not just I'm running faster, 
something faster than you, but 


1164
00:59:10,726 --> 00:59:15,168
but in a space like that. 
Yeah, no, I, I, I think what 
 

1165
00:59:15,176 --> 00:59:18,448
you're getting at is you, what 
we've seen from current machine 

1166
00:59:18,448 --> 00:59:20,951
learning 
 that it seems to do 
things that we don't fully 

1167
00:59:20,951 --> 00:59:23,771
understand, but 
 are like you, 
like you gave the original 

1168
00:59:23,771 --> 00:59:27,080
example of the time 
 series 
that an LLM that wasn't 

1169
00:59:27,080 --> 00:59:29,240
explicitly trained on it 
 
actually ended up being quite 

1170
00:59:29,240 --> 00:59:32,200
useful. 
 
So I guess we shouldn't 

1171
00:59:32,200 --> 00:59:35,720
underestimate the potential of 

how some of these things may do 

1172
00:59:35,720 --> 00:59:40,400
something that we're, it's not 

just about replicating what an 

1173
00:59:40,400 --> 00:59:43,240
existing simulation tool could 

do, but maybe it could help to 

1174
00:59:44,400 --> 00:59:47,320
merge different disciplines 
 
together and see the links and 

1175
00:59:47,320 --> 00:59:48,640
the, and the, and the things 
 
between them. 

1176
00:59:50,320 --> 00:59:54,480
I did want to pivot to a 
 
different topic, which is when 

1177
00:59:54,480 --> 00:59:58,440
some of the other people I've 
 
spoken to, I'm always really 

1178
00:59:58,440 --> 01:00:03,760
fascinated by the curiosity or 

the uniqueness or the 

1179
01:00:03,760 --> 01:00:09,040
peculiarity of academia and, and

 the sort of different people's

1180
01:00:09,040 --> 01:00:12,320
careers and, and, and 
 
progression through. 

1181
01:00:12,320 --> 01:00:16,360
And I think you have quite an 
 
interesting one, because you are

1182
01:00:16,400 --> 01:00:20,920
a very pure. 
 
Mathematician, academic, you 

1183
01:00:20,920 --> 01:00:25,000
know, you, you are very much in 
 that, but you also have, you 

1184
01:00:25,000 --> 01:00:27,760
know, the Amazon Scholar 
 
position; you're at, you know, 

1185
01:00:27,760 --> 01:00:35,000
Lawrence Berkeley. 
 
Was this a a strategic thing in 

1186
01:00:35,000 --> 01:00:39,520
the sense that you wanted to 
 
work on real problems? 

1187
01:00:40,240 --> 01:00:46,280
Was this just you find more 
 
interesting, like is there, How 

1188
01:00:46,280 --> 01:00:48,320
has it benefited you? 
 
And would you advise it to 

1189
01:00:48,320 --> 01:00:48,960
others? 
 
You know? 

1190
01:00:49,760 --> 01:00:52,040
Yeah. 
 
Yeah. 

1191
01:00:52,040 --> 01:00:54,160
I mean, it's there's a lot of, 

again, a lot of questions. 

1192
01:00:54,200 --> 01:00:57,040
I mean, at least the narrow 
 
questions, universities versus 

1193
01:00:57,040 --> 01:00:59,040
not, yeah. 
 
Yeah, there you go. 

1194
01:00:59,080 --> 01:01:03,240
I'll start. 
 
That but, but, but yeah, you 

1195
01:01:03,240 --> 01:01:05,320
tapped enough stuff into the 
 
question to make it forward- 

1196
01:01:05,320 --> 01:01:07,240
compatible with lots of answers 
 as it sounds like. 

1197
01:01:08,000 --> 01:01:12,032
I mean, so my, my, you know, OK,

 so my PhD's in physics, I 

1198
01:01:12,032 --> 01:01:14,480
never actually studied computer 
 science or statistics. 

1199
01:01:14,480 --> 01:01:16,880
And it was in computational 
 
statistical mechanics. 

1200
01:01:17,600 --> 01:01:21,960
And at the end of the 
 
dissertation I switched areas 

1201
01:01:21,960 --> 01:01:24,280
basically to theoretical 
 
computer science to work on 

1202
01:01:25,880 --> 01:01:30,040
originally Markov chain 
 
algorithms, some of the stuff 

1203
01:01:30,040 --> 01:01:32,720
I've been working on 
 
computationally, Markov chains 

1204
01:01:32,720 --> 01:01:34,320
and molecular dynamics. 
 
So it was a technical 

1205
01:01:34,320 --> 01:01:36,160
connection, but it was it was 
 
fairly loose. 

1206
01:01:37,000 --> 01:01:40,160
And then the randomization 
 
inside the Markov chains that 

1207
01:01:40,160 --> 01:01:42,600
the people I was working with 
 
were also working on these very 

1208
01:01:42,600 --> 01:01:44,360
early versions of these 
 
randomized linear algebra 

1209
01:01:44,360 --> 01:01:46,800
algorithms. 
 
So I got involved with that and 

1210
01:01:46,800 --> 01:01:49,204
I'd done enough coding that it 
was 
 clear certain ones are 

1211
01:01:49,204 --> 01:01:50,760
going to be useful and certain 
ones 
 weren't. 

1212
01:01:50,760 --> 01:01:54,960
And so at that point I decided 

to switch areas partly because 

1213
01:01:54,960 --> 01:01:59,000
of what I've been working on. 
 
Well, two reasons there was a 

1214
01:01:59,000 --> 01:02:01,120
forcing function. 
 
This is I just mentioned this 

1215
01:02:01,120 --> 01:02:03,720
because if you have younger 
 
people listening, it's good to 

1216
01:02:03,720 --> 01:02:06,720
know that sometimes older people

 have had chequered past, let's

1217
01:02:06,720 --> 01:02:08,760
say. 
 
So Long story short, I had 

1218
01:02:08,760 --> 01:02:11,028
career ending problems with the 
 dissertation advisor and and 

1219
01:02:11,028 --> 01:02:13,660
the supply-demand structure in 
the 
 natural sciences and 

1220
01:02:13,660 --> 01:02:15,640
engineering is very different 
than in 
 computer science. 

1221
01:02:16,160 --> 01:02:19,240
And so that was was motivation 

to, you know, jump off the 

1222
01:02:19,240 --> 01:02:25,240
Cliff. 
 
I also realized that that the 

1223
01:02:25,240 --> 01:02:27,320
particular work I've been doing 
 in computational statistical 

1224
01:02:27,320 --> 01:02:28,337
mechanics, this is a great area.

 

1225
01:02:28,345 --> 01:02:30,472
I loved it and and it's informed
a lot of the stuff that I've 
 

1226
01:02:30,480 --> 01:02:34,060
done since then. 
But you know, it was a great 
 

1227
01:02:34,068 --> 01:02:38,268
area to be in 1970 because you 
know, forward-compatible with a 

1228
01:02:38,268 --> 01:02:40,420
 huge expansion. 
There's going to be Nobel prizes

1229
01:02:40,420 --> 01:02:42,126

 given out. 
I mean, just lots of good stuff 

1230
01:02:42,126 --> 01:02:43,965
 going on. 
Not so much in 2000. 
 

1231
01:02:43,973 --> 01:02:47,622
The area, it's sort of sad, but 
it was clear that that the 
 

1232
01:02:47,630 --> 01:02:52,049
future is going to be data 
analysis and, and I joke that I 

1233
01:02:52,049 --> 01:02:54,635
 never knew more about 
algorithms or data than the day 

1234
01:02:54,635 --> 01:02:56,480
I switched. 
 
And then I've just been becoming

1235
01:02:56,480 --> 01:03:00,040
more ignorant since then. 
 
So I don't know quite what I 

1236
01:03:00,400 --> 01:03:05,520
meant by data analysis then, but

 but but clearly I was right. 

1237
01:03:05,520 --> 01:03:08,480
I mean, it was a fruitful area 

for the next couple decades. 

1238
01:03:08,920 --> 01:03:14,000
So I'd done work in theory of 
 
algorithms and, and with respect

1239
01:03:14,000 --> 01:03:17,880
to the academic question, if 
 
you're desperately wed to an 

1240
01:03:17,880 --> 01:03:21,920
academic career, you shouldn't 

do this because the academic 

1241
01:03:21,920 --> 01:03:24,480
hiring is fairly siloed and, and

 it's, it's not a good idea to 

1242
01:03:24,480 --> 01:03:27,000
switch areas after your 
 
dissertation and, and so on. 

1243
01:03:28,080 --> 01:03:32,160
If, if you, if there's a forcing

 function like I had, or you're

1244
01:03:32,160 --> 01:03:35,360
interested in sort of problems 

more broadly, then you can think

1245
01:03:35,360 --> 01:03:37,120
a little bit more broadly. 
 
So I've had a lot of students in

1246
01:03:37,120 --> 01:03:39,320
postdocs and I sort of give them

 this advice and you try and 

1247
01:03:39,320 --> 01:03:41,520
carve out various projects that 
 they're interested in. 

1248
01:03:41,520 --> 01:03:43,640
Some want an academic route, so 
 they should work on slightly 

1249
01:03:43,640 --> 01:03:45,400
more conservative things in the 
 general space. 

1250
01:03:45,400 --> 01:03:48,382
Some definitely want industry 
and 
 some, you know, could go 

1251
01:03:48,382 --> 01:03:50,414
with either and, and the ones 
that 
 could go with either, 

1252
01:03:50,414 --> 01:03:52,360
I've seen some go one way and 
some go the 
 other. 

1253
01:03:53,240 --> 01:03:56,480
As a general rule, this is going

 to be a marketable space and 

1254
01:03:56,480 --> 01:03:58,560
stuff I do. 
 
And so maybe it's a little bit 

1255
01:03:58,560 --> 01:04:03,080
less of an issue, but but, but 

that's something that they 

1256
01:04:03,080 --> 01:04:05,640
figure out. 
 
So I had done work in in 

1257
01:04:07,360 --> 01:04:09,440
theoretical computer science and

 it was clear that these 

1258
01:04:09,440 --> 01:04:13,120
algorithms would be useful more 
 broadly because the randomness 

1259
01:04:13,120 --> 01:04:17,000
entered in a very different way 
 than classical algorithms. 

1260
01:04:17,000 --> 01:04:20,520
And so then the question is, how

 does it evolve? 

1261
01:04:20,960 --> 01:04:23,560
And so it evolved. 
 
I was at Yale in the mathematics

1262
01:04:23,560 --> 01:04:25,960
department as in one of the 
 
junior faculty positions. 

1263
01:04:25,960 --> 01:04:29,520
I spent time at Yahoo and 
 
Stanford and moved to Berkeley. 

1264
01:04:29,520 --> 01:04:33,920
I'm in the statistics department

 there and at the International

1265
01:04:33,920 --> 01:04:36,360
Computer Science Institute and 

Lawrence Berkeley National Lab. 

1266
01:04:36,360 --> 01:04:39,920
And as you mentioned, a couple 

years ago started with as an 

1267
01:04:39,920 --> 01:04:42,560
Amazon Scholar working in the 
 
supply chain optimization 

1268
01:04:42,560 --> 01:04:46,040
technology group, where, you 
 
know, we work on better 

1269
01:04:46,040 --> 01:04:49,080
forecasting, demand forecasting 
 algorithms and supply chain 

1270
01:04:49,080 --> 01:04:51,640
decisions. 
 
So I think this is, this is 

1271
01:04:51,640 --> 01:04:53,320
something that's good for people

 to think about. 

1272
01:04:53,880 --> 01:04:56,840
And you know, there's at least, 
 I guess 4 hats there. 

1273
01:04:56,840 --> 01:04:59,480
And each hat has pros and cons 

and pluses and minuses. 

1274
01:04:59,480 --> 01:05:03,240
And so figuring out how to 
 
navigate that space is something

1275
01:05:03,240 --> 01:05:05,680
I try and encourage students and

 postdocs to think about sooner

1276
01:05:05,680 --> 01:05:08,280
rather than later. 
 
How? 

1277
01:05:09,760 --> 01:05:13,400
How much should people expect to

 have to move? 

1278
01:05:14,080 --> 01:05:19,270
You know, I always think this is

 one of the sad in a way, 

1279
01:05:19,270 --> 01:05:21,584
things about academia, at least 
from 
 what I've observed, that 

1280
01:05:21,584 --> 01:05:26,738
the fact that you do your PhD 
and 
 let's say you do a 

1281
01:05:26,738 --> 01:05:30,075
postdoc, how realistic is it 
that that person 
 in the US or 

1282
01:05:30,075 --> 01:05:32,640
the UK should assume that they 
can stay at 
 that same 

1283
01:05:32,640 --> 01:05:34,384
university and just move up a 
chain? 
 

1284
01:05:34,392 --> 01:05:37,304
Or how much should they expect 
that they're basically going to 

1285
01:05:37,304 --> 01:05:42,660
 have to move somewhere else in 
the country to find the 
 

1286
01:05:42,668 --> 01:05:44,945
position? 
Yeah. 
 

1287
01:05:44,953 --> 01:05:51,056
I mean, the short answer is that
when you do a postdoc, it's 
 

1288
01:05:51,064 --> 01:05:54,560
usually more targeted because 
you're working with a particular

1289
01:05:54,560 --> 01:05:59,318

 person, a particular group and
and so you could be at the same 

1290
01:05:59,318 --> 01:06:00,935
 place or not. 
It's usually a good idea to go 


1291
01:06:00,943 --> 01:06:03,556
elsewhere to get experience with
other people, the real filter's 

1292
01:06:03,556 --> 01:06:05,172
 at the assistant professor 
level. 
 

1293
01:06:05,180 --> 01:06:07,610
And you got to move there 
because, you know, except in 

1294
01:06:07,610 --> 01:06:10,308
rare 
 cases, you don't get a 
job at the same place. 
 

1295
01:06:10,316 --> 01:06:13,346
And that's not a statement about
whether the place where you are 

1296
01:06:13,346 --> 01:06:16,564
 at is good or bad or whether 
you're good or bad is just just 

1297
01:06:16,564 --> 01:06:19,377
 run the numbers, right. 
If there's 10, 20, 30 or 40 

1298
01:06:19,377 --> 01:06:22,495
places 
 hiring and that you're 
looking at and and the chances 

1299
01:06:22,495 --> 01:06:25,880
are one 
 and whatever that that
you get an offer, it's just 

1300
01:06:25,880 --> 01:06:28,244
unlikely to 
 happen at that 
same place even if you 

1301
01:06:28,244 --> 01:06:29,880
ultimately end up there. 
 
Sometimes people go and leave 

1302
01:06:29,880 --> 01:06:33,084
and come back. 
 
So moving, for better or worse, 

1303
01:06:33,084 --> 01:06:37,560
is part of the part of the 
 
equation, yeah. 

1304
01:06:38,760 --> 01:06:44,000
But did you, some of the 
 
questions I get is the should 

1305
01:06:44,000 --> 01:06:47,720
they go into industry or, or 
 
just let's put this two ways, 

1306
01:06:47,760 --> 01:06:52,560
does academia value industry? 
 
So if you've been an assistant 

1307
01:06:52,560 --> 01:06:55,240
professor or you've been a 
 
postdoc and you then say, you 

1308
01:06:55,240 --> 01:06:59,320
know what, I'm going to go into 
 industry, do they value then? 

1309
01:06:59,440 --> 01:07:05,152
And is the ability to come back 
 or have you been out of the 

1310
01:07:05,152 --> 01:07:08,240
game too long? 
 
You haven't sort of, you know, 

1311
01:07:08,240 --> 01:07:11,320
got the publications or the 
 
teaching experience, You know, 

1312
01:07:11,520 --> 01:07:14,440
is that a good piece of advice 

or would you say no, no, no, If 

1313
01:07:14,440 --> 01:07:16,160
you really want to be in 
 
academia, you basically need to 

1314
01:07:16,160 --> 01:07:23,200
stay in in academia. 
 
Yeah, I mean it, it's, it's more

1315
01:07:23,200 --> 01:07:25,480
textured than the following. 
 
But I, I think the short answer 

1316
01:07:25,480 --> 01:07:27,400
is that that that you need to 
 
stay. 

1317
01:07:28,840 --> 01:07:32,280
The slightly longer answer is it

 depends on the area, like 

1318
01:07:32,280 --> 01:07:35,040
computer science versus 
 
statistics versus engineering 

1319
01:07:35,320 --> 01:07:39,800
are rather different. 
 
In some cases having a startup, 

1320
01:07:41,320 --> 01:07:43,920
in some cases having a postdoc 

in industry is good and you can 

1321
01:07:43,920 --> 01:07:46,684
go back to universities and that

 tends to correlate with 

1322
01:07:46,684 --> 01:07:49,800
computer science. 
 
I don't know as much in 

1323
01:07:49,800 --> 01:07:52,880
engineering that may be changing

 over time, but less so. 

1324
01:07:52,880 --> 01:07:54,520
And maybe statistics and applied

 math. 

1325
01:07:56,280 --> 01:08:02,560
I think in a sense, on the one 

hand, people value the 

1326
01:08:02,560 --> 01:08:04,240
experience you might have, but 

not really. 

1327
01:08:04,240 --> 01:08:08,120
And by that I mean, you know, 
 
if, if you, if you check all, if

1328
01:08:08,120 --> 01:08:11,000
if you check all my boxes. 
 
And in addition, you have this 

1329
01:08:11,000 --> 01:08:13,160
other stuff, good, but you got 

to check all my boxes and the 

1330
01:08:13,160 --> 01:08:14,440
boxes are the things you alluded

 to. 

1331
01:08:15,680 --> 01:08:19,750
And so there's sometimes there's

 postdocs in industry that are 

1332
01:08:19,750 --> 01:08:21,640
basically academic post 
 docs, 
you write papers. 

1333
01:08:21,720 --> 01:08:24,279
So I'm not counting that because

 that's, that's effectively the

1334
01:08:24,279 --> 01:08:25,520
same sort of thing. 
 
But if you're out more than a 

1335
01:08:25,520 --> 01:08:29,600
couple years and you have fewer 
 papers, it just gets harder to 

1336
01:08:29,600 --> 01:08:32,840
publish. 
 
There are certainly cases where 

1337
01:08:32,840 --> 01:08:35,279
people do that, especially if 
 
they have some high profile ones

1338
01:08:35,960 --> 01:08:38,040
and, and can work their way back

 one way or the other. 

1339
01:08:38,040 --> 01:08:41,080
But it's, you know, you should 

know going in that you're, it's 

1340
01:08:41,080 --> 01:08:41,943
like a salmon swimming upstream.

 

1341
01:08:41,951 --> 01:08:44,145
And this is going to be a hard 
one. 
 

1342
01:08:44,153 --> 01:08:51,300
So how does he, this is maybe 
people in the US are maybe a 
 

1343
01:08:51,308 --> 01:08:53,285
little bit more familiar, but 
particularly people who are not.

1344
01:08:53,285 --> 01:08:54,800

 
So how does it work with the 

1345
01:08:54,800 --> 01:08:57,000
national labs? 
 
So you have a position at the 

1346
01:08:57,000 --> 01:09:00,840
national labs. 
 
How, how does that work between 

1347
01:09:00,840 --> 01:09:05,680
is this a, a quite common thing 
 in a way for these things to 

1348
01:09:05,680 --> 01:09:07,200
happen? 
 
I I can only imagine how you 

1349
01:09:07,200 --> 01:09:12,240
juggle the e-mail inboxes and 
 
the sort of meetings, but I 

1350
01:09:12,240 --> 01:09:14,319
guess it's valuable. 
 
And interestingly, after an hour

1351
01:09:14,319 --> 01:09:15,640
today and I haven't been hit by 
 e-mail. 

1352
01:09:15,640 --> 01:09:18,880
That's why I've turned it off. 

Except, e-mail is a little 

1353
01:09:18,880 --> 01:09:21,960
overwhelming. 
 
I mean, it's not so common. 

1354
01:09:22,240 --> 01:09:26,439
So, so I, I with the Berkeley 
 
hat, because I'm in the 

1355
01:09:26,439 --> 01:09:30,040
Statistics Department there, I'm

 at Lawrence Berkeley lab. 

1356
01:09:30,040 --> 01:09:31,278
It's literally just up the hill.

 

1357
01:09:31,286 --> 01:09:33,260
And so I can, you can walk 
there. 
 

1358
01:09:33,268 --> 01:09:37,724
And so they were interested in 
scientific machine learning and,

1359
01:09:37,724 --> 01:09:40,640
and I had 
 done work and was 
well known in machine learning 

1360
01:09:40,640 --> 01:09:42,890
and, and 
 algorithms and 
statistics, large scale 

1361
01:09:42,890 --> 01:09:45,200
statistics. 
 
And I had various projects over 

1362
01:09:45,200 --> 01:09:49,080
the years with people there and 
 elsewhere on scientific ML 

1363
01:09:49,080 --> 01:09:53,120
problems. 
 
And so started about two years 

1364
01:09:53,120 --> 01:09:56,240
ago. 
 
This is what we renamed, but the

1365
01:09:56,240 --> 01:10:01,840
scientific machine learning 
 
group, basically MLA: Machine 

1366
01:10:01,840 --> 01:10:05,836
learning and analytics. 
 
And so it was partly facilitated

1367
01:10:05,836 --> 01:10:11,468
by the fact that there there is 
a moderate amount of interaction

1368
01:10:11,468 --> 01:10:16,080

 between UC Berkeley campus and
LBNL Lawrence Berkeley National 

1369
01:10:16,080 --> 01:10:19,362
 Lab, Oak Ridge, Argonne. 
Some of the other labs have 
 

1370
01:10:19,370 --> 01:10:21,925
things like that, but it's not 
so, so, so common. 
 

1371
01:10:21,933 --> 01:10:25,488
Most people are there, meaning 
just there and 100% of their 

1372
01:10:25,488 --> 01:10:28,506
FTEs there and and there for 
long term but so it's not. 
 

1373
01:10:28,514 --> 01:10:33,238
I wouldn't say it's very common.
OK, which makes your experience 

1374
01:10:33,238 --> 01:10:36,908
 even more unique then to have 
done it. 
 

1375
01:10:36,916 --> 01:10:43,332
Yeah, what maybe is a sort of 
finishing off comment or, or 
 

1376
01:10:43,340 --> 01:10:51,524
topic is if you have somebody 
now who was wanting to get into 

1377
01:10:51,524 --> 01:10:55,509
 scientific machine learning, 
because this is I guess kind of 

1378
01:10:55,509 --> 01:10:59,467
 what we've been largely talking
about, what would you get them 


1379
01:10:59,475 --> 01:11:02,734
to focus on? 
They're doing a PhD or they're 


1380
01:11:02,742 --> 01:11:04,737
doing a postdoc. 
But is there any, doesn't have 


1381
01:11:04,745 --> 01:11:07,361
to be a very, you know, 
specific, specific, but what 

1382
01:11:07,361 --> 01:11:11,354
sort of area 
 we'd say, you 
know, this is something you 

1383
01:11:11,354 --> 01:11:13,520
should look at. 
 
This is an area that is ripe 

1384
01:11:13,920 --> 01:11:20,480
for, you know, progress. 
 
Yeah, that's a good question. 

1385
01:11:20,800 --> 01:11:30,520
There's certain things, I mean, 
 that are less ripe for progress

1386
01:11:30,520 --> 01:11:33,280
in scientific ML, even if 
 
they're important scientific 

1387
01:11:33,280 --> 01:11:36,600
problems. 
 
So what I, what I, what I would 

1388
01:11:36,600 --> 01:11:39,680
try and say is work. 
 
You should work on a problem 

1389
01:11:39,680 --> 01:11:41,160
that's of interest to both 
 
sides. 

1390
01:11:41,160 --> 01:11:43,480
Otherwise you're coming at it 
 
from one side or another. 

1391
01:11:43,480 --> 01:11:46,817
And this is true whether you're 
 coming from CS or stats 

1392
01:11:46,817 --> 01:11:49,150
learning scientific problems or 
coming 
 from one scientific 

1393
01:11:49,150 --> 01:11:51,780
area. 
You should work on trying to 
 

1394
01:11:51,788 --> 01:11:53,810
figure out how to frame what 
you're doing as something of 
 

1395
01:11:53,818 --> 01:11:56,704
interest to both sides. 
So I have projects that, you 
 

1396
01:11:56,712 --> 01:12:00,122
know, we, we, we construct 
projects, you know, a student or

1397
01:12:00,122 --> 01:12:04,469

 postdoc owns a piece of it. 
And, and a goal typically is we 

1398
01:12:04,469 --> 01:12:07,384
 want a publication in the 
particular scientific area, but 

1399
01:12:07,384 --> 01:12:10,506
 also in an ML venue. 
And it needn't be the same one. 

1400
01:12:10,506 --> 01:12:12,000
 
Sometimes it is, it's a long 

1401
01:12:12,000 --> 01:12:13,720
version, a short version. 
 
Sometimes it's just two 

1402
01:12:13,720 --> 01:12:16,960
different things, but work on a 
 method that's broad enough that

1403
01:12:16,960 --> 01:12:19,200
machine learning people are 
 
interested in it and that by the

1404
01:12:19,200 --> 01:12:20,960
definition of machine learning, 
 people are interested. 

1405
01:12:20,960 --> 01:12:22,640
You convince 3 reviewers to say 
 yes. 

1406
01:12:23,320 --> 01:12:25,800
And so it's an imperfect 
 
process, you know the review 

1407
01:12:25,800 --> 01:12:28,257
process, etcetera, but boom, you

 get it in a top machine 

1408
01:12:28,257 --> 01:12:30,840
learning venue. 
 
And similarly, on the scientific

1409
01:12:30,840 --> 01:12:33,630
side, you know, a statement that
it's 
 of value to the scientist

1410
01:12:33,630 --> 01:12:36,160
is you get three reviewers 
 to 
say accept and then it's 

1411
01:12:36,440 --> 01:12:38,370
accepted on the scientific side.

 

1412
01:12:38,378 --> 01:12:41,772
And oftentimes the way we try 
and scope our projects is to 
 

1413
01:12:41,780 --> 01:12:44,540
say, you know, we don't on the 
machine learning side, we don't 

1414
01:12:44,540 --> 01:12:47,564
 want to work on your problem if
it's only of interest to you. 
 

1415
01:12:47,572 --> 01:12:50,140
If I can't apply it to some 
other area, someone else, some 


1416
01:12:50,148 --> 01:12:52,622
other domain, Because if that's 
the case, I probably need to 
 

1417
01:12:52,630 --> 01:12:55,172
know so much about your area 
that I become, you know, an 
 

1418
01:12:55,180 --> 01:12:57,749
expert in your particular area. 
And similarly, on the scientific

1419
01:12:57,749 --> 01:13:01,120

 side, you know, if you have a 
particular method that is only 


1420
01:13:01,128 --> 01:13:03,902
that's so heavily tailored to 
your domain that it's not going 

1421
01:13:03,902 --> 01:13:06,220
 to be useful more broadly to 
machine learning people, that's 

1422
01:13:06,220 --> 01:13:08,605
 a much harder sell. 
Then you're not doing something 

1423
01:13:08,605 --> 01:13:11,442
 that satisfies both sides. 
So something that I knew about 


1424
01:13:11,450 --> 01:13:14,715
from working on years ago was 
like sequence alignment in in 
 

1425
01:13:14,723 --> 01:13:17,460
genetics, right? 
Clearly an important problem, 
 

1426
01:13:17,468 --> 01:13:19,800
not something that's portable 
the particular algorithm, not 
 

1427
01:13:19,808 --> 01:13:22,348
something that's portable to 
lots of other scientific areas, 

1428
01:13:22,348 --> 01:13:25,480
 but better machine learning 
methods for solving 
 

1429
01:13:25,488 --> 01:13:28,216
spatiotemporal forecasting 
problems with non-trivial 
 

1430
01:13:28,224 --> 01:13:31,528
boundary constraints. 
That's clearly of interest to a 

1431
01:13:31,528 --> 01:13:33,229
 pretty wide range of 
applications. 
 

1432
01:13:33,237 --> 01:13:37,196
And also to solve it, you're 
going to have to introduce 
 

1433
01:13:37,204 --> 01:13:39,540
technical solutions that are 
probably, you know, go, can you 

1434
01:13:39,540 --> 01:13:42,240
 come up with a differentiable 
optimizer for an end-to-end 
 

1435
01:13:42,248 --> 01:13:44,510
differentiable system where you 
have hard constraints? 
 

1436
01:13:44,518 --> 01:13:47,446
I mean, that's clearly something
of interest to machine learning 

1437
01:13:47,446 --> 01:13:50,584
 and optimization. 
So I guess I'd try and focus on 

1438
01:13:50,584 --> 01:13:52,730
 something there. 
And that is probably the way 
 

1439
01:13:52,738 --> 01:13:55,350
that you're going to be. 
And unless you know exactly what

1440
01:13:55,350 --> 01:13:58,700

 you want to be doing 30 years 
from now, that's probably the 
 

1441
01:13:58,708 --> 01:14:01,860
way to make yourself most 
forward-compatible with with 
 

1442
01:14:01,868 --> 01:14:04,270
however. 
That's a really, that's a really

1443
01:14:04,270 --> 01:14:06,628

 good point. 
I guess what you're saying is, 


1444
01:14:06,636 --> 01:14:08,790
yeah, it's true. 
Like later on in your career 
 

1445
01:14:08,798 --> 01:14:11,226
you've become established in a 
certain area, but most people 
 

1446
01:14:11,234 --> 01:14:14,730
during their PhD or the postdoc 
at that time have no real 
 

1447
01:14:14,738 --> 01:14:18,158
idea, necessarily. 
So you're saying if you do 
 

1448
01:14:18,166 --> 01:14:22,685
something broad enough that can 
get you exposed to different 
 

1449
01:14:22,693 --> 01:14:28,320
groups and and have a foundation
to your earlier comment, you can

1450
01:14:28,320 --> 01:14:31,072

 more easily go into verticals 
because you've built that 
 

1451
01:14:31,080 --> 01:14:32,998
foundation. 
Whereas I guess if you jump 
 

1452
01:14:33,006 --> 01:14:36,782
directly into a vertical right 
from the beginning for you to 
 

1453
01:14:36,790 --> 01:14:40,250
sort of reverse out that it's 
kind of harder, isn't it? 
 

1454
01:14:40,258 --> 01:14:42,246
So I I think that's what you're 
alluding to. 

1455
01:14:42,246 --> 01:14:45,498
It sets you up 
 better. 
Yeah, yeah, because it's not at 

1456
01:14:45,498 --> 01:14:46,872
 all. 
I mean, all the questions you're

1457
01:14:46,872 --> 01:14:48,915

 asking are fair ones. 
And it's not at all clear how 
 

1458
01:14:48,923 --> 01:14:53,616
the whole area will evolve. 
And so depending on how it 
 

1459
01:14:53,624 --> 01:14:55,836
evolves, having experience in 
one versus another. 
 

1460
01:14:55,844 --> 01:14:59,375
And, and I think I mean, having 
worked in a bunch of areas, you 

1461
01:14:59,375 --> 01:15:01,552
 know, it's, it's, but, but 
speaking substantially only 
 

1462
01:15:01,560 --> 01:15:04,928
English, but a but a little bit 
of other things I can imagine. 


1463
01:15:04,936 --> 01:15:08,000
And I see, you know, you learn a
second language, it's hard. 
 

1464
01:15:08,008 --> 01:15:11,088
You learn a third language, the,
the relative cost is a lot less.

1465
01:15:11,088 --> 01:15:11,960

 
You learn a fourth. 

1466
01:15:11,960 --> 01:15:15,600
I mean, so there's a diminishing

 returns in terms of the amount

1467
01:15:15,600 --> 01:15:17,382
of extra effort you need to do. 
 

1468
01:15:17,390 --> 01:15:19,895
So you work in one area, switch 
to the second. 
 

1469
01:15:19,903 --> 01:15:21,685
It's very hard switch to the 
third. 
 

1470
01:15:21,693 --> 01:15:25,580
You know, you can, you know, you
can, you can do that. 
 

1471
01:15:25,588 --> 01:15:29,332
And after that, one thing I've 
sort of been always struck by is

1472
01:15:29,332 --> 01:15:31,880

 it's amazing how little you, 
you need to know about a 

1473
01:15:31,880 --> 01:15:33,770
particular 
 area, especially if
you're working with good people 

1474
01:15:33,770 --> 01:15:34,880
who'll 
 complement what you're 
doing. 

1475
01:15:34,880 --> 01:15:36,080
You can learn from them, right? 
 

1476
01:15:36,088 --> 01:15:39,048
And so it's amazing what you 
don't need to know in order to, 

1477
01:15:39,048 --> 01:15:41,167
 to get interesting results in 
an area. 
 

1478
01:15:41,175 --> 01:15:44,538
And so, and so learning new 
things. 
 

1479
01:15:44,546 --> 01:15:46,890
By the time you know two or 
three things, it's easy to learn

1480
01:15:46,890 --> 01:15:49,250

 the fourth. 
Yeah, yeah, I know that that 
 

1481
01:15:49,258 --> 01:15:52,215
makes sense. 
Well, thank you so much. 
 

1482
01:15:52,223 --> 01:15:56,552
I it's the day after Labour Day,
which means that probably 
 

1483
01:15:56,560 --> 01:16:00,416
there's a whole bunch of emails 
and stuff that's started to come

1484
01:16:00,416 --> 01:16:02,227

 through on essentially your 
first day back. 
 

1485
01:16:02,235 --> 01:16:05,534
And I know we probably could 
have carried on for for another 

1486
01:16:05,534 --> 01:16:07,992
 couple of hours. 
But yeah, thank you. 
 

1487
01:16:08,000 --> 01:16:12,360
Really, really appreciate it and
look forward to catching up in 


1488
01:16:12,368 --> 01:16:14,894
person at some conference in the
future. 
 

1489
01:16:14,902 --> 01:16:17,418
Sounds great, thanks for having 
me, this has been fun. 
 

1490
01:16:17,426 --> 01:16:41,320
All right. 
Cheers, 
 Neil.

