1
00:00:00,040 --> 00:00:02,800
Welcome to the Deep Dive. 
We're the show that tries to cut

2
00:00:02,800 --> 00:00:06,640
through the hype and get to the 
core of complex research, making

3
00:00:06,640 --> 00:00:10,080
it clear and hopefully usable. 
Yeah, making it actionable. 

4
00:00:10,080 --> 00:00:12,720
Exactly. 
And today we're diving into 

5
00:00:12,720 --> 00:00:14,560
something really fascinating in 
AI. 

6
00:00:15,040 --> 00:00:18,520
Think about this for a second. 
So much of the world isn't neat 

7
00:00:18,520 --> 00:00:20,640
rows and columns. 
Oh, definitely not. 

8
00:00:20,960 --> 00:00:25,000
Think social networks, supply 
chains, even biological systems.

9
00:00:25,000 --> 00:00:27,320
Right. 
Or how chemicals bond together. 

10
00:00:27,320 --> 00:00:29,200
Or traffic flowing through a 
city. 

11
00:00:29,200 --> 00:00:32,400
It's all interconnected. 
It forms these intricate webs, 

12
00:00:32,400 --> 00:00:35,240
these graph structures. 
Precisely. 

13
00:00:35,240 --> 00:00:38,440
And that kind of data, this 
graph structure data, it's 

14
00:00:38,440 --> 00:00:41,080
historically been tricky for 
standard machine learning. 

15
00:00:41,240 --> 00:00:44,960
Yeah, models that love grids, 
like for images, or sequences 

16
00:00:44,960 --> 00:00:47,440
like for text. 
They kind of stumble when things

17
00:00:47,440 --> 00:00:49,680
get messy and interconnected 
like this, don't they? 

18
00:00:49,720 --> 00:00:52,160
You really do. 
It's like trying to understand a

19
00:00:52,160 --> 00:00:55,840
city by only looking at isolated
building blueprints. 

20
00:00:56,080 --> 00:00:58,200
You miss the roads, the 
utilities, the people moving 

21
00:00:58,200 --> 00:00:59,720
between them. 
You miss the connections. 

22
00:00:59,960 --> 00:01:02,920
That's a great way to put it. 
And that brings us right to our 

23
00:01:02,920 --> 00:01:06,480
topic graph neural networks, or 
GNS. 

24
00:01:06,480 --> 00:01:08,720
These are really a different 
breed of AI, aren't they? 

25
00:01:08,800 --> 00:01:10,880
Absolutely. 
They're a branch of AI and 

26
00:01:10,880 --> 00:01:14,960
machine learning specifically 
designed to understand and, 

27
00:01:15,080 --> 00:01:18,440
crucially, make predictions on 
this kind of interconnected 

28
00:01:18,440 --> 00:01:20,320
data. 
They finally let us make sense 

29
00:01:20,320 --> 00:01:23,440
of these complex networks. 
So our mission today for this 

30
00:01:23,440 --> 00:01:25,760
deep dive is to really unpack 
GNNS. 

31
00:01:25,760 --> 00:01:28,280
We want to get into not just 
what they are, but why they're 

32
00:01:28,280 --> 00:01:30,400
so powerful. 
Yeah, and how they actually work

33
00:01:30,400 --> 00:01:32,240
under the hood. 
We'll look at some real world 

34
00:01:32,240 --> 00:01:34,440
examples too, things you might 
already be using without 

35
00:01:34,440 --> 00:01:36,800
realizing it. 
And the goal here is clarity. 

36
00:01:37,080 --> 00:01:40,360
We want you to walk away with a 
solid conceptual grasp, maybe 

37
00:01:40,360 --> 00:01:43,040
even enough detail that if 
you're an engineer, you can 

38
00:01:43,040 --> 00:01:45,240
start thinking about how to 
apply these ideas. 

39
00:01:45,240 --> 00:01:46,600
Sounds good, where should we 
start? 

40
00:01:46,800 --> 00:01:48,280
Let's start right at the 
beginning. 

41
00:01:48,400 --> 00:01:51,400
We say graph neural network. 
What does the neural network 

42
00:01:51,400 --> 00:01:54,360
part actually do with the graph?
How do they come together? 

43
00:01:54,600 --> 00:01:58,520
OK, so AGNN at its heart is a 
neural network, but it's built 

44
00:01:58,520 --> 00:02:00,880
specifically for data structured
as a graph. 

45
00:02:01,200 --> 00:02:02,800
Remember, graph is just 
entities. 

46
00:02:02,800 --> 00:02:05,760
We call them nodes connected by 
relationships. 

47
00:02:05,800 --> 00:02:08,800
We call those edges. 
Like people and friendships on 

48
00:02:08,800 --> 00:02:10,720
Facebook or airports and flight 
paths. 

49
00:02:10,720 --> 00:02:12,560
Exactly. 
People are nodes. 

50
00:02:12,600 --> 00:02:15,280
Friendships are edges. 
Airports are nodes, flight 

51
00:02:15,280 --> 00:02:18,360
routes are edges. 
The neural network part comes in

52
00:02:18,360 --> 00:02:22,720
because GNNS learn 
representations of these nodes 

53
00:02:22,720 --> 00:02:24,840
and edges. 
They learn how to process these 

54
00:02:24,840 --> 00:02:27,320
complex, often non Euclidean 
structures. 

55
00:02:27,440 --> 00:02:30,360
Non Euclidean, meaning they 
don't fit neatly onto a flat 

56
00:02:30,360 --> 00:02:33,440
plane or a straight line like a 
grid or sequence does. 

57
00:02:33,520 --> 00:02:36,800
Precisely, they handle the 
irregular complex shapes of real

58
00:02:36,800 --> 00:02:39,000
world networks. 
OK, so the graph gives us the 

59
00:02:39,000 --> 00:02:41,720
structure. 
How does the GNN actually learn 

60
00:02:41,720 --> 00:02:43,640
from all those connections? 
What's the mechanism? 

61
00:02:44,160 --> 00:02:46,840
This is the really cool part. 
The core idea is usually called 

62
00:02:47,000 --> 00:02:50,440
message passing. 
Imagine each node in the network

63
00:02:50,440 --> 00:02:52,720
as having some initial 
information, some features. 

64
00:02:52,800 --> 00:02:55,560
Like a user's profile info or an
Adam's type. 

65
00:02:55,800 --> 00:03:00,600
Right then, in each step of the 
GNN, nodes essentially send 

66
00:03:00,600 --> 00:03:03,520
messages derived from their 
current features across the 

67
00:03:03,520 --> 00:03:06,320
edges to their direct neighbors.
OK, so they're talking to their 

68
00:03:06,320 --> 00:03:07,160
neighbors. 
Exactly. 

69
00:03:07,400 --> 00:03:10,800
And then each node receives all 
the messages from its neighbors 

70
00:03:10,800 --> 00:03:14,280
and it aggregates them, combines
them in some way, often with its

71
00:03:14,280 --> 00:03:17,960
own existing information. 
Aggregates like summing them up 

72
00:03:17,960 --> 00:03:19,680
or averaging? 
Yeah, it could be summing, 

73
00:03:19,680 --> 00:03:21,240
averaging, maybe something more 
complex. 

74
00:03:21,240 --> 00:03:22,360
We'll get into different ways 
later. 

75
00:03:22,600 --> 00:03:25,040
But the key is this iterative 
process. 

76
00:03:25,560 --> 00:03:28,880
Send messages, receive messages,
update your own state based on 

77
00:03:28,880 --> 00:03:30,960
those messages. 
And this happens over multiple 

78
00:03:30,960 --> 00:03:33,560
steps or layers. 
Typically, yes. 

79
00:03:33,840 --> 00:03:36,640
Multiple layers allow 
information to propagate further

80
00:03:36,640 --> 00:03:39,000
out across the graph. 
A node doesn't just learn from 

81
00:03:39,000 --> 00:03:41,520
its immediate neighbors, but 
eventually from its neighbors, 

82
00:03:41,520 --> 00:03:44,200
neighbors, and so on. 
So it builds up a richer picture

83
00:03:44,200 --> 00:03:46,280
of its place in the wider 
network. 

84
00:03:46,280 --> 00:03:48,680
Exactly. 
It learns A representation that 

85
00:03:48,680 --> 00:03:52,040
captures not just its own 
features, but also the structure

86
00:03:52,040 --> 00:03:55,560
of its local neighborhood and 
potentially even global 

87
00:03:55,560 --> 00:03:58,440
properties of the entire graph. 
It's context aware. 

88
00:03:58,640 --> 00:04:01,960
That makes intuitive sense. 
Why is this message passing 

89
00:04:01,960 --> 00:04:04,400
approach so important? 
Why not just use the node 

90
00:04:04,400 --> 00:04:07,200
features alone? 
Because in so many real world 

91
00:04:07,200 --> 00:04:10,480
problems, the connections are 
the crucial information. 

92
00:04:11,320 --> 00:04:15,120
Traditional models that ignore 
the graph topology, the specific

93
00:04:15,120 --> 00:04:18,839
pattern of connections, miss out
on a huge source of signal. 

94
00:04:19,000 --> 00:04:21,399
Like trying to predict if 
someone will click an ad without

95
00:04:21,399 --> 00:04:24,440
knowing what their friends like.
Perfect example, GN NS 

96
00:04:24,440 --> 00:04:26,680
explicitly leverage that 
topology. 

97
00:04:27,040 --> 00:04:30,240
They use the connections to 
inform the learning process and 

98
00:04:30,240 --> 00:04:32,800
this makes them incredibly 
effective for specific tasks. 

99
00:04:32,800 --> 00:04:35,080
Such as? 
Well, things like node 

100
00:04:35,080 --> 00:04:37,160
classification. 
You predict a property for each 

101
00:04:37,160 --> 00:04:39,320
node. 
Is this bank transaction 

102
00:04:39,320 --> 00:04:41,720
fraudulent? 
Is this protein involved in a 

103
00:04:41,720 --> 00:04:44,360
specific disease? 
OK, classifying individual items

104
00:04:44,360 --> 00:04:45,960
within the network. 
Right. 

105
00:04:46,200 --> 00:04:48,440
Or link prediction. 
Can we predict if a new 

106
00:04:48,440 --> 00:04:51,200
connection is likely to form? 
Will these two users become 

107
00:04:51,200 --> 00:04:53,160
friends? 
Will this drug interact with 

108
00:04:53,160 --> 00:04:55,600
this target protein? 
Predicting future relationships 

109
00:04:55,600 --> 00:04:57,160
makes sense for recommendations 
too. 

110
00:04:57,440 --> 00:04:59,760
Definitely, and also graph 
classification. 

111
00:05:00,080 --> 00:05:02,200
Here you're classifying the 
entire graph. 

112
00:05:02,520 --> 00:05:04,640
Is this molecule likely to be 
toxic? 

113
00:05:04,920 --> 00:05:08,000
Does this social network exhibit
community structure? 

114
00:05:08,160 --> 00:05:10,920
Yeah, typical of, say, political
polarization. 

115
00:05:11,040 --> 00:05:15,360
Wow, OK, so node link and graph 
level predictions, that's quite 

116
00:05:15,360 --> 00:05:17,400
versatile. 
Can we ground this in some real 

117
00:05:17,400 --> 00:05:19,720
world examples you mentioned? 
We might be interacting with 

118
00:05:19,720 --> 00:05:20,920
them already. 
Absolutely. 

119
00:05:21,040 --> 00:05:22,640
Let's talk about drug discovery 
first. 

120
00:05:22,640 --> 00:05:24,760
This is a huge one. 
Think about a molecule. 

121
00:05:24,760 --> 00:05:27,040
It's a natural graph, isn't it? 
How so? 

122
00:05:27,360 --> 00:05:30,080
Well, the atoms are the nodes, 
the chemical bonds between them 

123
00:05:30,080 --> 00:05:32,000
are the edges. 
I see. 

124
00:05:32,160 --> 00:05:33,720
And you could add features 
right? 

125
00:05:33,720 --> 00:05:36,760
Like the type of atom carbon, 
oxygen as node features. 

126
00:05:36,760 --> 00:05:39,880
Exactly, Atom type charge, 
things like that and even edge 

127
00:05:39,880 --> 00:05:41,960
features like is it a single 
bond or double bond? 

128
00:05:42,000 --> 00:05:43,920
Okay, so you have this molecular
graph. 

129
00:05:44,160 --> 00:05:48,560
What can a GNN do with it? 
It can learn to predict the 

130
00:05:48,560 --> 00:05:50,960
molecule's properties directly 
from its structure. 

131
00:05:51,280 --> 00:05:53,840
Will it dissolve in water? 
Will it bind to a certain 

132
00:05:53,840 --> 00:05:56,840
receptor? 
Will it kill bacteria? 

133
00:05:56,880 --> 00:05:58,840
And that leads to things like 
finding new drugs. 

134
00:05:59,000 --> 00:06:01,400
Precisely. 
The Halocen story is probably 

135
00:06:01,400 --> 00:06:04,760
the most famous example. 
Here, researchers use GN NS to 

136
00:06:04,760 --> 00:06:08,200
screen a massive library of 
existing chemical compounds. 

137
00:06:08,200 --> 00:06:10,480
OK. 
They train the GNN to predict 

138
00:06:10,480 --> 00:06:13,640
whether a molecule would inhibit
the growth of bacteria like E 

139
00:06:13,640 --> 00:06:16,720
coli, and the model flagged a 
compound called Halocen. 

140
00:06:17,000 --> 00:06:19,240
Which wasn't known as an 
antibiotic before. 

141
00:06:19,400 --> 00:06:21,000
No. 
It had been investigated for 

142
00:06:21,000 --> 00:06:24,000
diabetes I think, but it's 
antibiotic potential was 

143
00:06:24,000 --> 00:06:26,040
completely overlooked by 
traditional methods. 

144
00:06:26,040 --> 00:06:29,000
The GNN saw something in its 
structure and it turned out to 

145
00:06:29,000 --> 00:06:31,560
be a potent broad spectrum 
antibiotic. 

146
00:06:31,720 --> 00:06:34,880
Wow, so the GNN literally helped
discover a new antibiotic 

147
00:06:34,880 --> 00:06:38,560
candidate just by looking at the
graph structure of molecules. 

148
00:06:38,560 --> 00:06:42,240
It's a prime example of how GNNS
can accelerate scientific 

149
00:06:42,240 --> 00:06:44,320
discovery. 
It's really transforming parts 

150
00:06:44,320 --> 00:06:46,080
of chem, informatics and drug 
development. 

151
00:06:46,200 --> 00:06:48,800
That's genuinely game changing. 
What's another area you 

152
00:06:48,800 --> 00:06:50,120
mentioned? 
Something we use daily. 

153
00:06:50,120 --> 00:06:52,440
Yeah, navigation. 
Think about Google Maps or 

154
00:06:52,440 --> 00:06:54,600
similar apps. 
How do they estimate your 

155
00:06:54,600 --> 00:06:56,160
arrival time? 
Your ETA. 

156
00:06:57,040 --> 00:06:59,040
I guess they look at traffic 
speeds, distance. 

157
00:06:59,240 --> 00:07:02,440
Right, but how do they combine 
all that information across a 

158
00:07:02,440 --> 00:07:05,520
huge complex Rd. network that's 
constantly changing? 

159
00:07:05,640 --> 00:07:08,120
OK, I see where you're going. 
The road network is a graph. 

160
00:07:08,480 --> 00:07:10,680
Exactly. 
Intersections are nodes, roads 

161
00:07:10,680 --> 00:07:13,360
are edges, and you have features
on those edges. 

162
00:07:13,360 --> 00:07:17,040
Rd. length, speed limit, current
traffic speeds, maybe historical

163
00:07:17,040 --> 00:07:19,560
traffic patterns. 
So the GNN processes this 

164
00:07:19,560 --> 00:07:21,800
constantly evolving graph of 
traffic. 

165
00:07:21,800 --> 00:07:24,720
Precisely. 
Google actually deployed GN NS 

166
00:07:24,880 --> 00:07:27,560
to significantly improve their 
ETA predictions. 

167
00:07:28,120 --> 00:07:31,320
They found that GN NS could 
model the complex dependencies 

168
00:07:31,560 --> 00:07:34,920
how congestion on one Rd. 
effects surrounding areas much 

169
00:07:34,920 --> 00:07:37,600
better than previous methods. 
And did it make a difference? 

170
00:07:37,840 --> 00:07:40,200
A huge difference. 
They reported reductions in 

171
00:07:40,200 --> 00:07:43,440
negative user outcomes, 
basically wildly inaccurate ET 

172
00:07:43,560 --> 00:07:47,440
as by over 40% in cities like 
Sydney and Taichung after 

173
00:07:47,440 --> 00:07:51,240
deploying GN NS. 40% fewer bad 
ET as I definitely noticed that.

174
00:07:51,240 --> 00:07:52,600
That's impressive. 
It really is. 

175
00:07:52,880 --> 00:07:56,200
Shows how GNNS can handle 
dynamic large scale graph data 

176
00:07:56,400 --> 00:07:58,760
effectively. 
And I bet recommendation systems

177
00:07:58,760 --> 00:08:01,440
are another big one, like 
suggesting products or content. 

178
00:08:01,680 --> 00:08:04,320
You got it. 
Recommendations are a natural 

179
00:08:04,320 --> 00:08:06,840
fit. 
You have users, you have items, 

180
00:08:07,040 --> 00:08:08,680
products, movies, pins, 
whatever. 

181
00:08:08,840 --> 00:08:10,440
And you have interactions 
between them. 

182
00:08:10,440 --> 00:08:13,200
Purchases, ratings, clicks. 
That's a graph. 

183
00:08:13,280 --> 00:08:16,720
A bipartite graph, often users 
connected to items. 

184
00:08:17,040 --> 00:08:20,800
Yes, or even more complex graphs
involving user user connections 

185
00:08:20,800 --> 00:08:23,360
too. 
GN NS can learn embeddings for 

186
00:08:23,360 --> 00:08:26,760
users and items based on this 
graph structure, leading to 

187
00:08:26,760 --> 00:08:29,280
better recommendations. 
Any specific examples? 

188
00:08:29,440 --> 00:08:31,880
Pinterest's Pin Sage models a 
well known one. 

189
00:08:32,159 --> 00:08:34,960
They built a GNN to operate on 
their massive graph. 

190
00:08:35,360 --> 00:08:39,520
Billions of nodes, billions of 
edges connecting users, pins and

191
00:08:39,520 --> 00:08:40,840
boards. 
Billions. 

192
00:08:40,840 --> 00:08:42,799
How do they even compute on a 
graph that big? 

193
00:08:43,039 --> 00:08:44,720
That's where clever engineering 
comes in. 

194
00:08:44,840 --> 00:08:47,840
Pin Sage uses techniques based 
on something called Graph Sage. 

195
00:08:48,520 --> 00:08:51,120
The core idea is instead of 
trying to aggregate messages 

196
00:08:51,120 --> 00:08:53,560
from all of a node's neighbors, 
which could be thousands or 

197
00:08:53,560 --> 00:08:56,960
millions on Pinterest, you 
intelligently sample a smaller 

198
00:08:56,960 --> 00:09:00,000
subset of neighbors. 
So you get an approximation of 

199
00:09:00,000 --> 00:09:01,920
the neighborhood information 
without the massive 

200
00:09:01,920 --> 00:09:04,080
computational cost. 
Exactly. 

201
00:09:04,280 --> 00:09:07,040
It allows GNNS to scale to web 
size graphs. 

202
00:09:07,680 --> 00:09:11,120
Penn Sage uses this to learn 
embeddings that capture visual 

203
00:09:11,120 --> 00:09:14,440
and semantic similarity, 
powering their recommendations. 

204
00:09:14,560 --> 00:09:17,800
OK, these examples, drug 
discovery maps, recommendations 

205
00:09:17,800 --> 00:09:19,400
that they really drive home the 
why. 

206
00:09:19,960 --> 00:09:21,640
Now let's dig deeper into the 
how. 

207
00:09:21,640 --> 00:09:24,240
What's actually going on inside 
a GNN layer? 

208
00:09:24,760 --> 00:09:27,640
You mentioned message passing, 
but how does that translate 

209
00:09:27,640 --> 00:09:29,440
mathematically? 
Right, let's get a bit more 

210
00:09:29,440 --> 00:09:31,920
technical. 
A useful analogy is to think 

211
00:09:31,920 --> 00:09:35,680
about convolutional neural 
networks, CNNS, which are so 

212
00:09:35,680 --> 00:09:39,440
dominant in computer vision. 
OK, CN NS use kernels or filters

213
00:09:39,440 --> 00:09:41,040
that slide across an image. 
Exactly. 

214
00:09:41,040 --> 00:09:43,960
An image is basically a grid, a 
very regular graph. 

215
00:09:44,360 --> 00:09:47,560
ACNN kernel is a small matrix of
learnable weights. 

216
00:09:47,960 --> 00:09:50,640
It slides across the image 
looking at a small patch of 

217
00:09:50,640 --> 00:09:54,320
pixels, a local neighborhood at 
a time, and computes a feature 

218
00:09:54,440 --> 00:09:57,360
based on that patch. 
It explains the fact that the 

219
00:09:57,360 --> 00:09:59,240
same pattern might appear 
anywhere. 

220
00:09:59,240 --> 00:10:02,000
So it captures local patterns 
and because the grid is regular,

221
00:10:02,000 --> 00:10:03,560
the same kernel works 
everywhere. 

222
00:10:03,720 --> 00:10:06,400
Precisely. 
The challenge for GN NS is to 

223
00:10:06,400 --> 00:10:09,840
generalize that idea. 
Create an operator that acts on 

224
00:10:09,840 --> 00:10:12,440
a node's local neighborhood. 
But for graphs that are 

225
00:10:12,440 --> 00:10:16,360
irregular, node neighborhoods 
can vary wildly in size and 

226
00:10:16,360 --> 00:10:18,360
structure. 
So we can't just slide a fixed 

227
00:10:18,360 --> 00:10:21,000
size kernel around. 
Nope, We need something more 

228
00:10:21,000 --> 00:10:23,440
flexible. 
So if I'm an engineer designing 

229
00:10:23,440 --> 00:10:26,560
one of these layers, what 
properties am I aiming for? 

230
00:10:26,560 --> 00:10:29,360
What makes a good GNN layer? 
Good question. 

231
00:10:29,440 --> 00:10:34,120
You are a few key things. 
One, efficiency computation and 

232
00:10:34,120 --> 00:10:37,560
memory usage should ideally 
scale reasonably, maybe linearly

233
00:10:37,560 --> 00:10:40,360
or near linearly with the number
of nodes and edges. 

234
00:10:40,360 --> 00:10:42,880
Makes sense, you don't want it 
to blow up on large graphs. 

235
00:10:42,880 --> 00:10:46,680
Two parameter sharing. 
Like ACNN kernel, you want a 

236
00:10:46,680 --> 00:10:49,200
fixed number of learnable 
parameters regardless of the 

237
00:10:49,200 --> 00:10:51,320
graph size. 
The same layer should work on a 

238
00:10:51,320 --> 00:10:53,160
small graph or a massive 1. 
OK. 

239
00:10:53,160 --> 00:10:56,440
So it learns general principles 
of neighborhood interaction. 3 

240
00:10:56,760 --> 00:10:59,160
Locality. 
The operation should primarily 

241
00:10:59,160 --> 00:11:01,640
focus on nodes. 
Immediate neighborhood, say it's

242
00:11:01,640 --> 00:11:03,560
direct connections, one hop 
neighbors. 

243
00:11:03,920 --> 00:11:06,280
Deeper layers can then capture 
wider context. 

244
00:11:06,440 --> 00:11:10,600
Similar to how CNN's build up 
complexity layer by layer. 4 

245
00:11:11,200 --> 00:11:14,600
Ability to handle varying 
neighborhood sizes and maybe 

246
00:11:14,600 --> 00:11:16,800
assigned different importance to
different neighbors. 

247
00:11:16,920 --> 00:11:18,400
Not all connections are equal, 
right? 

248
00:11:18,400 --> 00:11:20,240
Some friends influence you more 
than others. 

249
00:11:20,280 --> 00:11:24,040
And five, inductive capability. 
This is crucial. 

250
00:11:24,600 --> 00:11:27,080
The model should be able to 
generalize to graphs it hasn't 

251
00:11:27,080 --> 00:11:29,160
seen during training. 
It should learn structural 

252
00:11:29,160 --> 00:11:31,200
patterns, not just memorize the 
training graph. 

253
00:11:31,200 --> 00:11:33,960
That's a demanding list. 
How do we actually achieve this 

254
00:11:33,960 --> 00:11:36,040
mathematically? 
What's the core operation? 

255
00:11:36,240 --> 00:11:40,440
The fundamental idea boils down 
to combining 2 things, the nodes

256
00:11:40,440 --> 00:11:42,040
features and the graph 
structure. 

257
00:11:42,040 --> 00:11:44,560
The connections. 
Let's represent all the node 

258
00:11:44,560 --> 00:11:48,960
features as a matrix, say H Each
row is a node, each column is a 

259
00:11:48,960 --> 00:11:51,520
feature dimension, and let's 
represent the graph structure 

260
00:11:51,520 --> 00:11:56,840
using the adjacency matrix A. 
This is a matrix where aij is 1 

261
00:11:57,080 --> 00:12:00,000
if node I is connected to node J
and 0 otherwise. 

262
00:12:00,000 --> 00:12:03,480
Standard graph representation. 
The simplest GNN like operation 

263
00:12:03,480 --> 00:12:06,200
might look something like this. 
You multiply the adjacency 

264
00:12:06,200 --> 00:12:08,520
matrix A by the feature matrix 
H. 

265
00:12:08,520 --> 00:12:12,160
What does AH do? 
Let's for each node I it sums up

266
00:12:12,160 --> 00:12:15,320
the feature vectors H of all 
nodes J that are connected to I.

267
00:12:15,600 --> 00:12:17,440
Exactly. 
It aggregates the features from 

268
00:12:17,440 --> 00:12:20,240
the neighbors. 
Then you typically multiply this

269
00:12:20,240 --> 00:12:23,440
result by a learnable weight 
matrix W to transform the 

270
00:12:23,440 --> 00:12:27,280
features and often apply a non 
linear activation function like 

271
00:12:27,280 --> 00:12:30,600
real U, So something like real 
UAHW. 

272
00:12:30,800 --> 00:12:34,040
OK, that seems like a basic 
message passing step neighbors 

273
00:12:34,040 --> 00:12:38,280
features summed and transformed.
But wait, you mentioned earlier 

274
00:12:38,280 --> 00:12:41,040
that a only defines connections 
to other nodes. 

275
00:12:41,160 --> 00:12:43,560
Does AH include the nodes own 
features? 

276
00:12:43,960 --> 00:12:46,320
Excellent point. 
Usually the standard adjacency 

277
00:12:46,320 --> 00:12:49,520
matrix A has zeros on the 
diagonal, meaning nodes aren't 

278
00:12:49,520 --> 00:12:53,000
connected to themselves. 
If you just use AH, a node only 

279
00:12:53,000 --> 00:12:55,360
gets information from its 
neighbors, potentially losing 

280
00:12:55,360 --> 00:12:57,640
its own original information 
after the transformation. 

281
00:12:57,800 --> 00:13:00,440
Which seems like a bad idea. 
You want the node to remember 

282
00:13:00,440 --> 00:13:02,000
itself. 
Definitely so. 

283
00:13:02,240 --> 00:13:05,280
A very common and important 
practical trick is to add self 

284
00:13:05,280 --> 00:13:07,760
loops to the graph before doing 
the aggregation. 

285
00:13:08,360 --> 00:13:09,680
Mathematically this is super 
simple. 

286
00:13:09,680 --> 00:13:12,800
You just add the identity matrix
I which has ones on the diagonal

287
00:13:12,800 --> 00:13:15,360
and zeros elsewhere, to the 
adjacency matrix A. 

288
00:13:15,560 --> 00:13:18,360
So instead of a you use a plus 
I. 

289
00:13:18,600 --> 00:13:22,840
Precisely. 
Now when you compute A+ IH, each

290
00:13:22,840 --> 00:13:26,280
node sums the features of its 
neighbors plus its own features.

291
00:13:26,920 --> 00:13:30,200
It's a simple tweak, but crucial
for stable learning in most GN 

292
00:13:30,200 --> 00:13:31,200
NS. 
Got it. 

293
00:13:31,840 --> 00:13:36,600
Add self loops using a plus I. 
So this basic real UA plus IHW 

294
00:13:37,120 --> 00:13:41,040
operation forms the foundation. 
Are there different variations 

295
00:13:41,040 --> 00:13:42,520
on this theme? 
Different ways to do the 

296
00:13:42,520 --> 00:13:43,880
aggregation. 
Absolutely. 

297
00:13:44,080 --> 00:13:47,280
That basic operation gives rise 
to the first type of GNN layer 

298
00:13:47,280 --> 00:13:49,640
we can discuss, which we could 
call some pulling. 

299
00:13:49,760 --> 00:13:52,800
OK some pulling because it 
literally just sums major 

300
00:13:52,800 --> 00:13:54,520
features plus self. 
Exactly. 

301
00:13:54,880 --> 00:14:00,400
He knew Ray LUA plus IHW 
Straightforward, but it has a 

302
00:14:00,400 --> 00:14:02,360
potential practical problem. 
What's that? 

303
00:14:02,440 --> 00:14:05,560
Imagine a node with thousands of
neighbors, like a celebrity on a

304
00:14:05,560 --> 00:14:07,800
social network. 
When you sum up all their 

305
00:14:07,800 --> 00:14:10,400
transformed features, the 
resulting vector, and he's new 

306
00:14:10,400 --> 00:14:12,720
for that node can become huge in
magnitude. 

307
00:14:13,200 --> 00:14:15,480
The scale can explode just 
because of the number of 

308
00:14:15,480 --> 00:14:16,360
connections. 
Right. 

309
00:14:16,360 --> 00:14:18,760
This can make the gradients 
during training very large or 

310
00:14:18,760 --> 00:14:21,720
unstable, making it hard for the
model to learn effectively. 

311
00:14:21,920 --> 00:14:25,200
So some pulling is simple but 
potentially unstable for graphs 

312
00:14:25,360 --> 00:14:28,880
with highly varied node degrees.
How do we fix that? 

313
00:14:29,280 --> 00:14:32,360
The most common fix leads to our
second type mean pulling. 

314
00:14:32,920 --> 00:14:34,640
Instead of just summing, we 
average. 

315
00:14:34,800 --> 00:14:37,920
OK, how do you implement 
averaging using matrix 

316
00:14:37,920 --> 00:14:40,000
operations? 
You normalize the A plus I 

317
00:14:40,000 --> 00:14:42,000
matrix. 
For each node, you divide the 

318
00:14:42,000 --> 00:14:45,640
contribution from its neighbors 
and a self by its degree, the 

319
00:14:45,640 --> 00:14:47,680
number of neighbors plus one. 
For the self loop, 

320
00:14:48,240 --> 00:14:51,840
mathematically, you multiply A +
I by the inverse of the degree 

321
00:14:51,840 --> 00:14:55,160
matrix. 
DD is a diagonal matrix with no 

322
00:14:55,160 --> 00:14:58,360
degrees on the diagonal, so the 
operation becomes something like

323
00:14:58,520 --> 00:15:04,360
real you D1A plus IHW. 
So D1 handles the division by 

324
00:15:04,400 --> 00:15:07,000
the degree for each node. 
That's right, it effectively 

325
00:15:07,000 --> 00:15:09,360
computes the mean of the 
features from the neighborhood, 

326
00:15:09,520 --> 00:15:11,800
including the node itself. 
This prevents the scale 

327
00:15:11,800 --> 00:15:14,080
explosion issue because you're 
averaging, not summing. 

328
00:15:14,280 --> 00:15:15,800
And this is generally more 
stable. 

329
00:15:15,800 --> 00:15:18,240
Much more stable. 
Mean pulling is a really solid, 

330
00:15:18,240 --> 00:15:20,520
simple baseline. 
It works well especially for 

331
00:15:20,520 --> 00:15:23,440
inductive problems where you 
need to generalize to new unseen

332
00:15:23,440 --> 00:15:24,520
graphs. 
OK. 

333
00:15:24,520 --> 00:15:27,240
Some pulling then mean pulling 
for stability. 

334
00:15:27,560 --> 00:15:30,120
You mentioned a really popular 
one earlier, GCM. 

335
00:15:30,160 --> 00:15:33,600
Yes, the Graph Convolutional 
Network GCN layer introduced by 

336
00:15:33,600 --> 00:15:36,080
Kipps and Welling. 
This is probably the most cited 

337
00:15:36,080 --> 00:15:38,160
and widely used GNN layer 
currently. 

338
00:15:38,360 --> 00:15:41,080
What's its specific innovation? 
Is it different from mean 

339
00:15:41,080 --> 00:15:43,240
pulling? 
It uses a slightly different, 

340
00:15:43,240 --> 00:15:46,240
more sophisticated normalization
of the adjacency matrix. 

341
00:15:46,920 --> 00:15:50,520
Instead of just D1A plus I, it 
uses what's called symmetric 

342
00:15:50,520 --> 00:15:56,080
normalization, D12A plus ID 12. 
D2 the minus half power. 

343
00:15:56,080 --> 00:15:58,360
What does that achieve? 
Intuitively, it bounces. 

344
00:15:58,360 --> 00:16:02,520
The influence mean pulling D1A 
normalizes based on the receiver

345
00:16:02,520 --> 00:16:04,680
nodes degree. 
Symmetric normalization 

346
00:16:04,680 --> 00:16:06,920
considers both the senders and 
receivers degrees. 

347
00:16:07,160 --> 00:16:10,520
It tends to work very well 
empirically, often slightly 

348
00:16:10,520 --> 00:16:12,440
outperforming mean pulling. 
Why do you think it's so 

349
00:16:12,440 --> 00:16:14,200
popular? 
It hits a sweet spot. 

350
00:16:14,520 --> 00:16:17,880
It's simple, computationally 
efficient, scalable and perform 

351
00:16:17,880 --> 00:16:19,680
strongly on many benchmark 
tasks. 

352
00:16:20,000 --> 00:16:22,680
It became a go to layer for many
researchers and practitioners. 

353
00:16:22,880 --> 00:16:25,680
OK, GCN uses that symmetric 
normalization. 

354
00:16:26,280 --> 00:16:29,880
What if the edges themselves 
have important features like the

355
00:16:29,880 --> 00:16:32,720
type of chemical bond or the 
strength of a connection? 

356
00:16:33,040 --> 00:16:37,360
Some MEAN and GCN seem focused 
on node features and 

357
00:16:37,360 --> 00:16:39,960
connectivity. 
Great point that brings to a 

358
00:16:39,960 --> 00:16:43,040
more general framework message 
passing neural networks and 

359
00:16:43,040 --> 00:16:43,920
PNNS. 
OK. 

360
00:16:44,000 --> 00:16:46,920
Sounds more fundamental. 
It is, and PNNS define the 

361
00:16:46,920 --> 00:16:49,320
process more explicitly. 
They typically have 3 

362
00:16:49,320 --> 00:16:53,080
components. 1A Message function.
This computes a message for each

363
00:16:53,080 --> 00:16:55,920
edge, potentially considering 
the features of the sending 

364
00:16:55,920 --> 00:16:59,080
node, the receiving node, and 
any features of the edge itself.

365
00:16:59,640 --> 00:17:03,400
2 An aggregation function. 
This takes all the incoming 

366
00:17:03,400 --> 00:17:06,319
messages for a node from its 
neighbors and aggregates them. 

367
00:17:06,480 --> 00:17:10,680
For example some mean Max three 
an update function or read out. 

368
00:17:11,040 --> 00:17:13,680
This combines the aggregated 
message with the nodes current 

369
00:17:13,680 --> 00:17:16,560
feature vector to produce the 
nodes new feature vector for the

370
00:17:16,560 --> 00:17:18,920
next layer. 
So it explicitly models message 

371
00:17:18,920 --> 00:17:21,560
creation based on sender, 
receiver and edge features. 

372
00:17:21,560 --> 00:17:22,880
That sounds much more 
expressive. 

373
00:17:22,880 --> 00:17:24,680
It is very expressive. 
If you have rich edge 

374
00:17:24,680 --> 00:17:26,880
information, MPN ends are often 
the way to go. 

375
00:17:27,079 --> 00:17:28,680
However, there's a potential 
downside. 

376
00:17:29,120 --> 00:17:31,800
Computing and storing 
potentially complex message 

377
00:17:31,800 --> 00:17:34,800
vectors for every single edge 
can become very memory 

378
00:17:34,800 --> 00:17:37,480
intensive, especially on large 
dense graphs. 

379
00:17:38,120 --> 00:17:40,600
They can also sometimes be prone
to overfitting if you don't have

380
00:17:40,600 --> 00:17:43,080
enough data to learn the complex
message functions well. 

381
00:17:43,160 --> 00:17:47,400
So powerful, but potentially 
resource hungry and data hungry.

382
00:17:47,720 --> 00:17:50,920
Best for maybe smaller graphs 
with really informative edges. 

383
00:17:50,920 --> 00:17:53,840
That's often where they shine, 
yes, like in molecular graphs 

384
00:17:53,840 --> 00:17:57,560
where bond types are critical. 
OK, we've covered summing, 

385
00:17:57,800 --> 00:18:00,480
averaging, symmetric 
normalization, and explicit 

386
00:18:00,480 --> 00:18:03,640
message passing. 
What about the idea of learning 

387
00:18:03,640 --> 00:18:06,880
the importance of neighbors? 
GCN and mean pulling? 

388
00:18:06,880 --> 00:18:10,360
Treat all neighbors after 
normalization somewhat equally, 

389
00:18:10,400 --> 00:18:12,040
right? 
That's a key limitation they 

390
00:18:12,040 --> 00:18:15,600
have, and that's exactly what 
Graph Attention networks Gats 

391
00:18:15,600 --> 00:18:18,360
were designed to address. 
Attention like in Transformers. 

392
00:18:18,360 --> 00:18:20,800
Very much inspired by the 
success of attention in sequence

393
00:18:20,800 --> 00:18:21,600
models. 
Yes. 

394
00:18:22,040 --> 00:18:25,200
The core idea in GAP is instead 
of using a fixed normalization 

395
00:18:25,200 --> 00:18:28,800
like mean or symmetric, let the 
model learn how important each 

396
00:18:28,800 --> 00:18:30,560
neighbor is for updating a given
node. 

397
00:18:30,680 --> 00:18:33,280
Just learn that. 
For each edge connecting node J 

398
00:18:33,280 --> 00:18:36,920
to node ijat computes an 
attention coefficient, let's 

399
00:18:36,920 --> 00:18:39,880
call it Alfie. 
This coefficient represents the 

400
00:18:39,880 --> 00:18:41,920
importance of node J's features 
to node I. 

401
00:18:42,400 --> 00:18:45,520
And how is Alfie calculated? 
It's typically calculated using 

402
00:18:45,520 --> 00:18:48,480
a small neural network and 
attention mechanism that takes 

403
00:18:48,480 --> 00:18:51,280
the features of node I and node 
J as input. 

404
00:18:52,160 --> 00:18:53,960
This mechanism is learned during
training. 

405
00:18:54,160 --> 00:18:57,640
So the network learns a function
to decide attention scores. 

406
00:18:57,640 --> 00:19:00,920
Exactly. 
Then these raw attention scores 

407
00:19:00,920 --> 00:19:04,560
for all of no dies neighbors are
usually normalized using a soft 

408
00:19:04,560 --> 00:19:08,240
Max function so they sum to one.
This gives you a weighted 

409
00:19:08,240 --> 00:19:11,640
average of neighbor features, 
where the weights Alfie are 

410
00:19:11,640 --> 00:19:15,080
learned by the model itself. 
So it's like mean pulling, but 

411
00:19:15,080 --> 00:19:17,520
instead of a simple average, 
it's a weighted average where 

412
00:19:17,520 --> 00:19:19,160
the weights are learned 
attention scores. 

413
00:19:19,200 --> 00:19:21,840
Precisely this give the model 
much more flexibility. 

414
00:19:22,320 --> 00:19:24,680
It can learn to pay more 
attention to relevant neighbors 

415
00:19:24,720 --> 00:19:27,320
and ignore irrelevant ones for a
specific task. 

416
00:19:27,600 --> 00:19:29,920
That sounds really powerful. 
Is it as computationally 

417
00:19:29,920 --> 00:19:33,080
expensive as MPN NS? 
Generally less so than the most 

418
00:19:33,080 --> 00:19:36,760
complex MPN NS. 
While it computes an attention 

419
00:19:36,760 --> 00:19:40,680
score per edge, that score is 
usually just a scalar value, not

420
00:19:40,680 --> 00:19:43,200
a high dimensional vector like 
some MPN messages. 

421
00:19:43,680 --> 00:19:46,720
This makes it more scalable. 
Often people use multi head 

422
00:19:46,720 --> 00:19:49,600
attention, computing several 
sets of attention scores in 

423
00:19:49,600 --> 00:19:52,920
parallel and combining them, 
which helps stabilize training 

424
00:19:53,040 --> 00:19:55,600
and capture different aspects of
neighborhood importance. 

425
00:19:55,600 --> 00:19:58,400
Multi head attention just like 
in Transformers. 

426
00:19:58,400 --> 00:20:00,400
That connection seems really 
strong now. 

427
00:20:00,600 --> 00:20:03,920
It is in fact you can view a 
transformer as a special kind of

428
00:20:03,920 --> 00:20:06,120
GNN operating on a fully 
connected graph. 

429
00:20:06,120 --> 00:20:09,920
Fully connected, meaning every 
input token like a word is 

430
00:20:09,920 --> 00:20:13,040
connected to every other token. 
Yes, if you think of words as 

431
00:20:13,040 --> 00:20:16,360
nodes, a transformer uses 
attention as the mechanism to 

432
00:20:16,400 --> 00:20:19,120
aggregate information from all 
other words nodes. 

433
00:20:19,760 --> 00:20:23,320
The main difference is that in 
text the order matters, but a 

434
00:20:23,320 --> 00:20:26,680
basic, fully connected graph 
doesn't have inherent order. 

435
00:20:26,840 --> 00:20:28,720
So Transformers need something 
extra. 

436
00:20:28,920 --> 00:20:31,320
Right, they need those 
positional embeddings added to 

437
00:20:31,320 --> 00:20:34,400
the input features to give the 
model information about the 

438
00:20:34,400 --> 00:20:37,040
sequence order GN NS operating 
on. 

439
00:20:37,040 --> 00:20:40,680
Naturally, sparse grass often 
don't need explicit positional 

440
00:20:40,680 --> 00:20:43,960
encoding because the structure 
itself provides context. 

441
00:20:44,360 --> 00:20:46,760
That's a fascinating link. 
It really frames Transformers 

442
00:20:46,760 --> 00:20:49,320
within this broader graph 
perspective. 

443
00:20:49,680 --> 00:20:51,920
OK, so we've covered a good 
range of architectures. 

444
00:20:52,160 --> 00:20:54,280
Let's shift gears slightly and 
talk about some cutting edge 

445
00:20:54,280 --> 00:20:57,000
research. 
NEC Labs America has been doing 

446
00:20:57,000 --> 00:20:59,320
some significant work in GN NS, 
right? 

447
00:20:59,320 --> 00:21:00,720
What are some of their 
contributions? 

448
00:21:00,960 --> 00:21:03,160
Yes, they've been tackling some 
really important practical 

449
00:21:03,160 --> 00:21:06,280
challenges. 
One area is making GN NS more 

450
00:21:06,280 --> 00:21:08,520
robust. 
Robust to what? 

451
00:21:08,960 --> 00:21:13,400
Noisy data attacks. 
Both real world graphs often 

452
00:21:13,400 --> 00:21:15,640
have noisy or even irrelevant 
edges. 

453
00:21:15,960 --> 00:21:18,720
Sometimes connections might be 
actively manipulated in 

454
00:21:18,720 --> 00:21:21,680
adversarial settings like 
finance or security gene and 

455
00:21:21,680 --> 00:21:24,200
performance can degrade 
significantly with this noise. 

456
00:21:24,200 --> 00:21:27,440
So how do you make them tougher?
NEC labs develop methods like 

457
00:21:27,440 --> 00:21:30,200
PTD NET, which stands for a 
Parameterized Topological 

458
00:21:30,200 --> 00:21:32,840
Denoising Network. 
That's a mouthful, but the idea 

459
00:21:32,840 --> 00:21:34,440
is clever. 
It essentially trains an 

460
00:21:34,440 --> 00:21:38,560
auxiliary network alongside the 
main GNN to learn which edges 

461
00:21:38,560 --> 00:21:41,520
are likely to be irrelevant or 
noisy for the task at hand. 

462
00:21:42,040 --> 00:21:45,400
It then learns to effectively 
down weight or even drop those 

463
00:21:45,400 --> 00:21:47,360
edges during the GN NS message 
passing. 

464
00:21:47,440 --> 00:21:50,880
So it learns to ignore the noise
in the graph structure itself. 

465
00:21:50,880 --> 00:21:53,680
Exactly. 
Wei Chang at NEC Labs put it 

466
00:21:53,680 --> 00:21:57,040
nicely, saying something like by
focusing on the most relevant 

467
00:21:57,040 --> 00:22:00,400
connections, we've been able to 
significantly enhance the 

468
00:22:00,400 --> 00:22:04,280
robustness of GN NS, especially 
against noisy data, which is 

469
00:22:04,280 --> 00:22:06,280
common in the real world. 
That makes sense. 

470
00:22:06,480 --> 00:22:09,040
If the model can clean up its 
own input graph it should 

471
00:22:09,040 --> 00:22:11,680
perform better. 
What about understanding why a 

472
00:22:11,680 --> 00:22:15,440
GNN makes a prediction? 
Explain ability is huge in AI 

473
00:22:15,440 --> 00:22:17,440
now. 
Absolutely critical, especially 

474
00:22:17,440 --> 00:22:19,680
if you're using GNNS for high 
stakes decisions. 

475
00:22:19,880 --> 00:22:22,480
NEC Labs has also focused 
heavily on advancing 

476
00:22:22,480 --> 00:22:24,920
explainability. 
How do you explain AGNN 

477
00:22:24,920 --> 00:22:27,880
prediction? 
Is it about which nodes or edges

478
00:22:27,880 --> 00:22:30,440
were most influential? 
That's the usual approach, 

479
00:22:30,560 --> 00:22:33,840
identifying important subgraphs 
or features, but a key challenge

480
00:22:33,840 --> 00:22:37,000
is reliably evaluating whether 
an explanation method is 

481
00:22:37,000 --> 00:22:38,960
actually faithful to what the 
model is doing. 

482
00:22:39,120 --> 00:22:40,440
Right. 
How do you know the explanation 

483
00:22:40,440 --> 00:22:42,920
is correct? 
NEC Labs developed new ways to 

484
00:22:42,920 --> 00:22:46,040
measure this fidelity using 
ideas from information theory. 

485
00:22:46,600 --> 00:22:48,960
They've worked on creating 
robust metrics to ensure that 

486
00:22:48,960 --> 00:22:51,520
when an explanation method 
highlights certain parts of the 

487
00:22:51,520 --> 00:22:55,640
graph, those parts are genuinely
driving the GN NS prediction. 

488
00:22:56,400 --> 00:22:59,440
Hyphen Chen at NEC Labs 
emphasized their goal is to make

489
00:22:59,440 --> 00:23:02,640
GN NS not just powerful, but 
also transparent and 

490
00:23:02,640 --> 00:23:04,120
trustworthy. 
Trustworthy. 

491
00:23:04,120 --> 00:23:06,240
Is the keyword there. 
OK, Robustness. 

492
00:23:06,240 --> 00:23:09,440
Explainability. 
What about graphs that change 

493
00:23:09,440 --> 00:23:12,440
over time, like social networks 
or communication networks? 

494
00:23:12,440 --> 00:23:15,240
That's another area anomaly 
detection and dynamic graphs. 

495
00:23:15,520 --> 00:23:18,640
Networks aren't static. 
They evolve new connections form

496
00:23:18,640 --> 00:23:21,480
old ones break. 
Detecting unusual changes can be

497
00:23:21,480 --> 00:23:24,880
critical for things like network
intrusion detection or 

498
00:23:24,880 --> 00:23:26,920
identifying emerging fraudulent 
patterns. 

499
00:23:26,920 --> 00:23:30,040
Standard GN NS might struggle 
with the time aspect, right? 

500
00:23:30,320 --> 00:23:34,280
They often process static 
snapshots, so NEC Labs develops 

501
00:23:34,280 --> 00:23:37,200
TRGNN Structural Temporal Graph 
Neural Networks. 

502
00:23:37,760 --> 00:23:40,400
These models are specifically 
designed to learn from sequences

503
00:23:40,400 --> 00:23:42,800
of graph stamp fats over time. 
So they look at both the 

504
00:23:42,800 --> 00:23:45,200
structure and how it changes. 
Exactly. 

505
00:23:45,440 --> 00:23:47,400
They integrate both the 
structural information from the 

506
00:23:47,400 --> 00:23:50,640
graph at each time step and the 
temporal dynamics of how edges 

507
00:23:50,640 --> 00:23:54,040
appear or disappear. 
Zheng Zheng Chen from the lab 

508
00:23:54,040 --> 00:23:57,840
highlighted that this 
integration allows CIGNN to 

509
00:23:58,040 --> 00:24:01,520
effectively pinpoint anomalies 
and dynamic systems, making it 

510
00:24:01,520 --> 00:24:04,240
valuable for proactive threat 
detection and security systems 

511
00:24:04,240 --> 00:24:05,560
where it's actually been 
deployed. 

512
00:24:05,640 --> 00:24:07,040
Wow. 
Deployed in real systems? 

513
00:24:07,040 --> 00:24:09,480
That's significant. 
One more challenge. 

514
00:24:09,840 --> 00:24:13,400
What happens when a trained GNN 
encounters types of nodes it 

515
00:24:13,400 --> 00:24:16,720
never saw during training? 
Say a new type of device 

516
00:24:16,720 --> 00:24:19,120
connects to your network. 
That's the outer distribution 

517
00:24:19,120 --> 00:24:21,120
problem, and it can really hurt 
reliability. 

518
00:24:21,600 --> 00:24:25,200
The GNN might make unpredictable
or poorly calibrated predictions

519
00:24:25,200 --> 00:24:28,080
for these unfamiliar nodes. 
NEC Labs tackled this with 

520
00:24:28,080 --> 00:24:30,200
calibration for out of 
distribution nodes. 

521
00:24:30,680 --> 00:24:32,480
How do you calibrate for 
something you haven't seen? 

522
00:24:32,800 --> 00:24:36,200
They developed a framework 
called GRDQ Graph Edge re 

523
00:24:36,200 --> 00:24:38,000
Weighting via Deep Queue 
Learning. 

524
00:24:38,440 --> 00:24:41,120
It uses reinforcement learning 
concepts, specifically queue 

525
00:24:41,120 --> 00:24:44,520
learning, to learn how to adjust
the importance the weights of 

526
00:24:44,520 --> 00:24:47,360
messages coming from potentially
out of distribution neighbors. 

527
00:24:47,640 --> 00:24:50,880
So it learns to dynamically 
adapt the message passing when 

528
00:24:50,880 --> 00:24:53,680
it encounters unfamiliar nodes. 
Essentially, yes. 

529
00:24:53,760 --> 00:24:56,560
It tries to mitigate the 
negative impact of these novel 

530
00:24:56,560 --> 00:24:59,960
nodes types, making the GNN's 
predictions more reliable and 

531
00:24:59,960 --> 00:25:02,480
calibrated even when the graph 
environment changes. 

532
00:25:02,760 --> 00:25:06,960
OK, so NEC Labs is pushing on 
robustness, explainability, 

533
00:25:06,960 --> 00:25:09,760
dynamic graphs and out of 
distribution handling. 

534
00:25:09,760 --> 00:25:11,600
That really covers a lot of the 
practical hurdles. 

535
00:25:11,920 --> 00:25:15,280
Now to tie all these 
architectural ideas, some mean 

536
00:25:15,280 --> 00:25:18,920
GCN attention together. 
Can we look at a concrete 

537
00:25:18,920 --> 00:25:21,240
example? 
Maybe a benchmark data set? 

538
00:25:21,240 --> 00:25:23,320
Absolutely. 
The Core Data set is a classic 

539
00:25:23,320 --> 00:25:25,360
benchmark, perfect for 
illustrating these differences. 

540
00:25:25,360 --> 00:25:27,640
Cora, what is it? 
It's a citation network data 

541
00:25:27,640 --> 00:25:29,000
set. 
Each node represents a 

542
00:25:29,000 --> 00:25:31,600
scientific paper. 
An edge exists from paper A to 

543
00:25:31,600 --> 00:25:33,880
paper B if paper A cites paper 
B. 

544
00:25:33,960 --> 00:25:35,960
OK, a network of papers and 
citations. 

545
00:25:35,960 --> 00:25:38,080
What's the task? 
Each paper comes with some 

546
00:25:38,080 --> 00:25:40,960
features, typically a bag of 
words vector representing the 

547
00:25:40,960 --> 00:25:42,600
words in the papers abstractor 
text. 

548
00:25:42,840 --> 00:25:46,520
The task is node classification,
predict the research topic like 

549
00:25:46,520 --> 00:25:49,520
neural networks, reinforcement 
learning, probabilistic methods 

550
00:25:49,520 --> 00:25:51,720
etcetera. 
There are 7 topics in Cora for 

551
00:25:51,720 --> 00:25:54,000
each paper. 
So classify papers based on 

552
00:25:54,000 --> 00:25:56,560
their content and who they cite 
or get cited by. 

553
00:25:57,080 --> 00:25:59,320
What makes Cora interesting for 
GNNS? 

554
00:25:59,480 --> 00:26:02,320
Two things mainly. 
First, the connections clearly 

555
00:26:02,320 --> 00:26:04,760
matter. 
Papers on similar topics tend to

556
00:26:04,760 --> 00:26:07,080
cite each other. 
Second, it's often used in a 

557
00:26:07,080 --> 00:26:09,560
semi supervised or low data 
setting. 

558
00:26:10,200 --> 00:26:13,680
You might have thousands of 
papers nodes, but maybe only 

559
00:26:13,680 --> 00:26:16,120
labels for a very small 
fraction, like 20 papers per 

560
00:26:16,120 --> 00:26:18,800
topic. 
So just 140 labeled nodes in 

561
00:26:18,800 --> 00:26:22,160
total for training. 
Only 140 training examples for 

562
00:26:22,160 --> 00:26:23,680
thousands of nodes? 
That's tiny. 

563
00:26:23,680 --> 00:26:26,000
Exactly. 
It forces the model to rely 

564
00:26:26,000 --> 00:26:28,880
heavily on the graph structure 
to propagate information from 

565
00:26:28,880 --> 00:26:31,400
the few labeled nodes to the 
many unlabeled ones. 

566
00:26:31,840 --> 00:26:34,800
You can't possibly succeed just 
using the node features alone 

567
00:26:35,000 --> 00:26:37,840
with so little labeled. 
Data OK So what happens if you 

568
00:26:37,840 --> 00:26:41,000
try what's the baseline 
performance using only the paper

569
00:26:41,000 --> 00:26:43,840
features the bag of words with a
standard classifier like a 

570
00:26:43,840 --> 00:26:48,120
simple multi layer perceptron 
MLP ignoring the citations if. 

571
00:26:48,120 --> 00:26:49,560
You do that, Just use the node 
features. 

572
00:26:49,560 --> 00:26:52,680
Your accuracy is typically 
pretty low, maybe around 5055%. 

573
00:26:53,280 --> 00:26:56,040
It struggles because it doesn't 
have enough labels and ignores 

574
00:26:56,040 --> 00:26:58,200
the rich relational information 
in the citations. 

575
00:26:58,440 --> 00:26:59,880
Right, just over chance 
basically. 

576
00:27:00,000 --> 00:27:01,680
So now let's bring in the graph 
structure. 

577
00:27:01,880 --> 00:27:05,000
What if we use that simple some 
pulling GNN layer we discussed 

578
00:27:05,000 --> 00:27:10,040
first re Lu, A+, IHW? 
OK, if you build a simple two 

579
00:27:10,040 --> 00:27:12,800
layer GNN using some pulling, 
the performance jumps 

580
00:27:12,800 --> 00:27:15,240
significantly. 
You'd likely see accuracy jump 

581
00:27:15,240 --> 00:27:19,120
up to around say 77.5%. 
Big improvement, but you 

582
00:27:19,120 --> 00:27:21,160
mentioned some pulling can be 
unstable. 

583
00:27:21,280 --> 00:27:23,600
Does that show up? 
It can, yes. 

584
00:27:23,840 --> 00:27:26,800
Depending on the initialization 
and hyper parameters, some 

585
00:27:26,800 --> 00:27:29,480
pulling results on CORIC and 
sometimes vary a bit more, 

586
00:27:29,720 --> 00:27:32,600
reflecting that potential 
instability from unnormalized 

587
00:27:32,600 --> 00:27:35,040
feature aggregation. 
OK, So what about mean pulling 

588
00:27:35,160 --> 00:27:39,560
where we normalize by no degree 
real UD1A plus IHW? 

589
00:27:39,960 --> 00:27:42,560
Mean pulling generally gives a 
more stable result and often a 

590
00:27:42,560 --> 00:27:46,320
slight boost in performance. 
You might get around 78.5% or 

591
00:27:46,320 --> 00:27:50,080
maybe close to 79% accuracy. 
It clearly shows the benefit of 

592
00:27:50,080 --> 00:27:52,680
that simple normalization for 
stability and performance. 

593
00:27:52,720 --> 00:27:55,480
Better and more reliable. 
Now what about the star player? 

594
00:27:55,560 --> 00:27:58,040
The GCN layer with symmetric 
normalization? 

595
00:27:58,040 --> 00:28:02,880
Real you D12A plus ID 12HW. 
This is where Core really shines

596
00:28:02,880 --> 00:28:06,880
as a benchmark for GCNS. 
A standard two layer GCN model 

597
00:28:07,080 --> 00:28:10,200
often pushes the accuracy 
significantly higher, achieving 

598
00:28:10,200 --> 00:28:14,520
results consistently above 80%, 
sometimes reaching 81% or even 

599
00:28:14,520 --> 00:28:17,800
slightly higher depending on the
specific setup and data split. 

600
00:28:18,080 --> 00:28:23,360
Wow, so going from mean pulling 
79% to GCNS 81% plus is a 

601
00:28:23,360 --> 00:28:26,560
noticeable step up just from 
changing the normalization 

602
00:28:26,560 --> 00:28:28,600
scheme? 
It is on this benchmark, yes. 

603
00:28:28,840 --> 00:28:31,160
It really highlighted the 
empirical effectiveness of that 

604
00:28:31,160 --> 00:28:33,880
symmetric normalization. 
When the GCN paper came out, it 

605
00:28:33,880 --> 00:28:35,280
became the number to beat for a 
while. 

606
00:28:35,480 --> 00:28:37,920
What about GAA? 
The graph Attention network? 

607
00:28:37,920 --> 00:28:39,960
Does learning neighbor 
importance help further on 

608
00:28:40,000 --> 00:28:42,440
Quora? 
Yes, GAT models also perform 

609
00:28:42,440 --> 00:28:45,120
very well on Quora, often 
achieving similar or even 

610
00:28:45,120 --> 00:28:47,520
slightly better results than 
GCM, perhaps getting into the 

611
00:28:47,520 --> 00:28:51,240
8283% range sometimes. 
It shows that allowing the model

612
00:28:51,240 --> 00:28:53,680
to learn which citations are 
more important for classifying a

613
00:28:53,680 --> 00:28:56,080
papers topic can provide an 
additional edge. 

614
00:28:56,200 --> 00:28:58,720
So that Quora example perfectly 
illustrates the journey. 

615
00:28:59,040 --> 00:29:03,040
Ignoring the graph gives poor 
results. 50% simple summing 

616
00:29:03,040 --> 00:29:07,520
helps. 77.5% averaging more of 
lane pulling is better and 

617
00:29:07,520 --> 00:29:11,760
stabler. 79% sophisticated 
normalization like GCN pushes it

618
00:29:11,760 --> 00:29:15,400
higher. 81% plus and learn at 
attention like geek can 

619
00:29:15,400 --> 00:29:17,960
potentially refine it even 
further. 82% plus. 

620
00:29:17,960 --> 00:29:19,920
Exactly. 
It makes the abstract concepts 

621
00:29:19,920 --> 00:29:22,160
very concrete. 
The way you leverage the graph 

622
00:29:22,160 --> 00:29:24,760
structure, especially how you 
normalize or weight the 

623
00:29:24,760 --> 00:29:27,480
information flowing between 
nodes, really matters for 

624
00:29:27,480 --> 00:29:29,600
performance. 
It's not just that nodes are 

625
00:29:29,600 --> 00:29:32,680
connected, but how the GNN 
processes those connections. 

626
00:29:33,040 --> 00:29:34,600
This has been incredibly 
insightful. 

627
00:29:34,600 --> 00:29:36,920
We've gone from the basic idea 
of graphs meeting neural 

628
00:29:36,920 --> 00:29:39,200
networks through the core 
mechanism of message passing, 

629
00:29:39,480 --> 00:29:42,320
explored a whole spectrum of 
architectures from simple sums 

630
00:29:42,320 --> 00:29:45,760
to sophisticated attention, seen
real world applications changing

631
00:29:45,760 --> 00:29:48,560
industries, touched on cutting 
edge research making them more 

632
00:29:48,560 --> 00:29:51,320
robust and trustworthy, and walk
through a concrete example 

633
00:29:51,320 --> 00:29:53,640
showing how these design choices
impact performance. 

634
00:29:53,640 --> 00:29:55,040
It really covers a lot of 
ground. 

635
00:29:55,360 --> 00:29:57,760
GNS provides such a powerful 
lens for looking at 

636
00:29:57,760 --> 00:30:00,480
interconnected data, moving 
beyond tables and sequences. 

637
00:30:00,640 --> 00:30:03,560
Absolutely. 
So thinking about this deep, the

638
00:30:03,560 --> 00:30:07,760
big take away seems to be that 
connections matter, and GNS give

639
00:30:07,760 --> 00:30:10,160
us the tools to finally learn 
from them effectively. 

640
00:30:10,600 --> 00:30:13,840
From finding new medicines to 
getting you places faster, to 

641
00:30:13,840 --> 00:30:17,080
suggesting things you might 
like, to even securing networks,

642
00:30:17,080 --> 00:30:19,320
it's all about understanding 
relationships. 

643
00:30:19,400 --> 00:30:21,840
Well said. 
The shift towards analyzing 

644
00:30:21,840 --> 00:30:25,120
relational data is a major 
trend, and GN NS are at the 

645
00:30:25,120 --> 00:30:27,600
forefront of it. 
So for everyone listening, the 

646
00:30:27,800 --> 00:30:31,600
final thought perhaps is this. 
Start looking for the graphs 

647
00:30:31,640 --> 00:30:35,040
hidden in plain sight. 
Think about your own work, your 

648
00:30:35,040 --> 00:30:37,320
own data, even just your daily 
life. 

649
00:30:37,320 --> 00:30:40,000
Where are the hidden networks? 
Yeah, customer interactions, 

650
00:30:40,000 --> 00:30:42,680
process flows, social dynamics, 
biological pathways. 

651
00:30:42,680 --> 00:30:45,760
Exactly identifying those 
relationships, those nodes and 

652
00:30:45,760 --> 00:30:48,720
edges, could be the first step 
towards unlocking completely new

653
00:30:48,720 --> 00:30:51,840
insights or building entirely 
new applications using the power

654
00:30:51,840 --> 00:30:55,240
of graph neural networks. 
This area is moving so fast too.

655
00:30:55,400 --> 00:30:57,800
What we covered is a solid 
foundation, but there's constant

656
00:30:57,800 --> 00:31:00,280
innovation happening. 
It's a really exciting feel to 

657
00:31:00,280 --> 00:31:03,800
follow or get involved in. 
Definitely, hopefully this deep 

658
00:31:03,800 --> 00:31:06,160
dive has given you that 
foundation and maybe sparked 

659
00:31:06,160 --> 00:31:09,000
some curiosity to explore GNS 
further. 

660
00:31:09,320 --> 00:31:12,640
It's a complex but incredibly 
rewarding area of AI. 

661
00:31:12,800 --> 00:31:14,240
Agreed. 
Thanks for joining us on the 

662
00:31:14,240 --> 00:31:15,680
deep dive. 
We'll catch you next time.

