1
00:00:00,040 --> 00:00:02,200
Welcome back to the Inference 
and Intelligence Lab. 

2
00:00:02,640 --> 00:00:06,040
Today is episode 2 of our 
Inference in the Wild series, 

3
00:00:06,480 --> 00:00:08,520
the Space where theory Meets 
Reality. 

4
00:00:09,080 --> 00:00:12,200
We dive into messy data, 
imperfect experiments, model 

5
00:00:12,200 --> 00:00:14,880
evaluation traps, and the 
day-to-day decisions that 

6
00:00:14,880 --> 00:00:17,280
determine whether an analysis is
actually useful. 

7
00:00:18,080 --> 00:00:21,160
Today we're focusing on a mantra
every data scientist needs to 

8
00:00:21,160 --> 00:00:23,800
hear. 
Running a statistical test is 

9
00:00:23,800 --> 00:00:27,000
just pressing the shutter. 
Designing A measurement system 

10
00:00:27,120 --> 00:00:29,600
is building the camera. 
Yeah, and honestly, building 

11
00:00:29,600 --> 00:00:32,759
that camera is where like almost
everyone messes up at some 

12
00:00:32,759 --> 00:00:33,600
point. 
Exactly. 

13
00:00:34,040 --> 00:00:37,000
So today's mission for you 
listening is really about how 

14
00:00:37,000 --> 00:00:40,480
measurement design fundamentally
guides your statistical choices.

15
00:00:40,640 --> 00:00:43,120
Right, because we're looking at 
the Super common problem where, 

16
00:00:43,280 --> 00:00:46,280
you know, true product 
improvements are just completely

17
00:00:46,280 --> 00:00:48,480
masked by the noise of the 
metric itself. 

18
00:00:48,480 --> 00:00:50,720
It's a classic, incredibly 
frustrating scenario. 

19
00:00:50,720 --> 00:00:54,280
So let's set the stage. 
Picture this, you run an AB test

20
00:00:54,280 --> 00:00:55,160
on new feature. 
Right. 

21
00:00:55,160 --> 00:00:57,720
OK, I'm with you. 
And it genuinely improves 

22
00:00:57,720 --> 00:01:01,040
session duration for most users,
like they are actually staying 

23
00:01:01,040 --> 00:01:03,600
longer on the platform. 
Which is, I mean, that's the 

24
00:01:03,600 --> 00:01:05,239
dream scenario for the product 
team. 

25
00:01:05,440 --> 00:01:08,200
Totally. 
The feature works perfectly, but

26
00:01:08,200 --> 00:01:11,200
then you run your standard T 
test and the result comes back 

27
00:01:11,200 --> 00:01:16,360
is P = .34. 
Oh man, the dreaded .34. 

28
00:01:17,200 --> 00:01:20,240
Not statistically significant, 
the test is just completely 

29
00:01:20,240 --> 00:01:23,280
blind to the improvement. 
I feel like my immediate 

30
00:01:23,280 --> 00:01:26,520
instinct there and probably the 
instinct for most analysts 

31
00:01:26,520 --> 00:01:30,080
listening is just like, do we 
just need a bigger sample size? 

32
00:01:30,080 --> 00:01:32,520
Are we underpowered? 
OK, let's unpack this. 

33
00:01:32,520 --> 00:01:35,400
Yeah, because that's the bubble 
we need to burst right now. 

34
00:01:35,960 --> 00:01:37,960
Adding more users won't fix 
this. 

35
00:01:37,960 --> 00:01:40,520
Wait, really? 
Not even if we let it run for 

36
00:01:40,520 --> 00:01:41,400
another month? 
Nope. 

37
00:01:41,960 --> 00:01:44,440
You can get millions of users 
and you'd still be stuck. 

38
00:01:44,840 --> 00:01:47,720
The root cause isn't your sample
size, it's that the session 

39
00:01:47,720 --> 00:01:50,160
duration metric is heavily right
skewed. 

40
00:01:51,320 --> 00:01:54,560
Right, because session duration 
is never just a nice neat bell 

41
00:01:54,560 --> 00:01:57,480
curve. 
Exactly conceptually, think 

42
00:01:57,480 --> 00:02:00,200
about what's actually happening.
You have this massive cluster of

43
00:02:00,200 --> 00:02:03,200
casual users who browse for just
a few minutes, right? 

44
00:02:03,200 --> 00:02:04,640
Yeah, they click a few things 
and bounce. 

45
00:02:04,680 --> 00:02:06,120
Right. 
But then you have a handful of 

46
00:02:06,120 --> 00:02:09,639
power users who stay for hours. 
We're talking 300 minutes, maybe

47
00:02:09,639 --> 00:02:12,720
five hours at a time. 
And those extreme outliers just 

48
00:02:12,720 --> 00:02:15,200
completely dominate the 
variance, yes. 

49
00:02:15,440 --> 00:02:18,920
The raw metric just amplifies 
this noise to such an extreme 

50
00:02:18,920 --> 00:02:21,440
degree. 
The massive distance between the

51
00:02:21,440 --> 00:02:24,960
casual user and the five hour 
power user blows U the variance 

52
00:02:25,400 --> 00:02:27,520
and it buries the actual signal 
of the feature. 

53
00:02:27,520 --> 00:02:30,480
So the test just can't see the 
casual users staying longer 

54
00:02:30,480 --> 00:02:33,960
because the math gets wrecked by
those five hour power users. 

55
00:02:34,000 --> 00:02:36,320
You nailed it. 
OK, well if sample size isn't 

56
00:02:36,320 --> 00:02:39,320
the issue, then the data itself 
needs to be handled differently,

57
00:02:39,320 --> 00:02:41,280
right? 
Like we need standard data 

58
00:02:41,280 --> 00:02:42,960
science tools to clean up that 
noise. 

59
00:02:43,000 --> 00:02:44,960
So what's your first instinct? 
What do you reach for? 

60
00:02:45,120 --> 00:02:47,360
Well, probably covariate 
adjustment, right? 

61
00:02:47,360 --> 00:02:50,360
I mean we condition on pre 
experiment data to reduce the 

62
00:02:50,360 --> 00:02:52,200
variance. 
It's a totally classic 

63
00:02:52,200 --> 00:02:54,440
practitioner move. 
And usually that is a great 

64
00:02:54,440 --> 00:02:56,720
tool. 
Yeah, here's the reality check. 

65
00:02:57,040 --> 00:03:00,600
Simulations on this specific 
heavy tailed metric show it 

66
00:03:00,600 --> 00:03:03,960
fairly improves the raw data. 
Wait, why wouldn't it if we 

67
00:03:03,960 --> 00:03:07,200
adjust for their past behavior? 
Because adjusting for past data 

68
00:03:07,200 --> 00:03:11,800
just isn't enough when a few 
massive extreme outliers still 

69
00:03:11,800 --> 00:03:13,840
dominate the current experiments
data. 

70
00:03:13,960 --> 00:03:16,320
Oh, I see. 
Even if we know they're power 

71
00:03:16,320 --> 00:03:19,800
users from historical data, 
their actual session length 

72
00:03:19,800 --> 00:03:23,200
today is still, you know, wildly
unpredictable. 

73
00:03:23,200 --> 00:03:24,880
Exactly. 
Last week they stayed 300 

74
00:03:24,880 --> 00:03:28,120
minutes, today they say 400. 
That residual variance is still 

75
00:03:28,120 --> 00:03:30,360
astronomical compared to the 
casual users. 

76
00:03:30,360 --> 00:03:33,000
OK, that is so frustrating. 
Well, in that case, searching 

77
00:03:33,000 --> 00:03:36,360
for another solution, my classic
data science reflex kicks in. 

78
00:03:36,800 --> 00:03:38,640
Let's just take the log of the 
metric. 

79
00:03:38,640 --> 00:03:40,760
And that right there is a very 
dangerous illusion. 

80
00:03:41,000 --> 00:03:45,240
Dangerous, but taking the log is
so standard to like pull in 

81
00:03:45,240 --> 00:03:48,080
those long tails. 
It's tempting, yeah, but you 

82
00:03:48,080 --> 00:03:51,120
cannot forget that any 
mathematical transformation 

83
00:03:51,400 --> 00:03:53,920
fundamentally alters the null 
hypothesis. 

84
00:03:54,400 --> 00:03:56,800
Taking the log changes what you 
were actually measuring. 

85
00:03:56,880 --> 00:03:59,520
Wait, how so? 
It's just reshaping the data. 

86
00:03:59,520 --> 00:04:02,280
It's shifting the estimate from 
the arithmetic mean to the 

87
00:04:02,280 --> 00:04:04,720
geometric mean. 
Right, and the geometric mean 

88
00:04:04,720 --> 00:04:06,680
handles spread differently. 
Very differently. 

89
00:04:06,760 --> 00:04:08,880
It inherently penalizes 
variance. 

90
00:04:08,880 --> 00:04:12,120
So let me give you the specific 
kind of shocking example here, 

91
00:04:12,120 --> 00:04:12,760
OK? 
Hit me. 

92
00:04:12,960 --> 00:04:16,640
Imagine your new feature engages
power users really heavily. 

93
00:04:16,800 --> 00:04:18,920
They love it, so they stay even 
longer. 

94
00:04:18,920 --> 00:04:20,519
That increases the overall 
variance, right? 

95
00:04:20,519 --> 00:04:21,839
Yeah. 
The tail gets even longer, the 

96
00:04:21,839 --> 00:04:25,520
spread widens, right? 
So for this exact data, the 

97
00:04:25,520 --> 00:04:29,120
arithmetic mean might go up by 
nearly 34.9%. 

98
00:04:29,160 --> 00:04:31,680
Wow, so a massive success. 
Exactly. 

99
00:04:31,880 --> 00:04:35,000
But because of that variance 
penalty, the geometric mean 

100
00:04:35,000 --> 00:04:37,360
actually goes down by 4.2%. 
Wait. 

101
00:04:37,520 --> 00:04:39,280
Are you serious? 
It goes down. 

102
00:04:39,280 --> 00:04:42,480
Yes, it's the exact same data, 
the exact same users, but 

103
00:04:42,480 --> 00:04:45,080
entirely opposite conclusions. 
That's insane. 

104
00:04:45,080 --> 00:04:48,000
The log transform test would 
literally report the future as 

105
00:04:48,000 --> 00:04:49,640
harmful. 
You would be answering a 

106
00:04:49,640 --> 00:04:51,800
question the business never 
asked, and you'd kill a 

107
00:04:51,800 --> 00:04:53,960
massively successful feature. 
Wow. 

108
00:04:54,320 --> 00:04:58,320
OK, so covariate adjustment 
fails and log transforms 

109
00:04:58,360 --> 00:05:00,440
actively lie to the business 
here. 

110
00:05:00,440 --> 00:05:03,000
If we can't use those, what 
actually works without 

111
00:05:03,000 --> 00:05:05,200
distorting the truth? 
This brings us to the proper 

112
00:05:05,200 --> 00:05:08,600
tool for this specific job, 
which is a rank transformation. 

113
00:05:08,760 --> 00:05:11,560
Rank transformation like just 
ranking the users. 

114
00:05:11,560 --> 00:05:13,640
Exactly. 
Instead of measuring raw 

115
00:05:13,640 --> 00:05:16,720
minutes, you replace each 
observation with its rank order 

116
00:05:16,720 --> 00:05:19,720
in the combined sample. 
Oh wow, let's just do that for 

117
00:05:19,720 --> 00:05:21,760
everything then. 
If it works, we should apply 

118
00:05:21,760 --> 00:05:24,040
this everywhere. 
OK, here's where it it's really 

119
00:05:24,040 --> 00:05:26,800
interesting. 
You cannot blindly apply 

120
00:05:26,800 --> 00:05:28,200
transformations. 
Fair enough. 

121
00:05:28,200 --> 00:05:30,480
Yeah, there's always a catch. 
So instead of just rushing to 

122
00:05:30,480 --> 00:05:34,040
code, we have to use the rank 
transformation to walk through 

123
00:05:34,040 --> 00:05:37,240
the ultimate four question 
checklist for measurement 

124
00:05:37,240 --> 00:05:38,400
design. 
OK, I'm ready. 

125
00:05:38,480 --> 00:05:41,240
What's question one? 
Question one, What is the 

126
00:05:41,240 --> 00:05:43,560
business question? 
Let's break down the math 

127
00:05:43,560 --> 00:05:47,040
conceptually. 
A standard T test essentially 

128
00:05:47,040 --> 00:05:49,680
asks, is the average mean the 
same? 

129
00:05:49,800 --> 00:05:53,280
Right, like how many exact 
minutes did session duration 

130
00:05:53,280 --> 00:05:54,480
increase? 
Exactly. 

131
00:05:54,760 --> 00:05:57,840
But the business actually wants 
to know does this feature make 

132
00:05:57,840 --> 00:05:59,880
users better off directionally? 
Ah. 

133
00:06:00,520 --> 00:06:03,000
OK, directionally. 
Yeah, and the rank test 

134
00:06:03,160 --> 00:06:05,280
mathematically answers that 
directional question. 

135
00:06:06,320 --> 00:06:09,880
It asks if a random draw from 
one group is equally likely to 

136
00:06:09,880 --> 00:06:11,720
exceed a random draw from the 
other. 

137
00:06:11,720 --> 00:06:14,120
Oh, so it's perfectly aligned 
with what the business is 

138
00:06:14,120 --> 00:06:16,640
actually trying to do. 
Rank passes Question one. 

139
00:06:16,840 --> 00:06:20,040
Perfectly so question two, what 
does the metric look like? 

140
00:06:20,280 --> 00:06:23,280
Well, we know it's heavily right
skewed with those crazy power 

141
00:06:23,280 --> 00:06:26,680
users blowing up the variance. 
Right, but rank neutralizes 

142
00:06:26,680 --> 00:06:30,360
those outliers, that 5 hour 
power user, They simply become 

143
00:06:30,360 --> 00:06:33,120
ranked number 10. 
I get it, their massive raw 

144
00:06:33,120 --> 00:06:35,840
minute count no longer destroys 
the variance because they're 

145
00:06:35,840 --> 00:06:38,320
just, you know, one rank above 
the next person. 

146
00:06:38,320 --> 00:06:42,760
Exactly ranks follow a clean, 
roughly uniform distribution. 

147
00:06:42,840 --> 00:06:45,960
That is so elegant. 
So what's question 3 then? 

148
00:06:46,040 --> 00:06:49,600
Question three. 
What approach best reduces noise

149
00:06:49,600 --> 00:06:51,760
while preserving what the 
business cares about? 

150
00:06:52,080 --> 00:06:55,520
Right, because we saw covariate 
adjustment failed to reduce the 

151
00:06:55,520 --> 00:06:59,200
noise and log transforms totally
flip the answer on us, yeah. 

152
00:06:59,680 --> 00:07:02,440
Rank transformation passes 
because it perfectly preserves 

153
00:07:02,440 --> 00:07:05,360
the ordering of users. 
Like if user A stayed longer 

154
00:07:05,360 --> 00:07:08,480
than user B in the raw data, 
they still outrank them in the 

155
00:07:08,480 --> 00:07:11,000
transformed data. 
The ordinal truth is completely 

156
00:07:11,000 --> 00:07:13,360
preserved. 
Exactly, while simultaneously 

157
00:07:13,360 --> 00:07:16,440
killing that outlier noise that 
was blinding our test. 

158
00:07:16,480 --> 00:07:18,880
OK, so that leaves question 4, 
which I have to ask as an 

159
00:07:18,880 --> 00:07:20,680
analyst. 
Does it actually help? 

160
00:07:20,960 --> 00:07:23,520
I need the proof. 
The proof is in the simulations,

161
00:07:23,960 --> 00:07:27,880
the rank transformation reaches 
80% statistical power at just a 

162
00:07:27,880 --> 00:07:29,880
10% effect size. 
Wait, really? 

163
00:07:30,000 --> 00:07:32,880
That matches the extreme 
sensitivity of the log 

164
00:07:32,880 --> 00:07:35,280
transform. 
It does match it, but here is 

165
00:07:35,280 --> 00:07:39,360
the kicker, it achieves this 
high power without inflating the

166
00:07:39,360 --> 00:07:41,920
false positive rate, the Type I 
era. 

167
00:07:41,920 --> 00:07:44,640
Oh wow, so it gives you the 
sensitivity you need without 

168
00:07:44,640 --> 00:07:46,840
tricking you into shipping duds?
Exactly. 

169
00:07:46,840 --> 00:07:49,080
You find the real wins without 
the fake ones. 

170
00:07:49,200 --> 00:07:51,520
That's incredible. 
If I synthesize this whole 

171
00:07:51,520 --> 00:07:54,880
checklist, it really just forces
the analysts to align the math 

172
00:07:54,880 --> 00:07:57,880
directly with stakeholder needs 
before touching any cone. 

173
00:07:58,040 --> 00:07:59,840
You can't just blindly run 
tests. 

174
00:07:59,920 --> 00:08:02,880
You hit the nail on the head, 
and it leads perfectly to this 

175
00:08:02,880 --> 00:08:04,520
final thought to Mull over 
today. 

176
00:08:04,920 --> 00:08:08,000
Most experimentation debates 
focus on which statistical tests

177
00:08:08,000 --> 00:08:11,040
to run, but the teams that ship 
better products focus on what to

178
00:08:11,040 --> 00:08:13,280
measure. 
That is such a powerful 

179
00:08:13,280 --> 00:08:16,160
perspective shift for anyone 
running AB tests. 

180
00:08:16,680 --> 00:08:18,920
Thanks for joining us for 
episode 2 of Inference in the 

181
00:08:18,920 --> 00:08:20,680
Wild. 
Make sure to hit subscribe to 

182
00:08:20,680 --> 00:08:23,320
the Inference and Intelligence 
Lab and we'll see you next time.

