1
00:00:00,040 --> 00:00:03,920
For this week, I am incredibly 
pleased to finally do an episode

2
00:00:03,920 --> 00:00:09,120
which isn't about macroeconomics
and the meta state of AI in the 

3
00:00:09,120 --> 00:00:12,720
world because of potential 
economic bubbles and moving on 

4
00:00:12,720 --> 00:00:15,000
to something that's a bit more 
timely and relevant just this 

5
00:00:15,000 --> 00:00:16,680
once. 
And this week we're going to be 

6
00:00:16,680 --> 00:00:22,040
talking about GPT 5.1. 
So for those that haven't yet 

7
00:00:22,040 --> 00:00:25,200
heard, Open AI just release a 
new language model. 

8
00:00:25,280 --> 00:00:29,720
GBT 5.1 was released three 
months after their big flagship 

9
00:00:29,720 --> 00:00:34,040
release of Five Point O, which 
is their fastest model upgrade 

10
00:00:34,040 --> 00:00:35,760
so far. 
Three months in isn't a very 

11
00:00:35,760 --> 00:00:38,920
long time, and they've released 
it with no benchmark data, 

12
00:00:39,400 --> 00:00:42,440
graphs and charts showing how 
good it is versus other models. 

13
00:00:42,600 --> 00:00:46,360
There's no technical metrics, no
performance metrics, and what 

14
00:00:46,360 --> 00:00:49,320
they're actually doing is now 
selling vibes. 

15
00:00:49,440 --> 00:00:52,480
So yeah, the company that's 
built its reputation on big 

16
00:00:52,480 --> 00:00:57,600
fancy charts and how clever it 
is compared to a PhD student is 

17
00:00:57,600 --> 00:00:59,720
now selling quite literally, 
vibes. 

18
00:00:59,880 --> 00:01:02,520
And I think the shift is going 
to tell us something about where

19
00:01:02,560 --> 00:01:05,600
AI is heading right now. 
And that is going to be the big 

20
00:01:05,600 --> 00:01:07,920
topic of today's episode. 
We're going to be covering the 

21
00:01:07,920 --> 00:01:11,720
five things Chachi BT 5.1 
actually does better than its 

22
00:01:11,720 --> 00:01:13,800
previous model and why it 
matters. 

23
00:01:13,960 --> 00:01:15,480
This is in the loop with Jack 
Horton. 

24
00:01:15,760 --> 00:01:35,030
I hope you enjoy the show. 
So First off, Chachi BT 5.1, 

25
00:01:35,030 --> 00:01:37,870
what does it do better? 
It turns out it actually follows

26
00:01:37,870 --> 00:01:42,750
instructions. 
So if you ask it to respond with

27
00:01:42,750 --> 00:01:46,070
only 6 words, it will actually 
respond in just six words. 

28
00:01:46,070 --> 00:01:49,920
For those that have tried to do 
any form of work with ChatGPT 5,

29
00:01:49,920 --> 00:01:52,040
you may have hit a wall of 
frustration. 

30
00:01:52,040 --> 00:01:53,800
When you ask it to do very 
simple things. 

31
00:01:53,800 --> 00:01:57,120
It says it's going to do that 
and seems to completely ignore 

32
00:01:57,120 --> 00:01:59,400
you. 
Write this thing and make sure 

33
00:01:59,400 --> 00:02:03,040
it has a word limit of say 500 
and it comes out with an 8000 

34
00:02:03,040 --> 00:02:06,200
word essay. 
Stop writing in short sentences 

35
00:02:06,200 --> 00:02:10,080
and it continues to use short 
sentences and open AI releases 

36
00:02:10,080 --> 00:02:12,120
prompting guide. 
So it's basically it's rules for

37
00:02:12,120 --> 00:02:17,720
how to work with ChatGPT 5.1 and
in its prompting guide, it's 

38
00:02:17,720 --> 00:02:20,800
asked the model to make 
assumptions about any missing 

39
00:02:20,800 --> 00:02:24,680
details and do not ask the user 
for any clarification unless 

40
00:02:24,720 --> 00:02:26,880
absolutely necessary. 
So it's going to be very 

41
00:02:26,880 --> 00:02:31,600
decisive rather than trying to 
gather information from you, 

42
00:02:31,960 --> 00:02:35,040
it's going to assume you already
know what you want it to do and 

43
00:02:35,040 --> 00:02:36,960
correctly prompted it to do that
thing. 

44
00:02:37,160 --> 00:02:38,840
So when you're really clear 
about what you want it to do, 

45
00:02:38,840 --> 00:02:42,760
that's going to be fantastic. 
When you're kind of context 

46
00:02:42,760 --> 00:02:44,760
dumping, it's not going to be 
that helpful. 

47
00:02:44,960 --> 00:02:49,440
Number 2 is decisiveness. 
So, so GPT 5.1 is going to try 

48
00:02:49,440 --> 00:02:51,560
and commit to recommendations 
more and more. 

49
00:02:51,720 --> 00:02:55,640
So previous models often hedged.
They said, on the one hand, do 

50
00:02:55,640 --> 00:02:57,360
this, but on the other hand, 
this. 

51
00:02:57,600 --> 00:02:59,960
And now it's just going to try 
and make recommendations much 

52
00:02:59,960 --> 00:03:01,640
faster. 
In fact, the prompting guide by 

53
00:03:01,640 --> 00:03:05,760
Open AI says quite literally 
have a bias for action. 

54
00:03:05,920 --> 00:03:08,400
So there's been lots of 
experiments online on Twitter or

55
00:03:08,400 --> 00:03:12,160
actual I say, and people have 
managed to get it to write 2425 

56
00:03:12,160 --> 00:03:16,520
word reports on strategies for 
improving supply chains just 

57
00:03:16,520 --> 00:03:18,920
from a single prompt. 
Now personally, that's not 

58
00:03:18,920 --> 00:03:22,120
really something I like, but I 
guess many people do. 

59
00:03:22,320 --> 00:03:25,880
And really this is where vibes 
over benchmark starts to get 

60
00:03:25,880 --> 00:03:29,240
quite interesting. 
A company called Surge AI ran an

61
00:03:29,240 --> 00:03:33,920
evaluation and found that in 
that study, 48% of people 

62
00:03:33,920 --> 00:03:38,520
preferred a G PT4-O style 
language model that asked 

63
00:03:38,520 --> 00:03:42,280
questions that didn't just 
assume you knew what you wanted 

64
00:03:42,280 --> 00:03:44,240
and tried to help and guide you 
to the decision. 

65
00:03:45,080 --> 00:03:48,240
And 52% of people wanted 
something that's just super 

66
00:03:48,240 --> 00:03:51,320
decisive and that tells us a lot
about society generally. 

67
00:03:51,760 --> 00:03:54,120
And they've clearly gone for 
that latter group. 

68
00:03:54,160 --> 00:03:56,800
Now you can go to settings and 
actually set personalized 

69
00:03:56,800 --> 00:04:00,320
instructions for your GPT and it
will follow them better. 

70
00:04:00,600 --> 00:04:05,560
But still they've gone for that 
52% of people that love decisive

71
00:04:05,560 --> 00:04:07,920
language models. 
I've one cancelled my license as

72
00:04:07,920 --> 00:04:11,160
a result of it constantly being 
direct and using constant short 

73
00:04:11,160 --> 00:04:14,320
sentences and just being 
infuriating. 

74
00:04:14,520 --> 00:04:17,120
And I actually switched to 
Claude as a result #3 is 

75
00:04:17,120 --> 00:04:20,200
planning. 
So it's now going to be more 

76
00:04:20,200 --> 00:04:24,360
verbose in its answers, which 
says a lot about how frustrating

77
00:04:24,360 --> 00:04:26,600
it was with its incredibly short
responses. 

78
00:04:27,440 --> 00:04:30,960
It's also going to be more 
explicit in how it got to a 

79
00:04:30,960 --> 00:04:32,960
result as well. 
So again, not much of A 

80
00:04:32,960 --> 00:04:36,080
technical improvement, but it's 
helping increase the feeling of 

81
00:04:36,080 --> 00:04:38,000
vibes. 
And I think really that is the 

82
00:04:38,000 --> 00:04:40,240
biggest pattern throughout this 
major releases. 

83
00:04:40,760 --> 00:04:44,760
Their GPT 5 just didn't have the
vibes to attract users. 

84
00:04:44,760 --> 00:04:48,800
I for one, like I, as I said, 
cancelled my membership and I 

85
00:04:48,800 --> 00:04:50,320
don't do things like that very 
lightly. 

86
00:04:50,320 --> 00:04:54,000
I found it infuriating to work 
with #4 is the try to improve 

87
00:04:54,000 --> 00:04:56,800
the writing. 
Now the big question here is 

88
00:04:56,800 --> 00:04:59,440
improve for whom? 
And I'll get on to what I mean 

89
00:04:59,440 --> 00:05:00,280
here. 
There's been a lot of 

90
00:05:00,280 --> 00:05:02,240
independent analysis in this as 
always. 

91
00:05:02,680 --> 00:05:05,960
And on the EQ benchmarks, 
Creative writing leaderboard, 

92
00:05:06,280 --> 00:05:09,840
which is basically a benchmark 
that evaluates a language models

93
00:05:09,840 --> 00:05:12,920
ability to generate characters 
and character development, 

94
00:05:13,200 --> 00:05:16,200
generate emotional engagement 
with readers or create plot 

95
00:05:16,200 --> 00:05:18,040
structures. 
And lots of reviewers have said 

96
00:05:18,040 --> 00:05:20,960
it no longer feels synthetic. 
And many people in the 

97
00:05:20,960 --> 00:05:23,520
communities have been saying 
that it's the first model 

98
00:05:23,520 --> 00:05:27,840
they've been excited to use for 
writing in a very long time. 

99
00:05:28,040 --> 00:05:31,200
Now, Claude, especially Claude 
Sonnet completely beats every of

100
00:05:31,200 --> 00:05:34,840
the model on creative tasks, for
example, in poetry or plot 

101
00:05:34,840 --> 00:05:39,480
development fiction. 
But GPT 5.1 has obviously tried 

102
00:05:39,480 --> 00:05:42,240
to gear towards the writers 
market. 

103
00:05:42,240 --> 00:05:44,880
Again, it's been noted that it's
especially good with technical 

104
00:05:44,880 --> 00:05:47,440
writing because it loves having 
short sentences. 

105
00:05:47,600 --> 00:05:49,360
Now, a little bit of caveat to 
this. 

106
00:05:50,440 --> 00:05:53,560
There's a study that was 
released by Cornell that 

107
00:05:53,560 --> 00:05:56,360
basically revealed something 
quite interesting about what 

108
00:05:56,360 --> 00:05:59,840
happens with language models and
the way it shapes our thought 

109
00:05:59,840 --> 00:06:03,040
process and our interactions as 
a society. 

110
00:06:03,080 --> 00:06:08,800
Researchers gave 118 people, so 
60 from I58 from the US at 

111
00:06:08,800 --> 00:06:13,400
writing tasks designed to 
understand and surface cultural 

112
00:06:13,400 --> 00:06:16,880
expression. 
So things like describe your 

113
00:06:16,880 --> 00:06:19,400
favorite festival or write about
your favorite food. 

114
00:06:20,120 --> 00:06:22,640
And now half wrote with AI 
assistance and the other half 

115
00:06:22,640 --> 00:06:25,360
didn't. 
And what they found was quite 

116
00:06:25,360 --> 00:06:28,640
interesting. 
When Indian participants used AI

117
00:06:28,640 --> 00:06:31,920
writing assistance, their 
writing suddenly shifted towards

118
00:06:32,200 --> 00:06:36,480
a more American cultural norms 
in the way they explained their 

119
00:06:36,480 --> 00:06:39,840
favorite food or festival. 
So a good example here was an 

120
00:06:39,840 --> 00:06:44,680
Indian participant wrote about 
Diwali and in it they wrote we 

121
00:06:44,680 --> 00:06:47,760
worship Goddess Laxmi pot 
crackers and eat sweets. 

122
00:06:48,520 --> 00:06:52,360
Yet with the AI assistance that 
was turned into each additional 

123
00:06:52,360 --> 00:06:56,360
breakfast items and have a day 
filled with happiness and 

124
00:06:56,360 --> 00:06:58,360
warmth. 
So those specific cultural 

125
00:06:58,360 --> 00:07:02,360
references, the goddess Laxmi 
pot crackers, sweets got 

126
00:07:02,360 --> 00:07:05,040
replaced for generic western 
descriptions. 

127
00:07:05,680 --> 00:07:08,720
So traditional Indian breakfast 
items and happiness and warmth 

128
00:07:09,560 --> 00:07:11,960
could be describing anything. 
And obviously AI has been 

129
00:07:11,960 --> 00:07:15,640
trained on a lot of Western 
content, and as a result, it's 

130
00:07:15,640 --> 00:07:20,600
been subtly pushing writing 
words with cultural norms in it.

131
00:07:21,760 --> 00:07:24,840
So I'm always a bit hesitant 
when they try to force changes 

132
00:07:24,840 --> 00:07:27,240
to language models in the way 
that it writes and expresses 

133
00:07:27,240 --> 00:07:31,720
itself because often it shapes 
the way that people express 

134
00:07:31,720 --> 00:07:34,880
themselves online. 
So yeah, when we say GPT 5.1 is 

135
00:07:34,880 --> 00:07:39,520
better than writing, I always 
wonder better for who, because 

136
00:07:39,520 --> 00:07:42,400
obviously it's always optimized 
towards a certain audience. 

137
00:07:42,600 --> 00:07:45,080
But anyway, that's just a small 
gripe with language models 

138
00:07:45,080 --> 00:07:46,920
generally. 
This is exactly the whole point 

139
00:07:46,920 --> 00:07:49,920
with vibes over benchmarks now, 
which is you can't really 

140
00:07:49,920 --> 00:07:53,920
measure writing quality on just 
technical dimensions like 

141
00:07:53,920 --> 00:07:56,360
clarity or structure. 
So yeah, vibes is going to 

142
00:07:56,360 --> 00:07:59,280
become increasingly important 
with all language models moving 

143
00:07:59,280 --> 00:08:01,480
forward. 
And finally, they had a warmer 

144
00:08:01,880 --> 00:08:05,120
personality added to GPT. 
Open AI published something 

145
00:08:05,120 --> 00:08:08,600
called a systems card. 
So essentially this is a safety 

146
00:08:08,600 --> 00:08:11,560
assessment that they publish and
it's a document for all major 

147
00:08:11,560 --> 00:08:15,680
models and buried in that system
card is key information about 

148
00:08:15,680 --> 00:08:20,600
the model and it's safety. 
So apparently GPT 5.1 showed 

149
00:08:20,600 --> 00:08:24,800
safety regression so got worse 
at safety benchmarks across 9 of

150
00:08:24,800 --> 00:08:28,880
13 categories compared to GPT 5.
Specifically, anything to do 

151
00:08:28,880 --> 00:08:32,400
with mental health metrics got 
worse and emotional resilience 

152
00:08:32,400 --> 00:08:35,080
got worse. 
And obviously they knew this 

153
00:08:35,080 --> 00:08:37,840
before they shipped it. 
You know, they documented it and

154
00:08:37,840 --> 00:08:40,880
they shipped it anyway because 
people wanted a model that felt 

155
00:08:40,880 --> 00:08:44,480
more human and warm. 
Now, what does emotional 

156
00:08:44,480 --> 00:08:46,760
resilience metrics got worse 
even mean? 

157
00:08:46,920 --> 00:08:50,680
Well, open eyes data shows that 
nought point nought 7% of weekly

158
00:08:50,680 --> 00:08:55,000
active users, which is about 
560,000 people per week show 

159
00:08:55,000 --> 00:08:59,720
signals of psychosis or mania 
related to their chachi BT use. 

160
00:08:59,720 --> 00:09:04,640
And another 0.15% which is about
1.2 million users shows 

161
00:09:04,640 --> 00:09:08,920
heightened emotional attachment.
So yeah, all those people are 

162
00:09:08,920 --> 00:09:12,200
those that their safety 
benchmarks have regressed in the

163
00:09:12,200 --> 00:09:13,600
people that are particularly 
vulnerable. 

164
00:09:14,240 --> 00:09:16,440
So there are pros and cons 
survives over benchmarks. 

165
00:09:16,600 --> 00:09:19,120
I personally prefer a model that
doesn't just use non-stop short 

166
00:09:19,120 --> 00:09:21,760
sentences like it's a robot. 
And the reason they've done this

167
00:09:21,760 --> 00:09:24,560
is probably quite clear because 
over this last period, their 

168
00:09:24,560 --> 00:09:29,360
enterprise market share fell 
from 50% to 25%, whilst Clawed 

169
00:09:29,600 --> 00:09:33,080
captured 32% of enterprise 
market share. 

170
00:09:33,160 --> 00:09:36,920
Claude code holds 42% versus 
open AI is 21%. 

171
00:09:37,640 --> 00:09:40,720
And most people are choosing 
Claude over open AI. 

172
00:09:40,920 --> 00:09:44,600
But ChatGPT has over 800 million
users who aren't doing big 

173
00:09:44,600 --> 00:09:47,840
evaluations into which package 
to choose. 

174
00:09:47,840 --> 00:09:50,000
They're doing it based on feel 
and vibes. 

175
00:09:50,160 --> 00:09:52,800
And so you could really say that
we've reached a level of 

176
00:09:52,800 --> 00:09:57,240
intelligence where things like 
personality matter more than 

177
00:09:57,240 --> 00:10:00,040
just raw intelligence. 
So yeah, I think Open Air made a

178
00:10:00,040 --> 00:10:03,880
very deliberate choice to 
optimize for what those 800 

179
00:10:03,880 --> 00:10:08,120
million people wanted, which is 
more warmth, personality, the 

180
00:10:08,120 --> 00:10:10,960
ability to follow instructions 
better, for it to be a bit more 

181
00:10:10,960 --> 00:10:12,960
verbose and explain its ideas 
better. 

182
00:10:24,240 --> 00:10:29,080
So yeah, to conclude those those
are the five things that GPT 5.1

183
00:10:29,080 --> 00:10:31,480
does better. 
And really it's quite obvious 

184
00:10:31,480 --> 00:10:33,800
what they're doing. 
It's selling vibes over 

185
00:10:33,800 --> 00:10:35,280
benchmarks. 
And I think we're going to start

186
00:10:35,280 --> 00:10:38,400
to see this as a pattern more 
and more and more. 

187
00:10:38,400 --> 00:10:41,360
We've definitely hit an 
inflection point where technical

188
00:10:41,360 --> 00:10:45,320
capability has basically 
equalised across all the major 

189
00:10:45,320 --> 00:10:47,960
providers and the competition 
has moved to something that's 

190
00:10:47,960 --> 00:10:52,800
much harder to measure. 
How does this thing feel when I 

191
00:10:52,840 --> 00:10:56,280
use it? 
Anyway, that's it for this week.

192
00:10:56,280 --> 00:10:59,120
I hope you enjoyed today's 
episode and I'll see you next 

193
00:10:59,120 --> 00:10:59,440
week.
