1
00:00:00,080 --> 00:00:03,680
So as we predicted last week, 
Open Eye just brought out a new 

2
00:00:03,680 --> 00:00:08,800
language model called GPT 5.2. 
What we described was something 

3
00:00:08,800 --> 00:00:12,560
called GPT Garlic, which is 
their code name for what is now 

4
00:00:12,560 --> 00:00:14,880
called 5.2. 
There's a lot of noise about it,

5
00:00:14,880 --> 00:00:18,760
especially because since writing
and researching this episode and

6
00:00:18,760 --> 00:00:21,760
testing out the new model, they 
brought out another language 

7
00:00:21,760 --> 00:00:24,640
model or a language model 
improvement, which was their 

8
00:00:24,720 --> 00:00:28,000
image model generation on top of
their 5.2 model. 

9
00:00:28,160 --> 00:00:30,080
So for this episode, we're going
to try and keep it short and 

10
00:00:30,080 --> 00:00:33,000
sweet for the final episode of 
2025. 

11
00:00:33,000 --> 00:00:36,040
We're just going to break down 
exactly what it means for you as

12
00:00:36,040 --> 00:00:39,200
an everyday user and the things 
you should be looking out for. 

13
00:00:39,400 --> 00:00:41,320
Now, this isn't a loop with Jack
Horton. 

14
00:00:42,080 --> 00:00:57,480
I hope you enjoyed the show. 
Where to start off, it feels 

15
00:00:57,480 --> 00:01:00,960
like I am dragging myself to the
end of this year. 

16
00:01:00,960 --> 00:01:04,280
It has been an incredible year, 
but it's been incredibly tiring,

17
00:01:04,280 --> 00:01:08,400
especially this last month. 
So I was going to be the last 

18
00:01:08,400 --> 00:01:10,720
episode of the year. 
I'm looking forward to a couple 

19
00:01:10,720 --> 00:01:14,560
weeks break and we'll be coming 
back better and stronger in 

20
00:01:14,560 --> 00:01:18,680
2026. 
So for context, on this episode,

21
00:01:18,680 --> 00:01:22,120
in early December, Sam Altman 
issued what was called a Code 

22
00:01:22,120 --> 00:01:25,400
Red, which is what last week's 
episode really discussed in 

23
00:01:25,400 --> 00:01:27,880
detail. 
And obviously we went through 

24
00:01:27,880 --> 00:01:33,320
the memo he shared and the real 
problems that Open AI really 

25
00:01:33,320 --> 00:01:36,960
face over the next 6/12/24 
months. 

26
00:01:37,080 --> 00:01:40,040
And as I said, we last week 
predicted that they were going 

27
00:01:40,040 --> 00:01:43,760
to bring out a new language 
model and it was codenamed at 

28
00:01:43,760 --> 00:01:48,400
the time GPT Garlic, which is 
now called GPT 5.2. 

29
00:01:59,800 --> 00:02:03,320
And to cut to the chase, really,
I think actually in many ways 

30
00:02:03,320 --> 00:02:07,320
they've matched a lot of 
Gemini's biggest model updates. 

31
00:02:07,480 --> 00:02:11,720
Whether that's something that 
people feel in the press, feel 

32
00:02:11,720 --> 00:02:13,840
in the market, I can't really 
say. 

33
00:02:13,840 --> 00:02:17,080
I don't think so yet. 
But from a benchmark perspective

34
00:02:17,080 --> 00:02:20,880
and an image generation model 
perspective, they've gone above 

35
00:02:20,880 --> 00:02:23,200
and beyond what I to be honest, 
expected. 

36
00:02:23,320 --> 00:02:26,160
So GPT 5.2 comes in three 
versions. 

37
00:02:26,320 --> 00:02:29,760
Instant is their super fast one.
As always, no thinking mode, 

38
00:02:30,480 --> 00:02:33,560
just rapid responses. 
What they've called thinking is 

39
00:02:33,560 --> 00:02:35,920
their standard model. 
As you come to expect from most 

40
00:02:35,920 --> 00:02:38,480
models these days, it has 
reasoning. 

41
00:02:38,480 --> 00:02:40,920
So it pauses, works through the 
problem, and then responds and 

42
00:02:40,920 --> 00:02:45,400
it creates a plan for itself. 
And now you've got Pro extended 

43
00:02:45,400 --> 00:02:47,720
thinking. 
So Matt Schumer, someone online 

44
00:02:47,720 --> 00:02:51,160
that there's a lot of commentary
on the space, had access to this

45
00:02:51,160 --> 00:02:55,880
model since November 25th, 
testing it out and apparently it

46
00:02:55,880 --> 00:02:58,800
was thinking for over an hour on
some hard problems. 

47
00:02:59,800 --> 00:03:03,080
But obviously Pro is only 
available on the $200 a month 

48
00:03:03,200 --> 00:03:06,960
subscription and frustratingly 
for many businesses, it's not 

49
00:03:06,960 --> 00:03:10,400
available via API. 
There's also a new reasoning 

50
00:03:10,400 --> 00:03:13,120
effort setting within the 
thinking mode, so you have 

51
00:03:13,120 --> 00:03:16,000
standard high or extra high, so 
that obviously the higher you 

52
00:03:16,000 --> 00:03:18,400
set it, the longer it thinks and
obviously you'd expect the 

53
00:03:18,400 --> 00:03:21,920
better the output to be. 
From a technical perspective, 

54
00:03:22,880 --> 00:03:26,040
the context window is now 
400,000 tokens. 

55
00:03:26,200 --> 00:03:29,640
To put that into context, GPT 
5.1 had a context window of 

56
00:03:29,800 --> 00:03:33,280
about 128,000. 
So it's a big substantial uplift

57
00:03:33,280 --> 00:03:35,800
here. 
And for those going, what the 

58
00:03:35,800 --> 00:03:38,400
heck is context windows? 
That is the amount of 

59
00:03:38,400 --> 00:03:43,120
information you can give to it 
and it be able to have within 

60
00:03:43,120 --> 00:03:45,000
its context. 
When it answers questions or 

61
00:03:45,000 --> 00:03:49,520
does tasks for you before it 
kind of gets poor and stupid and

62
00:03:49,520 --> 00:03:52,880
annoying as you can often find 
it doing or it just says sorry 

63
00:03:52,880 --> 00:03:56,240
limit reach, move to a new chat.
There's also now an auto setting

64
00:03:56,240 --> 00:04:00,080
that's supposed to and better 
choose between either instant or

65
00:04:00,080 --> 00:04:02,720
extended thinking. 
I'd recommend ignoring it. 

66
00:04:02,720 --> 00:04:05,560
From what I've read online and 
some limited testings I've 

67
00:04:05,560 --> 00:04:07,960
actually got rid of recently my 
ChatGPT license. 

68
00:04:08,080 --> 00:04:10,640
The model often would just 
automatically think for a couple

69
00:04:10,640 --> 00:04:13,200
of seconds and often get an 
answer that is poor or wrong. 

70
00:04:13,560 --> 00:04:16,279
And often you're going to need 
thinking for most professional 

71
00:04:16,279 --> 00:04:17,640
work. 
The second big area of 

72
00:04:17,640 --> 00:04:21,720
improvement is the area of 
professional work and activities

73
00:04:21,720 --> 00:04:23,960
and output. 
So let's talk spreadsheets and 

74
00:04:23,960 --> 00:04:26,640
presentations and things like 
that because this is where 

75
00:04:26,640 --> 00:04:29,520
Opening and Eye have clearly put
a lot of their marketing energy 

76
00:04:29,520 --> 00:04:31,560
into. 
And honestly, there's real 

77
00:04:31,560 --> 00:04:34,560
improvements made here. 
Another early access user called

78
00:04:34,560 --> 00:04:37,360
Simon Smith said this is the 
first time that ChatGPT had made

79
00:04:37,360 --> 00:04:40,160
spreadsheets and presentations 
that were actually presentable. 

80
00:04:40,800 --> 00:04:44,640
You know, on presentations, the 
YouTube reviewer Skill Leap gave

81
00:04:44,640 --> 00:04:47,880
a web link and asked it to 
create a full slideshow. 

82
00:04:47,880 --> 00:04:50,240
And it took 28 minutes. 
But the output for him was 

83
00:04:50,240 --> 00:04:53,480
really, really impressive. 
Really good layouts, information

84
00:04:53,480 --> 00:04:57,280
pulled correctly, slides that 
looked really professional and 

85
00:04:57,280 --> 00:05:00,440
his words were that it was 
shockingly good compared to 5.1 

86
00:05:00,440 --> 00:05:04,560
and another tester threw 10,000 
rows of spreadsheet data and 

87
00:05:04,560 --> 00:05:06,920
told it to make a PowerPoint 
with that and it created a 

88
00:05:06,920 --> 00:05:09,720
really good set of slides. 
For those of you who have to do 

89
00:05:09,720 --> 00:05:12,040
this sort of task all the time, 
this sort of thing is actually 

90
00:05:12,040 --> 00:05:14,640
music to your ears. 
However, there are many 

91
00:05:14,640 --> 00:05:17,320
fantastic tools like Gamma that 
do much the same thing. 

92
00:05:17,480 --> 00:05:20,880
Now open AI is benchmark for 
this is called the GDP Val 

93
00:05:21,280 --> 00:05:24,200
measuring essentially well 
specified knowledge work tasks 

94
00:05:24,200 --> 00:05:28,680
across over 44 occupations. 
And they claim that 5.2 thinking

95
00:05:28,680 --> 00:05:33,320
mode ties on human expertise 
around they claim that 5.2 

96
00:05:33,320 --> 00:05:38,840
thinking mode beats or at least 
ties with human experts on 70.9%

97
00:05:38,840 --> 00:05:43,520
of comparisons, up from 38.8%. 
So imagine it creating a task 

98
00:05:43,720 --> 00:05:47,200
and having experts review it. 
Essentially that's 70% of the 

99
00:05:47,200 --> 00:05:49,640
time it's either better or on 
par with humans. 

100
00:05:50,040 --> 00:05:52,520
Now, I think there's nuances 
here. 

101
00:05:52,920 --> 00:05:56,840
You know, well specified is 
doing a lot of work in that 

102
00:05:56,840 --> 00:05:58,320
sentence. 
You know, it means a model gets 

103
00:05:58,320 --> 00:05:59,680
handed everything it needs 
upfront. 

104
00:05:59,680 --> 00:06:04,000
So super clear instructions, all
the relevant context, define 

105
00:06:04,000 --> 00:06:07,360
success criteria. 
And I think real professional 

106
00:06:07,360 --> 00:06:09,240
work just isn't like that a lot 
of the time. 

107
00:06:10,240 --> 00:06:12,200
You know, you often have to 
figure out what information it 

108
00:06:12,200 --> 00:06:15,520
needs, you need to go find it. 
You need to make good judgement 

109
00:06:15,520 --> 00:06:17,480
calls, you need to make good 
prompting. 

110
00:06:18,240 --> 00:06:22,480
So that benchmark covers, I 
guess, well specified knowledge 

111
00:06:22,480 --> 00:06:25,800
work and the fact that they're 
giving the perfect prompt to do 

112
00:06:25,800 --> 00:06:28,600
that task to the language model,
whereas most people are not 

113
00:06:28,600 --> 00:06:32,160
giving such a well thought out 
structured prompt to a language 

114
00:06:32,160 --> 00:06:35,240
model. 
So yeah, 70.9% doesn't mean that

115
00:06:35,240 --> 00:06:39,000
GPT 5.2 can certainly do 71 
percent of a person's job. 

116
00:06:39,560 --> 00:06:42,480
It means for tasks where 
everything is perfectly 

117
00:06:42,480 --> 00:06:46,160
articulated, defined and 
essentially handed to a model on

118
00:06:46,160 --> 00:06:49,960
a plate, then it will perform at
expert level most of the time. 

119
00:06:50,240 --> 00:06:53,040
A third big areas of improvement
has been cogeneration. 

120
00:06:53,240 --> 00:06:56,160
As I said last last week, 
they're not trying to win in the

121
00:06:56,160 --> 00:06:59,600
code arena clearly. 
That being said, they've still 

122
00:06:59,600 --> 00:07:03,240
made fantastic leaps in coding. 
Maybe they were working on 

123
00:07:03,240 --> 00:07:05,840
making in a better coding model,
realise they'd made a better 

124
00:07:05,840 --> 00:07:09,440
model generally and just chucked
it out there because most of 

125
00:07:09,440 --> 00:07:11,840
their marketing is focused on 
professional work. 

126
00:07:12,040 --> 00:07:17,040
On the SWE Bench Pro, which is a
benchmark testing software 

127
00:07:17,040 --> 00:07:20,240
engineering across basically 
four different languages for 

128
00:07:20,320 --> 00:07:26,200
programming and 5.2 thinking 
scored 55.6% which is really a 

129
00:07:26,200 --> 00:07:28,280
new state-of-the-art. 
On another benchmark that's very

130
00:07:28,280 --> 00:07:33,640
similar on called SWE benchmark 
verified it hit 80% essentially 

131
00:07:33,640 --> 00:07:36,560
matching cords best model which 
was 80.9%. 

132
00:07:36,680 --> 00:07:39,880
Another area of improvement has 
been vision and long context. 

133
00:07:40,280 --> 00:07:44,200
So it's vision capabilities have
improved a lot, you know, on 

134
00:07:44,200 --> 00:07:46,960
chart understanding from 
scientific papers, for example, 

135
00:07:47,240 --> 00:07:51,440
it's accuracy jumped from 80% to
88% on user interface 

136
00:07:51,440 --> 00:07:53,680
understanding. 
So again that kind of agent mode

137
00:07:53,680 --> 00:07:56,560
of reading your screen and 
making clicks, it jumped from 

138
00:07:56,560 --> 00:08:01,040
64% to 86% and error rates have 
been cut in half. 

139
00:08:01,200 --> 00:08:03,280
On context windows, there's been
a massive improvement as I 

140
00:08:03,280 --> 00:08:08,320
mentioned, with 5.1 accuracy 
started to degrade as the amount

141
00:08:08,320 --> 00:08:11,040
of information you had given it 
throughout conversations or 

142
00:08:11,280 --> 00:08:15,960
system prompts for those using 
APIs got a lot longer, around 

143
00:08:15,960 --> 00:08:20,840
90% at about 8000 tokens, 
dropping to under 50% at about 

144
00:08:20,840 --> 00:08:25,400
256,000 tokens. 
Whereas with GPT 5.2, accuracy 

145
00:08:26,120 --> 00:08:30,840
stays at almost 100% across the 
entire context window, even when

146
00:08:30,840 --> 00:08:33,120
it's almost maxed out. 
This is one of the first models 

147
00:08:33,120 --> 00:08:38,039
ever to actually achieve near 
perfect accuracy levels on the 

148
00:08:38,039 --> 00:08:41,760
four needle challenge, which is 
essentially recalling 4 specific

149
00:08:41,760 --> 00:08:45,080
piece of information scattered 
across 200,000 words. 

150
00:08:45,240 --> 00:08:47,560
Another major improvement here 
has been hallucinations. 

151
00:08:47,840 --> 00:08:52,080
So Open AI claims that they've 
reduced hallucinations by 30% 

152
00:08:52,840 --> 00:08:58,280
from 8.8% in 5.1 to 6.2% in 5.2.
However, more independent 

153
00:08:58,280 --> 00:09:01,520
benchmark reviews have, let's 
give, let's say given them more 

154
00:09:01,520 --> 00:09:06,000
modest scoring. 
So Vectara found that GPT 5.2 

155
00:09:06,280 --> 00:09:11,640
had an 8.4% hallucination rate, 
which trails DeepSeek at 6.3%. 

156
00:09:12,320 --> 00:09:14,840
So a massive improvement, but 
still not a leading model. 

157
00:09:15,000 --> 00:09:17,160
Now let's move on to still 
things that are still not very 

158
00:09:17,160 --> 00:09:19,640
good. 
Speed is a still a real problem.

159
00:09:20,120 --> 00:09:23,240
You know, Matt Schumer said the 
standard 5.2 thinking is just 

160
00:09:23,240 --> 00:09:26,200
extremely slow. 
You know, very, very, very slow 

161
00:09:26,200 --> 00:09:28,880
for most questions, even 
straightforward ones, which 

162
00:09:28,880 --> 00:09:31,960
changes how he works. 
So quick questions means that 

163
00:09:32,040 --> 00:09:36,560
he'd go to Claude Opus. 
Deep reasoning now goes to 5.2 

164
00:09:36,560 --> 00:09:39,280
pro, which is quite interesting 
because it used to be or 

165
00:09:39,280 --> 00:09:41,320
certainly has been for me the 
other way around. 

166
00:09:41,440 --> 00:09:44,320
However, for those who do, a lot
of writing quality still lags 

167
00:09:44,320 --> 00:09:47,280
behind Claude. 
Dan Shippers publications called

168
00:09:47,280 --> 00:09:51,320
every ran a load of systematic 
tests and found that Claude Opus

169
00:09:51,320 --> 00:09:56,640
4.5 scored 80% in writing 
quality, whereas GPT 5.2 scored 

170
00:09:56,760 --> 00:10:01,040
74%, So still behind Claude. 
And many, many testers also 

171
00:10:01,040 --> 00:10:04,520
noticed big personality changes.
So Ali Miller, another big 

172
00:10:04,520 --> 00:10:07,880
commentator in the AI space, 
said that a simple question 

173
00:10:07,880 --> 00:10:10,840
turned into 58 bullet points and
numbered points. 

174
00:10:11,360 --> 00:10:15,040
And and then many people have 
been comparing 5.2 to a, a 

175
00:10:15,040 --> 00:10:18,000
brilliant freelancer. 
So yeah, as you know with me, 

176
00:10:18,000 --> 00:10:19,960
I've, I've said over and over 
again that benchmarks really 

177
00:10:19,960 --> 00:10:22,160
are, are not the be all and end 
all. 

178
00:10:23,000 --> 00:10:25,760
You know, comparing models is 
getting much harder and 

179
00:10:26,040 --> 00:10:28,600
performance is scaling very 
iteratively. 

180
00:10:28,800 --> 00:10:31,800
So as always, it's important to 
go and test this thing out and 

181
00:10:31,800 --> 00:10:34,400
see for yourself and experience 
it and see if you like the 

182
00:10:34,400 --> 00:10:47,520
improvements. 
But yeah, to conclude, I think 

183
00:10:47,880 --> 00:10:51,200
opening eyes messaging on this 
release has been very focused, 

184
00:10:51,200 --> 00:10:54,320
which is unusual because often 
they've been we're the best of 

185
00:10:54,320 --> 00:10:57,280
everything for everyone or would
you know that very scatter good 

186
00:10:57,280 --> 00:10:59,720
messaging. 
Now, every single executive and 

187
00:10:59,720 --> 00:11:03,560
interview and media has really 
focused on professional work for

188
00:11:03,560 --> 00:11:06,480
this release and economically 
valuable tasks. 

189
00:11:06,760 --> 00:11:09,440
So they're clearly not trying to
claim AGI breakthroughs, but 

190
00:11:09,440 --> 00:11:13,120
they are trying to win at 
professional work, clearly going

191
00:11:13,120 --> 00:11:15,240
for the enterprise market. 
And obviously, as I said, the 

192
00:11:15,240 --> 00:11:17,400
improvements are very real 
structured outputs. 

193
00:11:17,400 --> 00:11:20,280
It's the most capable model GPT 
has ever produced. 

194
00:11:21,320 --> 00:11:25,600
If your tasks involve slides and
spreadsheets, then 5.2 is a 

195
00:11:25,600 --> 00:11:28,880
serious upgrade. 
But I think if you zoom out, as 

196
00:11:28,880 --> 00:11:31,640
I mentioned, you're seeing 
incremental progress, not 

197
00:11:31,640 --> 00:11:34,320
massive leaps anymore. 
I think some people still hope 

198
00:11:34,320 --> 00:11:37,320
for this big flash of 
inspiration, you know, one model

199
00:11:37,320 --> 00:11:39,640
that conquers it all. 
But what we're starting to see 

200
00:11:39,640 --> 00:11:42,360
is a pattern of that not being 
quite true. 

201
00:11:42,520 --> 00:11:46,320
So 5.2 is a much better tool, 
but it's not a new era. 

202
00:11:46,600 --> 00:11:49,560
And the fact that Open AI is 
marketing better spreadsheets, I

203
00:11:49,560 --> 00:11:53,360
think, tells us a lot about 
where we are with AI right now. 

204
00:11:53,560 --> 00:11:56,480
Anyway, that's it for this week 
and also as I said, for this 

205
00:11:56,480 --> 00:11:58,200
year. 
I hope you find this episode 

206
00:11:58,280 --> 00:12:01,840
interesting and I look forward 
to spending 2026 with you. 

207
00:12:02,760 --> 00:12:04,400
Thank you and see you next year.
