1
00:00:00,120 --> 00:00:03,640
This week we have a new language
model announcement, and in fact,

2
00:00:03,640 --> 00:00:07,120
not even just one, but two 
announcements, and those two 

3
00:00:07,120 --> 00:00:10,400
have broken the Internet and 
they're both from Google. 

4
00:00:10,560 --> 00:00:13,360
Now, we know that Google, you 
could argue, has been not 

5
00:00:13,360 --> 00:00:15,960
winning the AI race. 
You could probably argue that 

6
00:00:15,960 --> 00:00:18,280
it's not losing, but it's 
definitely not winning for some 

7
00:00:18,280 --> 00:00:20,760
time. 
But obviously now everything 

8
00:00:20,760 --> 00:00:24,080
just changed to the point where 
Sam Altman has come out to his 

9
00:00:24,080 --> 00:00:27,640
entire company and said that he 
is worried. 

10
00:00:28,120 --> 00:00:31,600
He's worried about slower growth
and huge competition. 

11
00:00:31,720 --> 00:00:35,640
And for context, the Gemini app 
and the Nana Banana applications

12
00:00:35,640 --> 00:00:38,840
have hit #1 on the App Store, 
have generated hundreds of 

13
00:00:38,840 --> 00:00:42,120
millions of images and have gone
completely viral. 

14
00:00:42,320 --> 00:00:45,280
It's scored higher on almost 
every single benchmark than most

15
00:00:45,280 --> 00:00:49,000
language models out there. 
And it can do things that make 

16
00:00:49,200 --> 00:00:54,040
image generation finally truly 
useful for everyday users. 

17
00:00:54,160 --> 00:00:56,040
Anyway, in today's episode, I'm 
going to go through the 

18
00:00:56,080 --> 00:00:59,680
different things you can now do 
with these Gemini models. 

19
00:00:59,800 --> 00:01:02,280
This is in the loop with Jack 
Horton, and I hope you enjoy the

20
00:01:02,280 --> 00:01:16,490
show. 
So let's start with a little bit

21
00:01:16,490 --> 00:01:20,210
of background on this because on
November 18th, Google launched 

22
00:01:20,210 --> 00:01:26,090
Gemini 3 Pro, the main 
competitor model to the Open AI 

23
00:01:26,090 --> 00:01:30,520
Chachi BTS of the world. 
Two days later came Nano Banana 

24
00:01:30,520 --> 00:01:32,280
Pro, which is their image 
generation model. 

25
00:01:32,760 --> 00:01:36,280
Now officially it's called 
Gemini 3 Pro image, but 

26
00:01:36,280 --> 00:01:39,640
everyone's been calling it Nano 
Banana from the previous earlier

27
00:01:39,640 --> 00:01:41,160
versions that have been teased 
online. 

28
00:01:41,160 --> 00:01:45,440
And if you've been online 
recently in the AI community, 

29
00:01:45,760 --> 00:01:49,680
you will know that this model 
has been hyped by the community 

30
00:01:49,680 --> 00:01:51,760
in a way that we've probably 
never seen in a model since 

31
00:01:52,040 --> 00:01:54,600
early Chatchi BT days. 
Now, if you remember the episode

32
00:01:54,600 --> 00:01:58,400
I think five or six weeks ago 
where I talked about the new 

33
00:01:58,400 --> 00:02:02,680
Google Pixel phone having AI in 
it, it was actually overshadowed

34
00:02:02,680 --> 00:02:06,160
in some ways by the fact that 
everybody was hoping that Nano 

35
00:02:06,160 --> 00:02:08,960
Banana was going to be released 
and they didn't release it. 

36
00:02:09,360 --> 00:02:12,480
So that gives you a kind of 
insight into just how much hype 

37
00:02:12,480 --> 00:02:13,840
there's been around this model 
release. 

38
00:02:14,200 --> 00:02:16,120
And I mean, they're really 
building Google here because 

39
00:02:16,120 --> 00:02:19,160
they've got about 650 monthly 
users on their Gemini 

40
00:02:19,160 --> 00:02:22,600
application. 
They have 13 million developers 

41
00:02:22,600 --> 00:02:25,080
now building on top of their 
models, and now both of these 

42
00:02:25,080 --> 00:02:27,280
models they've just released are
available inside the main 

43
00:02:27,280 --> 00:02:31,200
applications and as we speak, 
being rolled out across every 

44
00:02:31,200 --> 00:02:34,240
single product feature you can 
possibly think of. 

45
00:02:34,400 --> 00:02:36,760
Now let's get into all the cool 
things you can now do with these

46
00:02:36,760 --> 00:02:39,960
models. 
OK, so I think of the two models

47
00:02:40,240 --> 00:02:42,800
I want to talk about today. 
So Gemini 3 Pro, which is kind 

48
00:02:42,800 --> 00:02:47,240
of the competitor to Open AI, 
ChatGPT and Cloud, and then the 

49
00:02:47,240 --> 00:02:50,640
Nano Banana, which is the 
competitor to the likes of Saura

50
00:02:51,320 --> 00:02:53,120
and other big image generation 
models. 

51
00:02:53,120 --> 00:02:55,640
I think I want to start with 
Nano Banana because it's 

52
00:02:56,160 --> 00:02:58,800
probably provided some of the 
most exciting new capabilities 

53
00:02:58,800 --> 00:03:14,370
I've seen in a very long time. 
So the first major capability 

54
00:03:14,370 --> 00:03:18,010
it's developed and built on is 
the way it deals with text. 

55
00:03:18,650 --> 00:03:22,090
So it's now supports every 
single language and deals with 

56
00:03:22,090 --> 00:03:25,880
text perfectly. 
So every other image model 

57
00:03:25,880 --> 00:03:27,920
breaks when it comes to text. 
You know? 

58
00:03:27,920 --> 00:03:32,120
Ask Dali or other providers to 
create a product label with 

59
00:03:32,120 --> 00:03:36,720
English, Korean and Arabic. 
You often get a lot of mistakes 

60
00:03:36,720 --> 00:03:41,360
and AI slop mid journey. 
Can't handle non Latin scripts 

61
00:03:41,360 --> 00:03:44,160
at all. 
Chinese characters might come 

62
00:03:44,160 --> 00:03:48,160
out as random strokes, Arabic 
reads right to left incorrectly.

63
00:03:48,160 --> 00:03:51,000
There's just a lot of errors, 
and it's not something that's 

64
00:03:51,000 --> 00:03:53,240
easy to do because there's a lot
of detail there. 

65
00:03:53,400 --> 00:03:55,400
Well, none of Nana completely 
flips that on its head. 

66
00:03:55,400 --> 00:03:57,160
And it enables you to do really 
cool things. 

67
00:03:57,160 --> 00:04:01,280
So first, it can generate 
multilingual text from scratch. 

68
00:04:01,800 --> 00:04:04,600
So someone tested, you know, 
generating product cans with 

69
00:04:04,600 --> 00:04:07,760
text on it across different 
languages, and it worked 

70
00:04:07,760 --> 00:04:10,120
perfectly. 
Second, big leap is the way it 

71
00:04:10,120 --> 00:04:13,640
is able to now translate text 
that's already in images. 

72
00:04:13,640 --> 00:04:17,320
This is genuinely new, you know,
take a photo of a friend's 

73
00:04:17,320 --> 00:04:20,000
restaurant menu and it'll be 
able to translate it to English.

74
00:04:20,519 --> 00:04:23,400
But going even further than 
that, it can match handwriting, 

75
00:04:23,800 --> 00:04:28,200
or it can take information and 
text from one image and generate

76
00:04:28,200 --> 00:04:31,720
it inside another. 
So one of the most amazing 

77
00:04:31,720 --> 00:04:35,880
examples so far I've seen is its
ability to even get this example

78
00:04:35,880 --> 00:04:39,360
up on my phone. 
Here, match your handwriting so 

79
00:04:39,360 --> 00:04:41,600
you can actually write something
on a notepad. 

80
00:04:41,600 --> 00:04:45,840
Imagine writing on a notepad, a 
maths equation or a problem and 

81
00:04:45,840 --> 00:04:48,600
then asking it to replicate that
handwriting but complete the 

82
00:04:48,600 --> 00:04:52,040
equation or write a page of 
notes and it will just do that 

83
00:04:52,040 --> 00:04:54,400
on the fly and it matches your 
handwriting. 

84
00:04:55,320 --> 00:04:57,520
So yeah, the way it deals with 
text is incredible. 

85
00:04:57,680 --> 00:05:00,520
Another one as a result is its 
ability to just generate entire 

86
00:05:00,520 --> 00:05:06,160
infographics and visual ways of 
representing information using 

87
00:05:06,160 --> 00:05:08,080
text. 
Again, this is something that 

88
00:05:08,200 --> 00:05:11,400
models just sucked up previously
and anyone that's tried to do 

89
00:05:11,400 --> 00:05:13,960
this has probably been 
disappointed over and over 

90
00:05:13,960 --> 00:05:16,240
again. 
So you can now just generate 

91
00:05:16,240 --> 00:05:20,040
entire amazing whiteboard and 
infographic images instantly. 

92
00:05:20,200 --> 00:05:23,040
So just imagine how this impacts
marketing teams, for example, 

93
00:05:23,040 --> 00:05:27,120
generating entire marketing 
campaigns for different regions 

94
00:05:27,840 --> 00:05:30,600
at scale. 
You can just imagine Facebook or

95
00:05:30,600 --> 00:05:36,680
meta ads that changes ads for 
not just a country, but a region

96
00:05:36,680 --> 00:05:41,240
or a street or so many different
layers of subculture. 

97
00:05:41,520 --> 00:05:44,400
Now, I, I personally find that 
kind of dark and terrifying, but

98
00:05:44,840 --> 00:05:49,160
you can see where this rolls out
to and how this impacts the way 

99
00:05:49,160 --> 00:05:51,760
things like marketing and 
information sharing go. 

100
00:05:51,920 --> 00:05:55,520
Now, I, for one, find this quite
a almost dark and scary concept.

101
00:05:55,720 --> 00:05:59,400
I always find big advertising 
and using data to generate new 

102
00:05:59,400 --> 00:06:02,480
images to persuade us to buy 
things a little bit. 

103
00:06:02,480 --> 00:06:05,920
Yeah, as I say, dark and evil. 
But you can see how much of an 

104
00:06:05,920 --> 00:06:08,280
impact this language model is 
going to have in that arena. 

105
00:06:08,480 --> 00:06:12,840
So another major upgrade here, 
number 2 is, is its ability to 

106
00:06:12,840 --> 00:06:16,080
maintain image identity across 
many different images. 

107
00:06:16,840 --> 00:06:20,600
So within say 48 hours of 
launch, you suddenly saw images 

108
00:06:20,600 --> 00:06:25,680
of sort of Elon Musk SAT with a 
cigar with the Google CEO, with 

109
00:06:25,680 --> 00:06:28,960
Jensen Twangs, Sam Altman, other
people like that and it looked 

110
00:06:28,960 --> 00:06:31,320
completely real. 
Now a lot of the time other 

111
00:06:31,320 --> 00:06:34,000
image providers, they might be 
able to take lots of single 

112
00:06:34,000 --> 00:06:36,680
images and combine them into a 
single image, but often it's 

113
00:06:36,680 --> 00:06:39,480
just not that good. 
With Nano Banana you can blend 

114
00:06:39,480 --> 00:06:43,720
up to 14 images at once whilst 
maintaining the appearance of up

115
00:06:43,720 --> 00:06:45,800
to five people. 
So now suddenly you've got 

116
00:06:46,640 --> 00:06:49,920
people being generated in 3D. 
People uploaded image selfies 

117
00:06:49,920 --> 00:06:52,800
and turning themselves into 
figurines. 

118
00:06:52,920 --> 00:06:56,080
Now the people from A16Z use 
this for their leadership teams 

119
00:06:56,080 --> 00:06:58,840
infographics. 
So they had five exact headshots

120
00:06:59,080 --> 00:07:01,840
with their actual appearances, 
family photos from individual 

121
00:07:01,840 --> 00:07:05,360
shots, and they blended it 
together with Nana Banana and 

122
00:07:05,360 --> 00:07:08,480
everyone looks perfect. 
But it's now in a completely 

123
00:07:08,480 --> 00:07:10,440
different context. 
That looks super professional, 

124
00:07:10,440 --> 00:07:14,680
but exactly like everybody now 
going on to #3 which really 

125
00:07:14,680 --> 00:07:18,920
blows my mind is how it can 
ground real time data with its 

126
00:07:18,920 --> 00:07:22,600
image generation. 
So previous image models might 

127
00:07:22,600 --> 00:07:24,560
only know what they learned 
during training. 

128
00:07:25,160 --> 00:07:27,840
So you asked Dali to create an 
infographic showing today's 

129
00:07:27,840 --> 00:07:30,440
weather. 
It'll invent A plausible looking

130
00:07:30,440 --> 00:07:33,960
weather icon or image, but it 
won't make any sense because it 

131
00:07:33,960 --> 00:07:36,360
won't be the actual weather. 
Or perhaps give it some 

132
00:07:36,360 --> 00:07:39,240
coordinates of an event that's 
just happened and ask it to 

133
00:07:39,240 --> 00:07:41,360
generate an image of that. 
What that event would have 

134
00:07:41,360 --> 00:07:43,720
looked like, it wouldn't have be
able to do that. 

135
00:07:43,720 --> 00:07:47,160
We'd have no clue. 
Well, Nana Banana now can use 

136
00:07:47,160 --> 00:07:50,640
Google search for injecting that
real time information about 

137
00:07:50,640 --> 00:07:53,640
what's happening in the world 
that second to help it generate 

138
00:07:53,640 --> 00:07:56,520
that information and image. 
So you can say generate me an 

139
00:07:56,520 --> 00:07:59,880
infographic or a visualization 
of an event that happened today.

140
00:07:59,960 --> 00:08:02,280
This search just happens 
automatically within Nano 

141
00:08:02,280 --> 00:08:03,360
banana. 
You don't have to use a 

142
00:08:03,360 --> 00:08:06,800
different model or see. 
I just think about the output of

143
00:08:06,800 --> 00:08:09,440
creative in the world that's 
going to that's going to be 

144
00:08:09,440 --> 00:08:11,120
online as a result. 
I mean, anyone that's been 

145
00:08:11,120 --> 00:08:14,960
online recently, you can just 
see so many Nano Banana images 

146
00:08:14,960 --> 00:08:17,280
and videos out there all the 
time right now. 

147
00:08:18,200 --> 00:08:20,880
And I mean, one of the things 
that I found amazing here is you

148
00:08:20,880 --> 00:08:24,040
can things like you can upload 
GPS coordinates and just ask 

149
00:08:24,040 --> 00:08:27,440
Nano Banana to generate images 
of what that event might have 

150
00:08:27,440 --> 00:08:30,400
looked like, both historic 
events and real time events 

151
00:08:30,400 --> 00:08:32,840
happening today. 
See, it's safe to say that I've 

152
00:08:32,840 --> 00:08:34,799
been blown away by this new 
model update. 

153
00:08:35,000 --> 00:08:38,640
It's really set the tone and 
expectation of things going 

154
00:08:38,640 --> 00:08:54,320
forward. 
Now moving on to Gemini 3 Pro 

155
00:08:54,320 --> 00:08:58,200
and it's upgrading capabilities 
and importantly how good it's 

156
00:08:58,200 --> 00:09:01,200
got at coding. 
So this is the fourth big 

157
00:09:01,200 --> 00:09:05,000
upgrade by Google. 
Here they claim that it's going 

158
00:09:05,000 --> 00:09:08,560
to have multi hour autonomous 
coding sessions because previous

159
00:09:08,560 --> 00:09:12,400
coding assistants, you know, 
like Cloud Code or others might 

160
00:09:12,400 --> 00:09:15,600
run for 30 minutes, 40 minutes, 
45 minutes without losing 

161
00:09:15,600 --> 00:09:19,280
context. 
Cursor can handle maybe an hour 

162
00:09:19,280 --> 00:09:21,920
of continuous work before 
needing to be restarted or 

163
00:09:21,920 --> 00:09:25,440
readjusted. 
Or they hit token limits, lose 

164
00:09:25,440 --> 00:09:27,480
track of what they're doing, or 
just start to make a lot of 

165
00:09:27,480 --> 00:09:29,280
errors without any human 
intervention. 

166
00:09:29,440 --> 00:09:33,600
Well, Google's what they call 
anti gravity agents apparently 

167
00:09:33,600 --> 00:09:36,240
can work for over 24 hours 
continuously. 

168
00:09:36,440 --> 00:09:40,280
So Google's documentation says 
in our internal evaluations we 

169
00:09:40,280 --> 00:09:44,560
observe Codex Max works on tasks
for more than 24 hours. 

170
00:09:45,240 --> 00:09:48,880
Now apparently it uses something
called compaction, which is a 

171
00:09:48,880 --> 00:09:51,520
technique that essentially 
prunes its history whilst 

172
00:09:52,080 --> 00:09:54,200
preserving contacts. 
So if you imagine a model going 

173
00:09:54,200 --> 00:09:57,120
out there and doing things, it's
constantly updating the 

174
00:09:57,120 --> 00:10:00,520
information about that task. 
Now if you can imagine you 

175
00:10:00,520 --> 00:10:04,120
working for three hours with 
this agent, it suddenly 

176
00:10:04,120 --> 00:10:07,840
collected 8 hours worth of 
information, contacts, things 

177
00:10:07,840 --> 00:10:09,240
that you said you do and don't 
like. 

178
00:10:09,240 --> 00:10:12,600
And after a while it's trying to
take too much in every time you 

179
00:10:12,600 --> 00:10:16,200
want it to do a new thing and so
it's performance can drop. 

180
00:10:16,320 --> 00:10:20,080
This innovation essentially is 
auto reducing, improving and 

181
00:10:20,080 --> 00:10:23,240
compressing the amount of 
context it has so it can be more

182
00:10:23,240 --> 00:10:26,040
efficient and therefore go on 
for a longer and longer amount 

183
00:10:26,040 --> 00:10:27,960
of time. 
So this means that Google claims

184
00:10:27,960 --> 00:10:30,920
that you could refactor an 
entire code base that might take

185
00:10:30,920 --> 00:10:34,320
weeks of time with a single 
prompt. 

186
00:10:34,320 --> 00:10:37,320
Now whether that's true or not, 
we'll see when we start to test 

187
00:10:37,320 --> 00:10:39,880
it, and I'm sure the team 
internally at Mindset will be 

188
00:10:39,880 --> 00:10:42,840
testing this very soon as well. 
So moving on to the fifth big 

189
00:10:42,840 --> 00:10:46,920
update by Google, you can now 
build complete games, which 

190
00:10:46,920 --> 00:10:49,480
although might not sound 
exciting, these games can have 

191
00:10:49,480 --> 00:10:52,600
complete physics and audio 
instantly. 

192
00:10:52,800 --> 00:10:55,720
So previous models could 
generate a game quite simply, 

193
00:10:55,720 --> 00:10:57,400
but it doesn't actually always 
work. 

194
00:10:57,760 --> 00:11:00,880
And you know, you could ask 
Claude or GBT to build a game 

195
00:11:00,880 --> 00:11:02,680
with physics. 
You might get code that's quite 

196
00:11:02,680 --> 00:11:05,040
plausible, a very simple 
prototype that will sometimes 

197
00:11:05,040 --> 00:11:07,520
will work OK, but often has a 
lot of bugs in it. 

198
00:11:07,680 --> 00:11:11,440
Well, Gemini 3 in their demos 
built entire games like 

199
00:11:11,440 --> 00:11:15,640
Ridiculous Fishing, the iOS game
from 2013, which had working 

200
00:11:15,640 --> 00:11:20,320
mechanics, sound effects and 
music like the full game, and it

201
00:11:20,320 --> 00:11:22,560
had accurate physics for 
catching fish. 

202
00:11:22,760 --> 00:11:25,000
And previous models really 
struggled with this types of 

203
00:11:25,000 --> 00:11:29,400
activity, whereas Gemini 3 can 
do it in just a few prompts. 

204
00:11:29,520 --> 00:11:31,760
Again, devil's in the detail 
here, and we'll be really 

205
00:11:31,760 --> 00:11:36,240
interested to see at what point 
you get stretched to its limits.

206
00:11:36,360 --> 00:11:41,240
OK, moving to number six, which 
is the ability for Gemini to 

207
00:11:41,240 --> 00:11:44,240
read very complex screenshots 
and take action on them. 

208
00:11:44,400 --> 00:11:47,560
So previous models are really 
good at describing what might be

209
00:11:47,560 --> 00:11:49,480
in a screenshot, but they can't 
interact with it. 

210
00:11:49,680 --> 00:11:53,200
So if you give it access to your
web browser, it can take 

211
00:11:53,200 --> 00:11:56,960
screenshots of things maybe. 
So this is agent mode, but it's 

212
00:11:56,960 --> 00:11:59,880
not very good at figuring out, 
you know, I need to press this 

213
00:11:59,880 --> 00:12:01,720
button and that button to do 
this thing. 

214
00:12:02,000 --> 00:12:04,920
So it could tell you, you know, 
the information on screen, but 

215
00:12:04,920 --> 00:12:06,920
not really good at figuring out 
where to press. 

216
00:12:07,080 --> 00:12:10,320
And that's why I think agent 
mode and autonomous agents, they

217
00:12:10,320 --> 00:12:14,120
get branded online have not been
hugely successful or popular. 

218
00:12:14,240 --> 00:12:18,440
Well, Gemini 3 has got 72% on 
the screenshot protest, which is

219
00:12:18,440 --> 00:12:21,160
double the previous best score 
of 36%. 

220
00:12:21,280 --> 00:12:24,040
And this test is only testing a 
model, whether it can look at a 

221
00:12:24,040 --> 00:12:28,080
screenshot, understand the UI it
sees and figure out where to 

222
00:12:28,080 --> 00:12:29,840
click and take the right action 
quickly. 

223
00:12:29,880 --> 00:12:32,720
And really interesting ways this
has been used is, for example, 

224
00:12:32,720 --> 00:12:37,000
giving it access to 3 1/2 hours 
of footage from a meeting and 

225
00:12:37,000 --> 00:12:40,200
getting it to produce 
transcripts, speaker names, time

226
00:12:40,200 --> 00:12:45,200
stamps, key sections, action 
items, all those types of 

227
00:12:45,280 --> 00:12:48,800
important parts, even where 
people disagreed and agreed, but

228
00:12:48,800 --> 00:12:51,640
doing so by watching the actual 
video and the actual footage as 

229
00:12:51,640 --> 00:12:53,720
well. 
So the type of innovations is my

230
00:12:53,720 --> 00:12:57,120
unlock is the ability for agents
to navigate software themselves.

231
00:12:57,400 --> 00:13:00,400
Again, this isn't an area that 
I've been particularly excited 

232
00:13:00,400 --> 00:13:03,400
by within AI because I think AP 
is so just, you know, 

233
00:13:03,440 --> 00:13:07,000
transferring information through
the back end is much quicker and

234
00:13:07,000 --> 00:13:09,920
easier. 
But it wouldn't surprise me if 

235
00:13:09,920 --> 00:13:12,040
there were really some 
interesting use cases emerging 

236
00:13:12,040 --> 00:13:15,520
because of this over time. 
You know, the personal shopping 

237
00:13:15,520 --> 00:13:20,200
assistance, that type of idea #8
is the ability for Gemini to use

238
00:13:20,360 --> 00:13:24,760
multiple valid approaches to 
figure out a solution to a 

239
00:13:24,760 --> 00:13:27,240
problem. 
So often a language model can 

240
00:13:27,240 --> 00:13:31,080
struggle when there's no single 
correct answer to a problem, 

241
00:13:31,360 --> 00:13:33,800
which is most problems in our 
world because it's a very 

242
00:13:33,800 --> 00:13:37,840
complex world. 
You know, ask GPT 5.1 for, you 

243
00:13:37,840 --> 00:13:40,240
know, three different approaches
to reduce our customer churn. 

244
00:13:40,320 --> 00:13:43,400
You'd often get different 
variations of the same, same or 

245
00:13:43,400 --> 00:13:45,840
similar approach. 
Ultimately, the model is pattern

246
00:13:45,840 --> 00:13:49,920
matching to common solutions 
rather than exploring very 

247
00:13:49,920 --> 00:13:53,360
different strategies. 
Now again, think about much of 

248
00:13:53,360 --> 00:13:56,040
these innovations not being 
suddenly made possible because 

249
00:13:56,040 --> 00:13:59,360
you can make a model do these 
things, but doing so within a 

250
00:13:59,360 --> 00:14:01,400
single prompt or just a couple 
of prompts. 

251
00:14:01,720 --> 00:14:04,800
So the speed to this raw 
intelligence and the ease to 

252
00:14:05,320 --> 00:14:08,800
access and utilize this raw 
intelligence in other models, 

253
00:14:08,800 --> 00:14:11,760
this might take a lot of, you 
know, twisting its arm, making 

254
00:14:11,760 --> 00:14:13,720
sure it's got the right 
information and context. 

255
00:14:14,000 --> 00:14:18,560
Whereas what Gemini is claiming 
to do is make this 10 times 

256
00:14:18,560 --> 00:14:20,600
easier. 
And a good example is Gemini 

257
00:14:20,880 --> 00:14:26,000
scored 45.1% on the ARC AGI two 
test, which is designed to have 

258
00:14:26,000 --> 00:14:29,480
multiple potential solutions. 
In this test, it shows you a few

259
00:14:29,480 --> 00:14:32,720
different examples of patterns 
and we'll ask the model to infer

260
00:14:32,720 --> 00:14:36,400
the rule, for example, with more
often than not more than a 

261
00:14:36,400 --> 00:14:41,840
single valid interpretation. 
PPT scored 17.6, Gemini 2.5 

262
00:14:41,840 --> 00:14:45,320
scored 4.9%. 
And again, as I said, Gemini 

263
00:14:45,320 --> 00:14:49,120
just scored, the Gemini 3 model 
just scored 45.1%. 

264
00:14:49,320 --> 00:14:52,360
That's that's a massive leap. 
And I think Matt Schumer said 

265
00:14:52,360 --> 00:14:56,600
this perfectly, which is he said
previous models often had a 

266
00:14:56,600 --> 00:14:59,200
certain spikiness. 
Their quality varied wildly 

267
00:14:59,200 --> 00:15:02,560
depending on different tasks. 
You could get brilliant at 1 

268
00:15:02,560 --> 00:15:06,920
task, but just OK results on 
another, whereas Gemini 3 feels 

269
00:15:06,920 --> 00:15:10,600
just more consistent across the 
board and when a model is 

270
00:15:10,880 --> 00:15:14,040
reliably good at most things 
instead of occasionally 

271
00:15:14,040 --> 00:15:17,520
brilliant, that's where you 
build entire systems around it. 

272
00:15:28,860 --> 00:15:31,420
So yeah, let's let's conclude 
for today's episode I mean, so 

273
00:15:31,420 --> 00:15:32,620
what's actually changed this 
week? 

274
00:15:33,380 --> 00:15:36,020
Well, Google has just released 2
models that have broken the 

275
00:15:36,020 --> 00:15:41,320
Internet and they've broken 
thresholds that people have 

276
00:15:41,320 --> 00:15:43,480
really struggled to break for a 
very long time. 

277
00:15:44,000 --> 00:15:47,760
Nana Banana now can merge 14 
images perfectly. 

278
00:15:48,200 --> 00:15:52,600
It can generate text perfectly. 
Gemini 3 can apparently do a 

279
00:15:52,600 --> 00:15:56,520
single task within a couple of 
prompts for over 24 hours. 

280
00:15:56,640 --> 00:16:00,000
CSA to say Google is super 
competitive again. 

281
00:16:00,880 --> 00:16:04,600
And as I said at the beginning, 
Sam Altman announced to open AI 

282
00:16:04,600 --> 00:16:09,560
that he is worried, and I mean I
would be too because Google is a

283
00:16:10,400 --> 00:16:13,880
huge Goliath as incredibly 
talented people and is now 

284
00:16:13,880 --> 00:16:18,040
making huge leaps into the 
future and being a market leader

285
00:16:18,040 --> 00:16:19,880
provider. 
Anyway, as always, the most 

286
00:16:19,880 --> 00:16:22,920
important thing I can now 
recommend is go try it for 

287
00:16:22,920 --> 00:16:25,720
yourself. 
Go play, go have fun, see if it 

288
00:16:25,720 --> 00:16:29,000
works in your day-to-day work. 
Anyway, that's it for today. 

289
00:16:29,000 --> 00:16:32,320
I hope you enjoyed the episode 
and I'll see you next week.

