Full transcript
0:00We need to talk about Gemini 3. Now, I
0:03know this video is coming out a little
0:04bit late. I was planning on getting this
0:06out last week, but then I just ended up
0:08playing with Gemini 3 more and more,
0:11as well as GPT 5.1 to try and compare
0:13that since that also came out recently.
0:16And I just found myself
0:18using Gemini 3 a lot, and I wanted to
0:20make sure I gave it a really thorough
0:22review. In order to do that, you really
0:24have to spend some time with it. You
0:25have to spend some time writing with it
0:27and and revising what it gives you for
0:29multiple chapters in my case. Uh because
0:32if you're new here, this chat this
0:33channel is all about writing with AI,
0:37especially for creative writing and what
0:39does well there. So, I'm sure that there
0:42are other use cases for GPT 5.1 or
0:44Gemini 3 that go beyond writing that one
0:48might be better than the other, but on
0:50this channel we're specifically talking
0:51about writing. And so, I spent some time
0:53writing
0:54about 10,000 words in general and
0:57reviewing all of that and going through
0:58it.
0:59And also testing other writing use cases
1:01like brainstorming and outlining and all
1:03of that to see which
1:06you know, how how good do these models
1:08measure up, especially in comparison to
1:10the other models that I was using in the
1:12past for each of the different steps of
1:14the writing process. And I'm going to
1:15walk you through all of that, but to
1:17keep the the long story short, I'm a
1:20little bit in love. Okay? Gemini 3 is an
1:23amazing
1:25platform. GPT 5.1 was not as good and I
1:29don't really see a huge upgrade from 5.1
1:32versus 5.
1:34And I did not end up using GPT 5.1 in
1:36any of the steps that I tested this on.
1:40So, let me just give you a brief
1:42overview of my process for testing
1:44models these days and what that looks
1:45like. First, we'll go to my automations.
1:50So, automations are how I do most of my
1:52writing these days. This is my
1:53automation that takes your outline, your
1:55characters, your world building,
1:56everything that you've developed about
1:58your story, and actually runs through
2:00the outline and writes the book for you
2:02chapter by chapter.
2:04And um
2:05there are four main steps to this
2:07writing
2:09workflow. The first is creating a scene
2:11brief, then it writes the first draft
2:13from that scene brief, then it creates
2:15an improvement plan,
2:16and then it does a rewrite of the
2:18chapter using that improvement plan to
2:20try and just make it a little bit more
2:22polished. In the [snorts] past, I've
2:23been using Sonnet 4.5 for the scene
2:26brief step. I've been using
2:28Opus 4.1 Claude Opus 4.1 for the first
2:31draft,
2:33which is an expensive model.
2:34And then I've been using
2:36Gemini 2.5 Pro for the improvement plan
2:40and the rewrite. Okay? And in my
2:42process, I tested each of these
2:44specifically with the the model that I
2:47have been using tested against Gemini 3
2:50and GPT 5.1. So, let's start with the
2:53scene brief
2:55step and we'll this one is interesting
2:58because I find that a any of the major
3:00models can do okay on this step.
3:03It's not really something that requires
3:06a whole lot of thinking. You think it
3:07would be, but it actually isn't because
3:10most of the information is present
3:12already in all of the outlining
3:14documents and everything. And so, all
3:16the scene brief step does is it kind of
3:18structures all of that information to
3:20collect all of the different things that
3:22the AI needs to know to write the first
3:24draft.
3:25And this is almost just a matter of
3:28summarization of certain things. And
3:30there's not a whole lot of thinking that
3:31needs to happen. And so, because of
3:32that, any of the decent models can do
3:35this this step pretty well. And I did
3:37find that Sonnet 4.5, which I was using
3:40before, Gemini 3, and
3:43um
3:44GPT 5.1 all did really well at this
3:48step.
3:49But I did end up giving the edge to
3:52Gemini 3.
3:54And the reason for this, this is
3:55actually really exciting. I was writing
3:58a scene with Gemini 3 and in the scene
4:01brief stage,
4:03I was, you know, having it create all of
4:05the different beats of the chapter and
4:06everything. And
4:08in [snorts] the story of this particular
4:10chapter,
4:11the main character visits an inn. This
4:14is like a medieval fantasy setting.
4:16And it's wintertime. And they go into
4:18this inn, they get a room, they go up to
4:20their room where
4:22it's no longer warm up there. There's no
4:24fire up there, obviously. Back then,
4:27they wouldn't have had insulation. And
4:29so, naturally, this room would be rather
4:32cold.
4:33And Gemini 3 in the scene brief stage
4:36added this extra detail that I was so
4:39impressed with.
4:40You know, like all of them did okay, but
4:43it was the adding of this detail that
4:45gave me gave the edge to Gemini 3 for
4:47me.
4:48Um
4:49in the the scene, she goes to like this
4:52washbasin and washes her face, right?
4:55It's like part of the scene. And that
4:57part was added by me in the outline. But
4:59what was not added in the outline was
5:01something that Gemini 3 added, which was
5:03the idea that there's a thin layer of
5:05ice over the water in the washbasin, and
5:08she had to crack it in order to get
5:10through to wash her face. And then, of
5:12course, it was freezing cold and all of
5:14this. That was something I did not put
5:16in there.
5:17And it's something that only a model
5:20that has exceptional reasoning would put
5:21in there because
5:23it knows from elsewhere that
5:27it's it's winter.
5:29It knows that
5:31you know, logically, it should know that
5:33this is a medieval setting, so there
5:34wouldn't be things like insulation or
5:36internal heating
5:37unless there's a fire, which there's
5:39not. It put all of those things together
5:41and knew that there would probably the
5:43room would probably be very cold enough
5:45to freeze water.
5:47Not a whole lot, but enough to like
5:49freeze a thin layer of water on top of
5:50the basin. And so, it was little details
5:53like that where you could see the
5:55advanced reasoning power of Gemini 3
5:58that really pushed it over the edge for
6:01me at the scene brief level. But like I
6:03said, the for the scene brief prompt,
6:05it's not really something that you need
6:07a a superpowered thing to do.
6:09Sonnet 4.5 and GPT 5.1 both did fine.
6:13But it wasn't you know, it was just that
6:15little thing. Anyway, so let's get back
6:18to this. All right, so the next step was
6:19the first draft. And this, in my
6:22opinion, is the most important step of
6:24these four because it sort of sets the
6:26tone. If you have a really poorly
6:28written first draft, this improvement
6:30plan and the rewrite step is not going
6:32to be able to fix it that well. It can
6:35improve it just a little bit, but really
6:36these two steps are really there to just
6:39give it a little bit more polish, but it
6:41doesn't really fix everything, you know?
6:44So, really, you need this first draft to
6:46be very well done. And I've been using
6:48Claude Opus 4.1 as my main workhorse of
6:51choice for this step.
6:54And it's very good. It's definitely, in
6:57my opinion, the best model out there for
6:59this particular step.
7:01However, it is one very expensive to
7:05run.
7:06Usually, it costs me between 70 and 90
7:09cents per chapter just for the step.
7:11With everything else combined, it ends
7:13up being a little over a dollar per
7:15chapter that I write. So, if you're
7:17writing a 40-chapter book, that's $40
7:19that you're forking over right there. If
7:21we look on OpenRouter at the actual
7:22costs for
7:25Opus
7:274.1, it's $15 for the input price. For
7:30that's for a million tokens of input and
7:33$75 for the output price on a million
7:36tokens of output. If we go and look up
7:39Gemini 3,
7:41we'll see that
7:43I mean, it depends on the provider a
7:45little bit, but in for the most cases,
7:47it's going to be $2 per million tokens
7:49of input and $12 per million tokens of
7:52output. So, significantly
7:54cheaper, less less than like five times
7:57cheaper
7:58than what Opus gives us, right?
8:02So, that has it's already got that in
8:04its favor, right? But also, I've now
8:07written probably the equivalent of two
8:09books worth of content, two like
8:12full-length books worth of content using
8:144.1 as my main
8:19model for this step. And [clears throat]
8:22as good as it is, I've noticed over time
8:24it starts to just feel the same, like
8:26the very same style. There's not a whole
8:28lot of
8:30um
8:31just difference in the style. Like I
8:33don't know exactly how to describe it,
8:35but it's only something that you will
8:36notice after you've started working with
8:38it for a long time. You're you're fixing
8:41a lot of the same things. And the the
8:43dialogue between two different
8:45characters might have a sort of similar
8:47rhythm to it as it goes along. You
8:49really can only do
8:52to figure it out as you spend more time
8:54with it. And that's why I spent a little
8:56bit of time
8:57testing Gemini 3 to to see if it I had
9:01similar And and to be fair, I think
9:03every model
9:04is going to have some issues like that
9:07that you're just going to pick up on
9:08them the more you use that model.
9:10But I did notice that with 4.1.
9:13This may be a good just as a side note
9:16here, this may be a good instance of um
9:18being able to understand which models
9:20have which kind of tone and being able
9:22to switch between them. For instance, I
9:24find Claude Opus 4.1 is exceptionally
9:27good at really climactic scenes.
9:31Decent for for action scenes as well.
9:35And so, maybe
9:37in in the future, we might be just
9:39picking between models for specific
9:41scenes. Like we know this scene this
9:43one's really good. I want like a sort of
9:45like higher level of tone to it, so I'll
9:50pick Opus 4.1. But then maybe I'll use
9:52Gemini 3 on the love scenes. I don't
9:54know. Like you just
9:56we're going to be learning which of
9:57these do which best cuz they're all
9:59going to be getting pretty good by
10:01themselves, okay? Anyway.
10:03>> [gasps]
10:04>> So I've been using that one
10:06and I switched to Gemini 3 and GPT 5.1
10:10to test out. First of all, let's just
10:13cover GPT 5.1. Um
10:16it it's not bad, but it's definitely not
10:18something I would use very often and it
10:20was exceptionally wordy. Uh it was
10:22spending like way too much time getting
10:24into the scene uh you know, in the past
10:26we had the opposite problem where these
10:28things wouldn't flush out the scene
10:30enough. Now with GPT 5.1, I was finding
10:33it was flushing out the scene too much
10:35to the point where I'm just like, okay,
10:37this doesn't actually make sense to be
10:39talking too much about this potato,
10:41right? Like um it doesn't really have
10:43the common sense to know where it needs
10:46to slow down and when it needs to speed
10:48up and it would just end up being wordy
10:50overall. I didn't like it. Gemini 3,
10:52however,
10:54>> [snorts]
10:56>> was exceptional. Okay? I'm going to give
10:58you a sample of its prose later, but
11:00suffice it to say it did exceptionally
11:03well at this stage. I actually found
11:05that a lot of the issues that I've had
11:06with Opus 4.1, like it being Opus 4.1 is
11:09also a little wordy
11:11and I find myself that the biggest type
11:14of edits that I do to prose written with
11:164.1 is just trimming it, right? I found
11:20I don't need to trim Gemini 3 hardly at
11:22all. Uh there were some other problems
11:24that came up with Gemini 3 which I'll
11:25get into, but um
11:28it was the cleanest prose I've ever
11:31seen.
11:32And
11:34that's why I spent some time working
11:36with it writing several chapters worth
11:37to be kind of sure that it was
11:40like I like it and and that I didn't
11:43come up with any major red flags in that
11:46time processing it. So I'll get back to
11:48the prose in a second. I'll give you
11:49examples, but let's just cover these
11:50other two steps. The improvement plan is
11:53also one I using Gemini 2.5 Pro for this
11:57because I felt like it would had the
11:58most logical
12:00analysis. And I tested this one This is
12:03another one like the scene brief like
12:04they all did okay.
12:06Um I found GPT 5.1 was giving advice
12:09that I was actually like, no, I actually
12:10don't want it to do that. [laughter]
12:13Um and so just looking at it, I was just
12:15like, yeah, I think we'll just upgrade
12:18this from Gemini 2.5 Pro to Gemini 3. I
12:20think it did the best job uh at giving
12:23me logical improvement ideas that
12:26actually made sense and fit what I had
12:28prompted it to do.
12:30And then the rewrite actually oddly
12:32enough
12:33uh I found that the the rewrite, you'd
12:35think this is a high skill task. So you
12:39want to have one of your best models do
12:41it. I've actually found that to be not
12:42the case.
12:43Because the rewrite step, if we if we
12:45look at the prompt for this,
12:47uh it says it's a very short prompt.
12:49It's It's just using the text of the
12:51original chapter and the improvement
12:52plan, I want you to implement the
12:54changes in the improvement plan. So
12:55that's all I'm doing. I'm not asking it
12:56to rewrite the chapter. I'm asking it to
12:59reproduce the chapter with the changes
13:01made
13:02that were in the improvement plan.
13:04And uh on the improvement plan step, I
13:07have it give specific examples of what
13:09to change and how to change it. So it's
13:10actually a really easy thing for AI to
13:12do. It just needs to look at the
13:14original chapter, look at the
13:15improvement plan and just implement the
13:17changes made in the improvement plan. So
13:18I actually found through this testing
13:20that I was probably using too too
13:23powerful of a model with Gemini 2.5 Pro.
13:26And yes,
13:27um Gemini 3 did it just fine,
13:30but so did Claude 4, so did GPT 5.1.
13:33They all did fine and so I actually
13:35tested this with Gemini 2.5 Flash using
13:39a less expensive, less powerful model,
13:41but cheaper.
13:43And Flash did it fine as well. So
13:47I ended up kind of going backwards on
13:49this one and being like, okay, let's
13:50actually use a cheaper model even if
13:52it's less powerful because
13:55um
13:56you don't need need a powerful step on
13:57this one. So you learn things by doing
13:59all this testing.
14:01Um regardless, um
14:03I mean Gemini 3 would probably be
14:05better,
14:06but you don't need it to do better here.
14:09Like this is
14:10it's just implementing what this step
14:14recommended. And Gemini 2.5 Flash is
14:17perfectly capable of doing that. As are
14:19the other
14:20less expensive models out there like GPT
14:225 Mini and Claude 4 Haiku. All of those
14:26do just fine. Anyway.
14:28So let's actually talk about the prose
14:30that I got out of this thing here. So I
14:31did actually pick up on a couple of
14:33quirks of Gemini 3 and I will show off
14:37one of those quirks here. The if we read
14:39this just this little section here from
14:41a book I was working on. Precise,
14:42legible, permanent. I took a deep breath
14:44of stale hot air, dust, and sweat. I sat
14:47there sweating, aching, heart hammering
14:48against the ribs. I leaned forward,
14:50forced the bot forced the body back into
14:52the hunch. I wasn't going home, not yet.
14:55This misery had a file number. Do you
14:57see the problem here?
15:00The main issue that I saw with the prose
15:03that it was giving me and this could be
15:06something it's getting from the style
15:07recommendations I've given in the past.
15:09You'll notice lots and lots of very
15:12short sentences all strung together
15:14one right after another, so it feels
15:16choppy. And this is this is a sign of
15:19bad writing. Short sentences are great,
15:21but they need to be kind of woven
15:22together with longer sentences
15:25to feel more natural.
15:26Uh if you do have more of an action
15:28scene, it can make more sense to have a
15:31a lot more shorter sentences because it
15:33feels more urgent. If there's like a
15:35little bit more anxiety in a lot of
15:37short sentences together. So it can work
15:39for certain instances, but generally
15:41speaking you want a mix of shorter and
15:44longer.
15:45And so I was looking through this and I
15:47thought, well, maybe because Gemini 3 is
15:49pretty smart, maybe if I just change the
15:51prompt and asked it to do more of the
15:54mixing sentences, it would do okay. And
15:56so I did and
15:58you can already see just like visually
16:00here there's a lot less white space
16:03uh and longer paragraphs here.
16:05Um but if we go and read this a little
16:07bit, the Styrofoam cup was hot against
16:09my palm. I squeezed it testing the give
16:11of the cheap material watching the dark
16:13liquid shift there inside. It smelled
16:14like burned beans and stale office air.
16:16It smelled like safety. All right, so
16:18this is a great use of a short sentence
16:20after a long sentence uh or in a
16:23paragraph to kind of add a little bit of
16:26punctuation to that thought.
16:29Um my lower back throbbed with a dull
16:30persistent ache that radiated down into
16:32my hips. This character's pregnant, by
16:33the way. Uh I shifted in the hard
16:36plastic chair trying to find a position
16:37that didn't make my spine feel
16:39compressed by a vice. The baby kicked a
16:41sharp protest against the adrenaline
16:42still flooding my system. Sit down,
16:44Sergio. My husband ignored me. He paced
16:47the length of the small kitchenette, his
16:48sneakers squeaking on the beige
16:50linoleum. Three steps forward, turn.
16:53Three steps back. He rubbed his wrist
16:55where the zip ties had been, the skin
16:56red and irritated. We shouldn't be here,
16:58he muttered. He wasn't looking at me. He
17:00was staring at the reinforced steel door
17:01with the intensity of a trapped animal.
17:03This is the box cat. We walked right
17:05into the box.
17:06Um by the way, I should note that this
17:08has not been edited by me at all. So
17:10obviously I could go through and
17:12probably improve some things.
17:13Um but uh
17:15this will just give you a sense for what
17:17it sounds like out of the box. And this
17:19is after going through those four steps.
17:21Um so this is not just
17:24the writing, but it's also
17:26uh drawing on a scene brief. It's having
17:28gone through that those two editing
17:31steps. So it's pretty polished at this
17:33point. Um ready for a human editor. This
17:36is a class A safe house, I said, my
17:38voice sending
17:39sounded scraping, exhausted. I took a
17:41sip of the coffee. It was bitter and
17:42lukewarm, exactly how it always tasted
17:44at the precinct.
17:45Steel reinforced walls, encrypted comms,
17:48three units paroling
17:50paroling the perimeter. Nobody gets in
17:51here unless Lieutenant Miller buzzes
17:53them in. It's a dead end, Sergio said.
17:55He stopped pacing and turned to me. His
17:56glasses were crooked, sitting askew on
17:58his nose. You saw what happened at the
18:00house. You saw the window. The glass
18:01didn't break, cat. It melted. It
18:03dissolved. Do you think a deadbolt is
18:04going to stop them? I think men with
18:06guns stop men with guns. I set the cup
18:08down on the Formica table with a hollow
18:10thud.
18:11Those things in the woods, that wasn't
18:12magic. It was tech. High-end military
18:14tech. Maybe DARPA, maybe foreign, but
18:16it's physical and if it's physical, the
18:17MPD can handle it. Obviously it's still
18:20AI, right? There's always going to be a
18:23lot of human guidance that is needed to
18:26get AI to a decent level. But I can say
18:29unequivocally
18:32unequivocally
18:34a Gemini 3 is better than any other
18:36model particularly for the cost cuz
18:39Claude 4.5 Opus is actually I I would
18:42almost put them neck and neck with each
18:43other.
18:45Um but in different ways.
18:47But for the cost, absolutely Gemini 3 is
18:50miles above what Claude can deliver.
18:53It's way above what I get from Claude
18:554.5 Sonnet, which is uh
18:58similar in price.
18:59Um and on par if not a little better
19:03with what I get out of Claude 4 Opus or
19:05Claude 4.1 Opus, which is five times
19:08more expensive, more than five times
19:10more expensive.
19:11And so
19:12as far as I'm concerned
19:14I as I was going through my automations,
19:16not this one, but not just this one,
19:18but also my outlining automations and my
19:21brainstorming automations and all of
19:23those other things that I've built,
19:25I found myself replacing the original
19:28model that I was using, whether that was
19:30a Sonnet model or a Gemini model, an
19:32older Gemini model, or a GPT model,
19:35replacing most of them with Gemini 3
19:38because as I was just going through and
19:39testing and looking at the difference
19:41between what the old model would give me
19:43and the new and the Gemini 3,
19:45um I was just finding like, yeah, Gemini
19:473 is better for this, too. And that's
19:50It's concerning to me because I don't
19:51like being so reliant on any one
19:52particular model. I like mixing the
19:54models in different ways and playing to
19:55their strengths, but to be completely
19:57honest, I think Gemini 3 is the best
20:01overall all around model
20:04that we have ever gotten for
20:06creative writing.
20:07That's my personal opinion. I know I
20:09said some good things about GPT-5 back
20:11when it came out, but after working with
20:13it a little bit more, I realized that I
20:15no longer agreed with those things
20:17uh because things just turned up you're
20:19like, "What is going on here?" That's
20:21why I spent more time working with
20:22Gemini 3 to try and see if that remained
20:26the case. And I can happily say that at
20:30least in the amount of testing that I've
20:31done, Gemini 3 is my favorite model on
20:35the market as of the recording of this
20:37video. Those are my thoughts. Hope this
20:39was a useful video for you, and I will
20:40see you in the next one.