Free YouTube Transcribe

Video transcript

How Well Do Gemini 3 and GPT 5.1 Write?

The Nerdy Novelist · 3,815 words · 18 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00We need to talk about Gemini 3. Now, I

0:03know this video is coming out a little

0:04bit late. I was planning on getting this

0:06out last week, but then I just ended up

0:08playing with Gemini 3 more and more,

0:11as well as GPT 5.1 to try and compare

0:13that since that also came out recently.

0:16And I just found myself

0:18using Gemini 3 a lot, and I wanted to

0:20make sure I gave it a really thorough

0:22review. In order to do that, you really

0:24have to spend some time with it. You

0:25have to spend some time writing with it

0:27and and revising what it gives you for

0:29multiple chapters in my case. Uh because

0:32if you're new here, this chat this

0:33channel is all about writing with AI,

0:37especially for creative writing and what

0:39does well there. So, I'm sure that there

0:42are other use cases for GPT 5.1 or

0:44Gemini 3 that go beyond writing that one

0:48might be better than the other, but on

0:50this channel we're specifically talking

0:51about writing. And so, I spent some time

0:53writing

0:54about 10,000 words in general and

0:57reviewing all of that and going through

0:58it.

0:59And also testing other writing use cases

1:01like brainstorming and outlining and all

1:03of that to see which

1:06you know, how how good do these models

1:08measure up, especially in comparison to

1:10the other models that I was using in the

1:12past for each of the different steps of

1:14the writing process. And I'm going to

1:15walk you through all of that, but to

1:17keep the the long story short, I'm a

1:20little bit in love. Okay? Gemini 3 is an

1:23amazing

1:25platform. GPT 5.1 was not as good and I

1:29don't really see a huge upgrade from 5.1

1:32versus 5.

1:34And I did not end up using GPT 5.1 in

1:36any of the steps that I tested this on.

1:40So, let me just give you a brief

1:42overview of my process for testing

1:44models these days and what that looks

1:45like. First, we'll go to my automations.

1:50So, automations are how I do most of my

1:52writing these days. This is my

1:53automation that takes your outline, your

1:55characters, your world building,

1:56everything that you've developed about

1:58your story, and actually runs through

2:00the outline and writes the book for you

2:02chapter by chapter.

2:04And um

2:05there are four main steps to this

2:07writing

2:09workflow. The first is creating a scene

2:11brief, then it writes the first draft

2:13from that scene brief, then it creates

2:15an improvement plan,

2:16and then it does a rewrite of the

2:18chapter using that improvement plan to

2:20try and just make it a little bit more

2:22polished. In the [snorts] past, I've

2:23been using Sonnet 4.5 for the scene

2:26brief step. I've been using

2:28Opus 4.1 Claude Opus 4.1 for the first

2:31draft,

2:33which is an expensive model.

2:34And then I've been using

2:36Gemini 2.5 Pro for the improvement plan

2:40and the rewrite. Okay? And in my

2:42process, I tested each of these

2:44specifically with the the model that I

2:47have been using tested against Gemini 3

2:50and GPT 5.1. So, let's start with the

2:53scene brief

2:55step and we'll this one is interesting

2:58because I find that a any of the major

3:00models can do okay on this step.

3:03It's not really something that requires

3:06a whole lot of thinking. You think it

3:07would be, but it actually isn't because

3:10most of the information is present

3:12already in all of the outlining

3:14documents and everything. And so, all

3:16the scene brief step does is it kind of

3:18structures all of that information to

3:20collect all of the different things that

3:22the AI needs to know to write the first

3:24draft.

3:25And this is almost just a matter of

3:28summarization of certain things. And

3:30there's not a whole lot of thinking that

3:31needs to happen. And so, because of

3:32that, any of the decent models can do

3:35this this step pretty well. And I did

3:37find that Sonnet 4.5, which I was using

3:40before, Gemini 3, and

3:43um

3:44GPT 5.1 all did really well at this

3:48step.

3:49But I did end up giving the edge to

3:52Gemini 3.

3:54And the reason for this, this is

3:55actually really exciting. I was writing

3:58a scene with Gemini 3 and in the scene

4:01brief stage,

4:03I was, you know, having it create all of

4:05the different beats of the chapter and

4:06everything. And

4:08in [snorts] the story of this particular

4:10chapter,

4:11the main character visits an inn. This

4:14is like a medieval fantasy setting.

4:16And it's wintertime. And they go into

4:18this inn, they get a room, they go up to

4:20their room where

4:22it's no longer warm up there. There's no

4:24fire up there, obviously. Back then,

4:27they wouldn't have had insulation. And

4:29so, naturally, this room would be rather

4:32cold.

4:33And Gemini 3 in the scene brief stage

4:36added this extra detail that I was so

4:39impressed with.

4:40You know, like all of them did okay, but

4:43it was the adding of this detail that

4:45gave me gave the edge to Gemini 3 for

4:47me.

4:48Um

4:49in the the scene, she goes to like this

4:52washbasin and washes her face, right?

4:55It's like part of the scene. And that

4:57part was added by me in the outline. But

4:59what was not added in the outline was

5:01something that Gemini 3 added, which was

5:03the idea that there's a thin layer of

5:05ice over the water in the washbasin, and

5:08she had to crack it in order to get

5:10through to wash her face. And then, of

5:12course, it was freezing cold and all of

5:14this. That was something I did not put

5:16in there.

5:17And it's something that only a model

5:20that has exceptional reasoning would put

5:21in there because

5:23it knows from elsewhere that

5:27it's it's winter.

5:29It knows that

5:31you know, logically, it should know that

5:33this is a medieval setting, so there

5:34wouldn't be things like insulation or

5:36internal heating

5:37unless there's a fire, which there's

5:39not. It put all of those things together

5:41and knew that there would probably the

5:43room would probably be very cold enough

5:45to freeze water.

5:47Not a whole lot, but enough to like

5:49freeze a thin layer of water on top of

5:50the basin. And so, it was little details

5:53like that where you could see the

5:55advanced reasoning power of Gemini 3

5:58that really pushed it over the edge for

6:01me at the scene brief level. But like I

6:03said, the for the scene brief prompt,

6:05it's not really something that you need

6:07a a superpowered thing to do.

6:09Sonnet 4.5 and GPT 5.1 both did fine.

6:13But it wasn't you know, it was just that

6:15little thing. Anyway, so let's get back

6:18to this. All right, so the next step was

6:19the first draft. And this, in my

6:22opinion, is the most important step of

6:24these four because it sort of sets the

6:26tone. If you have a really poorly

6:28written first draft, this improvement

6:30plan and the rewrite step is not going

6:32to be able to fix it that well. It can

6:35improve it just a little bit, but really

6:36these two steps are really there to just

6:39give it a little bit more polish, but it

6:41doesn't really fix everything, you know?

6:44So, really, you need this first draft to

6:46be very well done. And I've been using

6:48Claude Opus 4.1 as my main workhorse of

6:51choice for this step.

6:54And it's very good. It's definitely, in

6:57my opinion, the best model out there for

6:59this particular step.

7:01However, it is one very expensive to

7:05run.

7:06Usually, it costs me between 70 and 90

7:09cents per chapter just for the step.

7:11With everything else combined, it ends

7:13up being a little over a dollar per

7:15chapter that I write. So, if you're

7:17writing a 40-chapter book, that's $40

7:19that you're forking over right there. If

7:21we look on OpenRouter at the actual

7:22costs for

7:25Opus

7:274.1, it's $15 for the input price. For

7:30that's for a million tokens of input and

7:33$75 for the output price on a million

7:36tokens of output. If we go and look up

7:39Gemini 3,

7:41we'll see that

7:43I mean, it depends on the provider a

7:45little bit, but in for the most cases,

7:47it's going to be $2 per million tokens

7:49of input and $12 per million tokens of

7:52output. So, significantly

7:54cheaper, less less than like five times

7:57cheaper

7:58than what Opus gives us, right?

8:02So, that has it's already got that in

8:04its favor, right? But also, I've now

8:07written probably the equivalent of two

8:09books worth of content, two like

8:12full-length books worth of content using

8:144.1 as my main

8:19model for this step. And [clears throat]

8:22as good as it is, I've noticed over time

8:24it starts to just feel the same, like

8:26the very same style. There's not a whole

8:28lot of

8:30um

8:31just difference in the style. Like I

8:33don't know exactly how to describe it,

8:35but it's only something that you will

8:36notice after you've started working with

8:38it for a long time. You're you're fixing

8:41a lot of the same things. And the the

8:43dialogue between two different

8:45characters might have a sort of similar

8:47rhythm to it as it goes along. You

8:49really can only do

8:52to figure it out as you spend more time

8:54with it. And that's why I spent a little

8:56bit of time

8:57testing Gemini 3 to to see if it I had

9:01similar And and to be fair, I think

9:03every model

9:04is going to have some issues like that

9:07that you're just going to pick up on

9:08them the more you use that model.

9:10But I did notice that with 4.1.

9:13This may be a good just as a side note

9:16here, this may be a good instance of um

9:18being able to understand which models

9:20have which kind of tone and being able

9:22to switch between them. For instance, I

9:24find Claude Opus 4.1 is exceptionally

9:27good at really climactic scenes.

9:31Decent for for action scenes as well.

9:35And so, maybe

9:37in in the future, we might be just

9:39picking between models for specific

9:41scenes. Like we know this scene this

9:43one's really good. I want like a sort of

9:45like higher level of tone to it, so I'll

9:50pick Opus 4.1. But then maybe I'll use

9:52Gemini 3 on the love scenes. I don't

9:54know. Like you just

9:56we're going to be learning which of

9:57these do which best cuz they're all

9:59going to be getting pretty good by

10:01themselves, okay? Anyway.

10:03>> [gasps]

10:04>> So I've been using that one

10:06and I switched to Gemini 3 and GPT 5.1

10:10to test out. First of all, let's just

10:13cover GPT 5.1. Um

10:16it it's not bad, but it's definitely not

10:18something I would use very often and it

10:20was exceptionally wordy. Uh it was

10:22spending like way too much time getting

10:24into the scene uh you know, in the past

10:26we had the opposite problem where these

10:28things wouldn't flush out the scene

10:30enough. Now with GPT 5.1, I was finding

10:33it was flushing out the scene too much

10:35to the point where I'm just like, okay,

10:37this doesn't actually make sense to be

10:39talking too much about this potato,

10:41right? Like um it doesn't really have

10:43the common sense to know where it needs

10:46to slow down and when it needs to speed

10:48up and it would just end up being wordy

10:50overall. I didn't like it. Gemini 3,

10:52however,

10:54>> [snorts]

10:56>> was exceptional. Okay? I'm going to give

10:58you a sample of its prose later, but

11:00suffice it to say it did exceptionally

11:03well at this stage. I actually found

11:05that a lot of the issues that I've had

11:06with Opus 4.1, like it being Opus 4.1 is

11:09also a little wordy

11:11and I find myself that the biggest type

11:14of edits that I do to prose written with

11:164.1 is just trimming it, right? I found

11:20I don't need to trim Gemini 3 hardly at

11:22all. Uh there were some other problems

11:24that came up with Gemini 3 which I'll

11:25get into, but um

11:28it was the cleanest prose I've ever

11:31seen.

11:32And

11:34that's why I spent some time working

11:36with it writing several chapters worth

11:37to be kind of sure that it was

11:40like I like it and and that I didn't

11:43come up with any major red flags in that

11:46time processing it. So I'll get back to

11:48the prose in a second. I'll give you

11:49examples, but let's just cover these

11:50other two steps. The improvement plan is

11:53also one I using Gemini 2.5 Pro for this

11:57because I felt like it would had the

11:58most logical

12:00analysis. And I tested this one This is

12:03another one like the scene brief like

12:04they all did okay.

12:06Um I found GPT 5.1 was giving advice

12:09that I was actually like, no, I actually

12:10don't want it to do that. [laughter]

12:13Um and so just looking at it, I was just

12:15like, yeah, I think we'll just upgrade

12:18this from Gemini 2.5 Pro to Gemini 3. I

12:20think it did the best job uh at giving

12:23me logical improvement ideas that

12:26actually made sense and fit what I had

12:28prompted it to do.

12:30And then the rewrite actually oddly

12:32enough

12:33uh I found that the the rewrite, you'd

12:35think this is a high skill task. So you

12:39want to have one of your best models do

12:41it. I've actually found that to be not

12:42the case.

12:43Because the rewrite step, if we if we

12:45look at the prompt for this,

12:47uh it says it's a very short prompt.

12:49It's It's just using the text of the

12:51original chapter and the improvement

12:52plan, I want you to implement the

12:54changes in the improvement plan. So

12:55that's all I'm doing. I'm not asking it

12:56to rewrite the chapter. I'm asking it to

12:59reproduce the chapter with the changes

13:01made

13:02that were in the improvement plan.

13:04And uh on the improvement plan step, I

13:07have it give specific examples of what

13:09to change and how to change it. So it's

13:10actually a really easy thing for AI to

13:12do. It just needs to look at the

13:14original chapter, look at the

13:15improvement plan and just implement the

13:17changes made in the improvement plan. So

13:18I actually found through this testing

13:20that I was probably using too too

13:23powerful of a model with Gemini 2.5 Pro.

13:26And yes,

13:27um Gemini 3 did it just fine,

13:30but so did Claude 4, so did GPT 5.1.

13:33They all did fine and so I actually

13:35tested this with Gemini 2.5 Flash using

13:39a less expensive, less powerful model,

13:41but cheaper.

13:43And Flash did it fine as well. So

13:47I ended up kind of going backwards on

13:49this one and being like, okay, let's

13:50actually use a cheaper model even if

13:52it's less powerful because

13:55um

13:56you don't need need a powerful step on

13:57this one. So you learn things by doing

13:59all this testing.

14:01Um regardless, um

14:03I mean Gemini 3 would probably be

14:05better,

14:06but you don't need it to do better here.

14:09Like this is

14:10it's just implementing what this step

14:14recommended. And Gemini 2.5 Flash is

14:17perfectly capable of doing that. As are

14:19the other

14:20less expensive models out there like GPT

14:225 Mini and Claude 4 Haiku. All of those

14:26do just fine. Anyway.

14:28So let's actually talk about the prose

14:30that I got out of this thing here. So I

14:31did actually pick up on a couple of

14:33quirks of Gemini 3 and I will show off

14:37one of those quirks here. The if we read

14:39this just this little section here from

14:41a book I was working on. Precise,

14:42legible, permanent. I took a deep breath

14:44of stale hot air, dust, and sweat. I sat

14:47there sweating, aching, heart hammering

14:48against the ribs. I leaned forward,

14:50forced the bot forced the body back into

14:52the hunch. I wasn't going home, not yet.

14:55This misery had a file number. Do you

14:57see the problem here?

15:00The main issue that I saw with the prose

15:03that it was giving me and this could be

15:06something it's getting from the style

15:07recommendations I've given in the past.

15:09You'll notice lots and lots of very

15:12short sentences all strung together

15:14one right after another, so it feels

15:16choppy. And this is this is a sign of

15:19bad writing. Short sentences are great,

15:21but they need to be kind of woven

15:22together with longer sentences

15:25to feel more natural.

15:26Uh if you do have more of an action

15:28scene, it can make more sense to have a

15:31a lot more shorter sentences because it

15:33feels more urgent. If there's like a

15:35little bit more anxiety in a lot of

15:37short sentences together. So it can work

15:39for certain instances, but generally

15:41speaking you want a mix of shorter and

15:44longer.

15:45And so I was looking through this and I

15:47thought, well, maybe because Gemini 3 is

15:49pretty smart, maybe if I just change the

15:51prompt and asked it to do more of the

15:54mixing sentences, it would do okay. And

15:56so I did and

15:58you can already see just like visually

16:00here there's a lot less white space

16:03uh and longer paragraphs here.

16:05Um but if we go and read this a little

16:07bit, the Styrofoam cup was hot against

16:09my palm. I squeezed it testing the give

16:11of the cheap material watching the dark

16:13liquid shift there inside. It smelled

16:14like burned beans and stale office air.

16:16It smelled like safety. All right, so

16:18this is a great use of a short sentence

16:20after a long sentence uh or in a

16:23paragraph to kind of add a little bit of

16:26punctuation to that thought.

16:29Um my lower back throbbed with a dull

16:30persistent ache that radiated down into

16:32my hips. This character's pregnant, by

16:33the way. Uh I shifted in the hard

16:36plastic chair trying to find a position

16:37that didn't make my spine feel

16:39compressed by a vice. The baby kicked a

16:41sharp protest against the adrenaline

16:42still flooding my system. Sit down,

16:44Sergio. My husband ignored me. He paced

16:47the length of the small kitchenette, his

16:48sneakers squeaking on the beige

16:50linoleum. Three steps forward, turn.

16:53Three steps back. He rubbed his wrist

16:55where the zip ties had been, the skin

16:56red and irritated. We shouldn't be here,

16:58he muttered. He wasn't looking at me. He

17:00was staring at the reinforced steel door

17:01with the intensity of a trapped animal.

17:03This is the box cat. We walked right

17:05into the box.

17:06Um by the way, I should note that this

17:08has not been edited by me at all. So

17:10obviously I could go through and

17:12probably improve some things.

17:13Um but uh

17:15this will just give you a sense for what

17:17it sounds like out of the box. And this

17:19is after going through those four steps.

17:21Um so this is not just

17:24the writing, but it's also

17:26uh drawing on a scene brief. It's having

17:28gone through that those two editing

17:31steps. So it's pretty polished at this

17:33point. Um ready for a human editor. This

17:36is a class A safe house, I said, my

17:38voice sending

17:39sounded scraping, exhausted. I took a

17:41sip of the coffee. It was bitter and

17:42lukewarm, exactly how it always tasted

17:44at the precinct.

17:45Steel reinforced walls, encrypted comms,

17:48three units paroling

17:50paroling the perimeter. Nobody gets in

17:51here unless Lieutenant Miller buzzes

17:53them in. It's a dead end, Sergio said.

17:55He stopped pacing and turned to me. His

17:56glasses were crooked, sitting askew on

17:58his nose. You saw what happened at the

18:00house. You saw the window. The glass

18:01didn't break, cat. It melted. It

18:03dissolved. Do you think a deadbolt is

18:04going to stop them? I think men with

18:06guns stop men with guns. I set the cup

18:08down on the Formica table with a hollow

18:10thud.

18:11Those things in the woods, that wasn't

18:12magic. It was tech. High-end military

18:14tech. Maybe DARPA, maybe foreign, but

18:16it's physical and if it's physical, the

18:17MPD can handle it. Obviously it's still

18:20AI, right? There's always going to be a

18:23lot of human guidance that is needed to

18:26get AI to a decent level. But I can say

18:29unequivocally

18:32unequivocally

18:34a Gemini 3 is better than any other

18:36model particularly for the cost cuz

18:39Claude 4.5 Opus is actually I I would

18:42almost put them neck and neck with each

18:43other.

18:45Um but in different ways.

18:47But for the cost, absolutely Gemini 3 is

18:50miles above what Claude can deliver.

18:53It's way above what I get from Claude

18:554.5 Sonnet, which is uh

18:58similar in price.

18:59Um and on par if not a little better

19:03with what I get out of Claude 4 Opus or

19:05Claude 4.1 Opus, which is five times

19:08more expensive, more than five times

19:10more expensive.

19:11And so

19:12as far as I'm concerned

19:14I as I was going through my automations,

19:16not this one, but not just this one,

19:18but also my outlining automations and my

19:21brainstorming automations and all of

19:23those other things that I've built,

19:25I found myself replacing the original

19:28model that I was using, whether that was

19:30a Sonnet model or a Gemini model, an

19:32older Gemini model, or a GPT model,

19:35replacing most of them with Gemini 3

19:38because as I was just going through and

19:39testing and looking at the difference

19:41between what the old model would give me

19:43and the new and the Gemini 3,

19:45um I was just finding like, yeah, Gemini

19:473 is better for this, too. And that's

19:50It's concerning to me because I don't

19:51like being so reliant on any one

19:52particular model. I like mixing the

19:54models in different ways and playing to

19:55their strengths, but to be completely

19:57honest, I think Gemini 3 is the best

20:01overall all around model

20:04that we have ever gotten for

20:06creative writing.

20:07That's my personal opinion. I know I

20:09said some good things about GPT-5 back

20:11when it came out, but after working with

20:13it a little bit more, I realized that I

20:15no longer agreed with those things

20:17uh because things just turned up you're

20:19like, "What is going on here?" That's

20:21why I spent more time working with

20:22Gemini 3 to try and see if that remained

20:26the case. And I can happily say that at

20:30least in the amount of testing that I've

20:31done, Gemini 3 is my favorite model on

20:35the market as of the recording of this

20:37video. Those are my thoughts. Hope this

20:39was a useful video for you, and I will

20:40see you in the next one.

More from The Nerdy Novelist

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.