Full transcript
0:00All right. Today we're showing off
0:01another automation that I've put
0:03together. And this one will be a little
0:04bit interesting because it's not the
0:06kind of thing that a lot of people would
0:07think about, but it's actually extremely
0:09useful if you are interested in
0:11publishing any kind of public domain
0:12content. I've had a fair amount of
0:14success with public domain content. At
0:16least, you know, nothing huge, but I
0:19usually get a couple hundred dollars a
0:20month from it. And it's a relatively
0:21simple thing to be able to create. But
0:24there are a few little parts of the
0:27process that are actually really
0:28difficult. And the biggest of these is
0:31kind of creating a clean version of your
0:34text. Because if you go online looking
0:36for a public domain text to publish, you
0:39will find a number of different
0:40versions. But there are often a lot of
0:42problems with the manuscript, especially
0:44if you're trying to copy and paste it
0:46into a formatter like Vellum or Attacus
0:49because there's often like little
0:50footnotes and things or misspelled
0:53words. Sometimes it was created using
0:55OCR software uh where it was looking at
0:59literally scanning a physical page and
1:01then trying to come up with the text
1:03from that. and there's all kinds of
1:04errors in it. Regardless, there are a
1:07ton of reasons why it's difficult to
1:09have a clean version of the manuscript.
1:13And so, I created this automation to
1:15essentially negate that problem and make
1:19the public domain text perfectly suited
1:22and ready for formatting. So, let me
1:24show you exactly what that looks like.
1:26All right, so this is the automation
1:27here. Fairly very fairly simple looking.
1:30Um, but the way this works is we grab
1:32two documents. First is going to be the
1:35document where the clean version, the
1:36cleaned up version is going to go. And
1:38when I say clean, I mean clean of little
1:41artifacts and footnotes and other things
1:43that we don't want in there, not like a
1:46clean like content clean uh type of
1:49document. And uh and then we grab the
1:52original public domain document. So,
1:53what I do is I go out and I grab a I
1:56just kind of copy and paste the whatever
1:59I can find out there for the public
2:01domain book that I'm wanting to publish.
2:03I just sort of copy and paste it into a
2:04Google doc. And I don't worry about it
2:06too much. The only sort of
2:07post-processing that I will do manually
2:10is I will go through the chapter
2:12headings for these and I will often give
2:15it an H1 or H2 heading. However, you
2:18don't really have to do that with this
2:20particular automation because it's not
2:23going to be going chapter by chapter.
2:24It's going to be going by every thousand
2:26words. Sometimes chapters are so big
2:28that I actually don't trust the AI to be
2:30able to handle the whole thing at once
2:32and be able to reproduce it verbatim uh
2:34with everything I need it to. I'm I'm
2:37worried that it might like accidentally
2:39shorten it or something like that. So, I
2:42go by a thousand word chunks instead.
2:44And this is for this particular public
2:47domain book. This is one that I've
2:49already published. Uh I published an
2:51annotated version of it and I published
2:52a modern pros version of it. All of
2:54which was achieved with the help of AI.
2:57U but as you can see this is a this is
3:00an epic poem that was written in uh I
3:02think the 16th century or or maybe early
3:0517th century. Um, and uh, so it's been
3:09about 400 years since and there's a lot
3:12of really weird
3:14spellings and stuff. Uh, like if you see
3:17here, unfitter task, but the the U is a
3:20V. And this was pretty common in a lot
3:23of, uh, older texts like that, you have
3:27U's and V's are a little bit
3:28interchangeable. You also see Fs and S's
3:32um, often interchangeable. And so I
3:34wanted to clean this up, not to change
3:36the words because I want the words to
3:38say the same, but just to update the
3:41spelling so it's a little less
3:42confusing. Um, as well as remove any
3:45kind of artifacts. I mean, it was mostly
3:48spelling for this one. Uh but sometimes
3:50you do get a lot of things that you
3:52don't want in there like um like page
3:54numbers uh which sometimes get in there
3:58uh or or paragraph symbols or footnotes
4:02uh endotes that kind of thing. We want
4:04to strip all of that out so we just have
4:06the text itself very nicely formatted.
4:09And so this automation, what it does is
4:12it pulls in those two documents, the the
4:14original public domain document that I
4:16just showed you and a blank document
4:18where the cleaned up version is going to
4:20go. And then this is a little bit of
4:22code that splits the original document
4:24into 1,00 word chunks. Okay, so this is
4:27a really handy thing to have um because
4:31it makes it into a little bit more of a
4:33manageable thing that we can then
4:35process through an AI text. However, I
4:39didn't want this to just split into
4:42random uh thousandword chunks because
4:45what if that ends up in the middle of a
4:47sentence, right? So, what this code
4:48does, and once again, I am not a coder,
4:51but I use tools like Perplexity to vibe
4:54code a little bit of what I need here.
4:56This is what the code looks like. Um,
4:58it's fairly simple. Not too much to it.
5:01Um, but what this does, it is it divides
5:03it into a thousand word chunks, but then
5:06it searches for the next period. so that
5:08it actually divides it a few words after
5:11the thousand words at the end of the the
5:13next sentence. That way it kind of keeps
5:15everything together. So we're not
5:17dividing the chunks in the middle of a
5:20sentence cuz that would be kind of
5:22confusing for the AI later. So I have it
5:25do that and then it runs this loop on
5:27each of the thousandword chunks. And we
5:29have just one thing here um which
5:31currently is being done by Gemini 2.5
5:34Flash. Although now I could probably uh
5:36update this to Gemini 3 flash because
5:39that is out now. That's going to be a
5:41little bit more effective. And the
5:43prompt looks something like this. Here's
5:45the section of public domain text you
5:46will analyze and it pulls in the text
5:49here. Your task is to create a cleaned
5:51up version of the above text. Here are
5:53some things to watch out for. Number
5:54one, look for spellings that are
5:56different than modern English. Change
5:57these spellings to more modern English
5:59spellings, but do not change the word
6:00itself. Make sure the original text is
6:02still outputed just with more modern
6:04spellings so as not to confuse the
6:06readers. Number two, remove any
6:07footnotes or references to footnotes
6:09within the text. Number three, format
6:11grammar appropriately according to
6:12modern Chicago manual of style. However,
6:14do not change the sentence structure or
6:16change the meaning of the original text
6:18in any way. If there is a conflict, air
6:19on the side of sticking to the original
6:21style or grammar. Number four, if
6:23formatting poetry, use the same line
6:25breaks and original stanza length. Uh,
6:27always utilize the same line breaks,
6:28paragraph breaks, etc. as the original.
6:31Number five, make sure to maintain the
6:33same capitalization used in the original
6:35text. Number six, the text you receive
6:37may be poorly transcribed text from OCR
6:39software. Like the [snorts] older text,
6:41make sure the spellings and grammar
6:42match the original intended words,
6:44fixing the errors from the OCR software.
6:46Uh, just in case that's what we're
6:48running into. Note, if your text
6:50includes a chapter or section title,
6:52examples include, but are not limited
6:53to, chapter, blank, prologue, epilog,
6:56preface, kanto, etc. then format these
6:58on a separate line using markdown for an
7:00H1 header examples here. So that's where
7:04I tell it to basically if it finds a
7:06chapter header at some point in the in
7:08the story in the um the text to make
7:12that a heading and that's going to be
7:14important. Uh, so then I just give it
7:17this information about how to format the
7:20thing. And it will now run this on each
7:22of the thousandword chunks. And then
7:25once it's done that, it will add it to
7:27the um cleaned up document that we
7:31pulled in here. It will add it to that
7:33and then continue the loop and continue
7:36the process for the whole thing. And so
7:38to give you an idea of what this looks
7:39like, here is the cleaned up version of
7:42that same text. Um, so you'll notice
7:44there were a couple of things in here
7:46that it was getting wrong, uh, or that
7:48just had weird spelling like containing
7:51with a Y. That is now just containing.
7:54Um, we have shepherds.
7:57Um, but this unfitter task here is what
8:00I'd like to look at. Let's look at where
8:02that is. And yep, we now it now says
8:04unfitter task. So the words are still
8:07accurate words, but it's now cleaned up
8:10and looking a little bit now. it will
8:12sometimes leave certain words in there
8:15if it kind of if there's not really a
8:17good real world equivalent like if it
8:20would have to change the word it leaves
8:22the word there. So, you'll see a few of
8:24those in here, but on the whole, this is
8:25a much more cleaned up version that is
8:28really useful. And you'll notice that
8:29the uh title here is an H1 heading. And
8:33if I move on to Kanto one, uh that is an
8:35H1 heading. And this makes it super easy
8:39for me to just download this Google Doc
8:42and as a as a docs file and then upload
8:45it to a formatting tool like Attekus.
8:47Makes it so much easier to just go
8:50through and go through that process. And
8:52I'll actually show you what this uh what
8:54it looks like in Attekus or rather I
8:57will show you what it looks like in its
8:59fully formatted PDF version. So here I
9:01have it all with the table of contents
9:03and everything. And here we have the
9:05preface, the first book of the fairy
9:06queen containing the legend of the night
9:08of the red cross at war of holiness.
9:10There's the unfitter task right there.
9:12So it just puts everything here. And
9:15then I actually have some
9:18um notes and annotations and things that
9:20I added in there myself. But that's a
9:22separate automation. Uh, so you can see
9:25it looks very nice and I was very happy
9:27with this. So this is this saved me so
9:29much time having to rather than having
9:32to do a lot of this work myself, just
9:34being able to clean everything up. And
9:35if you know, you know, like if you've
9:37tried to do any uh public domain
9:39publishing in the past, getting that
9:41manuscript from a random place on the
9:43internet to the point where it is
9:45publishable and clean and looking nice
9:48is a bit of a headache. So, this really
9:51has been a extremely helpful automation
9:54for me personally. Now, I realize not
9:56everybody's going to be into this, but I
9:58have found it to be a really good time.
10:00So hopefully this has been a useful
10:02video for you and I will see you in the
10:03in the next