Full transcript
0:00The last few months have been some of
0:01the most interesting and
0:02industrydefining moments in tech and in
0:05copyright history. With the case of
0:07Barts versus Anthropic, which is set to
0:09really determine just how legal is it
0:12for these big AI companies to train
0:14their large language models on
0:16copyrighted material. Now, I've been
0:18monitoring this court case since it got
0:20started, and it's now pretty much over,
0:22and the results have been nothing short
0:24of fascinating. And believe it or not,
0:26it's actually a big win for authors and
0:28creatives, whether you use AI or not.
0:35Now, if you're new here, my name is
0:37Jason. I wrote 14 books the traditional
0:39way. Worked at Kindlepreneur for a
0:41number of years, which by the way, they
0:42have a really impressive article on this
0:44exact topic, so you should go check them
0:46out. I'll link to it down below. But I
0:47started this channel as a way to show
0:49authors how to use AI as a productivity
0:52tool without actually compromising your
0:54creativity or ethics. I should also
0:56point out that I am not a lawyer and
0:58none of this should be taken as actual
1:00legal advice. I have tried to do my due
1:02diligence to actually find credible
1:04sources for this though. So hopefully
1:06this will be useful in just breaking
1:07down this case for you. So what exactly
1:10is this lawsuit all about? Well, it was
1:12filed in late 2023 when three authors,
1:14Andrea Barts, Charles Greyber, and Kirk
1:17Wallace Johnson, filed a class action
1:19lawsuit against Anthropic, which is the
1:21company behind Claude. And if you have
1:23been a fan of this channel for a while,
1:24you know I really enjoy the Claude
1:26models. They tend to be one of the best
1:28ones at creative writing. But they
1:29claimed that Anthropic used their books
1:31along with millions of others to train
1:33its AI without permission. Something
1:35they thought was wrong and
1:37understandably so. Now, just as a
1:39reminder of what copyright actually
1:41protects, I like to think of it in terms
1:43of the three Ds. So, you have the
1:46display of a text, the distribution of
1:49that text, and the derivation of that
1:52text. Now, in this case, derivation is
1:54not really applicable as it's next to
1:56impossible to actually get AI to create
1:59derivative work without the express
2:01instruction of the user, in which case
2:04the user would actually be the one
2:05culpable for feeding it with these
2:07derivative instructions. You know, like
2:09if I asked it to create Star Wars
2:11fanfiction, for example, and then I went
2:13and tried to sell that fanfiction,
2:14that's not really the fault of the AI.
2:16That's more my issue. And a lot of these
2:19AI models actually have safeguards in
2:21place to sort of prevent any sort of
2:23derivative type of copyright
2:25infringement. It's not perfect, but they
2:27do try to do that. Now, the display of a
2:29copyrighted text is also something that
2:31large language models can't really do
2:33because once again, there safeguards to
2:35prevent that sort of thing. It's really,
2:37really difficult to try and get a large
2:39language model to give you more than a
2:41tiny quote or two here from other books.
2:44often those quotes will be wrong and
2:46trying to get it to produce a more
2:48substantial piece of the text is just
2:51not something it can do because remember
2:53these models were not they don't have a
2:55database that they're drawing from when
2:57they were trained on these models. The
2:58training happens and then the training
3:00data is discarded. It it isn't needed
3:03anymore. And so all it's done is it's
3:05learned the probability of language from
3:07that training data. Which means yes, in
3:09some cases it might be able to from a
3:11probability standpoint get something
3:13that's close to the original text, but
3:15it's actually really really difficult
3:17from the way large language models work
3:19to display the text. And so any lawsuits
3:21that have come trying to argue that
3:24these large language models are
3:25displaying text of copyrighted material
3:28to the end user essentially get thrown
3:30out because there's absolutely nothing
3:31that can really be brought as evidence
3:33for that. However, we have the third D
3:35and that is distribution and this was
3:38the main issue around this lawsuit and
3:40the one for which the large settlement
3:43occurred. So what do we get out of this
3:45lawsuit? Put simply, there are a number
3:48of things that got really definitively
3:51uh decided on by the judge on what this
3:54gives us. The first is that training
3:57claude on legally purchased books was
4:00officially deemed fair use. We're not
4:02likely to see that change because one of
4:03the things that Anthropic was doing is
4:05they were purchasing books and then
4:07scanning them and then using those
4:09scans, the digitization of those books
4:12in their training data. And
4:13additionally, the judge found that
4:14digitizing those print books that
4:16Anthropic bought and using those for
4:18training, that was also deemed fair use.
4:20And for good reason, honestly. Like, if
4:22I buy a book and then I digitize it for
4:25myself, just for my own purposes. And
4:28I'm not distributing it to anybody else.
4:30I'm not giving it away to anybody. I'm
4:31just using it for my own purposes. And I
4:33actually have done this. I actually have
4:35a book scanner. I'll show you one
4:37example. I recently bought this and a
4:39bunch of other little pamphlets like
4:40this from Alex Rormozi. And there's no
4:42ebook version of this. And I wanted to
4:44have an ebook version that I could carry
4:46around on my Kindle. I cannot do that
4:49and then give away the digitization that
4:51I have created. But I can 100% scan this
4:54and then run it through a prompt locally
4:57on my computer that transcribes it and
5:00then from there I can put that on my own
5:02Kindle and that is fair use. Uh because
5:04I'm only using it for myself. I am not
5:06creating derivative work from it. I am
5:08not displaying it to other people
5:10publicly and I am not distributing it in
5:13any way. So as long as I want to have
5:15that right to be able to do that for
5:16myself and as long as you want to have
5:18the right to be able to do pretty much
5:20whatever you want with your the stuff
5:22that you have purchased, it makes sense
5:23that this would be fair use and that
5:25anthropic would also have the same
5:27legality assigned to it because a
5:30business is treated legally much in the
5:32same way as an individual is. And if we
5:34were to change that, it would create a
5:36pretty dangerous precedent for what we
5:38could do with our own work that we
5:42purchased. So those first two things,
5:44the purchasing, scanning, and
5:47digitization of those books was
5:48considered fair use by the judge, which
5:50is actually really a handy thing to know
5:53that that is okay. However, Claude was
5:55also trained on databases that included
5:58pirated training data, aka data that was
6:00pulled from pirated websites where those
6:03websites were illegally distributing the
6:05work. Now, I think we all know that
6:06pirating is bad and is definitely
6:09illegal and against copyright
6:10protection. And remember that copyright
6:12protects against the display and
6:14distribution of a work. And when
6:16Anthropic pulled from those pirated data
6:19sets, they were essentially aiding in
6:21the illegal distribution of that work.
6:23So while buying and digitizing books to
6:25use in a data set is deemed fair use and
6:28completely legal, using pirated books in
6:30the same data set was not and this is
6:33what the judge ruled and this is why
6:35Anthropic eventually settled this
6:36lawsuit with the plaintiffs. So what
6:38does the settlement look like and can we
6:41as authors actually benefit from it?
6:42Well, Anthropic has agreed to settle for
6:44a whopping $1.5 billion to be paid to a
6:48settlement fund which can then be
6:50distributed to authors who qualify.
6:52That's approximately $3,000 per book for
6:54the estimated count of 500,000 books.
6:57And if there are more than 500,000 books
6:59found, Anthropic will pay an additional
7:02$3,000 per book to the fund. Hey, Jason
7:04from the future here. Um, just as I was
7:06about to post this video, some news came
7:08out that the judge had actually rejected
7:10that $1.5 billion settlement, saying it
7:14was quote nowhere close to complete. Um,
7:17and I think his main concern was just uh
7:19inefficiency in uh the detail and
7:22transparency. He thought it uh risked
7:24actually unfair payouts to authors and
7:28um uh he just generally wasn't happy
7:31with it. So that just shows that like a
7:33lot of this stuff is in flux and uh
7:35things may change and we don't really
7:37know what exactly this is going to look
7:39like when it's finally finalized. Um
7:43most of what I uh the other things I say
7:45in this video are still relevant though.
7:46So, we'll just go back to that. Now, as
7:48a side note, for those of you who are
7:49hoping that Anthropic and AI's demise
7:51from cases like this, I have bad news
7:53for you because Anthropic just recently
7:56raised their valuation by 13 billion to
7:58a total of 183 billion. 1.5 billion
8:01actually isn't. It's kind of a drop in
8:02the bucket considering the investments
8:04and the investment money that they have.
8:06And the same goes for other AI companies
8:08are even bigger than Anthropic. But
8:09Anthropic must also delete all pirated
8:12books from their training data set going
8:14forward. That was also part of the
8:15settlement. This is only relevant for
8:17future training materials as this
8:18lawsuit does not work retroactively to
8:21past models that were trained on this
8:23data. And as I mentioned before, you
8:25don't have to have like once a model has
8:27been trained on the training data, it
8:29doesn't need that training data anymore.
8:30So just by throwing out the pirated
8:32data, it's not actually going to make
8:34our current cloud models perform any
8:37worse because they've already been
8:38trained. That part is over. And the way
8:40legal cases like this work is that it
8:42doesn't work retroactively. they don't
8:44have to stop using the cloud models just
8:47because those cloud models were trained
8:48on pirated data. All this does is it
8:50prevents Anthropic from using pirated uh
8:54work in the future for their future data
8:57sets and for their future large language
8:59models that they train. So just a
9:00clarifying note there. So how can
9:02authors take advantage of this if you
9:04want to get that $3,000 per work for
9:06example? Now as of right now the
9:08comprehensive list of effective works is
9:10still in development but that will be
9:12available eventually. And in the
9:13meantime, you can visit a tool that I'll
9:15link to below from the Atlantic. And
9:16that will show you if your book was
9:18included in LibGen, which is one of the
9:20big databases that was that Anthropic
9:22used and was in question here. So, that
9:24should give you an idea of whether or
9:26not you qualify to earn some kickback
9:28from this. However, there are a few
9:30qualifications authors must have in
9:32order to qualify for these damages. And
9:34unfortunately, this is going to rule out
9:36a lot of you. The first is that this
9:37only applies to those in the US as this
9:40was a US copyright case. and so only
9:43applies in the US. You must also have a
9:45registered copyright for the book in
9:48question. And that copyright needs to
9:49have been filed within 5 years of having
9:51published it and before Anthropic used
9:53it to train your models. So you can't
9:55just go out and copyright your book
9:57right now and then expect to get
10:00compensated for it. It needs to had that
10:02copyright applied to it within 5 years
10:04of the publishing of that book. Now,
10:06unfortunately for self-published
10:07authors, most of us do not actually
10:09copyright every single book because it
10:11costs money and there's a hassle behind
10:13it. And in my case, I only had one book
10:15that I've copyrighted that is also on
10:18the libgen database. So, the most I
10:20could possibly get from this maybe would
10:22be $3,000 for the one book. And we
10:24actually found out because of this case
10:26that a lot of traditional publishers
10:27were also not filing copyright for some
10:30of the books that they published because
10:32a lot of people were looking into this
10:33and realizing that their book was not
10:34published even though in most cases
10:37traditional publishers are contractually
10:39obligated to file copyright for the
10:41book. So if you are a traditionally
10:42published author make sure to check that
10:45your book has been filed in the
10:47copyright office for a US copyright.
10:50Now, the third qualification is that
10:51your book must also have an
10:53international standard book number or
10:54ISBN. This is less likely to be a
10:57problem because even if you didn't
10:58purchase an ISBN for yourself, if you
11:01published it on a platform like Amazon,
11:03you will have gotten one anyway. So,
11:05your book will have an ISBN as long as
11:07you published it somewhere. And it
11:09doesn't have to be an ISBN that you
11:11purchased. So, all that said, um, if you
11:14think that you might qualify, if you
11:15check the LibGen database and you find
11:17out that you have books in there and
11:19those books were copyrighted at the time
11:21that Anthropic used them, your book has
11:23an ISBN number, then you might qualify.
11:26And so, there's a website I will link to
11:28down below,
11:28anthropiccopyrightslement.com,
11:30where you can go and submit your
11:31copyright information and your contact
11:33information. And I think there will be
11:35other updates on this as we go along
11:38where you might be able to get more
11:39clear instructions about what to do.
11:41Don't worry, this is not something, you
11:43know, these cases take a long time, so
11:45this is not something you need to rush
11:46for just yet. Uh, we don't even have the
11:49official list of affected works yet. So,
11:52at the very least, we can wait until
11:53that's out before starting to figure out
11:55how can we get there our damages for
11:57this. So, why do I say that this is a
11:59good thing for authors, even if you are
12:01a pro AAI author? Well, first of all,
12:04after years of arguing and moaning and
12:07groaning, we finally have some solid
12:09legal precedent for what is and isn't
12:11legal around using copyrighted works for
12:13training data. And in other words, uh,
12:15legally acquiring a copy of a book and
12:16using that seems to be fine while using
12:18pirated data is not. And this will
12:20likely lead to a lot of AI companies
12:22being a little bit more careful about
12:24the data that they include to make sure
12:26that there is no pirated work in there.
12:28Second, this means the authors will get
12:30at a minimum at least one legally bought
12:33copy of their book for training data
12:35purposes. And that's not going to sound
12:37like a lot and definitely not enough for
12:38most people, but I think it's actually
12:40better than what I would have expected,
12:42honestly. And I think we'll see more
12:43companies giving higher licensing
12:45options in order to sell exclusive
12:47rights to training data in the future.
12:49So, for instance, if they want to train
12:51on your book and other AI companies
12:53aren't allowed to do so, they want to
12:55buy that exclusive right from you, that
12:57could potentially uh drive more
12:59competition for training data, that sort
13:01of thing. So, having more of these
13:02licensing deals will also make their
13:04legal right to those books far more
13:06ironclad, which I'm sure is a security
13:08that most of these companies would want.
13:11Now, unfortunately, what this doesn't do
13:13is provide authors with a way of opting
13:15out of training LLMs on their books.
13:17That might be something we see in the
13:18future. But since these companies can
13:20now clearly just buy a copy of your book
13:22and digitize it without your permission
13:25because that is considered fair use.
13:26Clearly, there are lots of things that I
13:28can do with your book without permission
13:30because once I've bought that book, I
13:32have fair use to do a lot of things with
13:33it. As long as I don't distribute,
13:35display publicly or create derivative
13:37works from that book, which are the only
13:39things that copyright protects against.
13:41And copyright doesn't protect me against
13:42photocopying your book for my own
13:44purposes as long as I don't distribute
13:46it or for using it as toilet paper if I
13:49so desire. Well, you know, people do
13:51that. But ultimately, I think having
13:53clear rules around training data is
13:54nothing but a good thing. And this is
13:57just the start. I think we're going to
13:58start actually seeing a lot of these
14:00gray areas get much more well-
14:01definfined from a legal standpoint, from
14:03a licensing standpoint, and all of that.
14:06And if you're a pro AI author, this is
14:08good news for you as well. It doesn't
14:10mean that these LLMs are going to get
14:11dumber because they can't train on
14:13pirated material. If anything, now that
14:15AI companies know for sure that they can
14:17legally buy a copy of a book and
14:18digitize it, we might actually see more
14:21books join the data sets because they
14:23can now add books that either weren't
14:25digital or weren't in pirated data sets.
14:27And it's a heck of a lot cheaper to buy
14:29one copy and digitize it than to pay
14:31$3,000 in legal fees per book, which as
14:33I established, even 1.5 billion is a
14:36drop in the bucket for these huge tech
14:38companies. The last reason why I think
14:40this is a good thing is that now that
14:41this precedent has been set, we're going
14:44to see more lawsuits coming against
14:46other AI companies who have also trained
14:48their models on copyrighted works.
14:50There's already several of them in
14:51process and we're going to probably see
14:52a lot more settlements of this type. And
14:54this case will make it much easier for
14:56the plaintiffs to argue whether or not
14:58there was infringement that occurred.
15:00And I'll keep you up to date as this
15:01case continues unfolding, including
15:03where you'll be able to go once the full
15:05listing of infringing works is
15:06officially published. And if this video
15:08was useful for you, I would be happy if
15:10you could just do all the YouTube stuff.
15:12We talk a lot about AI and creative
15:13writing on this channel. I also have uh
15:15some ways to work with me a little bit
15:17more directly. We do challenges in my
15:20gold group like a book in a month
15:21challenge. Right now, we're writing a
15:23short story without AI as a sort of
15:25creative writing exercise to try and
15:27actually improve our skills. So, we're
15:29very serious about actually not just
15:31writing with AI, but building our own
15:33creativity in the process and using AI
15:35as a productivity tool. So, if that's
15:36interesting to you, links are all down
15:38below and I will see you in the next