Free YouTube Transcribe

Video transcript

I Built a SillyTavern Alternative — Local AI Chat With Voice

Fran Solo · 7,407 words · 34 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Hello everybody. In this video, I'm

0:01going to show you how you can create a

0:04local AI chart on your computer using

0:08just two pieces of software and how you

0:11can customize it very easily like you

0:14would do with any other charts that you

0:16tried online such as text generation web

0:19UI, silly tavern and similar. Now, I

0:23created this chat mainly for me and

0:26while I was creating it and developing

0:28it, I thought to record a

0:31behind-the-scene episodes, actually more

0:33than one, to share how I go through the

0:37process of creating something from

0:39scratch using LLMs. In this case, I use

0:41Quen Studio, but I also use Chart GPT as

0:44well. In this video, I'm going to show

0:45you the final work about the interface,

0:48how it works, and then I'm going to

0:50share with you the interface itself. You

0:53can just uh click on the Patreon link at

0:57the bottom or whatever links is

0:59available at the bottom of the

1:01description and you can have it as well.

1:05Okay. So, let's share the screen and

1:08let's have a look at this app. The first

1:11things you see when you open it is this

1:13welcome page. And you see this is called

1:16Aura AI and the version is 1.26 because

1:19I did roughly 26 or 27 iterations of

1:24this uh interface. This is an HTML file

1:28essentially and that's all it is. And

1:31what it does it allows you to chat

1:34locally with your AI. So you could do

1:36that using LM Studio. You can use that

1:39uh using any other local LLM, but this

1:43one also gives you the option to

1:45customize the voice. So you have a

1:48speechtoext feature and you also have a

1:51texttospech feature as well. So you can

1:53talk, you can converse with the AI and

1:56it will reply you back with a voice

1:58almost like having your own private

2:00Jarvis. In terms of responsiveness,

2:04it's going to be fast. if you load a

2:07small model or at least a model that

2:10fits your graphic card. So here you see

2:13all the things that you can do. So you

2:16can chat you can create charts and

2:18folders very similar to projects in chat

2:20GPT and in uh Quen Studio. And if you

2:23want to see it, if I go to my chat GPT,

2:26you'll see I've got all my list of chats

2:28here on the left hand side. I've got my

2:31login account down below. I've got also

2:34my projects and uh for each projects

2:37I've got chats like a list of chats etc.

2:41When you click on the dotted icon here

2:43you have option to you have the option

2:45to share the project to rename it

2:47project to to see the homepage of the

2:50project

2:52all the things that you can do here. So

2:54I pretty much copied or at least

2:57personalized my own interface to work

3:00the best. Now this is not a finalized

3:03interface necessarily but it's one of

3:06iterations one of the many iteration I

3:08did in terms of uh seeing comparing it

3:11to Quen Studio. You'll see here on the

3:13left hand side is very similar again you

3:15have the option to do new charts to

3:17search for chats. You also have the

3:20options to see the libraries with all of

3:22the media etc. And uh you can organize

3:26your chats into folders or projects. My

3:29interface here does pretty much the same

3:31things. Other thing you can do, you can

3:33personalize your the voice that you can

3:37use to answer your back. So, the AI can

3:40have a custom voice, and I'm going to

3:42show you how that works in a second. You

3:44have also a narrator voice that you can

3:46customize. You can change the character

3:50personality as well. If you want to, you

3:52can do that. and also add other

3:55graphically appealing widgets in the

3:58interface. In terms of saving your

4:01charts or exporting your charts, I made

4:03it in such a way that you can export it

4:05retaining all the media that you use in

4:07that specific thread. Exporting the same

4:10thing and then when you click on next

4:12launching and everything launching

4:15everything you have the uh instructions

4:17on how to use it. Now this chat this

4:20local AI chat works with LM studio in

4:23the background which works as a brain

4:26and with all talks uh TTS this is the

4:30one that I use so far and then I tested

4:33I tested some others which I wasn't

4:36really happy about uh like such as

4:38Google Microsoft Azour then I also

4:42tested um Piper and others as well and

4:46the voices were horrible.

4:48at least in my opinion. So, all is the

4:51one that I chose at the end. Of course,

4:54I've got omnivoice in confi, but there

4:57was no easy way to interface to link

5:00confui with an interface like this or at

5:04least I wasn't able to to actually do

5:06it. So, I prefer to have this

5:10configuration instead. So all you need

5:12to do to make this work is to have LM

5:15Studio installed with a model downloaded

5:19in it and then you need to have the

5:22launcher which is a file that I created

5:25which will install alls TTS into your

5:28machine and will also launch the ST

5:31server as well. Now the old talk to

5:33talks was a software was developed up to

5:38December 2025

5:40and after that uh it wasn't I don't

5:43think it was even touched by the

5:45developers I think for what I saw but it

5:49is a bit buggy when you install it if

5:51you go to the old talks DTS on a face or

5:54yeah I think it's on a face or github

5:57you might find some uh issues when you

6:00install it I think It's because of

6:03Python version. I had some several

6:07issues installing this and I found a

6:09solution by installing the right Python

6:13version which was used when they develop

6:16all talks DTS and that worked.

6:19So I also discovered a bug through an

6:21LLM. I'm not a genius. I use NLM. I use

6:24Quen Studio for that. And by doing it

6:27back and forth, I eventually discovered

6:30that the, you know, it discovered there

6:32was a glitch on the original code. And

6:36we fixed it. He fixed it for me. And now

6:39it works fine. And um then again, this

6:44launcher is going to be just a double

6:46click for you. So you all you need to do

6:47is simply double click on it and it will

6:49launch. I'm going to show you exactly

6:50how I do it. And also you need to just

6:54open the interface on your browser or

6:56any browser you prefer and you're ready

7:00to go. So you can of course customize

7:04your own voice and easily upload it or

7:08put it into the folder that I'm going to

7:09show you in a second and you can back up

7:11your conversations etc. If you change

7:14your mind, you want to use some other

7:16interface. Then I'm going to click on

7:17start chatting and this is how the

7:20interface looks like. It looks quite

7:21plain but you can customize quite a few

7:23things. So the first things we're going

7:25to do here on the left hand side is of

7:29course creating new chat if you want to.

7:31Of course on top you have the logo here

7:33if you want to find out about the

7:34instructions. You can just click on it

7:36and it will it's going to open this

7:38window again the instruction window and

7:40the features list. Then you can click on

7:42close if you want to. Then you can go to

7:44new chat. This will uh will open a new

7:46chart down here which you can rename by

7:49just hovering the cursor over the three

7:51dotted icon here. You can rename, you

7:53can clone the chart if you want to. You

7:55can export it as a HTML. You can move it

7:58into a project or into a folder in this

8:01case or you can delete.

8:04And you can also search for chats by

8:06keywords and it will automatically

8:08filter out this list down here. And you

8:11can choose a media folder. This will

8:14allow you to save all the uploaded

8:16images that you use during the chat if

8:19you want to. If you use a NLM model who

8:22has uh that has a vision in it, you can

8:25basically uh gather them all together in

8:28one place. So all I'm going to do, I'm

8:29going to click on media and I'm going to

8:31choose my folder. This case here is

8:33under the desktop under the Aura AI

8:36folder and it's called media chat. But

8:39you can choose any folder you like and

8:41click on select folder. And now it's

8:43going to ask me if I want to allow the

8:45site to uh to edit it. I'm going to

8:47click on allow. It says here media

8:50folder linked media chart. So that is a

8:52confirmation that it works. Now the

8:55other thing we need to make sure is the

8:58connection. At the moment the connection

9:00here has three LED. They are all green

9:03means that the chat is well connected

9:07with LM Studio and all talks. But I'm

9:09going to close this and I'm going to

9:11show you this from scratch how it works.

9:13So, first things we're going to do,

9:14we're going to close this and we're

9:16going to actually close the entire

9:19browser and we're going to also close

9:21the the PowerShell and LM Studio as

9:24well. I'm going to just start from

9:25scratch completely. And so when you

9:29install this uh folder, I'm going to uh

9:32you're going to find the zip file. After

9:33you unzip it, you're going to see a file

9:36called start local voice chat or

9:41something similar. But I will basically

9:42give you the step-by-step instructions

9:44when I create the the guide. But you see

9:48you got start local voice chat. This

9:50will when you double click on it, it

9:51will open two powershell. One will

9:54launch the old talk TTS and the other

9:56one will launch will launch the driver

9:59for the speech to talk to allow you to

10:02talk through the mic and dictate what

10:04you want. So you don't need to do

10:06anything else. Once you double click on

10:08it, you just minimize these two windows

10:10and forget about them. The second thing

10:12you need to do is simply open LM Studio.

10:15So, open LM Studio if you don't have it

10:18downloaded and then go into the to the

10:21search and just search a model that you

10:25want to chat with. At the moment, I've

10:27got already a few models in my list. And

10:31I'm going to choose a couple. One is a

10:34small model so we can have a quick

10:37response and have the feel that we

10:39chatting with our Jarvis AI. And this is

10:43going to be the the 12 GB model which

10:46has also a vision. And also we have the

10:49GPT OSS which is 22 GB. And we have also

10:54the Gemma 412B obliterated which is 12.7

10:59GB which is another small LLM that fits

11:02in my graphic card. By the way, I'm

11:04using an RTX 5070 Ti with 16 GB of RAM.

11:08I also have 96 GB of system RAM for

11:12those of you who are wondering what I've

11:13got and a Ryzen 7 7700 3.8 GHz as CPU.

11:19So that's what I've got. But in order to

11:21make Aura AI chat with you and have a

11:24pretty responsive and quick reply is

11:28good to have a model that fits into your

11:30graphic card and does not offload

11:33anything on your system RAM. So, I'm

11:35going to go into this tab here on the

11:37left hand side, which is developer. And

11:39here, I'm going to load that model. It's

11:41already loaded, but I'm going to load my

11:43model from here. All you need to do is

11:44simply go to load model up here. I'm

11:47going to choose something small. I'm

11:48going to choose this one here, 11 GB.

11:50It's going to load model fairly quickly.

11:53The other thing you need to make sure is

11:55under server settings, you make sure

11:58that enable cores is enabled. So, make

12:01sure that one is on. And that basically

12:03allows you to to share LM Studio with

12:06third party software. In this case, our

12:08TTS and SDT.

12:11Then after we done that, we're going to

12:13go to status

12:15here and running. Make sure that this is

12:17green. And that's it. That's all you

12:20need to do here in LM Studio. Once you

12:23launch this start uh local voice chat

12:27and LM studio and you load your model,

12:30now we are ready to open our Aura AI and

12:34chat with it. So I'm going to go into my

12:36desktop. I'm going to go to Aura AI and

12:39click on my HTML. And you can see now

12:43that the connection here are all green.

12:45Now for each of these elements if you

12:47don't know what they do you have a

12:49little question mark next to it and it

12:51will tell you what that lead is about

12:53LED is about or well the element is

12:55about in this case the first here is LM

12:57studio and the second one is our old

13:00talks TTS and the third one is our

13:03microphone ear so is ST server which

13:07allows us to actually use the mic. So

13:09now we're going to test the the AI and

13:11the mic if it works. So I'm going to go

13:13into the mic down here. But before I do

13:15that, I've got a couple of check boxes

13:17here. So the first one here is called

13:19stream. If you don't know what that

13:20does, make sure you hover the cursor

13:22over the the question mark and it tells

13:25you when on the AI replies appears word

13:29by word like a like a conversation. When

13:31off, the chat waits until the full

13:34answer is ready before showing anything.

13:37I wanted to have that feature. So while

13:39I'm talking it type something just to

13:41have some sort of a satisfaction of see

13:43if it's responsive or if actually is

13:46listening properly of what I'm saying

13:49and also a way to test your mic levels

13:51as well. The other one here is a is the

13:55auto send response. So if I want after I

13:58talk I can click on the mic again and

14:00that will auto send the text into the AI

14:05conversation. If I don't want to do

14:07that, I can uncheck that box and still

14:10see if everything has been typed

14:12correctly before I send it to the AI.

14:15I'm going to click on the mic and I'm

14:16going to start to talk. And of course,

14:18every time you uh click on the mic, you

14:20have to agree or to allow the browser to

14:24access your mic. And this will happen

14:26each time you use the mic because it's a

14:29local HTML file. And I think all the

14:33browsers do this unless you have a

14:35browser that automatically accept that.

14:37But I think after you allow this allow

14:40while busying the site, it won't bother

14:42you anymore once you retype the

14:44microphone. But each time you relaunch

14:46Aura AI, you would have to allow it at

14:48least once. I'm going to click on the

14:50first one. Hi, how are you? I am testing

14:55this local AI and wanted to just check

15:00if what I'm saying is been typed

15:03correctly. Let me know what you think

15:06about this interface. Now I clicked on

15:09the mic and now everything is typed

15:12here. Probably one thing I have to fix

15:14here. I wanted to be able to see the

15:16whole paragraph because it looks like it

15:18does this all in one line. But this is

15:20something I'm going to fix probably

15:21later on. So, it looks like it's typed

15:24everything correctly. If I'm happy with

15:25that, I'm going to click on send. We

15:27also have an option to add a an emoji or

15:31just to make a little bit the

15:32conversation a little bit more fun and

15:34click on send. So now the LM Studio that

15:40works in the background is loading the

15:43my reply my text and you can see how

15:46many tokens is taking. This is a small a

15:50relatively small model 12.67 GB. It

15:53shouldn't really bother my RTX, but we

15:56also recording. So now it's basically

15:59cooking that in the background. It's

16:01going to now type the response. And I'm

16:04going to hear also the response as well.

16:07And I'm going to show you also how to

16:09customize the voices here on the right

16:12hand side of the interface.

16:15The first time you launch this, it might

16:17take a little bit longer, but as you can

16:19see, after that, it should be fairly

16:21quick. As you can see, now it's typing

16:23everything. And it's not going to type

16:25everything all in one chunk. It just

16:27does this uh gradually.

16:30And it says, "Hello, I'm doing well.

16:32Thank you for asking to answer your

16:34questions." And it's going to basically

16:35answer. So, this is a quite long

16:38response. And you can also customize the

16:42response here on the right hand side

16:44under reply length. This will basically

16:47give you the option to go all the way up

16:51to 8192. I think if you go and if you

16:54want to find out how to use all these uh

16:57options, just need to go on the question

17:00mark and it will tell you range from 64

17:02to 81 192. Now it's about to talk. So

17:05I'm going to shut uh shut my mouth and

17:07let it talk. But as you can see here,

17:10you have all the explanation how to use

17:11it.

17:12>> Hello, I am doing well. Thank you for

17:15asking. To answer your questions, one,

17:18is what you typed correct? Yes,

17:22everything you typed is perfectly

17:24correct and readable. Your sentence

17:26structure is clear, the grammar is

17:28accurate, and there are no typos in your

17:31message. Hi, how are you? I am

17:34>> Okay, I'm gonna stop there because

17:35otherwise gonna read the whole thing.

17:37Um, and by the way, you can also choose

17:40to stop the speech if you decide to. You

17:43can also do a bunch of things here at

17:46the bottom. You can read aloud if you

17:48want that to reread. You can edit the

17:50response. You can regenerate. You can

17:52branch in new chats. And you can delete

17:54as well. And you can do the same thing

17:56on your reply on your text. When you

17:59hover over the uh, you know, on the

18:01message, you see the option to edit and

18:03to delete. You can also change the image

18:06avatar for you and your AI bot and I'm

18:10going to show you how to do that in a

18:11second. But for the voice, you have here

18:13a list of voices here on the right hand

18:15side. You have female Z01 Wave. So when

18:20you click on it, you see you have a list

18:22of all available voices. You can also

18:25add your own voices. And by the way, if

18:27you want to add your own voice, all you

18:29need to do is simply go to the folder,

18:31the old talks TTS folder, which you're

18:35going to have after you install it.

18:36Under voices, you can simply drag and

18:39drop all the W file that you like. Of

18:41course, don't import something that is 1

18:45hour long. Trying to make it that it's

18:47either 10 seconds, 30 seconds, one

18:50minute maximum. I think two minutes or

18:52something like that. I had that. I added

18:55some of them quite long and it works

18:58fine so far. So, make sure you add them

19:00here. After you add them here, you need

19:03to just refresh the page by clicking on

19:06the refresh button and they will appear

19:08on this list. If I want to select any

19:11other any other voice, I can just go

19:13here and I can go for instance to let's

19:17say male01 and I can test the voice by

19:20clicking on test voice here at the

19:21bottom. Hello. This is a test of your

19:23local text speech voice. It works.

19:27>> Okay. So, if I want to choose a

19:28different one, I can go to let's say

19:30female 01.

19:32>> Hello. This is a test of your local text

19:35to speech voice. It works.

19:37>> So, I quite like this voices. They're

19:40not bad. And the fact that I can

19:42customize mine and I can put mine in,

19:44it's quite nice. The very reason why I

19:46created this uh local chat interface is

19:49because I hated using Silly Tavern. I

19:51hated using those uh interface. I have

19:54to open text generation web UI and then

19:56I have to open silly tavern and then it

19:58they weren't linking properly the web

20:00bugs. So I decided I'm going to do my

20:03own local AI. And of course you need to

20:07launch at least LM Studio and you need

20:09to double click on that uh file that

20:11will launch both servers. But after you

20:15done that, it's the easiest way. You

20:16double click there, you open LM Studio

20:18and then you open this and it's ready to

20:20go. You see the connection already going

20:23and make I made it as easy as possible

20:25for you to use. In terms of linking LM

20:29Studio with this uh just I want to

20:32reiterate that you'll find LM Studio

20:34here on top which is already have the

20:37HTTP uh address which is this one here.

20:40Make sure at the end you have a forward

20:42slashv1 which you probably going to find

20:44already in the installation. And on the

20:47lm studio you also see here the uh it

20:50says is reachable at 18 1271234

20:55but it doesn't mention slashv1. So that

20:58it's quite important to actually add

21:00that into your interface in order for

21:02them to talk. The other thing you can do

21:04you can customize your character. You

21:05can add here a personality. Let's say I

21:08would say something like uh be brief in

21:12your answers and to the point. So I can

21:18do that. Um I can also select and

21:22narrate. If I want to have a narration,

21:24let's say if they answer an expl and you

21:26have like an explanation like a scene

21:28narration and you want to have a

21:30different voice, you can just simply

21:33tick this box and choose the voice you

21:35want to have as narrator. And the voice

21:38on top would be the voice of your

21:40Jarvis, your local AI assistant. In

21:44terms of uh other options here, you got

21:46down below is voice cap. Voice cap is is

21:50basically how many characters the voice

21:52will read after generate a response. So

21:56I put it zero. So it means it will read

21:58the whole paragraph. But if you want to

21:59only have for instance the first

22:01paragraph or so all you need to do is

22:03simply choose any many any characters

22:06from here and you can just do that

22:08creativity. So all these options here I

22:11saw some of them in silly tabvern which

22:13has like hundreds of those and have some

22:17crazy terminology which I don't know I

22:19need a translator to understand that

22:21interface. So I thought let me just make

22:24one which is super clear for everyone

22:26and easy to use. So the creativity

22:30allows you to to range from zero to two.

22:33So when is zero is dry, factual and

22:37predictable. When is for instance 07 is

22:39balanced as a default is 07. When is 1.5

22:43and up is going to be wild, poetic,

22:45imaginative, etc. So I put 1.8. This is

22:48going to be more like a test to see how

22:51it behaves. In terms of reply length, I

22:54put 20 48, but you can choose a range

22:57that goes from 64 all the way to 81 92.

23:02So this is something you can have a very

23:05thorough answer if you want to or a very

23:07short answer by just limiting this. The

23:10variety will let you choose between zero

23:13to one. And this is going to be how

23:16adventurous the AI is with a word

23:19choice. So if you choose something like

23:2209 is going to be if you want the the

23:26the chat to be balanced. I put one

23:29because I want the AI to be varied when

23:32it talks. The repetition also same

23:34thing. If you have a range from zero to

23:37two, how hard you want the AI to try not

23:41to repeat itself. I put 1.1. And in

23:45terms of memory, this will uh I would

23:47probably suggest you to leave this at

23:48zero. But this will remember all the

23:51conversation within the same thread. But

23:54if you want that to be less to remember

23:56only the last 10 or 15 replies, you can

23:59just choose 10 or 15 here on this

24:01number. In terms of look, and this is

24:03the fun part, you can choose your own

24:06avatar. In this case, I put my own uh

24:08image here. If I want to change my

24:10avatar, I can just go to my avatar and

24:13choose a different image. So, if I go to

24:17let's say downloads, I'm going to go to

24:20image. I'm going to choose an image of

24:23me. Let's say this one here. And this

24:25will change the image avatar for me. If

24:28I want to have my AI avatar as well, I

24:31can click on AI avatar and I can choose

24:34an image. Let's say on a desktop. Let's

24:36say I've got this image here. And by the

24:38way, you can choose either images with

24:40with or without the background. And I

24:42show you why would be beneficial to have

24:45that. So I'm going to choose let's say

24:47this one here with background and click

24:50on open. This will open the the image

24:53avatar here on the uh left hand side as

24:56well as here on the circle.

24:58Now if I want to customize this, if I

25:00want to move it around, I can just go

25:02down here and uh I can basically change

25:06the portrait size. So this will change

25:08the portrait size and as you can see

25:10goes beyond my chat. I can move it

25:13around. So I can click on uh move mode

25:16and this will allow me to move this

25:18image wherever I want in the chat. Let's

25:20say I want to have it here and I want to

25:23have it a little bit smaller. And you

25:26also have the option to flip it. If I

25:27click on flip, it will flip it as well.

25:30And this is purely graphic. If you want

25:32to have like a more interactive

25:34experience with your AI, you want to

25:36have it like this. You want to have

25:38whatever image you like. After you

25:40finish, you can click on move mode off

25:42and the portrait and the fade. This will

25:46allow you to fade your charts in case

25:48you have a background. And I'm going to

25:50show you how to choose a background. I'm

25:52going I'm going I'm going to go to my

25:54background here and I'm going to go into

25:57pictures and choose one of these

25:58background here and click on open. And

26:01this will open the background here. Now

26:03I can have the charts to be completely

26:05solid here, solid bubbles, or I can have

26:08a little bit shaded. It's up to you how

26:11you want to have it. This is purely

26:13graphic. After you finish that, you can

26:16decide to takeick it off if you want to

26:18have it solid or have it faded. It's up

26:20to you. And you can also reset the

26:22portrait position. So if I click on

26:24that, it's going to reset the position

26:25of the portrait there. If you want to

26:27move it again, you tick the box, you

26:29move it around, and you tick the box

26:31again once you finish the to position

26:33it. So, another thing you can do is back

26:36up everything. This will back up all of

26:38your chats in a zip file and it will

26:41organize them into conversation. You can

26:44restore a backup by clicking on restore.

26:47Select the restore zip file and it will

26:50restore everything in your chat

26:52interface. and the test voice. We

26:54already saw it before how it works. So,

26:57as you can see, this is something that

27:00you can use for voice. I'm going to test

27:03again with another reply. And this is

27:06going to be more like just a short

27:09conversation. I'm going to tick this

27:11box. This will basically auto send my

27:15reply. And click on the microphone

27:17again. Hi, how are you today? So, I

27:20clicked on the red button so that it

27:24will automatically send it for me.

27:27>> I'm doing well. How can I help you with

27:29your testing today?

27:30>> I'm testing the mic and how the

27:33interface behaves with the text. Can you

27:37help me understand about a product that

27:41I just bought? So, I'm just, you know,

27:44improvising here just to see if he can

27:46answer me fairly quick. Again, I'm using

27:49an 11 GB model, so the response is

27:53fairly quick.

27:54>> Yes, I can certainly help with that.

27:57Please let me know the name of the

27:58product and any specific features or

28:00concepts you would like explain.

28:02>> I've got an Herman Miller chair that I

28:04just purchased secondhand, and I would

28:07like to know where I could find the

28:08cylinder replacement for it. So this uh

28:12specific model that I loaded, I think

28:15it's uh updated up to the December 2025.

28:18So it will probably come up with

28:20whatever knowledge he has about it. But

28:23as you can see, he's already replying

28:26and uh and basically it's already

28:27telling me it's asking me about the what

28:30model.

28:32Now, it depends on how long the response

28:35is, but if you want to have a back and

28:37forward response, you probably just want

28:39to have different kind of conversations.

28:41But for this kind of stuff, you can just

28:44uh let it run. To identify the correct

28:46gas lift cylinder replacement for your

28:48Herman Miller chair, follow these steps.

28:51One, identify the chair model. Herman

28:54Miller uses different cylinders

28:56depending on the model, eg Aeron, Mirror

29:002. You must know which model you have

29:02before purchasing a replacement. Two,

29:04where to buy Owen part's original

29:06manufacturer.

29:07>> Okay, I'm going to stop there, but as

29:09you can see, it's answering you. It's uh

29:12fairly quick. Now, I can also attach

29:14images. Now, as you can see here, I've

29:17got media. When I click on media, I'm

29:19going to select my uh media folder. At

29:23the moment, when I open the chat for the

29:25very first time, after let's say one day

29:28or two, I reopen it. I need to go and

29:30select the media because apparently it

29:32forgets it. But I will fix this uh

29:35probably later on. After you select your

29:37media now, if I go to upload an image,

29:41um I'm going to ask to describe the

29:42image. I can do that. So for instance,

29:45I'm going to go into LM Studio. At the

29:46moment, this model does not have vision.

29:49So I would have to make sure that the

29:51model I load here has a vision feature.

29:54I'm going to go for this one here, 20

29:57gigabyte. Okay. After I done that, I'm

30:00going to go here. I'm going to go to the

30:02attachment here and I'm going to select

30:04an image from the desktop. Let's say I'm

30:08going to choose an image of I'm going to

30:10choose this image.

30:12So, I attach the image. I'm going to

30:15dictate this. Please describe the image

30:18I attached. And as you can see now, the

30:21this interface is going to do the job.

30:23It's going to analyze the image and it's

30:26going to describe it for me. I wanted to

30:28show you here under the desktop under

30:32Aura AI under media chart. When I double

30:35click there, as you can see, the image

30:36has been saved here as well. I've done

30:38it this before. I was testing it. So

30:41each time you upload an image in your

30:42thread, whatever charts you are in, it's

30:45going to be saved in that folder. I

30:46found it useful to have that um attached

30:50and well organized in the same chat. If

30:53you have any idea of any features that I

30:55haven't thought about, please leave it

30:57in the comment down below. It would be

30:58very useful to uh to add those additions

31:02if uh if I found it useful. If you find

31:04it useful, I would uh try to implement

31:07them into future versions of this

31:10interface. So another thing you can do

31:12here in the interface uh is to save the

31:15character. So if you are changing the

31:17personality, if you uh have let's say

31:20done you have done all these changes

31:22here and you want to save it, all you

31:23need to do is simply click on save

31:25character, give it a name. I'm going to

31:26call it let's say I don't know Jarvis or

31:30whatever and click okay. I'm going to

31:32have that saved into this list and I can

31:36choose this next time I open the

31:38interface or if I want to uh delete it,

31:40I can click on delete and this will

31:42delete the character for me. So that is

31:46all I've done so far with this. I spend

31:49a couple of days at the moment and 27 or

31:51more iterations to you know to to write

31:54this HTML. Of course I didn't write it.

31:58The LLM did it but I basically test it

32:00and amend it so far. If you want to know

32:03how I did this, please click the link

32:06down below which I basically share all

32:08my um thought process and all the what I

32:11went through in order to develop this.

32:14So you see exactly behind the scene on

32:16how I did it if you are curious to know.

32:19Basically you're also going to get the

32:21interface itself with that package. It's

32:24just a small little uh contribution to

32:27support the channel and hopefully you're

32:29going to use this chat as well. So let's

32:32have a look what's going on here with

32:33the with the model. Bear in mind that

32:36when you open the model for for the very

32:38first time it's going to take a little

32:39bit longer to answer your queries. But

32:42as you can see, it's almost done here.

32:45And but after it loads the uh the model,

32:49the subsequent replies should be fairly

32:52quick or reasonably quick. At the

32:54moment, I chosen a 20 GB model which now

32:58goes over my graphic card which is 16

33:00GB. So some of them is offloaded into

33:02the memory system. Let's see if it's

33:05actually working and describing the

33:07image. And as you can see now is

33:09answering the question is describing the

33:12image. The portrait features an African

33:14woman's face in closeup against a warm

33:16cream background and etc. is doing the

33:19whole explanation and this is done you

33:23know word by word at the moment because

33:24I asked the you know the LLM to uh

33:27develop this HTML in such a way that

33:30when the LM studio is analyzing the

33:34answer it will write you know straight

33:36away without me waiting for very long

33:38time. So I can actually start to read

33:41the answer before it finishes

33:44completely.

33:46Hello everybody. In this video, I'm

33:47going to show you how to install all

33:50talks DTS into your machine and how to

33:52use my patcher that you find in the zip

33:55file. So, first things first, let's go

33:58into the desktop and we're going to go

34:00into Google and type all talks DTS on

34:03screen here. The first search should be

34:06the one that we're looking for, which is

34:08all talks DTS. Going to click on that.

34:11And here you'll see we are under GitHub.

34:14So, all you need to do is simply go into

34:16the code here on top and download the

34:19zip file. So, after you download the zip

34:22file, you're going to make this a little

34:24bit smaller.

34:27The zip file should be on your

34:29downloads, which is this one here. We're

34:31going to extract that. Uh, let me just

34:33minimize this for a second. We're going

34:35to go into right click, extract all, and

34:39we're going to extract it here on the

34:42download folder. Now, we're going to

34:44double click that. And this is the

34:45folder that we need to drag into your

34:48into your AI folder or into your system.

34:50I'm going to drag it into my AI folder.

34:55Move it there. So, it's here. I'm going

34:57to click on it. I'm going to name it

35:00test_all

35:04talk. This is just for you guys that you

35:07can see it. Now, when I double click

35:09here, now you have all these files. We

35:12need to basically start the setup here.

35:14We're going to double click on at

35:16setup.bat.

35:17I'm going to double click there. I'm

35:19going to run it.

35:21And now the choice here are two. We're

35:24going to use it as a standalone

35:26application. We're going to press two

35:28and press return. And then we're going

35:30to press one here to install the the

35:33files into our system. This is going to

35:35take around 5 to 10 minutes depending of

35:38your internet speed. and it's going to

35:41install all the files needed for all

35:44talks to work and also the environment

35:47as well. Be patient and uh let's wait

35:50until this is done.

35:55Now, if it the installation failed or

35:57failed the the AV to run properly here

36:01to build a wheel for AV, do not worry

36:04because my patch is going to do the job

36:06and uh fix these issues. So I spend uh

36:09two or three days only to sort this out.

36:12So as you can see now it's ready to be

36:15launched. What we need to do is simply

36:16just press any key to close the

36:21PowerShell and then we're going to press

36:23nine here to escape and go back here.

36:26Now we need to use our patcher. When we

36:28use our patcher, we're gonna go into the

36:31zip file that you find in the download.

36:35And

36:37we're going to go into our desktop.

36:39That's what my Aura installer is. And

36:42there we go. So, these are the file you

36:44find into the zip file I gave you. And

36:47uh the one that we need to look at is

36:50Aura AI_installer.bot.

36:53So all you need to do is simply go back

36:55into the main folder called test all

36:58talk or whatever name you gave it and

37:01basically we drag this folder and drop

37:03it into aura installer.bat.

37:06This will start to install

37:10all the files that that are needed for

37:12uh text to speech. Actually, it's going

37:14to fix texttospech files from the

37:17original installation and it's going to

37:20install ST the speechtoalk server as

37:24well. So, everything is going to run

37:26hopefully smoothly here. And after this

37:30is finished, you're going to see a

37:34launcher popping up on your desktop and

37:36there's going to be a file that you need

37:38to double click on. So you are ready to

37:40go and use the local AI directly. Let's

37:44wait until this one is done and then we

37:46are going to be back. The installation

37:48has finished. Now it says here, press

37:50any key to continue. I'm going to press

37:52return to close it. And on the desktop

37:55now we have start aura AI.but. But so

37:59before we double click on that, we need

38:01to open LM Studio and of course load our

38:04model which is our brain. So we're going

38:06to go into LM Studio here. So make sure

38:09you download LM Studio. The interface is

38:12going to look like this. And uh all you

38:14need to do is simply go to the search

38:16here and download the model you want. I

38:19downloaded already a few models here.

38:21I'm going to my developer here on the

38:25left hand side on the top tab and I'm

38:28going to load my model. Going to eject

38:30this one because it's already loaded.

38:32All you need to do here, you're probably

38:34going to see this page. And then you

38:36need to go to load model here on the top

38:38right hand corner. I'm going to choose a

38:40small model so we can test it quickly.

38:42The Gemma 412B, the the first one here,

38:46which is going to be fairly quick to

38:47load. The other thing you need to make

38:49sure is that you go into server settings

38:53and make sure that enable cause is on.

38:56This enable enables uh LM studio to

39:00share the model with other servers in

39:04this case our ST and TTS server. So make

39:07sure that you turn that on and then once

39:09this is loaded you need to just the uh

39:12click on the status to run it. That's

39:15it. You don't have to do anything else

39:17here. Okay, just need to minimize this

39:18and then we are ready to double click on

39:20start aura.bat.

39:23We're going to double click there. This

39:26is going to open all talks TTS and our

39:29ear as well and ST. And this is how the

39:34interface looks like at the time of this

39:37recording. And this is how the interface

39:39going to look like. and make sure that

39:41this connection here on top are all

39:43green at the moment is still loading our

39:46shell here. The power shell is still

39:48loading. The ST the speech to talk is

39:52already loaded. You can simply um

39:55minimize that and all talks now will

39:58start and if you uh start all talks for

40:01the very first time is going to download

40:04the models. Next time you relaunch it uh

40:08you don't have to do this. is going to

40:09be much faster of course, but the very

40:11first time you have to wait until all

40:13these models are downloaded. Here you

40:16got 1.8 GB to go through and another

40:18couple as well. So, make sure you wait

40:21for that. And then on the interface,

40:24you're going to have all of the uh

40:26traffic lights here in green. As you can

40:28see, my ST server now is loaded in port

40:327852.

40:34And the TTS now is waiting until this

40:37one is finished. And now you be ready to

40:40go and use your interface on your

40:44machine locally, privately. So let's

40:47wait just a few seconds until this is

40:49finished and then we're going to be

40:51back. That's it. So now we need to just

40:53uh minimize this and uh let's check if

40:56our chart now is all green. If it's not

40:59green yet, do not worry. Just refresh

41:02here. this little blue button on the

41:04side. And if it hasn't done it yet,

41:08don't worry. It's still loading. So,

41:10let's go back into our PowerShell and

41:12see if it's still cooking something. It

41:14says now ready.

41:17And now we going to minimize this. And

41:18as you can see now, the traffic light is

41:20all green. And if you want to check the

41:23voices, if you don't see them here,

41:24click on the refresh button. And now we

41:27got all the voices that you can just

41:30click and test. For those of you who are

41:32interested to know on how I developed

41:34this interface from scratch and you want

41:36to watch over my shoulder on how I did

41:38it, sort of like behind the scene, watch

41:41my following uh videos here on YouTube.

41:44I'm going to be releasing episode one,

41:46two, and three. This is actually episode

41:484, the one you watched, and I showcased

41:51the interface. But in episode one, two,

41:54and three, you're going to see exactly

41:57how I developed this uh this app, this

42:00little HTML. All the nitty-gritty

42:03details that you wanted to see are shown

42:06in the following episode. So stay tuned,

42:08subscribe, and I'll see you very soon in

42:11the next video. Bye.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.