Free YouTube Transcribe

Video transcript

OpenAI's Team Runs One Chat for Months. Here's How.

Dylan Davis · 3,792 words · 18 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00People inside of Open AI, the company

0:02that brought you ChatGPT, now run one AI

0:05conversation for months and it gets more

0:07useful the longer it goes. If you're new

0:09here, I'm Dylan. I run an AI

0:11consultancy. And for the past year, I've

0:13told my coaching clients the exact

0:15opposite advice. Start fresh chats early

0:18and often.

0:19That advice still holds most of the

0:21time, but the tools have changed. And

0:23now there's one kind of work where it

0:25backfires. I'll show you what changed,

0:27the two ways your AI now remembers, and

0:30the test that matches each one to the

0:32right job. So, let's get into it. Now,

0:34there are two key things that have

0:36improved both in Claude and ChatGPT

0:39around memory and how an AI can have a

0:40longer conversation with you. The first

0:42thing here is compaction. So, you've

0:44probably seen when you're using Claude

0:45or ChatGPT, once you've had a really

0:47long conversation, it summarizes. And

0:50the summarizing effect, what happens, is

0:52simply looking at the previous

0:53conversation, the AI determines what it

0:55feels is most important. It summarizes

0:57that to itself, not showing you, and

0:59then you can keep having that

1:01conversation ongoing. And that's why you

1:02can have longer conversations with AI

1:04today. Its intelligence doesn't

1:06automatically degrade. But it's

1:08important to note that this benefit is

1:10more obvious in different tools. And

1:12I'll talk more about that later. But in

1:13the past, probably not even a year ago,

1:15but probably even 6 months ago, when

1:17you're using these tools, if you had a

1:18long conversation with an AI, it got

1:21dumb fast. Its intelligence dropped like

1:23a rock after around 50% of filling up

1:25its head. And that's the reason this

1:26happens is the AI only has so much space

1:28in its head. So, when you fill the head

1:30too much, it doesn't have that much

1:31space to actually think about the task

1:33at hand. So, this is the first

1:34improvement that we've seen around

1:35memory and long conversations. The

1:37second improvement is native memory in

1:39the tools you're using. So, both Claude

1:41and ChatGPT have a native memory feature

1:44where it can remember things about you.

1:45And right now still, this is very

1:47surface level. So, most of what it can

1:48remember about you are basic facts. So,

1:51your name, your location, what you do,

1:53and some general preferences that you

1:54have that likely map across multiple

1:56activities. So, this isn't that

1:58detailed. Both of the providers,

2:00Anthropic and Open AI, are working

2:01heavily to improve these, but right now

2:03they're still not that great. And that's

2:04the second thing, or the second thing

2:06that's improved in the last couple of

2:07months. Now, to the point around how

2:09these benefits are not equally

2:10distributed. So, if you're inside the

2:12browser using Claude or ChatGPT, you're

2:15still likely not going to see massive

2:17uplifts in long conversations and the

2:18value associated to those. Reason being

2:21is that in desktop agents like Claude

2:23Co-work and Codec's, the ability to

2:25compact conversations and save memories

2:27more effectively is much better. So, you

2:29can have much longer conversations. Now,

2:31it's important to know if you already

2:32have a subscription with ChatGPT or

2:34Claude, you already have access to these

2:36tools. All you have to do is download

2:37them. So, Claude Co-work is for Claude

2:40and Codec's is for ChatGPT. Both of

2:42which are extremely easy to set up and

2:43easy to use. And that's my caveat here.

2:45So, if you're not using a desktop agent,

2:47I wouldn't recommend watching the rest

2:48of this video. But, if you're willing to

2:50try it out and or download it, then you

2:52can keep going. So, I'm going to walk

2:54you through two different setups. Both

2:55of which have different configurations

2:57of memory. And there are two primary

2:58forms of memory that we're going to mix

3:00and match for both of the setups. Quick

3:02pause. If you're enjoying this, you're

3:03going to enjoy two other things. First

3:05off, below is a 30-day AI insight

3:07series, completely free. You'll get 30

3:09insights in your inbox of how you can

3:10apply AI to your business and your work.

3:13The second thing is if you'd like to

3:14work with me, below are a series of

3:16offerings to see if there's a good fit

3:17between the two of us. Now, let's get

3:18back to the video. The first form of

3:20memory is written memory. So, this is

3:22the long-term memory the AI will have.

3:24And the reason it's written is the AI is

3:26going to externalize these memories into

3:28files that it can access later on. So,

3:30you can think of this kind of like a

3:31filing cabinet where really important

3:33detailed facts go. The other thing is

3:35working memory. So, you can think about

3:36this as short-term memory. And this is

3:38the memory the AI has for ongoing

3:41conversations. So, if each one of these

3:43dots on this line are representative of

3:45a compaction of a conversation, the AI

3:48will have enriched context of previous

3:50conversations

3:51when you interact with it in future

3:53iterations. So, this is the running

3:55conversation memory. Those are the two

3:57things we have. We're first going to

3:58start with setup A. And the reason we're

4:00starting with setup A is this is the

4:01setup I'd recommend most people start

4:03with and most people use for most use

4:05cases. So, 95% of the things you're

4:07going to do with a desktop agent to

4:08extend its memory are going to be in

4:10setup A. And this is where we're

4:11actually having many fresh conversations

4:13with the AI. So, we're still maintaining

4:15the old advice that I gave, which is

4:17start fresh conversations early and

4:18often

4:19with one caveat that you're actually

4:21going to externalize the AI's activities

4:23and memories so it can then reference

4:25them in the future. So, in this case,

4:26let's say that we have a folder

4:28dedicated to writing proposals for us.

4:30So, this is a finite task. And inside of

4:32this folder, we're going to have an

4:34instruction file. So, depending on the

4:35tool you're using, if it's CodeX, it'll

4:37be agent.md. If it's Claude Co-worker,

4:39it'll be claude.md. But, these are

4:40basically the instructions AI's going to

4:42look at.

4:43I'll show you what a simple version of

4:44those look like in a second. In addition

4:46to this file, we're going to have the AI

4:47create another file, which is a memory

4:49file. And this is basically a file that

4:51contains all the lessons it learned over

4:52time that it feels are meaningful enough

4:54to save for future conversations. The

4:56reason this setup is so powerful is you

4:58get both the strength of having a fresh

5:01chat with the AI, so you get the maximum

5:03intelligence to achieve a task, but also

5:05you get the benefit of previous

5:06conversations and memories associated to

5:09that, so you don't have to keep

5:10correcting the AI over and over again on

5:12certain things you prefer about a given

5:13task. So, this is our setup A, what it

5:15looks like at a high level. The prompt

5:17I'd recommend starting with for that

5:19claude.md or agent.md is here. And

5:21again, this is a simple version, so

5:22you're probably going to add to this,

5:24but this is a template you can begin

5:25with. So, at the very top of this, we

5:27start with the purpose. And the reason

5:29we start with the purpose is that when

5:30you use a desktop agent, you have to

5:32open it through a folder. So, you have a

5:34bunch of folders in your computer. You

5:35open up that folder with a desktop agent

5:37and it works inside that folder. When it

5:38opens that folder, it's going to look at

5:40these instructions. When it looks at the

5:41instructions, it first needs to

5:42understand, what is my purpose? Why am I

5:44here? That's what this line does. So,

5:46we're saying this folder is for blank.

5:48So, you fill in the fill in the blank

5:49activity here for yourself. In this

5:51case, we could say writing client

5:52proposals. In addition to the purpose,

5:54we then have the files that are inside

5:56of this folder. So, when the AI reads

5:57the instructions here, it's going to

5:59read them every single time we start a

6:00new conversation with it, and it's going

6:01to know that there's a memory file

6:03inside this subfolder. And we're saying

6:05that I want you to read that file every

6:06time you start any task. The reason

6:08being is that it holds many of the

6:09previous lessons learned and preferences

6:11from other conversations. And finally, a

6:13bit of a maintenance part of the prompt

6:15is AI can actually update that memory

6:16file for us. We don't have to tell it to

6:18do it. It'll do it itself. And that's

6:20what this prompt here does. We're simply

6:21telling the AI, "If I corrected you,

6:23where you learned something worth

6:24keeping, I want you to add a short dated

6:27line to this memory file. And make sure

6:29you don't rewrite over old lines." This

6:31part here is important. Having it be

6:32short and dated. The reason we want it

6:34to be short, and I'll reference this in

6:36the future, but we want this memory file

6:38to be below probably 150 to 200 lines.

6:41Because AI is going to reference that

6:42file likely every time you interact with

6:44it. So, if it's too long, we're going to

6:46immediately fill up the AI's head and

6:48degrade its intelligence too quickly,

6:50and it kind of counteracts the whole

6:51point of doing this. So, we want to make

6:52sure that file is somewhat minimal. So,

6:54this is the prompt you're going to have

6:55for your instructions. And as a

6:56reminder, the type of tasks that fit in

6:58setup A are going to be tasks that are

7:00finite. There's end to it. And that's

7:02going to be like 95% of the tasks that

7:04you do. So, that could be writing

7:05proposals, generating reports,

7:07contracts, etc. There's a clear finish

7:09line. And that's setup A. So, now we go

7:11to setup B. So, setup B is when you have

7:13a pinned thread that runs on for weeks

7:15or months. And this is what most people

7:17do accidentally, and they degrade the

7:19AI's intelligence. But if you do this

7:20intentionally with the right setup, you

7:21can get benefit from this, and the

7:23memories compound. Before we get into

7:25the setup, there are two questions you

7:27need to ask yourself. If you say no to

7:29either of these, you should immediately

7:30go back to setup A. In addition to that,

7:32if you're even hesitant, you're like,

7:33"I'm not even sure if I should do setup

7:35B," just start with setup A, and then

7:36you can try the infinite thread later

7:38on. But the two questions you want to

7:40ask is the task I want to have a ongoing

7:43thread for, is there a finish line? If

7:45there isn't a finish line, it might be

7:46good for this setup B. Second question,

7:48is it beneficial to understand what

7:50happened yesterday in the thread for

7:52today's work? If so, there's a good

7:53chance that this might be suitable for

7:55setup B, which is the ongoing thread.

7:57Now, I've mentioned ongoing thread a few

7:58times. What does this actually mean and

8:00what does it look like? I'll show you

8:01both in Codex and Co-work how you pin a

8:03conversation so you can have that

8:05ongoing thread and not lose it. So,

8:06here's an example of Codex. So, on the

8:08left-hand side I have a bunch of

8:09projects and a series of chats inside

8:11those projects. Up here I have pinned,

8:13so you can see we have pinned listed

8:15here. And then over here we have a

8:17conversation that's pinned there. Now,

8:18if I wanted to unpin something, I would

8:20simply just select this pin and it would

8:22unpin that chat. If I wanted to pin a

8:23chat, I would just go to that

8:24conversation, select it, and then select

8:26pin. When I do that, you can see it

8:28automatically goes up to the pin

8:29section. That's how we do it inside of

8:30Codex. And in Co-work it's very similar.

8:33So, here we have Claude Co-work and

8:34actually a conversation I was having for

8:36building this presentation. And if I

8:37wanted to pin this, there's two ways I

8:39can do it. I can drag and drop it, so

8:41above here you can see we have the pin

8:42section. So, if I drag this up here and

8:45hover it, it's going to be pinned. If I

8:47bring it back down, drag it and let it

8:48go, it's unpinned. And another way I can

8:50do it is simply going to the three dots

8:52and selecting pin and it'll move it up

8:54there. So, that's how you connect a

8:56conversation to the pin section so you

8:57don't lose it and you can have a really

8:58long going conversation there. Now, what

9:00does this look like for certain use

9:01cases? Well, there are two primary use

9:03cases I've seen a lot of people get

9:04benefit from for ongoing conversations

9:07that are optimal for setup B. The first

9:08one is an inbox assistant. So, this is

9:10simply having an AI act as your

9:12assistant in triaging, researching, and

9:14drafting replies. The reason this is

9:16optimal is because it's the type of task

9:18that has no ending. You're going to keep

9:19on sending emails and drafting emails.

9:21And also, sometimes it's relevant for

9:23the AI to know what has happened

9:24previously in previous emails you've

9:25drafted to other people that might have

9:27interconnected topics or tasks. This is

9:29also the use case I've seen discussed

9:31most often from people at OpenAI. That's

9:33one use case. Another one is monitoring

9:35projects. So, if you're monitoring an

9:36ongoing project that's going to be maybe

9:386 months to 12 months, you can set up a

9:40pin thread for this and have an AI

9:42automatically check a series of systems

9:44and give you updates on that project on

9:46a daily or weekly basis. Now, as a

9:48reminder, if you have a task like

9:49writing proposals, which most of you

9:50will likely have finite tasks like this,

9:52it's optimal for setup A. And then

9:54another thing I've seen a lot of people

9:55run into issues with is they'll say,

9:57"Oh, I have this huge client that I want

9:58to have an ongoing thread for." That's a

10:00bad idea because for a given client,

10:02that's a topic not a task. So, what you

10:04would do in this situation is you would

10:05have a parent folder for that client,

10:07that important client, you'd put all

10:08your information in there. You have a

10:09series of subfolders, and each subfolder

10:11would have a given task for that client.

10:13So, answering questions, drafting

10:15proposals, reviewing contracts, etc. So,

10:17that's what this looks like in an

10:18applied way. And now for a pin

10:20conversation, it's important to note

10:21that within Codex and Co-work, I still

10:23recommend opening it through a folder.

10:25So, for those things that I showed you,

10:26those specific threads that we're going

10:27to have ongoing, you would start that

10:29conversation off in a folder so you can

10:31still get the benefit of externalizing

10:32memories. So, for example, if you're

10:33doing the inbox assistant, you would

10:35create a folder somewhere on your

10:36computer and call it email buddy or

10:38inbox assistant or something like that.

10:40You will open the thread in there and

10:41have that continued conversation. When

10:43you do that, you're going to create an

10:44instructions file like we did last time

10:45for setup A, but this is for setup B.

10:47So, this is going to be inside of either

10:48your cloud.md or your agents.md

10:50instructions files that the AI

10:51references every single time. And in

10:53these instructions, when the AI starts

10:55the conversation, it's always going to

10:56look at this file. It's going to see

10:57that your job is blank. So, in this

10:59case, it could be acting as an inbox

11:01assistant. So, you're supposed to

11:03triage, research, and draft replies for

11:05me. And then we lay out a series of

11:07rules that I recommend you copy. And

11:08this is around storing memory and open

11:10loops. So, the first one is stating that

11:12if there's anything worth keeping like

11:14decisions, prices, etc., I want you to

11:16write them to this memory file

11:17immediately. That's the first memory

11:19piece. And the second one is open loops.

11:21So, if you have an assistant that's

11:22monitoring your inbox, and there are

11:23open loops of previous conversations

11:25you've had or things you're waiting to

11:26follow up on, the AI can track that as

11:28well. So, we're saying I want you to

11:29track open loops, who owes me what or

11:31who do I owe something, and also what

11:33are we waiting on? And then add that to

11:35this loops file and update that specific

11:37loop file at the end of every single

11:38session. So, these simple additions to

11:40the prompt are ways for the AI to keep

11:42track of its own work and notify you on

11:44what's going on. And it's important to

11:45note with threads is that they will

11:48eventually have to be refreshed. So,

11:50even if you've had a thread for months,

11:52you'll start to see a degraded

11:53intelligence associated to the AI,

11:55irrelevant of the model you're using.

11:56And this will manifest in different

11:58things where the AI forgets obvious

11:59facts and starts to contradict itself in

12:01a variety of other things. When you

12:03notice this degradation of intelligence,

12:05it's time to start a new long thread.

12:07And the way that you would do this is

12:08you would just open up a new chat in

12:10that same folder. And the reason you

12:11open that up a new chat in that same

12:12folder is remember, you have the memory

12:14file and the loops file there. So,

12:16that's already some of the memory that's

12:17been externalized. The second and really

12:20important piece is that at the end of

12:21the old conversation, before you retire

12:24it, you want to ask the AI to give you a

12:26handoff document. So, this could be

12:27probably one to two pages summarizing

12:30any of the context it feels is suitable

12:32for a new AI to pick up where it left

12:34off based off of the existing

12:35information inside that thread. You'll

12:37get that document, you'll pass it off to

12:38the new AI in that new chat, and then be

12:41able to continue that long conversation

12:43in just a few minutes. And now, one

12:44important thing I want to call out for

12:45both setup A and setup B is that for

12:48those externalized memory files, again,

12:49we need to make sure they're minimal.

12:51They're not too long because the AI is

12:52going to look at them frequently and we

12:54don't want to blow the context window.

12:55So, I'd recommend running this on a

12:57weekly or monthly basis depending on how

12:59fast you're filling up that memory file

13:00for the tasks you're giving to the AI.

13:02And what this audit or pruning process

13:04is going to do is we're going to ask the

13:06AI to review our memory file line by

13:07line. As it's reviewing that file, it's

13:09going to flag anything that's stale,

13:11repeated, or no longer true. If it

13:12notices any of those, it's going to flag

13:15those to be removed and placed into an

13:17archive folder. So, we're not deleting

13:19these memories completely. Instead,

13:21we're putting them to archive folder

13:22that the AI then can get access to in

13:24the future if needed. And we're then

13:25telling the AI exactly why we're doing

13:26this because we need to keep the memory

13:28file below a certain line count. And

13:30then at the end we're telling the AI, "I

13:32want you to show me what you want to

13:33change in the memory file or delete or

13:35whatever else before you do anything,

13:37because we want to be the person that

13:38approves it before it actually happens."

13:40And that's specific for this pruning or

13:41audit process. And now as a quick recap,

13:44to extend the AI's memory to have really

13:45long conversations and get the benefit

13:47of compounding knowledge over time, we

13:48need to match the container to the job.

13:50That's probably the most important

13:51piece. And the two containers we can

13:53choose from are setup A and setup B.

13:54Most of the things you're doing with AI

13:56is going to be setup A, because most of

13:57the tasks you're outsourcing to AI has a

14:00finish line. So, that's going to be a

14:01dedicated folder skill to a given task

14:03that has externalized memories. But if

14:05you do have a task that you feel like

14:06falls into setup B, the long

14:08conversation, then you can ask yourself

14:10the two questions we asked previously.

14:11Does today's conversation benefit from

14:13yesterday's? And is there no finish line

14:15to this task? If the answer is yes to

14:17both, then there's a chance it might be

14:18setup B. If it is setup B, you still

14:20want to create a folder for that task.

14:22So, if it's an inbox assistant, you

14:24would create a folder for that. And you

14:25would open up that thread inside that

14:27folder, because you still want to

14:28externalize files for the memory. So,

14:30certain facts that need to be remembered

14:31that are important and detailed, you put

14:33into the memory file and don't rely on

14:35the conversation, because remember it

14:36gets compacted over time and the

14:38summaries get distilled down further and

14:39further and further. And for that memory

14:41file you're creating for either setup A

14:42or setup B, you want to make sure that

14:44it's dense and no longer than 150 to 200

14:46lines, because the AI is going to

14:47reference it frequently. We don't want

14:49to contradict ourselves and blow the

14:50context window too quickly. And that's

14:52it. So, as a reminder, two things. First

14:55off, below is a 30-day AI insight

14:57series, completely free. You'll get 30

14:59insights in your inbox if I can apply AI

15:01to your business and your work. The

15:02second thing is if you'd like to work

15:04with me, below are a series of offerings

15:06to see if there's a good fit between the

15:07two of us. Now, before you go, it's

15:09important to be aware that none of this

15:11works if you don't know what to write

15:13into those memory files.

15:15Every lesson worth keeping starts as the

15:17decision your AI made without you

15:18asking.

15:20In this video right here, I'll show you

15:21the one-line prompt that logs those

15:23decisions, so your files fill up with

15:26the right lessons and not guesses. I'll

15:28see you next time, internet.

More from Dylan Davis

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.