Free YouTube Transcribe

Video transcript

Stanford Webinar - Agentic AI: A Progression of Language Model Usage

Stanford Online · 6,708 words · 31 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:09INSOP SONG: My name is Insop.

0:11So today, I'd like to go over agentic AI, agentic language

0:17model as a progression of language model usage.

0:20So here is the outline of today's talk.

0:24We'll go over the overview of language model and how we use,

0:29and then the common limitations, and then some of the methods

0:33that can prove towards this common limitation.

0:37And then we'll transition into what is the agentic language

0:41model and its design patterns.

0:45So a language model is a machine learning

0:49model that predicts the next coming word given the input

0:53text.

0:54As in this example, if the input is the students open their,

0:59then language model can predict what's

1:04the most likely word coming next as a next word.

1:08So if the language model is trained with a large corpus,

1:13it is predict--

1:14it is generating the probability of next coming word.

1:18In this example, you can see, books and laptops

1:23have a higher probability than other words in the vocabulary.

1:28So the completion of this whole sentence

1:32could be the students open their books.

1:34And then if you want to keep generating the what's

1:37coming next, then we can turn them in as an input

1:41and then put it into the language model,

1:44and then language model continuously generating

1:47the next coming word.

1:50Then how these language models are trained?

1:54Largely two parts-- pre-training part and then post training

1:59part.

2:00And then first pre-training portion

2:03is the one that language models are

2:06trained with large corpus, texts collected from internet

2:12or books, or different types of texts, publicly available text,

2:16and then trained with the next token or next word prediction

2:20objectives.

2:21So once the models is finished in this pre-training stage,

2:26models are fairly good at predicting

2:29any words coming next as a word given the inputs.

2:35However, the pre-trained model itself is not easy to use.

2:42So hence, the post training steps are coming.

2:46And then these post training stage

2:50would include instruction following training,

2:53as well as the reinforcement learning with human feedback.

2:56And what this training stage means

3:00is we could prepare a data set in such

3:05a way that specific instruction or question.

3:08And then the answers or the generated output

3:12that is what the user would expect or more

3:19related to the questions and answers.

3:22So that's how the models are trained

3:25so that it's easier to use.

3:26And then also, it will respond to a specific styles.

3:30And then once this is done, and then additional training method

3:36is aligning to human preference by using reinforcement learning

3:42with human feedback, which is using human preference to align

3:48the model by using rewards schemes.

3:53And let's take a quick look, really

3:56quick look on the instruction data set.

3:58This is the template that we will

4:02use to train the model in instruction following training

4:06phase.

4:06As you can see, there is a specific instructions

4:10will be substituted in.

4:11And then expected output will be substituted in.

4:14And then this is fed to the model.

4:16And the model is only trained on the response

4:20part that is generating the output based on the given

4:26instructions.

4:29So language model that is trained on pre-training stage,

4:33as well as post-training stage, is

4:36quite capable of generating text given instruction.

4:43Essentially, it has a lot of world knowledge

4:46that could easily generate the outputs.

4:50So these are rapidly developing.

4:53And then these models are used in various application domains

4:58that we use day-to-day work such as AI coding

5:04assistance or domain-specific AI copilots

5:08or most widely known ChatGPT and related

5:12conversational interfaces.

5:14And then in order to use these type of models

5:18as for your applications or specific tools,

5:26you could use the cloud-based API calls towards the model

5:33provider or the model servers, or some other ways

5:38that you could also host the models on your local machines

5:44or even mobile machines for the models

5:47that are small enough to host on this compute

5:52constrained environments.

5:54So what does it mean by using API calls?

5:59So we step back.

6:01The language model is taking the input, natural language input

6:06text, and then generating the output.

6:09So that means we need to prepare a certain form

6:14of free-form text, natural language text

6:17as an instruction or a question, and then

6:20put them in a specific format that you could make an API call

6:24towards the model provider.

6:27And then the model provider takes that usually

6:30on a cloud environment, and then generate the output

6:33and then respond to your API calls with the generated output.

6:38Then your software around, that software

6:42that based on this model, will parse the output.

6:45And then use it as is, or maybe you

6:49could make a follow up LLM API calls

6:54to further generate the output.

7:00So the input to the model is again free-form text.

7:07So how you prepare your input, also known as a prompting

7:11is critical.

7:14So there are well-known best practices strategy,

7:20how you prepare your prompt.

7:22And here are some of them.

7:23And such as write a clear and very descriptive and detailed

7:29instructions that will help the model

7:32to generate the output that you want.

7:34And you could include a couple of examples.

7:38The form that you want to see as in the style or form.

7:43And also, you could provide the provide the references

7:48or context such a way that model rely on that context

7:53that you provide.

7:55And instead of just ask the model to answer right away,

8:00you could ask model to give model

8:03to time to think about it, such as reasoning,

8:08enable reasoning, or using chain of thought or COT method.

8:13And the next one is instead of asking model

8:19really complex task, you could break them down and then

8:23ask them chain in sequence, also chain the complex prompts.

8:30And the last one is something that

8:33is a good engineering practice.

8:36Have a good way of systemic trace and logging will help you.

8:41And also, automated evaluation is always

8:44helpful to develop your essential,

8:47to develop your progress on your application.

8:50So let's take a quick look at on each items, what that means,

8:56to get more familiar with.

8:58So write clear and descriptive instruction.

9:03As an example on the left, instead of asking short request,

9:09you could describe in detail so that model

9:13knows what you are asking because model

9:16doesn't understand.

9:17Model cannot read your mind so that means you need to describe

9:21what you want the model to generate the output for you.

9:27So this is always useful for using language model in general.

9:35And include few shot examples.

9:38So meaning that give keep model the example input and output

9:44that you would expect.

9:46As in this case, you have some type of consistent style output.

9:51Then what is the consistent style?

9:54You provide an input as an example input and example output

9:59that you would ask.

10:00And then finally, you ask your original questions.

10:05Then it's going to be--

10:06the model will generate the output based on your input.

10:10So few shot examples is always helpful to generate the output

10:15that you would want to generate.

10:19So provide relevant context and references.

10:24This is really helpful for many of the cases that

10:28are related to generating text-based

10:32on some factual information.

10:35So LLM can easily generate some incorrect output, also known

10:41as a hallucination.

10:43For those topics that it doesn't know or is not

10:48too confident about, so for those type of cases,

10:52providing context or references would always help.

10:56So here is the example prompt template

11:01that you would want to use in cases such as retrieval

11:05augmented generation that we'll look at in the following slides.

11:10Saying that only answer based on the input, based on the article

11:16that you provide, you could substitute your related

11:20reference.

11:20And then only answer based on these references.

11:25If the model cannot find the answer,

11:27you could just say that cannot find the answer.

11:31Then model will likely generate the answer

11:33based on your references.

11:40So this is important part that keep models time to think.

11:46In other words, instead ask direct question,

11:49you ask model to think through it

11:52or come up with its own solution.

11:55And then finally, compare and generate the output.

12:00So this is also known as a chain of thought.

12:04So here is one example that might not work

12:10in some medium-sized model.

12:15You could ask model saying that evaluate the student's solution

12:21is correct or not, and then provide a solution description.

12:25And then finally give student solution.

12:28And since the system prompt or your original request

12:34is saying that just answer it is correct or wrong.

12:38Model might not get it right.

12:40However, for the same model, that

12:43might not get the right answer for this

12:45if you prepare your prompt in a way that you could ask.

12:51First, work out your own solution to the problem

12:54first, and then compare your solution

12:56to the student solution.

12:58Then by doing this, model will generate its own solution,

13:02and then as it does, it will have an opportunity

13:06to provide good attention to all these inputs

13:10that it's from the original input as well as the output

13:15that it generated, and then turn towards the right answers.

13:20So reasoning chain of thought is always

13:23helpful to generate the generate output that you want to see.

13:31So here's an interesting one, probably

13:35easy to implement for your application.

13:38So instead of asking your request, that

13:41includes, say, multiple tasks in one request,

13:44you could prepare your prompt in small, simple stages.

13:51So how you do is prepare a simple prompt

13:56and then generate the output.

13:58And then prepend the output to the next stage two prompt.

14:03And then generate the output.

14:05And then again, prepend the output from the previous stage.

14:10And then generate the output in a third stage, like here.

14:13And then finally, generate the output

14:15that you would want to see.

14:16So by doing this, you may need to do it manually,

14:21or this can be done by LLM as we'll

14:26see in the following slides.

14:28But having a simple, clear task per each request

14:33would be a good way to do it.

14:41So this may not be obvious.

14:44But because many type of--

14:50as with many engineering application or development

14:55having a good way of keeping, tracing,

14:58logging will definitely be helpful for your development

15:02for debugging as well as auditing.

15:04So same principle applies to language

15:08model-based development.

15:09So keep track of the log is always good.

15:12And so that also relates to having

15:19a automated evaluation from the early stage of your development

15:24will definitely help you.

15:25In other words, you need to prepare question and answer

15:30pair, ground truth answer pair so that you could compare that

15:35against the generated output.

15:38And you could use a human to evaluate that.

15:42But that's usually a costly and time-consuming.

15:47So you may use to use--

15:49you may use a language model as a judge.

15:52Meaning that you could ask language model to evaluate

15:57model-generated output as well with ground truth output

16:03so that model can score the generated output's quality so

16:07that you can use that against your currently developed,

16:11your own applications.

16:14This will help.

16:15This will be very important because the language models are

16:20continuously and rapidly improving,

16:23as well as the methodology and the tools

16:26that you are using for developing language model is

16:30also rapidly developing.

16:32In other words, without clear evaluation,

16:36it is hard to make forward progress

16:39or even hard to change the model,

16:42change the different type of models.

16:44Because the models are rapidly developing also

16:49means that some models are rapidly deprecated,

16:53which means you may need to force to change your language

16:57model that you are using in your application.

17:00So having a good evaluation methodology upfront

17:05from the beginning will definitely help.

17:10So this is simple idea, but will be

17:15helpful for many applications.

17:19So instead of taking your input prompt as is and then process

17:23it, you may have some software or some model

17:28that you could detect the intention

17:30and then send it to a different prompt handlers.

17:33Or this is also known as prompt router.

17:37So based on the input query type,

17:41you may need to use simple prompt together

17:44with simple language model.

17:46And then this will both help in terms of operation costs

17:52as well as generate a more appropriate output

17:56together with a more relevant prompt with the language

18:03model that is more capable of that type of query.

18:12So maybe Petra, if this may be a good moment that we could take

18:20question if you could see.

18:24PETRA: Thank you so much, Insop, and such an inspirational

18:28talk already.

18:29Thank you so much.

18:30We will get to more details about agentic AI just shortly.

18:34We did want to provide you with some background

18:37and the progression what has been going on in the field.

18:41But I think maybe let's ask one question that came up.

18:46It's a little bit specific, but it

18:48might be what more people are wondering about.

18:51And is there any optimal amount of data

18:53to perform a good training or anything that you

18:57could advise people around the data available

19:01or data being used?

19:05INSOP SONG: Maybe I'll be short on this.

19:08So I assume the training here is meaning that fine

19:13tuning the LLMs in additional training on top of open source

19:19language model.

19:22Yeah, depends on your task.

19:26It would definitely vary because it's

19:29hard to say one or the other.

19:31But if you have enough data set or text

19:38that you would want to see, then you

19:40may come up with a simple question

19:42and answer pair or instruction following data set format.

19:46And then you could also make use of language model

19:49to further generate more if you need it more.

19:53But I would think you would start

19:56with, say, tens of data-- samples of data set first,

20:02and then see whether that makes the model behave what

20:05you actually see it to behave.

20:09And then you could add based on the result or signals

20:13from the initial quick test.

20:15Then you could additionally add more data set

20:18or data set samples possibly used

20:22language model to augment it or synthetic data

20:24that you could create.

20:26Great question.

20:27PETRA: Thank you so much.

20:29I see questions started coming in, which is wonderful.

20:32I think we will pause the question for now,

20:35and we will try to get to as many as we can at the end.

20:37But please keep them coming.

20:39Definitely makes the session more engaging

20:41and also let us know what you are interested in.

20:43Thank you.

20:44INSOP SONG: Thank you, Petra.

20:46So so far, we've been looking at overview of language model.

20:52Great.

20:54Very powerful models that are out there,

20:56many models out there, and then how we use.

20:59However, even there are still limitations

21:05for models that are available, that

21:09are listed here, such as hallucination

21:11is a well-known issue that models can sometimes oftentimes

21:20generate incorrect or incorrect information,

21:24particularly if it's related to some computation

21:27or some other specific area.

21:29So this is a problem that we want to avoid

21:36in your application domain.

21:39And other thing is there's always a knowledge cut up

21:42in data set preparation.

21:44So model creator prepare data set.

21:48However, they need to, at some point,

21:51cut up their data set collection and then use it.

21:55So model may have not seen the recent information

22:01or news as part of their pre-training data set.

22:06Lack of attribution-- so model can

22:09answer a lot of world knowledge questions

22:12and can answer those type of general questions.

22:17However, it may not tell--

22:18it's not going to tell you where they

22:21drew this, the answer from particular specific data source.

22:28Data privacy is one fact that model creator,

22:34you prepare the data set using publicly available data source.

22:41That means model have not seen your proprietary data set

22:46from your organization or particular domain.

22:51And limited context length--

22:54although it's a rapidly increasing,

22:58however, it's fine balance because providing longer context

23:08will give more context information to the model.

23:12However, it comes with operational cost as well as

23:17the speed of the latency in text generations.

23:24So in order to address these common limitations,

23:32retrieval augmented generation is

23:35one way to handle this, such as it could

23:38reduce the hallucination by using

23:41the actual relevant reference.

23:44It also address the citation because it knows where

23:51this reference is coming from.

23:53And then this will allow you as an application developer

23:59or a system developer to prepare systems

24:04so that you could use your own proprietary data

24:07set or the text.

24:08And then you could use good use of small number of context links

24:16because it only select relevant data set.

24:20So how it works is you could pre-index your own data set

24:28or your own text by turn them into smaller chunk, chunk

24:35of text, and then convert them into an embedding space using

24:40embedding model, and then stage them

24:44as part of your database or vector database.

24:47And then when the request or query came,

24:51you could turn this query into embedding space

24:56so that you could do nearest neighboring search and then

24:59select top K relevant information, relevant chunk,

25:05text chunk, and then place them as part of your prompt.

25:09Some of the slide that we see previously

25:12is you put the reference as part of the prompt,

25:15and then use that as the model.

25:19Only make use of this reference.

25:22So this is one good way to make use

25:26of your own proprietary data.

25:28And then similar method can be used in the actual AI search

25:36so that instead of using index data set,

25:40you could also rely on web search or different type

25:42of search so that you could provide the information as part

25:46of the index.

25:48And one of the thing also mentioned here

25:52is there are many methods or ideas for retrieval augmented

25:59generation.

26:00More commonly used method is something

26:04that we've just mentioned.

26:06We've just talked about, meaning turn the text

26:09chunk into embedding space.

26:11And then do a nearest neighbor search.

26:13However, there are many methods.

26:14And then you could also use knowledge graph base.

26:19So if you could generate a knowledge graph from your text

26:24source, and then that could also help to extract

26:28the more relevant information.

26:31Also known as graph RAG is one part of it.

26:35However, there are many methods and then

26:37you may need to look into the right method or if right method,

26:44make use of those.

26:48Tool usage-- so language model being most widely used form

26:54is a text in and text output, which means that it could

26:58answer many type of queries.

27:02However, it's not going to execute or extract information

27:08from the external.

27:09So that's where this tool usage came

27:13to rescue or also known as a function calling.

27:16So with this method, you could get real-time information,

27:22or you could actually do a computation

27:26by generating software or computer code.

27:30So what does it mean is--

27:32let's look at an example here.

27:35So if you have AI chatbot that if you

27:42ask what is the weather in, say, San Francisco,

27:45then model will not know it in itself.

27:48So however, if you tell model previously,

27:54as part of the prompt, saying that if you

27:57ask weather-related question generated output form

28:02that the software that parse the output can make an API call.

28:08So as in this example, model will generate the output

28:13in a way that, hey, this is the case for tool usage.

28:18So it generates an output as in the form here, get weather,

28:25and then input argument to this API call or function call

28:29is the place that we ask.

28:31So then software receives this text output from the model,

28:36and then parse.

28:37And then actually make an API call

28:40towards the weather provider, and then get

28:43the weather information.

28:45And then again, provide back to langpack back

28:49to the language model.

28:50And then language model will generate

28:53more human-friendly or helpful output based

28:57on this API-based result. And for some cases,

29:04model can also generate a software code that

29:09can be executed as part of the sandbox outside of the language

29:14model by the software that is coordinating

29:17all these activities.

29:22So agentic language model--

29:26so there can be many definitions.

29:30One definition is it could interact with environment.

29:35So compared to simple language model usage,

29:38generally use simple language model usage

29:44as you seen here, text input and text output.

29:48Agentic language model usage could

29:52be language model could do something with the environment

29:57by generating tool usage or retrieval request.

30:01And then from the environment, anything outside

30:05of language model could provide an output,

30:10could provide an information that can be fed back to language

30:16model as an observation.

30:18And then the whole thing, the agentic language model,

30:23which includes language model at its core

30:26with the software around it, will process it,

30:30and then also put them in a memory,

30:33and as with its conversational history,

30:38which can be taken as a memory.

30:41So this is one way of definition of agentic language model.

30:48The other way to look at it is this.

30:53Agentic language model usage can be

30:58defined as it could reason as well as it could action.

31:02It could do an action, so also called ReAct, reason and action.

31:07So reasoning part is something that you could encourage model

31:12to reason about by using a method such as chain of dots,

31:18and then doing an action using a method that we have seen

31:23in previous slides, as in retrieval or search engine,

31:27or actually using calculator by making an API call

31:31or different type of API calls, such as weather API that we have

31:35seen, and also generating a Python code

31:39so that you could run as part of your sandboxes.

31:43So by combining these reasoning and action,

31:47model can do a lot more complex task

31:51than simple input and output type of interaction.

31:57So let's look at a little more detail on this.

32:03What does it mean by reasoning and action?

32:06So reasoning part, instead of doing the task that is asked,

32:14you could prepare your prompt in such a way that ask

32:19model to break down the task and then make a plan.

32:22So instead of breaking down the task by yourself,

32:28as we've seen in the previous chaining the prompt slide,

32:32you could ask model to break down and then

32:36prepare your task so in other words, plan the actions.

32:42And then based on that breaking down,

32:45model can generate different actions

32:48by making API calls or tool usage

32:53so that it could extract or collect additional information

32:59from the external world.

33:00And then by combining all these, put them in a memory

33:04so that it knows, model knows what's been happening.

33:09And then based on that, finally draw an answer for you.

33:13So let's look at a concrete example here.

33:18So if you have a customer support AI agent,

33:23then how it might look?

33:25How it might work?

33:26So as a customer, ask, can I get a refund for product full?

33:35Then agentic system will break down this task,

33:39break down this request task into the following four

33:43different type of actions.

33:46Check the refund policy, check the customer information,

33:49and then check the products.

33:51And then finally, collect and then decide what to do.

33:56In each step language model will spit out the API calls

34:02so that it could collect the information.

34:05For example, check the refund policy,

34:08a language model could ask a retrieval system

34:13against the pre-indexed company policy, refund policy.

34:17And from there, it could retrack the information,

34:20and then put them in its own context.

34:22And then using that, also request the customer order

34:27information.

34:28It could either ask customer, in the chat format,

34:32collect more information.

34:34Or it could look it up in the system

34:37because it depends on how this chat system is prepared.

34:43The same thing for the product so that it

34:46could collect more information.

34:47And then finally, draw the conclusion

34:50based on the policy, and then product information,

34:53as well as the customer order information.

34:56And then send the request to the follow-up system as an API call

35:04as well as the send, prepare, say, response draft.

35:09And then that's going to be handled by final approval.

35:17So workflow is generally like this.

35:21So in a sense, agentic language model system

35:25is generally language model is making iterative calls

35:32by reviewing the document or task

35:35and then making external tool calls.

35:40An example, if you want to do some research

35:45of certain matters, you could prepare your agent

35:49to do a research, web search or different type of search,

35:53and then summarize them iteratively, and then

35:58finally prepare the report to you or to your system.

36:03Another example could be software assistant agent

36:08that you could ask this software agent or free agent that

36:16ask the issues of certain type of software bug or issues.

36:20Then this agent will look it up and review this issue,

36:25and then collect the relevant piece of code or files,

36:30and then review them and then propose the output.

36:33Or it could also execute in its sandbox environment,

36:38and then test the fix, and then get the output.

36:44And then iteratively try to find the fix.

36:46And then finally, pulls the pull request or the changes

36:52to users or developers.

36:57These are the ways that we can use language model

37:01in agent format and by doing interactive language model

37:07calls.

37:10So the main reasons and difference

37:18why these agentic language model usage is getting more widely

37:25used is given that if you have the same model, if you ask just

37:32direct request to the model, model

37:35may not be able to handle it.

37:38However, if you put your task in this type of agentic format

37:44or patterns, then model will do more complex tasks,

37:49even using a model that may not be able to do it

37:52if you don't do this way.

37:55So that's one of the reasons that agentic language model

37:58is pushing the boundaries, so that the things that we

38:03can do with AI agent is more complex or different domains

38:10that we can rely on.

38:15So here are the real-world applications--

38:19software development, code generation or bug fixing

38:23or this type of development is widely investigated

38:30or being researched by different organizations

38:34as well as there are companies that

38:36are trying to provide these services as decision

38:39on the right side, and research and analysis that

38:43gather information, synthesize it, and then provide

38:46a summary for the users.

38:49And then task automation is one of the areas

38:52that agentic method could be used.

39:00So to make it more clear, here are some of the design patterns

39:08that you could use agentic language model.

39:13Planning is critical because by asking a model

39:19to break down the task to make it simpler task or clear task

39:26so that language model later can make an API call

39:30or use the tool usage.

39:32So planning is critical.

39:35Then reflection is something that model

39:38can generate in the cell.

39:40And then the next model call can criticize

39:49the output that actually came from the same model.

39:53So by doing this, the output could be improved.

39:57And then two usages is something that something that outside

40:03of the language model that you need,

40:04the real-time information or different type of information,

40:09then you can use this.

40:10And then multi-agent collaboration

40:12is one way to handle this.

40:16Reflection is a pattern that quick to implement

40:20and then leads to a good performance.

40:23And let's use a concrete example here.

40:27So if you want to refactor a programming code,

40:30instead of asking model to improve it right away,

40:34if you do this pattern, as in this example,

40:37saying that you asked ask a model

40:40saying that here is the code, and then check the code

40:42and provide constructive feedback.

40:46And then take that feedback to the second prompt,

40:51as in this example.

40:53You could also prepare prompts saying

40:56that here is the code and the feedback,

40:58which came from the model itself.

41:01And then ask the model to refactor it.

41:04And then this way of reflection will likely

41:08generate better output or better fix for the code

41:13that you are asking to the model.

41:17Tool usage is something that we've seen before, ask model

41:22to generate the API patterns so that you

41:26can use this API function prototype to make

41:29an actual code.

41:31Or if the task is related to actual computation

41:36or some different form, you could also

41:40ask the model to generate a program as an output,

41:43and then you can run that on a safe sandbox environment

41:48that your software or software scaffolding around the language

41:54model can execute and then provide an input,

41:57provide the execution output back to the model

42:01so that model can synthesize it.

42:06So multi-agent is an interesting way

42:10to implement or accomplish your complex task.

42:14So you could split up the task--

42:19or you could split up your task and then

42:21assign those tasks in a different agent that are

42:25dedicated for specific task.

42:28And then this agent, in this case, in this context,

42:34could be just as a different prompt or different persona.

42:39So the prompt usually consists of helpful AI agent.

42:46You could change that into a different persona

42:49to a different agent.

42:50And also, you may or may not use the same model

42:54or the different model based on the task.

42:57So let's use a concrete example here.

43:00So if you build a multi-agent system for smart home

43:04automation, you could create a different agent, climate control

43:09agent, lighting control agent, and so on.

43:14And then these are the software piece

43:16that includes a different prompt with a persona

43:20as well as handling external triggers.

43:23And then those are the ones that work internally, and then

43:28that coordinate these agents essentially a model prompt

43:33together with software scaffolding around it

43:36coordinates the whole activity.

43:40So that brings up to our summary.

43:43So the agentic language model usage

43:51is a progression or extension to a existing language modal usage

43:56method.

43:57So for the best practice that you

44:00have used in language model for simple cases, most of them

44:07are applicable.

44:09However, you could use different additional methods

44:12such as more retrieval search tool usage,

44:16and then prepare different type of prompts.

44:19And then workflow so that you could

44:22use language model in its core as a reasoning or smart intern.

44:28And then you could use a tool usage or other retrieval method

44:34to interact with the external world,

44:37and then combine these results such a way

44:42that you could achieve a complex task instead of simple input

44:50and output type of language model usage.

44:53And that said, Petra?

44:56PETRA: Thank you so much, Insop.

44:58It has been really great.

45:00So much information.

45:02Hopefully, this is useful for everybody.

45:04We keep collecting the questions.

45:06We got so many.

45:07We will try to get as many as we can.

45:09But yeah, please feel free to keep them coming.

45:13Maybe the first question for you, Insop,

45:16and let's focus on the agentic AI.

45:19It's about the evaluation.

45:21And do you have some recommendations

45:23for a good strategy for evaluating agents beyond just

45:27using an LLM as a judge?

45:29It seems it should be a little bit more complicated to do

45:32the evaluation on agents.

45:33At least, that's the general notion, the questions,

45:35and people are wondering how that could be done.

45:40INSOP SONG: I think this is a great question.

45:43So just to make a quick context, LLM as a judge

45:48is commonly used method that you actually

45:51use LLM to evaluate the model-generated output

45:55against the ground truth answer or some type of reference

45:59information, which works great.

46:01And then why do we use--

46:02and also, why we use that as well for the--

46:07I think one thing that I've recently tried

46:12is agentic judging method, meaning

46:18that I use reflection type of pattern

46:23that we have seen previously.

46:27Instead of just ask one question right away to LLM as a judge,

46:31I ask first LLMs, provide initial reference

46:37and then feedback.

46:39And then I also ask again another LLM call

46:44or different prompt saying that, hey, this

46:47is a feedback from your junior engineer.

46:51If you are a senior engineer, how

46:53would you compare the junior engineer's evaluation

46:57against the output that you are evaluating?

47:00So I find that this reflection pattern

47:04was helpful to get the better evaluation instead of just one

47:09shot LLM as a judge output.

47:12However, I think there can be more creative way

47:16to improve your evaluation stage using agent

47:21patterns because the evaluation, it is really, really important.

47:27I can't emphasize more because that

47:32will help you to advance fast or change models

47:35and different type of changing the prompt

47:39and so on and so forth.

47:40So I think that was a great question.

47:43PETRA: Thank you so much, Insop.

47:45We got a few questions about augmenting the AI agents

47:51for specific uses and making sure

47:54that they get shaped in a way you need for your application.

47:58Like on the technological side, what is there to do?

48:02What is there to use?

48:02We got many questions going to this idea.

48:06So if you have some kind of information

48:08that would be helpful.

48:10INSOP SONG: Again, I think this is a great question.

48:18So first thing came up to my mind

48:21is if you have a task which is simple,

48:28then you just use language model, simple use cases.

48:31However, if you see a little more involved or complex tasks,

48:36you could experiment with simple agentic tasks.

48:40Even if it's a agentic language model usage could be--

48:48You can define agentic language model

48:50in many, many different ways.

48:51But if you have a little more involved task,

48:55you could do an iterative language model call.

48:59Even that could improve your output.

49:02So I think it all depends on your actual application domain.

49:11But instead of trying to look for the task that you could

49:17apply language model, turn that around,

49:20that how this task cannot be solved with a simple cases,

49:25so simple language model usage.

49:27So try to apply simple usage first, and then

49:33try to improve it using different patterns.

49:38Same for I think a slightly tangential

49:43but fine-tuning cases.

49:45Also, if you could have a task--

49:48if you have a task that you want to solve,

49:53try to use the existing model first whether it makes sense

49:56or not.

49:57And then from there, you could decide how to do it.

50:00Then you could prepare small data samples

50:04and then try it and then make a progress.

50:07And then really quick iteration instead of trying

50:10to invest upfront too much.

50:15PETRA: Thank you, Insop.

50:18A few questions came in.

50:20And this is a giant question on its own.

50:24So I will let you choose how you want to respond to that.

50:28But a lot of questions about ethical considerations, how

50:31to avoid hallucinations, how to avoid

50:34using data that might be unethical or there

50:39might be something kind of behind them.

50:41What would be your recommendation?

50:43Again, this question is so big, but maybe looking for something

50:47that you would say about this.

50:50INSOP SONG: Yeah, another great question.

50:54Yes, due to its probabilistic generation in nature,

51:02hallucination is always there, although a lot of people

51:06are working on it.

51:07So it's a problem As well as the contents

51:14output could be concerning.

51:19So I think a model provider themselves

51:22is checking the generated output in terms of different categories

51:28as well as your application, provide application builder

51:32yourself.

51:34Also add some guardrails.

51:36Guardrails being checking the output, using small language

51:42model so that you could quickly check

51:44the output or some type of I guess

51:48criteria using classifier or some type of thing

51:51so that you could actually filter them out to see.

51:54It could be either on the final generation stage,

51:58or it could be in the input stage

52:02that the query comes in so that you could actually avoid it.

52:07This can backfire.

52:09But I think as an enterprise or business,

52:14you may be on the safer side so they're making sure

52:18generated input, generated output, as well

52:21as the inputs queries that are being requested

52:26could be more safe or reasonable cases.

52:31So this is evolving.

52:33But I think fundamentally, it is something

52:36that check the output using type of classifier or even decoder

52:41type model to check it.

52:45PETRA: Thank you so much.

52:46And a little plug-in.

52:48We do have generative AI program that

52:50also covers a lot of the ethical aspects of LLMs.

52:54It's not a technical program, but I'm

52:56reaching to people who are asking about this very

52:59important questions.

53:02Maybe last question Insop for you.

53:05And we are not going to be endorsing,

53:08of course, at Stanford.

53:09But a question is like, how to get started?

53:12Are there any open source models you recommend?

53:14Is there anything people can do to start testing this out,

53:18like start playing with this?

53:21INSOP SONG: Great question.

53:23So one thing that as a take home message here

53:30is start simple, and then experiment it,

53:34and then iterate it.

53:35And also, agentic language model sounds complex, fancy,

53:41but it is a progression and extension of language model

53:45usage.

53:46So to start, I think you could--

53:51there are many language model usage framework and also

53:56agentic framework that you could use.

54:00However, to start, I would suggest

54:02to use playground type of environment.

54:08Say, model provider generally have their playground

54:11that you could type your prompt or input so that you

54:16can see the output right away so that you can experiment it

54:20really quickly.

54:21And then once you get familiar with it,

54:24and then use the API to make a call

54:27from your code, your program, and then see what's going on.

54:32In that way, you gain insights as well as

54:36practice best practices for prompt preparing.

54:41So once you have that, you could intelligently

54:46decide whether it makes sense for you

54:48to actually continue your own code base

54:50or make use of widely available libraries.

54:54So in short, start simple, work on a playground

54:58first, and then make a simple API calls,

55:00and then decide whether you want to use more extended libraries

55:04or just continue your own way.

55:07This applies for language model usage as well as

55:11agentic language model usage as well

55:15because there's also many agentic language

55:18model framework out there.

55:21PETRA: Thank you so much.

55:23And to close us out in this session,

55:27in the last minute, Insop, the field is progressing so fast.

55:31There is so much happening.

55:33Basically, every week, we see news

55:34about something new coming up in any of these fields.

55:38Are there any resources you study, you follow,

55:42anything you would recommend people

55:43to keep track of, to stay up to date in LLMs agentic

55:47AI in this field?

55:53INSOP SONG: Difficult question, but great question.

56:00I think it's a little hard, but I picked some experts

56:08who are known in this field.

56:09I follow them either in formerly Twitter or YouTube channel,

56:16and then from there, I get more updated information

56:20and then do my own digging.

56:23But I think find out a good starting point is a good thing.

56:28And then I think reference here, you can screenshot it.

56:33Then this can be a good starting point

56:36because it includes some agentic usage courses

56:42as well as a good type of courses,

56:44either from Stanford as well as different places.

56:48So, yeah.

56:51PETRA: Thank you so much, Insop.

56:53This has been really great.

56:54So much helpful information.

56:56And you're absolutely the best person to hear this from.

56:59Thank you so much for taking the time.

57:01Thank you also everyone who joined us live.

57:04Thank you for your time.

More from Stanford Online

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.