Full transcript
0:09INSOP SONG: My name is Insop.
0:11So today, I'd like to go over agentic AI, agentic language
0:17model as a progression of language model usage.
0:20So here is the outline of today's talk.
0:24We'll go over the overview of language model and how we use,
0:29and then the common limitations, and then some of the methods
0:33that can prove towards this common limitation.
0:37And then we'll transition into what is the agentic language
0:41model and its design patterns.
0:45So a language model is a machine learning
0:49model that predicts the next coming word given the input
0:53text.
0:54As in this example, if the input is the students open their,
0:59then language model can predict what's
1:04the most likely word coming next as a next word.
1:08So if the language model is trained with a large corpus,
1:13it is predict--
1:14it is generating the probability of next coming word.
1:18In this example, you can see, books and laptops
1:23have a higher probability than other words in the vocabulary.
1:28So the completion of this whole sentence
1:32could be the students open their books.
1:34And then if you want to keep generating the what's
1:37coming next, then we can turn them in as an input
1:41and then put it into the language model,
1:44and then language model continuously generating
1:47the next coming word.
1:50Then how these language models are trained?
1:54Largely two parts-- pre-training part and then post training
1:59part.
2:00And then first pre-training portion
2:03is the one that language models are
2:06trained with large corpus, texts collected from internet
2:12or books, or different types of texts, publicly available text,
2:16and then trained with the next token or next word prediction
2:20objectives.
2:21So once the models is finished in this pre-training stage,
2:26models are fairly good at predicting
2:29any words coming next as a word given the inputs.
2:35However, the pre-trained model itself is not easy to use.
2:42So hence, the post training steps are coming.
2:46And then these post training stage
2:50would include instruction following training,
2:53as well as the reinforcement learning with human feedback.
2:56And what this training stage means
3:00is we could prepare a data set in such
3:05a way that specific instruction or question.
3:08And then the answers or the generated output
3:12that is what the user would expect or more
3:19related to the questions and answers.
3:22So that's how the models are trained
3:25so that it's easier to use.
3:26And then also, it will respond to a specific styles.
3:30And then once this is done, and then additional training method
3:36is aligning to human preference by using reinforcement learning
3:42with human feedback, which is using human preference to align
3:48the model by using rewards schemes.
3:53And let's take a quick look, really
3:56quick look on the instruction data set.
3:58This is the template that we will
4:02use to train the model in instruction following training
4:06phase.
4:06As you can see, there is a specific instructions
4:10will be substituted in.
4:11And then expected output will be substituted in.
4:14And then this is fed to the model.
4:16And the model is only trained on the response
4:20part that is generating the output based on the given
4:26instructions.
4:29So language model that is trained on pre-training stage,
4:33as well as post-training stage, is
4:36quite capable of generating text given instruction.
4:43Essentially, it has a lot of world knowledge
4:46that could easily generate the outputs.
4:50So these are rapidly developing.
4:53And then these models are used in various application domains
4:58that we use day-to-day work such as AI coding
5:04assistance or domain-specific AI copilots
5:08or most widely known ChatGPT and related
5:12conversational interfaces.
5:14And then in order to use these type of models
5:18as for your applications or specific tools,
5:26you could use the cloud-based API calls towards the model
5:33provider or the model servers, or some other ways
5:38that you could also host the models on your local machines
5:44or even mobile machines for the models
5:47that are small enough to host on this compute
5:52constrained environments.
5:54So what does it mean by using API calls?
5:59So we step back.
6:01The language model is taking the input, natural language input
6:06text, and then generating the output.
6:09So that means we need to prepare a certain form
6:14of free-form text, natural language text
6:17as an instruction or a question, and then
6:20put them in a specific format that you could make an API call
6:24towards the model provider.
6:27And then the model provider takes that usually
6:30on a cloud environment, and then generate the output
6:33and then respond to your API calls with the generated output.
6:38Then your software around, that software
6:42that based on this model, will parse the output.
6:45And then use it as is, or maybe you
6:49could make a follow up LLM API calls
6:54to further generate the output.
7:00So the input to the model is again free-form text.
7:07So how you prepare your input, also known as a prompting
7:11is critical.
7:14So there are well-known best practices strategy,
7:20how you prepare your prompt.
7:22And here are some of them.
7:23And such as write a clear and very descriptive and detailed
7:29instructions that will help the model
7:32to generate the output that you want.
7:34And you could include a couple of examples.
7:38The form that you want to see as in the style or form.
7:43And also, you could provide the provide the references
7:48or context such a way that model rely on that context
7:53that you provide.
7:55And instead of just ask the model to answer right away,
8:00you could ask model to give model
8:03to time to think about it, such as reasoning,
8:08enable reasoning, or using chain of thought or COT method.
8:13And the next one is instead of asking model
8:19really complex task, you could break them down and then
8:23ask them chain in sequence, also chain the complex prompts.
8:30And the last one is something that
8:33is a good engineering practice.
8:36Have a good way of systemic trace and logging will help you.
8:41And also, automated evaluation is always
8:44helpful to develop your essential,
8:47to develop your progress on your application.
8:50So let's take a quick look at on each items, what that means,
8:56to get more familiar with.
8:58So write clear and descriptive instruction.
9:03As an example on the left, instead of asking short request,
9:09you could describe in detail so that model
9:13knows what you are asking because model
9:16doesn't understand.
9:17Model cannot read your mind so that means you need to describe
9:21what you want the model to generate the output for you.
9:27So this is always useful for using language model in general.
9:35And include few shot examples.
9:38So meaning that give keep model the example input and output
9:44that you would expect.
9:46As in this case, you have some type of consistent style output.
9:51Then what is the consistent style?
9:54You provide an input as an example input and example output
9:59that you would ask.
10:00And then finally, you ask your original questions.
10:05Then it's going to be--
10:06the model will generate the output based on your input.
10:10So few shot examples is always helpful to generate the output
10:15that you would want to generate.
10:19So provide relevant context and references.
10:24This is really helpful for many of the cases that
10:28are related to generating text-based
10:32on some factual information.
10:35So LLM can easily generate some incorrect output, also known
10:41as a hallucination.
10:43For those topics that it doesn't know or is not
10:48too confident about, so for those type of cases,
10:52providing context or references would always help.
10:56So here is the example prompt template
11:01that you would want to use in cases such as retrieval
11:05augmented generation that we'll look at in the following slides.
11:10Saying that only answer based on the input, based on the article
11:16that you provide, you could substitute your related
11:20reference.
11:20And then only answer based on these references.
11:25If the model cannot find the answer,
11:27you could just say that cannot find the answer.
11:31Then model will likely generate the answer
11:33based on your references.
11:40So this is important part that keep models time to think.
11:46In other words, instead ask direct question,
11:49you ask model to think through it
11:52or come up with its own solution.
11:55And then finally, compare and generate the output.
12:00So this is also known as a chain of thought.
12:04So here is one example that might not work
12:10in some medium-sized model.
12:15You could ask model saying that evaluate the student's solution
12:21is correct or not, and then provide a solution description.
12:25And then finally give student solution.
12:28And since the system prompt or your original request
12:34is saying that just answer it is correct or wrong.
12:38Model might not get it right.
12:40However, for the same model, that
12:43might not get the right answer for this
12:45if you prepare your prompt in a way that you could ask.
12:51First, work out your own solution to the problem
12:54first, and then compare your solution
12:56to the student solution.
12:58Then by doing this, model will generate its own solution,
13:02and then as it does, it will have an opportunity
13:06to provide good attention to all these inputs
13:10that it's from the original input as well as the output
13:15that it generated, and then turn towards the right answers.
13:20So reasoning chain of thought is always
13:23helpful to generate the generate output that you want to see.
13:31So here's an interesting one, probably
13:35easy to implement for your application.
13:38So instead of asking your request, that
13:41includes, say, multiple tasks in one request,
13:44you could prepare your prompt in small, simple stages.
13:51So how you do is prepare a simple prompt
13:56and then generate the output.
13:58And then prepend the output to the next stage two prompt.
14:03And then generate the output.
14:05And then again, prepend the output from the previous stage.
14:10And then generate the output in a third stage, like here.
14:13And then finally, generate the output
14:15that you would want to see.
14:16So by doing this, you may need to do it manually,
14:21or this can be done by LLM as we'll
14:26see in the following slides.
14:28But having a simple, clear task per each request
14:33would be a good way to do it.
14:41So this may not be obvious.
14:44But because many type of--
14:50as with many engineering application or development
14:55having a good way of keeping, tracing,
14:58logging will definitely be helpful for your development
15:02for debugging as well as auditing.
15:04So same principle applies to language
15:08model-based development.
15:09So keep track of the log is always good.
15:12And so that also relates to having
15:19a automated evaluation from the early stage of your development
15:24will definitely help you.
15:25In other words, you need to prepare question and answer
15:30pair, ground truth answer pair so that you could compare that
15:35against the generated output.
15:38And you could use a human to evaluate that.
15:42But that's usually a costly and time-consuming.
15:47So you may use to use--
15:49you may use a language model as a judge.
15:52Meaning that you could ask language model to evaluate
15:57model-generated output as well with ground truth output
16:03so that model can score the generated output's quality so
16:07that you can use that against your currently developed,
16:11your own applications.
16:14This will help.
16:15This will be very important because the language models are
16:20continuously and rapidly improving,
16:23as well as the methodology and the tools
16:26that you are using for developing language model is
16:30also rapidly developing.
16:32In other words, without clear evaluation,
16:36it is hard to make forward progress
16:39or even hard to change the model,
16:42change the different type of models.
16:44Because the models are rapidly developing also
16:49means that some models are rapidly deprecated,
16:53which means you may need to force to change your language
16:57model that you are using in your application.
17:00So having a good evaluation methodology upfront
17:05from the beginning will definitely help.
17:10So this is simple idea, but will be
17:15helpful for many applications.
17:19So instead of taking your input prompt as is and then process
17:23it, you may have some software or some model
17:28that you could detect the intention
17:30and then send it to a different prompt handlers.
17:33Or this is also known as prompt router.
17:37So based on the input query type,
17:41you may need to use simple prompt together
17:44with simple language model.
17:46And then this will both help in terms of operation costs
17:52as well as generate a more appropriate output
17:56together with a more relevant prompt with the language
18:03model that is more capable of that type of query.
18:12So maybe Petra, if this may be a good moment that we could take
18:20question if you could see.
18:24PETRA: Thank you so much, Insop, and such an inspirational
18:28talk already.
18:29Thank you so much.
18:30We will get to more details about agentic AI just shortly.
18:34We did want to provide you with some background
18:37and the progression what has been going on in the field.
18:41But I think maybe let's ask one question that came up.
18:46It's a little bit specific, but it
18:48might be what more people are wondering about.
18:51And is there any optimal amount of data
18:53to perform a good training or anything that you
18:57could advise people around the data available
19:01or data being used?
19:05INSOP SONG: Maybe I'll be short on this.
19:08So I assume the training here is meaning that fine
19:13tuning the LLMs in additional training on top of open source
19:19language model.
19:22Yeah, depends on your task.
19:26It would definitely vary because it's
19:29hard to say one or the other.
19:31But if you have enough data set or text
19:38that you would want to see, then you
19:40may come up with a simple question
19:42and answer pair or instruction following data set format.
19:46And then you could also make use of language model
19:49to further generate more if you need it more.
19:53But I would think you would start
19:56with, say, tens of data-- samples of data set first,
20:02and then see whether that makes the model behave what
20:05you actually see it to behave.
20:09And then you could add based on the result or signals
20:13from the initial quick test.
20:15Then you could additionally add more data set
20:18or data set samples possibly used
20:22language model to augment it or synthetic data
20:24that you could create.
20:26Great question.
20:27PETRA: Thank you so much.
20:29I see questions started coming in, which is wonderful.
20:32I think we will pause the question for now,
20:35and we will try to get to as many as we can at the end.
20:37But please keep them coming.
20:39Definitely makes the session more engaging
20:41and also let us know what you are interested in.
20:43Thank you.
20:44INSOP SONG: Thank you, Petra.
20:46So so far, we've been looking at overview of language model.
20:52Great.
20:54Very powerful models that are out there,
20:56many models out there, and then how we use.
20:59However, even there are still limitations
21:05for models that are available, that
21:09are listed here, such as hallucination
21:11is a well-known issue that models can sometimes oftentimes
21:20generate incorrect or incorrect information,
21:24particularly if it's related to some computation
21:27or some other specific area.
21:29So this is a problem that we want to avoid
21:36in your application domain.
21:39And other thing is there's always a knowledge cut up
21:42in data set preparation.
21:44So model creator prepare data set.
21:48However, they need to, at some point,
21:51cut up their data set collection and then use it.
21:55So model may have not seen the recent information
22:01or news as part of their pre-training data set.
22:06Lack of attribution-- so model can
22:09answer a lot of world knowledge questions
22:12and can answer those type of general questions.
22:17However, it may not tell--
22:18it's not going to tell you where they
22:21drew this, the answer from particular specific data source.
22:28Data privacy is one fact that model creator,
22:34you prepare the data set using publicly available data source.
22:41That means model have not seen your proprietary data set
22:46from your organization or particular domain.
22:51And limited context length--
22:54although it's a rapidly increasing,
22:58however, it's fine balance because providing longer context
23:08will give more context information to the model.
23:12However, it comes with operational cost as well as
23:17the speed of the latency in text generations.
23:24So in order to address these common limitations,
23:32retrieval augmented generation is
23:35one way to handle this, such as it could
23:38reduce the hallucination by using
23:41the actual relevant reference.
23:44It also address the citation because it knows where
23:51this reference is coming from.
23:53And then this will allow you as an application developer
23:59or a system developer to prepare systems
24:04so that you could use your own proprietary data
24:07set or the text.
24:08And then you could use good use of small number of context links
24:16because it only select relevant data set.
24:20So how it works is you could pre-index your own data set
24:28or your own text by turn them into smaller chunk, chunk
24:35of text, and then convert them into an embedding space using
24:40embedding model, and then stage them
24:44as part of your database or vector database.
24:47And then when the request or query came,
24:51you could turn this query into embedding space
24:56so that you could do nearest neighboring search and then
24:59select top K relevant information, relevant chunk,
25:05text chunk, and then place them as part of your prompt.
25:09Some of the slide that we see previously
25:12is you put the reference as part of the prompt,
25:15and then use that as the model.
25:19Only make use of this reference.
25:22So this is one good way to make use
25:26of your own proprietary data.
25:28And then similar method can be used in the actual AI search
25:36so that instead of using index data set,
25:40you could also rely on web search or different type
25:42of search so that you could provide the information as part
25:46of the index.
25:48And one of the thing also mentioned here
25:52is there are many methods or ideas for retrieval augmented
25:59generation.
26:00More commonly used method is something
26:04that we've just mentioned.
26:06We've just talked about, meaning turn the text
26:09chunk into embedding space.
26:11And then do a nearest neighbor search.
26:13However, there are many methods.
26:14And then you could also use knowledge graph base.
26:19So if you could generate a knowledge graph from your text
26:24source, and then that could also help to extract
26:28the more relevant information.
26:31Also known as graph RAG is one part of it.
26:35However, there are many methods and then
26:37you may need to look into the right method or if right method,
26:44make use of those.
26:48Tool usage-- so language model being most widely used form
26:54is a text in and text output, which means that it could
26:58answer many type of queries.
27:02However, it's not going to execute or extract information
27:08from the external.
27:09So that's where this tool usage came
27:13to rescue or also known as a function calling.
27:16So with this method, you could get real-time information,
27:22or you could actually do a computation
27:26by generating software or computer code.
27:30So what does it mean is--
27:32let's look at an example here.
27:35So if you have AI chatbot that if you
27:42ask what is the weather in, say, San Francisco,
27:45then model will not know it in itself.
27:48So however, if you tell model previously,
27:54as part of the prompt, saying that if you
27:57ask weather-related question generated output form
28:02that the software that parse the output can make an API call.
28:08So as in this example, model will generate the output
28:13in a way that, hey, this is the case for tool usage.
28:18So it generates an output as in the form here, get weather,
28:25and then input argument to this API call or function call
28:29is the place that we ask.
28:31So then software receives this text output from the model,
28:36and then parse.
28:37And then actually make an API call
28:40towards the weather provider, and then get
28:43the weather information.
28:45And then again, provide back to langpack back
28:49to the language model.
28:50And then language model will generate
28:53more human-friendly or helpful output based
28:57on this API-based result. And for some cases,
29:04model can also generate a software code that
29:09can be executed as part of the sandbox outside of the language
29:14model by the software that is coordinating
29:17all these activities.
29:22So agentic language model--
29:26so there can be many definitions.
29:30One definition is it could interact with environment.
29:35So compared to simple language model usage,
29:38generally use simple language model usage
29:44as you seen here, text input and text output.
29:48Agentic language model usage could
29:52be language model could do something with the environment
29:57by generating tool usage or retrieval request.
30:01And then from the environment, anything outside
30:05of language model could provide an output,
30:10could provide an information that can be fed back to language
30:16model as an observation.
30:18And then the whole thing, the agentic language model,
30:23which includes language model at its core
30:26with the software around it, will process it,
30:30and then also put them in a memory,
30:33and as with its conversational history,
30:38which can be taken as a memory.
30:41So this is one way of definition of agentic language model.
30:48The other way to look at it is this.
30:53Agentic language model usage can be
30:58defined as it could reason as well as it could action.
31:02It could do an action, so also called ReAct, reason and action.
31:07So reasoning part is something that you could encourage model
31:12to reason about by using a method such as chain of dots,
31:18and then doing an action using a method that we have seen
31:23in previous slides, as in retrieval or search engine,
31:27or actually using calculator by making an API call
31:31or different type of API calls, such as weather API that we have
31:35seen, and also generating a Python code
31:39so that you could run as part of your sandboxes.
31:43So by combining these reasoning and action,
31:47model can do a lot more complex task
31:51than simple input and output type of interaction.
31:57So let's look at a little more detail on this.
32:03What does it mean by reasoning and action?
32:06So reasoning part, instead of doing the task that is asked,
32:14you could prepare your prompt in such a way that ask
32:19model to break down the task and then make a plan.
32:22So instead of breaking down the task by yourself,
32:28as we've seen in the previous chaining the prompt slide,
32:32you could ask model to break down and then
32:36prepare your task so in other words, plan the actions.
32:42And then based on that breaking down,
32:45model can generate different actions
32:48by making API calls or tool usage
32:53so that it could extract or collect additional information
32:59from the external world.
33:00And then by combining all these, put them in a memory
33:04so that it knows, model knows what's been happening.
33:09And then based on that, finally draw an answer for you.
33:13So let's look at a concrete example here.
33:18So if you have a customer support AI agent,
33:23then how it might look?
33:25How it might work?
33:26So as a customer, ask, can I get a refund for product full?
33:35Then agentic system will break down this task,
33:39break down this request task into the following four
33:43different type of actions.
33:46Check the refund policy, check the customer information,
33:49and then check the products.
33:51And then finally, collect and then decide what to do.
33:56In each step language model will spit out the API calls
34:02so that it could collect the information.
34:05For example, check the refund policy,
34:08a language model could ask a retrieval system
34:13against the pre-indexed company policy, refund policy.
34:17And from there, it could retrack the information,
34:20and then put them in its own context.
34:22And then using that, also request the customer order
34:27information.
34:28It could either ask customer, in the chat format,
34:32collect more information.
34:34Or it could look it up in the system
34:37because it depends on how this chat system is prepared.
34:43The same thing for the product so that it
34:46could collect more information.
34:47And then finally, draw the conclusion
34:50based on the policy, and then product information,
34:53as well as the customer order information.
34:56And then send the request to the follow-up system as an API call
35:04as well as the send, prepare, say, response draft.
35:09And then that's going to be handled by final approval.
35:17So workflow is generally like this.
35:21So in a sense, agentic language model system
35:25is generally language model is making iterative calls
35:32by reviewing the document or task
35:35and then making external tool calls.
35:40An example, if you want to do some research
35:45of certain matters, you could prepare your agent
35:49to do a research, web search or different type of search,
35:53and then summarize them iteratively, and then
35:58finally prepare the report to you or to your system.
36:03Another example could be software assistant agent
36:08that you could ask this software agent or free agent that
36:16ask the issues of certain type of software bug or issues.
36:20Then this agent will look it up and review this issue,
36:25and then collect the relevant piece of code or files,
36:30and then review them and then propose the output.
36:33Or it could also execute in its sandbox environment,
36:38and then test the fix, and then get the output.
36:44And then iteratively try to find the fix.
36:46And then finally, pulls the pull request or the changes
36:52to users or developers.
36:57These are the ways that we can use language model
37:01in agent format and by doing interactive language model
37:07calls.
37:10So the main reasons and difference
37:18why these agentic language model usage is getting more widely
37:25used is given that if you have the same model, if you ask just
37:32direct request to the model, model
37:35may not be able to handle it.
37:38However, if you put your task in this type of agentic format
37:44or patterns, then model will do more complex tasks,
37:49even using a model that may not be able to do it
37:52if you don't do this way.
37:55So that's one of the reasons that agentic language model
37:58is pushing the boundaries, so that the things that we
38:03can do with AI agent is more complex or different domains
38:10that we can rely on.
38:15So here are the real-world applications--
38:19software development, code generation or bug fixing
38:23or this type of development is widely investigated
38:30or being researched by different organizations
38:34as well as there are companies that
38:36are trying to provide these services as decision
38:39on the right side, and research and analysis that
38:43gather information, synthesize it, and then provide
38:46a summary for the users.
38:49And then task automation is one of the areas
38:52that agentic method could be used.
39:00So to make it more clear, here are some of the design patterns
39:08that you could use agentic language model.
39:13Planning is critical because by asking a model
39:19to break down the task to make it simpler task or clear task
39:26so that language model later can make an API call
39:30or use the tool usage.
39:32So planning is critical.
39:35Then reflection is something that model
39:38can generate in the cell.
39:40And then the next model call can criticize
39:49the output that actually came from the same model.
39:53So by doing this, the output could be improved.
39:57And then two usages is something that something that outside
40:03of the language model that you need,
40:04the real-time information or different type of information,
40:09then you can use this.
40:10And then multi-agent collaboration
40:12is one way to handle this.
40:16Reflection is a pattern that quick to implement
40:20and then leads to a good performance.
40:23And let's use a concrete example here.
40:27So if you want to refactor a programming code,
40:30instead of asking model to improve it right away,
40:34if you do this pattern, as in this example,
40:37saying that you asked ask a model
40:40saying that here is the code, and then check the code
40:42and provide constructive feedback.
40:46And then take that feedback to the second prompt,
40:51as in this example.
40:53You could also prepare prompts saying
40:56that here is the code and the feedback,
40:58which came from the model itself.
41:01And then ask the model to refactor it.
41:04And then this way of reflection will likely
41:08generate better output or better fix for the code
41:13that you are asking to the model.
41:17Tool usage is something that we've seen before, ask model
41:22to generate the API patterns so that you
41:26can use this API function prototype to make
41:29an actual code.
41:31Or if the task is related to actual computation
41:36or some different form, you could also
41:40ask the model to generate a program as an output,
41:43and then you can run that on a safe sandbox environment
41:48that your software or software scaffolding around the language
41:54model can execute and then provide an input,
41:57provide the execution output back to the model
42:01so that model can synthesize it.
42:06So multi-agent is an interesting way
42:10to implement or accomplish your complex task.
42:14So you could split up the task--
42:19or you could split up your task and then
42:21assign those tasks in a different agent that are
42:25dedicated for specific task.
42:28And then this agent, in this case, in this context,
42:34could be just as a different prompt or different persona.
42:39So the prompt usually consists of helpful AI agent.
42:46You could change that into a different persona
42:49to a different agent.
42:50And also, you may or may not use the same model
42:54or the different model based on the task.
42:57So let's use a concrete example here.
43:00So if you build a multi-agent system for smart home
43:04automation, you could create a different agent, climate control
43:09agent, lighting control agent, and so on.
43:14And then these are the software piece
43:16that includes a different prompt with a persona
43:20as well as handling external triggers.
43:23And then those are the ones that work internally, and then
43:28that coordinate these agents essentially a model prompt
43:33together with software scaffolding around it
43:36coordinates the whole activity.
43:40So that brings up to our summary.
43:43So the agentic language model usage
43:51is a progression or extension to a existing language modal usage
43:56method.
43:57So for the best practice that you
44:00have used in language model for simple cases, most of them
44:07are applicable.
44:09However, you could use different additional methods
44:12such as more retrieval search tool usage,
44:16and then prepare different type of prompts.
44:19And then workflow so that you could
44:22use language model in its core as a reasoning or smart intern.
44:28And then you could use a tool usage or other retrieval method
44:34to interact with the external world,
44:37and then combine these results such a way
44:42that you could achieve a complex task instead of simple input
44:50and output type of language model usage.
44:53And that said, Petra?
44:56PETRA: Thank you so much, Insop.
44:58It has been really great.
45:00So much information.
45:02Hopefully, this is useful for everybody.
45:04We keep collecting the questions.
45:06We got so many.
45:07We will try to get as many as we can.
45:09But yeah, please feel free to keep them coming.
45:13Maybe the first question for you, Insop,
45:16and let's focus on the agentic AI.
45:19It's about the evaluation.
45:21And do you have some recommendations
45:23for a good strategy for evaluating agents beyond just
45:27using an LLM as a judge?
45:29It seems it should be a little bit more complicated to do
45:32the evaluation on agents.
45:33At least, that's the general notion, the questions,
45:35and people are wondering how that could be done.
45:40INSOP SONG: I think this is a great question.
45:43So just to make a quick context, LLM as a judge
45:48is commonly used method that you actually
45:51use LLM to evaluate the model-generated output
45:55against the ground truth answer or some type of reference
45:59information, which works great.
46:01And then why do we use--
46:02and also, why we use that as well for the--
46:07I think one thing that I've recently tried
46:12is agentic judging method, meaning
46:18that I use reflection type of pattern
46:23that we have seen previously.
46:27Instead of just ask one question right away to LLM as a judge,
46:31I ask first LLMs, provide initial reference
46:37and then feedback.
46:39And then I also ask again another LLM call
46:44or different prompt saying that, hey, this
46:47is a feedback from your junior engineer.
46:51If you are a senior engineer, how
46:53would you compare the junior engineer's evaluation
46:57against the output that you are evaluating?
47:00So I find that this reflection pattern
47:04was helpful to get the better evaluation instead of just one
47:09shot LLM as a judge output.
47:12However, I think there can be more creative way
47:16to improve your evaluation stage using agent
47:21patterns because the evaluation, it is really, really important.
47:27I can't emphasize more because that
47:32will help you to advance fast or change models
47:35and different type of changing the prompt
47:39and so on and so forth.
47:40So I think that was a great question.
47:43PETRA: Thank you so much, Insop.
47:45We got a few questions about augmenting the AI agents
47:51for specific uses and making sure
47:54that they get shaped in a way you need for your application.
47:58Like on the technological side, what is there to do?
48:02What is there to use?
48:02We got many questions going to this idea.
48:06So if you have some kind of information
48:08that would be helpful.
48:10INSOP SONG: Again, I think this is a great question.
48:18So first thing came up to my mind
48:21is if you have a task which is simple,
48:28then you just use language model, simple use cases.
48:31However, if you see a little more involved or complex tasks,
48:36you could experiment with simple agentic tasks.
48:40Even if it's a agentic language model usage could be--
48:48You can define agentic language model
48:50in many, many different ways.
48:51But if you have a little more involved task,
48:55you could do an iterative language model call.
48:59Even that could improve your output.
49:02So I think it all depends on your actual application domain.
49:11But instead of trying to look for the task that you could
49:17apply language model, turn that around,
49:20that how this task cannot be solved with a simple cases,
49:25so simple language model usage.
49:27So try to apply simple usage first, and then
49:33try to improve it using different patterns.
49:38Same for I think a slightly tangential
49:43but fine-tuning cases.
49:45Also, if you could have a task--
49:48if you have a task that you want to solve,
49:53try to use the existing model first whether it makes sense
49:56or not.
49:57And then from there, you could decide how to do it.
50:00Then you could prepare small data samples
50:04and then try it and then make a progress.
50:07And then really quick iteration instead of trying
50:10to invest upfront too much.
50:15PETRA: Thank you, Insop.
50:18A few questions came in.
50:20And this is a giant question on its own.
50:24So I will let you choose how you want to respond to that.
50:28But a lot of questions about ethical considerations, how
50:31to avoid hallucinations, how to avoid
50:34using data that might be unethical or there
50:39might be something kind of behind them.
50:41What would be your recommendation?
50:43Again, this question is so big, but maybe looking for something
50:47that you would say about this.
50:50INSOP SONG: Yeah, another great question.
50:54Yes, due to its probabilistic generation in nature,
51:02hallucination is always there, although a lot of people
51:06are working on it.
51:07So it's a problem As well as the contents
51:14output could be concerning.
51:19So I think a model provider themselves
51:22is checking the generated output in terms of different categories
51:28as well as your application, provide application builder
51:32yourself.
51:34Also add some guardrails.
51:36Guardrails being checking the output, using small language
51:42model so that you could quickly check
51:44the output or some type of I guess
51:48criteria using classifier or some type of thing
51:51so that you could actually filter them out to see.
51:54It could be either on the final generation stage,
51:58or it could be in the input stage
52:02that the query comes in so that you could actually avoid it.
52:07This can backfire.
52:09But I think as an enterprise or business,
52:14you may be on the safer side so they're making sure
52:18generated input, generated output, as well
52:21as the inputs queries that are being requested
52:26could be more safe or reasonable cases.
52:31So this is evolving.
52:33But I think fundamentally, it is something
52:36that check the output using type of classifier or even decoder
52:41type model to check it.
52:45PETRA: Thank you so much.
52:46And a little plug-in.
52:48We do have generative AI program that
52:50also covers a lot of the ethical aspects of LLMs.
52:54It's not a technical program, but I'm
52:56reaching to people who are asking about this very
52:59important questions.
53:02Maybe last question Insop for you.
53:05And we are not going to be endorsing,
53:08of course, at Stanford.
53:09But a question is like, how to get started?
53:12Are there any open source models you recommend?
53:14Is there anything people can do to start testing this out,
53:18like start playing with this?
53:21INSOP SONG: Great question.
53:23So one thing that as a take home message here
53:30is start simple, and then experiment it,
53:34and then iterate it.
53:35And also, agentic language model sounds complex, fancy,
53:41but it is a progression and extension of language model
53:45usage.
53:46So to start, I think you could--
53:51there are many language model usage framework and also
53:56agentic framework that you could use.
54:00However, to start, I would suggest
54:02to use playground type of environment.
54:08Say, model provider generally have their playground
54:11that you could type your prompt or input so that you
54:16can see the output right away so that you can experiment it
54:20really quickly.
54:21And then once you get familiar with it,
54:24and then use the API to make a call
54:27from your code, your program, and then see what's going on.
54:32In that way, you gain insights as well as
54:36practice best practices for prompt preparing.
54:41So once you have that, you could intelligently
54:46decide whether it makes sense for you
54:48to actually continue your own code base
54:50or make use of widely available libraries.
54:54So in short, start simple, work on a playground
54:58first, and then make a simple API calls,
55:00and then decide whether you want to use more extended libraries
55:04or just continue your own way.
55:07This applies for language model usage as well as
55:11agentic language model usage as well
55:15because there's also many agentic language
55:18model framework out there.
55:21PETRA: Thank you so much.
55:23And to close us out in this session,
55:27in the last minute, Insop, the field is progressing so fast.
55:31There is so much happening.
55:33Basically, every week, we see news
55:34about something new coming up in any of these fields.
55:38Are there any resources you study, you follow,
55:42anything you would recommend people
55:43to keep track of, to stay up to date in LLMs agentic
55:47AI in this field?
55:53INSOP SONG: Difficult question, but great question.
56:00I think it's a little hard, but I picked some experts
56:08who are known in this field.
56:09I follow them either in formerly Twitter or YouTube channel,
56:16and then from there, I get more updated information
56:20and then do my own digging.
56:23But I think find out a good starting point is a good thing.
56:28And then I think reference here, you can screenshot it.
56:33Then this can be a good starting point
56:36because it includes some agentic usage courses
56:42as well as a good type of courses,
56:44either from Stanford as well as different places.
56:48So, yeah.
56:51PETRA: Thank you so much, Insop.
56:53This has been really great.
56:54So much helpful information.
56:56And you're absolutely the best person to hear this from.
56:59Thank you so much for taking the time.
57:01Thank you also everyone who joined us live.
57:04Thank you for your time.