Full transcript
0:01Hi everybody, and thank you for joining
0:02Project Neutron ME onboarding.
0:05Today we're going to have an overview of
0:07our Slack channels, tasking overview,
0:09and common tasking errors.
0:11In this project, we're training AI
0:13models to get better at expert work by
0:16letting them practice inside a
0:17real-estate reinforcement learning
0:19environment. This is a sandbox for
0:21models take real action with real
0:23software.
0:25Project Neutron focuses on one domain,
0:27which is mechanical engineering.
0:29As I give a high-high overview of the
0:30project, you upload a PDF-formatted
0:33technical drawing and let the models try
0:36to create a 3D model from it.
0:38Your goal is to have a drawing complex
0:40enough that it stumps the models.
0:43The 2D drawing must be fully
0:45dimensioned,
0:46as this gives the models a fair shot at
0:48creating it.
0:50I would also like to note that on this
0:52project, there is a zero-tolerance
0:54policy for LLM usage. So,
0:58there's only one task type on the
1:00project currently, which is PDF
1:03technical drawing to 3D CAD design.
1:06On the task on the platform home home
1:09screen, you can um sort by unclaimed and
1:13claim a task with this type. As soon as
1:16you claim, then all all work is done on
1:19platform.
1:21So, I would also I'll start right here
1:24real really quick. Um Neutron has a
1:27couple different channels here. Um most
1:29of your questions can be directed to the
1:32Slack channel, where um
1:34they will be posted every day the
1:36different issues that you're you may be
1:38facing. We also have a social group.
1:41People uh
1:42just for everyone to chat with. Feel
1:43free to to hop in there, drop a meme,
1:45have some fun.
1:46And then also all of our announcements
1:48come through here.
1:50So, after you claim a task, you will see
1:53this screen. As you can see, PDF
1:55technical drawing to 3D CAD design.
1:58You hit start timer.
2:00I would like to know on this project, it
2:01is currently pay per task, so whether it
2:04takes you 1 hour or 10 hours to complete
2:06a task, your payout is is still the
2:08same. So, you don't have to worry if you
2:11go over the 5-hour limit.
2:14From here, you can one, click
2:16instructions and, you know, get more
2:19informed, or two, continue.
2:22This is uh a pledge that you are you are
2:26making that you are not using LLM on
2:29this project at all. You check the box,
2:31may continue.
2:33Also on this project, we have a
2:34requirement that every asset that you
2:36upload, you own the rights to it. It is
2:40yours to distribute and share.
2:43So, after you check that, you can click
2:46continue.
2:48So, before we get into the details of
2:50the prompt, I'm going to give a kind of
2:52a
2:53overview of how a task works. So, I
2:57already mentioned, you attach a
2:59technical drawing, PDF format only, and
3:03you write a prompt.
3:04Number two, the model will read your
3:07prompt, and
3:09I'm sorry, the model will read your
3:10prompt
3:12and the drawing that you have attached,
3:14and it will attempt to build a CAD file
3:16from it.
3:17This happens twice, as two independent
3:19rollouts. So, you will have two model
3:22responses.
3:23Number three, after the models run,
3:27you inspect them. You ensure that they
3:30did fail, that it did a bad job at
3:32creating your part.
3:34Then, you go to the golden answer.
3:36And the golden answer has two parts,
3:38which is a
3:40a two written parts. One is a
3:42description of the CAD file, and two is
3:45the steps to solve.
3:48Number four, after and in the golden,
3:50you need to upload your golden file,
3:52which is a step file, a CAD file in step
3:56file format that is the perfect answer
3:59to the
4:00drawing and prompt and in the prompt
4:04input area.
4:07After this, you go to the rubric. The
4:09rubric is auto-generated from your
4:11golden answer.
4:12Um
4:13but I will say the rubric that is
4:15auto-generated is truly a starting point
4:17and only a starting point. You will need
4:20to
4:21modify it
4:23drastically to bring bring it up to up
4:27to par. Um it's a whole lot better than
4:29writing it from scratch.
4:31Um
4:32but just keep that in mind it will need
4:34to be edited.
4:36So, after you write your rubric, then
4:38you run the evaluation. And what this
4:40does, it scores the rubric against your
4:42two against the two model responses. If
4:46that scores below 40%,
4:48then you are good to submit the task.
4:51That 40% is the average between the two
4:54results.
4:56So, here on the prompt, I'm going to
4:57pause this and we're going to actually
4:59look at a task that we have in the
5:00review layer right now.
5:05So, as you can see, this task here they
5:07accepted that accepted that. And this is
5:10the prompt area. So, in the prompt they
5:14they stated, "I am an a mechanical
5:17engineer who just received a drawing of
5:18a mount from an intern. Please review
5:21the attached PDF of the drawing and
5:23generate a step file of the model.
5:26Assume units of millimeters along with a
5:28tolerance of plus or minus 0.5
5:30millimeters or 1° unless specified
5:33otherwise." Additionally, they uploaded
5:36their PDF drawing.
5:40The drawing needs to be fully
5:41dimensioned
5:44to ensure that the models have a fair
5:46chance at trying to exactly replicate
5:48it.
5:52In this prompt, your goal is to stump
5:55the model.
5:56The
5:57everything must be realistic, not a
6:00trick, and and unambiguous.
6:03It must be hard enough
6:05to stump the models. Engineering is
6:07difficult. These models are smart, you
6:09will find out. And so, um
6:12you do need to use somewhat complex
6:15drawings to stump these models.
6:17There's one rule that we go by big time
6:19on MEK-E, and that is if 50 competent
6:22mechanical engineers solved your task
6:24independently, they should all land on
6:26the same answer.
6:28So, that means that there should be zero
6:30room for interpretation from your prompt
6:33and from your drawing.
6:36So, the five parts of a prompt that we
6:37have, um five key parts. Um one is the
6:41context. Say who's asking and why they
6:44need the work.
6:45So, mechanical engineer just received
6:48the drawing of the mount from an intern.
6:51Oop, we don't want that.
6:54What they need, the main thing you want
6:56done in plain terms. We state right
6:58here, generate a step file of the model.
7:01Number three, where the data is. Point
7:03to the PDF file you attached.
7:06We state it right here, attached PDF.
7:09Number four, what to produce. On this
7:11project, we only want one file back, and
7:14that is step file, a step file.
7:18And so, here they stated generate a step
7:19file of the model.
7:21And number five, action to take, the
7:23specific modeling work to do, if any.
7:25So, this is not always the case, um but
7:28sometimes, if you want to want to add
7:30complexity to your prompt, you can have
7:34a prompt like this, um have your
7:36drawing, and then require additional
7:39additional modeling features to be
7:41completed.
7:43So, maybe
7:45increase the size of the part by 50% and
7:47add a
7:4930 30 thou radi- or a
7:53a 30 thou chamfer to all sharp edges.
7:56So, you can require additional work to
7:58be done.
8:01So, here in our golden answer. So,
8:03before I get there, you will have two
8:04model responses, response one and
8:07response two. As you can see, the models
8:09will often give several different um
8:12output files, but as clearly stated in
8:15the prompt here, a step file,
8:18it did it output a step file. That is
8:20what we want to inspect.
8:23Both models will generate a step file.
8:28So, before you typically go to this part
8:31of the task, what I would do is open
8:33these step files and at this point you
8:36should have your golden file, which is
8:38the perfect answer to your prompt, and
8:41you can compare it to this step file. If
8:44these look drastically different, much
8:46different, there is a good chance you
8:48have stumped the model. If they look
8:51pretty identical,
8:52your task, your prompt likely was not
8:55complex enough before you proceed
8:57further in the task, I would recommend
8:58adding complexity to your prompt,
9:01whether that's
9:02um
9:03altering your original file or
9:06additional modeling uh features to be
9:09added here.
9:13If it all checks out, you've got a
9:15winner. You've stumped the model, then
9:17you come down to the golden.
9:19In the golden,
9:20you want to upload first your step file.
9:24It needs to be zipped. Uh that is a
9:26platform requirement. So, take your step
9:28file that is the perfect answer to your
9:30prompt, put it in a zip file, and upload
9:32it here.
9:34In the written portion of the golden,
9:36there are two parts.
9:38One is the description of the CAD file.
9:41So, this states all the key values with
9:44tolerances and units.
9:48Two,
9:49um
9:53we don't want necessary this response is
9:57a little bit long. However,
10:00we we don't want a text wall. Sometimes
10:03this will take up several screens worth
10:06of scrolling. That is not ideal. What we
10:09would prefer is you to just call out the
10:10key values that make up your part and
10:13state them here.
10:15So, that is the description of the CAD
10:17file.
10:18Step number two is the steps to solve.
10:21So, in this part, you don't have to um
10:25we we we want you to keep it high level.
10:27So, how to build the part, no need to
10:29state every value or describe every
10:32minor feature. This is a high-level
10:34overview of how you would build this
10:36part.
10:38So,
10:39um
10:42one other note on this I would like to
10:43say,
10:45it may seem redundant, but this is a
10:47requirement. So, make sure you keep
10:49these two sections separate. So, you
10:52have the finished part, which is the
10:54description of the CAD file, and then
10:57section two, steps to solve steps to
10:59solve. Please keep them separate.
11:02So, after you do this, then you will
11:05come down to the rubric. Um when you
11:08submit this part, the golden, then the
11:12the rubric will auto generate. That auto
11:15generation can sometimes take between
11:2020 minutes to an hour.
11:22Um
11:23it is getting worked on right now as far
11:25as speed. They're trying to speed it up,
11:27but we do ask for your patience during
11:29this time.
11:30Additionally, after you submit your
11:32prompt right here, these model rollouts
11:35will sometimes take
11:38up to an hour. We have seen them take
11:40longer. It doesn't typically, but please
11:42be patient during that time. Typically,
11:45if you think that your drawing is
11:46complex enough, especially after your
11:48first task and you get a good feel of
11:50the project, what it's going to take to
11:52stump the models, then as this is
11:54running, you can either be working on
11:56another task or be preparing this part
11:59of of the task.
12:03So, both
12:04model rollouts and the rubric do take a
12:07little bit to generate.
12:08Um
12:10So, here on the rubric, you open it up
12:13and
12:16I need to
12:19refresh.
12:27So, you scroll down here. You guys won't
12:29have to do that. You all.
12:32But if you look at the rubric,
12:34so
12:36it like I said, it is auto-generated.
12:38This The rubric um this is where you
12:40create criteria that grade the model
12:43responses. The goal is for your task to
12:46be complex enough that when the two
12:48model rollouts are graded against your
12:50rubric, that the average score below the
12:53between the two responses are below 40%.
12:56If they are not, then you need to go
12:59back to the prompt and add complexity.
13:01Um I would like to say though, keep in
13:03mind any edits that you make upstream
13:06erases downstream data. So, if you edit
13:09the prompt, your golden answer and your
13:11rubric will be wiped out. For that
13:13reason, always click here and download a
13:17copy of the CS of the rubric in CSV
13:19format. Additionally, copy this and save
13:23it to a local document. That way, if you
13:25do have to update the prompt, uh you can
13:28easily paste back in and make quick
13:29modifications accordingly.
13:32So,
13:33um
13:34hopefully you have a complex complex
13:36enough task you don't have to do that.
13:40So, I would also like to drive home the
13:42point again that this rubric is a
13:45starting point only a starting point.
13:47This one has been worked on. I have not
13:49audited it yet. Um but quick glance, it
13:53does not look too bad. So,
13:55what you need to
13:57uh ensure is that the rubric meets these
14:02uh that these five principles. One, uh
14:05it is self-contained. So, it is gradable
14:08from the criterion text alone. A grader
14:10with no access to the prompt or files
14:13can evaluate it. Name the final value,
14:16never a file or a stores.
14:19Two, needs to be objective. Two graders
14:21reach the same pass/fail. No vibe words,
14:24no judgment calls.
14:26I will say one of the biggest things
14:27that we hear back from taskers is that
14:30they disagree with the response that the
14:32rubric the score that the rubric is
14:34given.
14:36Typically, 99% of the time, the reason
14:39the tasker is getting that is because
14:40they're being too vague in their in
14:43their grading on a criterion. When you
14:46explicitly state what you want targeted,
14:50getting extremely detailed,
14:5299% of the time, the evaluation will
14:56score
14:57the correct way.
14:59So, just keep that in mind.
15:02Number three,
15:03um
15:04your criterion need to be verifiable.
15:06So, point to evidence invisible in the
15:07model's output. Um four, atomic. Each
15:11criterion must check only one thing. For
15:14example, don't check delivers a
15:18a base that is
15:2150 mm wide by 50 mm tall. That is not
15:24atomic. You check one
15:26feature per criterion. So, that would be
15:29need to be divided into two.
15:32And then number five is making mutually
15:34exclusive, collectively exhaustive.
15:37Um the sum of all criteria cover every
15:40aspect of a perfect response without any
15:43without any redundantly related
15:45criteria, which penalize twice for the
15:47same mistake.
15:50All criteria need to be verb led. They
15:53start with a verb.
15:55And then number seven, no double
15:57negatives. If you grade something here,
16:00um and it's positive, don't have an
16:02equal negative at the bottom that
16:04penalizes the model.
16:08Also on our rubrics,
16:09um we need between 20 and 100 criterion.
16:13Um
16:14every weight um
16:16can be one of six values, which is 9 5 1
16:20-1 -5 -9. No other numbers, zeros, no
16:25decimals.
16:27I would also like to say
16:29Well, I'll get to that in a second. So,
16:32of your
16:34uh waiting, we do have distribution
16:36requirements. So, 35 to 45% of your
16:40total criterion account uh criterion
16:43count must be at the plus or minus nine
16:47value. 45 to 60 must be at the plus or
16:50minus five value. And 10 to 20 must be
16:53at the plus or minus one value.
16:56Additionally,
16:58at most 35% of items can be negative.
17:03Your rubric needs to be fair
17:06and a real a real realistic request
17:09that an engineer would actually make.
17:12Additionally, every value in a criterion
17:15must call out the units and the
17:16tolerances.
17:18So, as you can see here, measuring 70 mm
17:21tall plus or minus 0.5 in the delivered
17:24step model.
17:28So, there's going to be some questions
17:29about the
17:31weighting here. Um
17:33so, I would highly recommend before you
17:37do a task, you need to go through these
17:40instructions very thoroughly. I mean, a
17:43couple of times. There's a lot of
17:45information here, but it is all good
17:47information that will only lead to
17:49successful successful tasking.
17:52So, this page will show you exactly how
17:54your weights
17:56um are decided. Nines are decisive
17:59must-have. Um fives are important
18:02should-have, and one is supportive
18:04nice-to-have.
18:06So, if you get everything lined out,
18:09you can you'll run the evaluation
18:11results. If you are below 40% between
18:14the two rollouts when averaging them
18:17them together, 40% does not pass.
18:20Anything below 40 does, then you can
18:22submit the task.
18:27So, at this point, I would say,
18:32let's look at the top five error
18:33categories and then we should be done.
18:36So, these are
18:39in a task that we review.
18:42Um so, that this is not one of my tasks.
18:44This is one in the queue right now that
18:46needs to be reviewed, but it does go
18:48through two review cycles. So,
18:52if it passes R1, they don't find any
18:54issues, then it goes to R2. If R2
18:56doesn't find any issues, then it goes to
18:59ready to deliver.
19:01At that point, you should see payment
19:03pop up in your
19:05in your um
19:07platform dashboard, cuz that means your
19:09task has passed all of the checks.
19:12If your task does not pass, don't worry.
19:15Um
19:16R1 or R2, if there are any issues, they
19:19will catch it. They will politely send
19:21it back to you with detailed feedback on
19:23what needs to be updated. If you will
19:25make those changes and resubmit, um then
19:29we will review again until we we get it
19:31right. And then once the task has met
19:34all uh standards of the project, then it
19:37will go to the RTD layer.
19:40So, common tasking errors, top five
19:42error categories. Um number one, the
19:45rubric criteria are not verbal-led.
19:48So, if we go down here, I know I
19:50mentioned it.
19:52Make sure every one of these criteria
19:55start with a verb.
19:59The task doesn't stump the model.
20:02Um
20:02sometimes we
20:04you we will get tasks that the average
20:06score is only 61%.
20:09It has to be below 40 for us to accept
20:12it. If it is not, um we will send it
20:14back and ask you to add complexity to
20:16the task.
20:18The golden answer is missing the
20:19critical steps.
20:21So, this critical critical steps is the
20:23steps to solve. Um but like I said, the
20:26golden answer has two parts. Well,
20:28actually three if you include you have
20:30to upload your golden step file.
20:33Must be zipped. Um
20:35then the two written portions of that uh
20:38golden answer is the description of the
20:40CAD file and the steps to solve.
20:43Number four, prompt asset is low
20:45quality. We see this quite a bit. And we
20:48do not want to try to trick the model.
20:51We need to give the model a fair shot at
20:54trying to create your drawing. Create a
20:583D model based on your drawing. And so,
21:00that means it needs to be fully
21:02dimensioned. It There does not need to
21:05be room for interpretation. Doing this
21:07allows the model to have a fair shot.
21:10And number five, this is very common,
21:13the rubric is unbalanced. Please ensure
21:16that your rubric does meet these
21:18requirements. Nines 35 to 45%, fives 45
21:22to 60%, ones 10 to 20%, and then
21:26negatives 35% or less.
21:29That is all I have for today.
21:32Um
21:32I wish everyone the best of luck on this
21:34project. We have a really good team.
21:36Um everyone's eager to work together and
21:40um
21:41good luck to everybody out there. Thank
21:42you.