Free YouTube Transcribe

Video transcript

Task Failed Successfully - Jeremiah Lowin

PyData · 6,984 words · 32 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00good afternoon my name is Jeremiah

0:01thanks so much for coming to my talk I

0:04hope everyone here knows this is a talk

0:06about the perils of workflow management

0:08and not one of the machine learning

0:09talks is taking place right now so if

0:11you want to take a moment reevaluate the

0:13room you're in that's totally fine for

0:15the rest of you I'm really happy to be

0:16here as one of the representatives of

0:18our DC Data community I have spent most

0:24of my career in various forms of data

0:26science usually in finance usually in

0:29risk management a few years ago I became

0:31a PMC of Apache airflow and this year I

0:34formed a software company called prefect

0:36here in DC I also want to mention I've

0:39had a long very friendly relationship

0:41with a company called quanto peein some

0:42of you may have attended James's talk

0:44this morning on using air flow in

0:46productions James is an SRE Encanto peon

0:49he's a really fantastic resource if you

0:51want to talk about deploying air flow in

0:54fact the first person who ever saw our

0:56prefect work was the head of data

0:57science at quant opions I think

0:59incredibly highly of their company

1:00however I do want to give some fair

1:03warning that this is going to be a very

1:04different type of talk than James and

1:07the and the inspiration for today's talk

1:09actually came from an error message task

1:13failed successfully the more I thought

1:17about this first I thought it was funny

1:18and then and then I I really liked it

1:20because this is sort of what I do I'm a

1:22risk manager and more recently I built

1:24workflow software so my job is to make

1:27sure that when failure happens it is

1:29handled correctly but even so what does

1:32it actually mean to fail

1:34successfully and and and maybe more

1:36importantly for all of you how do you

1:38turn that into actionable advice so I

1:40think we should start with the basics

1:42what is workflow management it's a good

1:45question and here's what I think in a

1:47very simple sense workflow management is

1:50when you have things you want done and

1:51you put a system in place to make sure

1:54they happen and that's it our team uses

1:56the terms positive and negative

1:58engineering to differentiate between

2:00these two types of work positive

2:02engineering is the work you actually

2:03want to do negative engineering is

2:06everything you have to do to make sure

2:07it gets done that's the error trapping

2:10and handling the type coercion the data

2:11serialization

2:12infrastructure all of that stuff the

2:15most data scientist data engineers want

2:16to be writing positive code and I

2:18believe it's the role of a good workflow

2:20system to support them by taking on the

2:22negative side in that sense negative

2:25engineering is a form of risk management

2:29now negative engineering requires you to

2:32anticipate and defend against all the

2:35possible ways your system could fail or

2:37in a more general sense behave some

2:40pessimists might say that's impossible

2:42and others missed the point by assuming

2:45it's easy they say well if you don't

2:47want errors just write better code and

2:49in my past life that would have been

2:51like saying oh you don't want to lose

2:52money just buy better stocks it's easy

2:55to say very very hard to do and it's

2:58frankly silly either way as a risk

2:59manager I can tell you that this is very

3:01doable in the right framework so to

3:04illustrate what negative engineering

3:05actually entails let's walk through an

3:08example

3:08so here's finally some Python this is a

3:10simple made-up task that I want to do

3:14and since it's what I want to do it's a

3:16form of positive engineering but now we

3:19are putting on our risk management hats

3:20and we start thinking about negative

3:21outcomes how do we defend against

3:24failure so the first thing that most

3:27people do is they trap the error right

3:29it's very easy it's very straightforward

3:31do this if it doesn't work raise a

3:32helpful error now a cynic might note

3:34that this just replaces one error with

3:37another and while that's true it could

3:39also be a case of better the error I

3:41know

3:41than the error I don't in either case we

3:44can do better and so most people will

3:46move on to retrial and so here we've

3:49added some logic to retry the task up to

3:51five times and it's gonna wait ten

3:54minutes between each attempt in case the

3:56system needs to cool down so we feel

3:58good about this we put it into

3:59production and immediately it starts

4:01failing something's wrong with the

4:03function arguments we don't have a way

4:05to replicate it locally so let's debug

4:08it in production and the way we're going

4:09to do this is with very specific errors

4:12when our arguments have the wrong type

4:13so have we actually solved a problem

4:16no but do we feel better about it maybe

4:20what we have done is taken a situation

4:23we already knew was bad

4:25and added an explicit hook to find out

4:27when it happens it may seem silly but

4:29that is a totally valid negative

4:31engineering exercise no one has ever

4:33complained about having too much

4:35information but now it turns out that

4:37our task is hogging resources and when

4:39it sleeps for 10 minutes it is still

4:41running so our full failure cycle takes

4:43almost an hour our team wants us to do

4:45better so we're gonna retry

4:47asynchronously and now in addition to

4:49writing error trapping code we're gonna

4:51write error trapping infrastructure

4:53which I'm obviously not gonna bother to

4:55write here but our function now returns

4:57these AP likes API like status messages

5:00so that some other system presumably can

5:02manage it and and retry it appropriately

5:04and it turns out we're not even done

5:06with that sometimes you might remember

5:09that this function which I don't even

5:10know where it is here anymore was doing

5:121 divided by a result and oh sorry

5:16and every now and then it turns out that

5:18that result can be 0

5:19so we need to introduce a special

5:21handler for that zero case lest that

5:24error take down our whole system so

5:26let's take a step back and see where we

5:27started and where we ended up this is

5:29what we wanted to do this is what we had

5:32to write to actually do it this is the

5:34positive engineering that's the negative

5:36engineering so yes my example is

5:39extremely contrived but if you think

5:41it's unrealistic then you need to talk

5:44to more data engineers for this example

5:46we worked with a single task to control

5:48issues that might arise in only its

5:50execution and that's great but even that

5:52doesn't justify a full workflow

5:54management system workflow managers

5:56become much more important when we start

5:58to deal with the dependencies between

5:59tasks and I would argue that what takes

6:03a series of function calls and graduates

6:06it to a workflow with a capital W is the

6:09fact that some process is monitoring and

6:11managing the full state of the system

6:13because while any task can in theory

6:16look out for its own state and its own

6:18failures we need something to ensure the

6:20integrity of the system as a whole and

6:23in order to do that we have to know how

6:25to work with failure the simple question

6:29is if a task has an error and no one is

6:32around to depend on that task and does

6:33anyone really care did it actually fail

6:37but if someone does depend on that task

6:39whether it's a person or another task

6:40then all of a sudden handling that

6:42failure becomes critical and so this

6:44idea of failing successfully is not just

6:47about reporting that an error happened

6:49it's about defining the behavior of the

6:51system once that error takes place and

6:53this is super important because task

6:55failure should not automatically mean

6:57system failure

6:59failing successfully means that failure

7:02is a part of our system it is a state

7:04just like success that we can work with

7:07that we can react to and that can drive

7:09business logic if you the designer want

7:13to apply meaning to that failure then

7:15that's totally fine maybe you want the

7:17workflow to stop as soon as it has a

7:19problem maybe you wanted to stop only

7:21after 10% of tasks fail

7:23maybe you expect failure and want to

7:25react to it every single time

7:27all of those are totally valid it's not

7:29for the workflow system to decide

7:31whether failure is necessarily good or

7:33bad it's job is to reveal it promote it

7:38and make it available for any business

7:40purpose and that is how you fail

7:43successfully now what happens when we

7:47don't embrace failure when we treat it

7:49as an outlier workflow systems are often

7:52victims of something that in my finance

7:53days we called wrong-way risk wrong-way

7:56risk is the idea that your exposure to a

7:59bad outcome is exaggerated when the bad

8:02outcome occurs in the classic example is

8:04that you buy insurance against some

8:06natural disaster and one day the

8:08disaster actually happens and you go to

8:10file a claim and you find out that your

8:12insurance company only sold policies to

8:14people in your neighborhood so

8:16everybody's making a claim and the

8:18insurance company is bankrupt and you

8:19wind up with nothing so you had

8:21wrong-way risk you thought you were

8:22insured but your insurance turned out to

8:24be negatively correlated with the thing

8:26that you are insuring against workflow

8:29management systems introduced wrong-way

8:31risk when they treat failure as an

8:33outlier or an afterthought by not taking

8:36it seriously or not providing

8:38first-class tools for working with it

8:39they themselves at exactly the

8:42moment that they should be most useful

8:43and the proof of that is simple if you

8:45could guarantee that your workflow would

8:47run perfectly you don't need a workflow

8:49manager at all

8:51just put in a script and run it whenever

8:52you need it it's only as we start to

8:54introduce some probability of failure

8:56the workflow system becomes really

8:58important it's how we're gonna deal with

8:59those failed states so therefore if the

9:02workflow system treats failure is

9:04unexpected it is least equipped to deal

9:07with the situation that we need it the

9:10most and that is wrong-way risk now a

9:13lot of people assume they can control

9:15this by building robust data pipelines

9:18but usually they were exclusively about

9:19the data and you talk about data lineage

9:21and data latency and data storage but

9:23you never hear about the workflow system

9:25itself collapsing and that's because of

9:27a rather massive assumption that we all

9:29make that our workflow systems are

9:30perfect that they handled themselves the

9:32right way and that they do not fail and

9:34I wish I could live in that world but I

9:36have touched almost every line of

9:37airflows code and I can assure you that

9:38there's nothing magic about a workflow

9:40system it has a strength it has a

9:42weakness and it doesn't care how much

9:44time you spent on your ETL checks it

9:45will fail if you don't use it the right

9:47way

9:47so you haven't set yourself up to fail

9:50successfully you haven't set yourself up

9:52to succeed in any way at all if you

9:54don't play to your workflow system

9:55strengths we can see how this will play

9:57out this is what a workflow system is

10:00supposed to do it's supposed to protect

10:02against problems that's the negative

10:04engineering aspect it's supposed to

10:06promise that work will get done that's

10:07the positive engineering side and of

10:09course it needs to provide us with the

10:10tools that we need for working with the

10:12results of those tasks good or bad but

10:14what really happens is quite different

10:17first workflow systems compromise in

10:21order to make good on their promise to

10:23protect against failure they limit what

10:25they let you do because you can't do

10:28very much they don't have to defend

10:29against very much and they minimize the

10:31number of things that can go wrong so

10:32the first sin is the compromise the

10:36second sin is the cheat the cheat is how

10:39we get around the compromise if your

10:41workflow system is compromised because

10:43it doesn't handle retries it's no

10:45problem just write a script like I did

10:46before that does the retry yourself at

10:49the end of the day it's all code and you

10:52can do whatever you want the only

10:53question is whether your code complies

10:55with the expectations of your workflow

10:57system so that might not seem like a

10:59problem but it can be really really

11:00dangerous you end up

11:02effectively running two systems neither

11:04aware of the other and without any

11:06common vocabulary for understanding what

11:08the other ones doing and that leads to

11:10the third sin corruption all workflow

11:14systems are designed with a philosophy

11:15of how workflow should behave if it's

11:17not written down then it's expressed in

11:19what kind of tasks they can run and what

11:20the contract is that those tasks fulfill

11:22the reason the previous sin was called

11:24cheating is because we were forcing the

11:26workflow manager to do something

11:28contrary to its philosophy contrary to

11:30what it was designed to do and that

11:32corrupts it how do you record a tasks

11:35retry attempt if you don't know what a

11:36retry is how do you decide if a task is

11:39ready to run based on data if you don't

11:40track data dependencies how do you

11:42resume from a failed state if your

11:44system was designed to exit on failure

11:46this is the most common type of airflow

11:49bug we see it happens when you try to do

11:51something that airflow doesn't really

11:52understand but lets you do anyway the

11:54system gets corrupted recovery can be

11:56impossible how would you like to see

11:59some examples first up something called

12:04X coms one of air flows compromises is

12:08that tasks communicate purely by a state

12:11in this case success or failure and

12:14therefore there is no way for them to

12:15exchange data or information people kept

12:18coming up with ways to get around this

12:19so a few years ago we cooked up

12:21something called X coms which stands for

12:23cross communication to a user X comes

12:26look like a first class way to move data

12:29between tasks but from airflows point of

12:31view air flows are excuse me ex-cons are

12:34a hack they're definitely a cheat under

12:38the hood and don't worry too much about

12:39this and the XCOM is just a way for

12:41tasks to impersonate an admin user and

12:44write arbitrary data to the airflow

12:45database tasks push or pull information

12:48from the database which creates strong

12:49data dependencies between tasks and

12:52these data dependencies can exist

12:53between two tasks between two dads or

12:56even between a single task and a past or

12:58future version of itself it is a very

13:01rigid and error-prone form of dependency

13:04and the kicker is that airflow the

13:06system we use to manage dependencies has

13:08no idea it exists it has no vocabulary

13:10to describe or enforce a database

13:13dependency

13:14and so the single most common airflow

13:16bug we see is users who have used X

13:18comms to hack strong data dependencies

13:21into a workflow system that doesn't

13:22track data dependencies what this means

13:25is that if a task ever fails and

13:27therefore doesn't push the XCOM the next

13:30task is also gonna fail an airflow has

13:32no idea why because we just we went

13:34outside the system it leads to a very

13:36frustrating chain of failures even worse

13:39X comms are serialized into the airflow

13:42database with no way to delete them and

13:44no exploration so if you're processing a

13:4610 gigabyte data frame and you pass it

13:48through 100 tasks then every single run

13:51is generating a terabyte of permanent

13:52storage which is just it's sort of

13:54insane and it all begs the question

13:56should I use X coms and the answer is

14:00probably not and the punchline

14:02is that that's me okay I wrote X coms

14:05about 3 years ago because I was

14:06frustrated by this problem and I own

14:09this cheat and I regret it and the truth

14:12is I don't think airflow is usable for

14:13data science without X comments I think

14:15you need to move data around data

14:17dependencies are just too important in

14:19the modern data stack but when I

14:21designed these X comes I was using them

14:23to pass around tiny little bits of

14:24transient information and I remember

14:26Maxime he's the original developer of

14:28airflow telling me that the system was

14:30prone to abuse but I didn't listen I

14:32just I thought this was too important to

14:34include in the system but sadly as I

14:37said it's become one of the largest

14:38source of largest sources of untraceable

14:41bugs in the air flow world so watch out

14:43for your ex comms

14:44the next thing I'd like to talk about is

14:47something called the branch operator

14:49just by truth the branch operator allows

14:52us to include branching logic which is

14:55pretty fantastic we need that a lot of

14:57the time and it's also pretty amazing

14:59because it's not possible in air flow

15:00the same compromised functionality that

15:03led me to make X comms says that there's

15:06no way for a single task to tell

15:08multiple downstream tasks that one of

15:10them should run and the other shouldn't

15:11all it can do is return success or

15:13failure but branching is important so so

15:15what do we do and of course we cheat we

15:20don't need to look too closely at the

15:21code but but if you do look in the code

15:23you'll see that the task is calling a

15:24method called skip

15:26and that seems kind of reasonable that

15:28it would call this method that lives on

15:30itself until you ask well why is this

15:31task skipping itself and the answer is

15:33it's not it's actually skipping the

15:36tasks that goes after it

15:37so the Skip method is another backdoor

15:40into the air flow metadata this task is

15:42manipulating the state of a task that

15:44hasn't been evaluated yet and from a

15:46data integrity standpoint that should

15:48horrify you right airflow the system

15:51that tracks dependencies and enforces

15:53contracts between tasks is letting a

15:55task impersonate an admin and mess

15:57around with other tasks it's a

15:59completely untraceable form of

16:01dependency and it's the most pure

16:02violation of air flows philosophy that

16:04I'm aware of but it's incredibly useful

16:08so we forgive it and in the interest of

16:11time I'm going to move through some of

16:12the some of the texts but these will be

16:14available later for you to read the last

16:16example that I want to run through in

16:18detail is about the intersection of air

16:19flow state and business logic so this

16:23isn't so much a compromise actually as

16:24it is a choice remember that a good work

16:27flow system is supposed to track state

16:28but be agnostic to how it's used for a

16:30business purpose to agnostic to what it

16:33signifies and that's what lets us fail

16:35successfully we elevate state to a place

16:37where you can make decisions based on it

16:39so this is an interesting situation this

16:41is a work flow that has two tasks one of

16:44them is very important it does something

16:45that you care about and the other one

16:47just cleans up if the first one has a

16:49problem so on the left this is a

16:51situation where the important task was

16:53successful and we didn't even bother

16:55running the cleanup and I think we all

16:57agree with air flow that that is a

16:58successful run of this workflow but on

17:00the right we have a situation where the

17:02important tasks failed and the cleanup

17:05task was successful now air flow thinks

17:07that this run was successful and to be

17:10honest I don't know if it was or wasn't

17:12that's up to the person who designed it

17:14maybe they only care that the workflow

17:16ran and so yes they would like to see

17:17this as a green light on their dashboard

17:19but I can easily imagine the situation

17:21where someone cares deeply about the

17:23important first task running in which

17:25case they'd like to see this as a

17:27failure but air flow does not give you a

17:30choice

17:31air flow has it hard-coded the task

17:34State excuse me dag State is always

17:36determined by the last tasks in the

17:39Dagon so this is a design failure this

17:42means that our workflow system is not

17:44capable of reflecting business logic and

17:47therefore we cannot fail successfully

17:50the system is corrupt again skip some

17:53texts in the interest of time and we'll

17:54just move very quickly if you're

17:56interested through a couple of other

17:57quick examples people use airflows

18:00global variables in place of parameters

18:02which Justin creates a very staple mess

18:05people abuse the fact that airflow

18:07reruns a dag every single time it runs a

18:09task to introduce tags that change while

18:12they're running which I don't even have

18:14to tell you how horrifying that should

18:15be off-schedule runs air flow used to

18:18not support off schedule runs we

18:20eventually sort of found a way to make

18:22it do it and we found that this this

18:24assumption was so baked into the system

18:26that both the scheduler and the UI broke

18:28as soon as we ran something off schedule

18:30in particular the schedule was working

18:32by adding an interval to the most recent

18:35run so if you had something ticking

18:36along every hour it was totally fine

18:38then all of a sudden if you ran

18:39something on the half-hour the scheduler

18:41would start adding the hour to the

18:42half-hour time and and keep going from

18:44there so again the idea is that these

18:47these are not impossible you can make

18:49the system do whatever you want but if

18:50you don't respect the assumptions of the

18:52system you're going to end up in an

18:53unexpected place and the last one I'll

18:56talk about is that air flows database

18:57has a unique constraint on a dag in its

18:59execution time which means that you

19:01cannot run a dag twice at the same time

19:04it doesn't support the concept of

19:06simultaneous runs of the same dag now if

19:08you live in air flows world that's

19:10totally fine that makes sense if you

19:11live in my world where I want to run

19:12workflows all the time no matter what

19:14for any circumstance that makes no sense

19:16at all so that's a that's a big problem

19:17to me I think it's extremely important

19:22for me to note that these problems did

19:23not start with air flow and they are not

19:25unique to air flow I just happen to know

19:26them very well in an air flow context

19:28from solving users problems I want to

19:31repeat that all work flow managers have

19:33strengths and weaknesses and identifying

19:35these weaknesses doesn't automatically

19:37mean it's bad it's just a benchmark for

19:39you to measure your appropriate use of

19:42the tool against and we can see how this

19:44has played out through history

19:45all right so in the beginning there was

19:47cron and in fact cron I was very

19:49surprised to learn they saw the way back

19:50to 90

19:5187 and I'll bet that every single person

19:53in this room has written a cron cheat

19:55because cron makes the compromise of

19:57course that there are no dependencies so

19:59if you want two things to run in

20:00sequence you schedule one for midnight

20:02you schedule one for 1:00 a.m. and you

20:04hope they work that's the crunchy and

20:06after cron we can move pretty pretty

20:08quickly this is Uzi which is cron in

20:10Java and XML

20:12this is Azkaban which is cron with

20:14dependencies this is Luigi with

20:16dependent which is cron with

20:17dependencies that are based on side

20:18effects SPARC is an interesting one to

20:21mention it doesn't feel like it fits in

20:24this pattern of data engineering

20:25workflow managers but if you have spent

20:27enough time in both the data engineering

20:28world and the data science world you'll

20:29quickly remember that spark and other

20:31tools like it including beam task even

20:33tensor flow these are just ways of

20:35describing computational graphs and

20:37moving data or some state through them

20:40so it just happens that SPARC passes

20:41data instead of States but it's the same

20:44idea and last but not least of course we

20:46have air flow which I would basically

20:47describe this cron with a database so

20:51the question is what have we learned

20:52from this history of work flow

20:54innovation which by the way I think

20:56stopped in 2015 this is a quote from the

21:00creator of Luigi said all workflow

21:02engines kind of suck

21:04including Luigi the question is well why

21:06why do they have to second and one

21:08answer is that they were originally

21:10created to solve a specific problem at a

21:13specific company which means that while

21:15we can today point back and say yeah

21:17they were compromised

21:18they weren't compromised in the moment

21:19they were appropriate to the task at

21:21hand it's only once they were open

21:22sourced and attempted to be widely

21:24applied that these compromises became

21:26evident and in the ahta in the absence

21:28of a modern alternative they've become

21:30weighed down by all of these cheats that

21:32I've been enumerated a common response

21:35from the authors of these platforms

21:37including myself is that users should

21:39just write idempotent tasks does

21:42everyone know what idempotency is it's

21:44it's it's basically the idea that if you

21:46run something multiple times it'll only

21:47cause an action once so this is sort of

21:51like the get-out-of-jail-free card of

21:53data engineering and data engineering

21:54you know workflow authors because of

21:56course if you have an idempotent task

21:58it'll work that you could use cron and

22:01just spam a server every two seconds

22:02with a bunch of item potent fast you

22:03don't need a workflow manager

22:05so it's like yeah this is true but does

22:07it really matter is it practical and the

22:09answers is no it's not so I'd like us

22:11you know you know to look to be clear

22:13idempotency is good but we need to live

22:15in the real world where it's unlikely

22:16and possibly impossible and perhaps

22:19impossible and so instead of pretending

22:21that people can write perfect code let's

22:24acknowledge that they write bad code and

22:26see what we can do to support them in

22:31today's talk I've shown that these

22:33workflow managers they commit these same

22:35basic sins over and over and because

22:37they implement compromise limited

22:39functionality they end up cheating to

22:41implement the features they want how can

22:42we overcome this we need to start with

22:44what users actually want to do and if

22:47you're the users in this room what you

22:49want to do is you want to write Python

22:51so this is what you know the the

22:54simplest ETL script I can imagine looks

22:56like in Python and this is an example of

22:58positive engineering this is what you

23:00want to do how would we design an

23:03idealized workflow system around this

23:05code well if you think about it almost

23:07everything we need is already here right

23:09we know what the what the tasks are we

23:11know what the workflow definition is we

23:13just need a way to communicate this to

23:15of this magic system that we're

23:17inventing and on and the way we need to

23:19do that is by creating entry points for

23:21our negative engineering hooks so here's

23:24what that might look like and if you

23:25don't see the difference I'll highlight

23:27it all we've done is introduce these

23:29decorators which will tell our system

23:32these are the tasks they're gonna be

23:34they're gonna run keep an eye on them

23:36manage their state manage their data

23:37dependencies and then down here we've

23:39created a workflow we're saying hey this

23:41is my computational graph it's any

23:43Python and I want you to pay attention

23:45to what happens this is going to be the

23:46workflow that we run in the future and

23:48that's it I don't have to explain

23:50anything else in theory it's all here so

23:52there's one last thing that we need and

23:55that's to import prefect so this is real

23:58this works today this is actually our

24:00hello world and this is the first time

24:01that we've ever shown it publicly which

24:03is kind of cool for us my team has

24:05worked incredibly hard to build this API

24:07not to mention the system that powers it

24:09and we're very excited that we're going

24:11to be open sourcing it in the very near

24:13future and I promise that it fails more

24:16successfully than anything you've

24:18scene in fact the most important thing

24:23that we did with prefect was that we

24:25allowed it to fail successfully prefect

24:27will track the state and result of every

24:29task in your workflow and let them

24:30depend on each other in a variety of

24:32ways we use these rich state objects the

24:34currency of the system and we allow

24:36users to apply semantics to those

24:38objects don't worry we have sensible

24:40defaults so you don't actually have to

24:41do anything more than what I showed a

24:43moment ago if you don't want to so a

24:45traditional data pipeline implemented in

24:47Prefect would ignore state and just pass

24:49data a more complex pipeline might use

24:51the state to make decisions maybe to

24:54kick off a cleanup job and oh really out

24:55there workflow could actually use could

24:58actually let failed tasks return data

25:01for example their own exception and use

25:04that to drive Diagnostics downstream as

25:06we think these abstractions become

25:07incredibly powerful and they are only

25:09possible because the system is agnostic

25:11to our applicants to its application it

25:13is failing successfully and I want to

25:16mention that under the hood we are using

25:18desc to drive pretty much all the

25:20execution we think that Basques is one

25:22of the best tools out there and was

25:23actually when I really got my hands on

25:25desk as a data scientist that I thought

25:27for the first time that something like

25:29prefect might even be possible and will

25:31have a lot more to say about that very

25:33soon so one last thing since I've been

25:36talking a lot about air flow today I

25:37spent years on air flow it's a very

25:39important tool to me it's a very

25:41important tool to the data engineering

25:42community I do think it has some flaws

25:44prefect does not share any code with air

25:47flow but it will run air flow for you

25:50and so we have some partners who are

25:51using prefect to modernize their air

25:53flow stack without losing any of the

25:55work that they put into building it so I

25:57want to thank you all for coming to my

25:59talk about failure if you'd rather talk

26:01about success we're working with a

26:03number of partners to develop prefect

26:05and if you're in this room today we

26:06would love to work with you you can

26:08reach us at this address which is hello

26:10at prefect IO or come find us outside

26:12we're the guys who have been handing out

26:13free compass coffee all day and since I

26:16think we have a little bit of time left

26:18um do we yeah then I'm happy to take

26:20some questions

26:21if there are any

26:26[Music]

26:28[Applause]

26:31or if there aren't yes one second hey

26:40Arno media capital one so the question I

26:43have is when we look at some of these

26:45workflows obviously there's all the

26:48problems that happen operationally and

26:49then there's all of the things than the

26:50challenges that we face what we're

26:52trying to develop them trying to figure

26:53it out debug it piece it through how can

26:56you compare and contrast what you're

26:57promising so far with Prefect

26:59in terms of as I am building a workflow

27:02what the debugging process looks like

27:05getting into the state being able to

27:06develop locally versus sort of that

27:08deployment stage so this will run on

27:12your local machine if you want to run it

27:15in a synchronous process so you can

27:16actually put breakpoints and inspect

27:18them that's totally fine we do that for

27:19our own testing if you want to run it in

27:21a local desk cluster or even using a

27:23desk synchronous executor that's totally

27:25fine and then you can also deploy this

27:26at you know into a cluster very easily

27:28but these are these are just functions

27:31there they're exposed in this form to

27:33you and they will do whatever it is that

27:34you want one of the one of the core

27:35design principles of prefect is that the

27:38tasks themselves are black boxes we we

27:40don't really care what's in them we're

27:42agnostic to it we actually don't even

27:43receive the code from our partners sort

27:46of an important privacy barrier that

27:47we've created so you can do whatever you

27:49want interact with your systems however

27:51you want we really don't care frankly so

27:55if you need to debug this locally go you

27:56just call it as a function you don't

27:57even have to run the flow it's it's not

27:59tied to anything unless you want it to

28:00be so this is actually we worked very

28:17hard on this so I like to show it off

28:18this is sort of our functional API it's

28:20not actually running the functions it's

28:21just building a computational graph

28:22similar to you know any deep learning

28:24tool that you might use today we have a

28:26fully class-based you know system behind

28:29it just like air flow if you want to

28:30call set upstream if you want to

28:31manipulate tasks and edges directly you

28:33you absolutely can if I had added a line

28:36here that was flowed out run

28:38you would receive back a dictionary with

28:40all of the state objects from all of the

28:42tasks in that run you could manipulate

28:45them and examine them and do whatever

28:46you want with them you can even crack

28:47open the hood and see how we're working

28:49with them so to be clear yeah we have a

28:54functional API we also have what we call

28:56our imperative API for more programmatic

28:58access but this is more fun could you

29:04talk more about integration with airflow

29:07I kind of made it you sounded like you

29:11made a promise if I keep working in air

29:12flow then you promise me something this

29:14is better later on I'm not like hosts in

29:17terms of like six months of development

29:18of like bunch of tags we we I challenged

29:23my team to did I do that I challenged my

29:27team to translate airflow into prefect

29:29which I thought would be rather doable

29:33based on my knowledge of prefect based

29:34on my knowledge of air flow we decided

29:37not to do that because there are some

29:38features of air flow that we don't

29:40support with a direct analog X comms

29:43being one of them we have what we think

29:44is a much stronger data dependency

29:47feature that we're that we're showing

29:48here and we didn't want to write

29:50translation logic that we would then

29:52have to maintain so we're working with

29:56some of our partners some of whom are

29:58actually other airflow PMC's who are

30:01using prefect to instrument their

30:04airflow deploys so you keep your airflow

30:06database running you keep air flow

30:07running we have flow in prefect which

30:12creates a copy a one four one copy of

30:14every task in your dag loads it into

30:16prefix so that you can use all of our

30:18instrument respectin tools you can use

30:20our API you can use our UI to examine it

30:22and run it but at the end of the day

30:24it's calling back to airflow to run your

30:26existing system so we've abandoned the

30:28idea of translating it because we just

30:30refuse to support some features

30:34maybe this is a weird thing to say but

30:36so I use a Z in my career before and I'm

30:38using air flow right now and you know

30:41I've never been a JavaScript programmer

30:43but you see the like the rise and death

30:44of like JavaScript frameworks and

30:47visualizations of that one of the

30:50reasons well we chose the Oracle flow

30:53because of the utility to us but we also

30:55were excited as someone on Google cloud

30:57platform that they kind of made it a

30:59first-class citizen with cloud composer

31:00do you see this world is like the coming

31:04and dying of the same way we have

31:05JavaScript frameworks like what's hot

31:07right now in terms of workflow

31:08orchestration and it's like next year

31:10like I'm gonna be applying for air flow

31:11jobs if you're like I don't care about

31:13that at all I in that time is it just

31:17like are there still like as you were

31:19saying people make these things and they

31:21open source and then everybody makes a

31:22ton of demands about what they should do

31:24and then they're like it wasn't designed

31:25to do that and they're like okay now on

31:27the next one I think that's the first

31:30time anyone's ever described the

31:31workflow manager space is hot to me I

31:33had I had to beg people to take me

31:36seriously when I said that this was even

31:37possible I mean we've had conversations

31:40with people who deny that it's possible

31:41until we show them how we how we did it

31:43air flow is close to four years old now

31:46and nothing has unseated it and so while

31:49it has gained momentum and to be clear I

31:50am an air flow PMC I am extremely

31:52excited to see that I think I have the

31:54third most commits of anyone in the

31:55world on air flow it's an amazing thing

31:58but it's not a hot space it's not a

32:01crowded space I'm frankly amazed that we

32:05see this emergence of an air flow

32:09developer like I'm not sure that's a

32:11good thing to be honest with you sort of

32:12my argument here is that the workflow

32:13system should fade right you shouldn't

32:16be an air flow developer you should be a

32:17developer who happens to use air flow

32:24any other questions yes we have we've

32:41designed our compute to be modular at

32:43the moment we preferred ask and we used

32:45ask running in kubernetes like how do

33:00you sorry I mean how do you had let's

33:04suppose suppose you have one step that's

33:06got a lot of different software than the

33:08next step like how would you address

33:10that a lot of different dependencies yes

33:12so software dependencies yeah no it's a

33:14good question today the the workflow is

33:16built as a container with a single set

33:17of dependencies you put them all in and

33:19we're working towards supporting an

33:21individual container per task yeah yeah

33:23okay okay is that a hint are we supposed

33:30to get out okay I'm happy to take

33:33questions as long as there are yes

33:42there you go thank you my apologies if

33:45you already answer this but I know many

33:47people have found frustrations with

33:51distributing airflow drops using celery

33:55just with the complications of a celery

33:58and things what's the story for prefect

34:01surround running a cluster where jobs

34:04are run across it sure so that's what we

34:08used ask for tasks is hands-down the

34:11best way to do that can you describe

34:13that a little bit further assuming

34:15people know I know about tasks and

34:18assuming others do can you describe how

34:20it yeah we're just best voice one of the

34:22easiest ways I've found to describe desk

34:24is actually a sort of a pure Python and

34:26friendly version of SPARC so desk is I

34:28think just an extremely easy way to

34:30distribute computation across any number

34:33of nodes whether it's your local laptop

34:35or something running in the cloud

34:36it'll take a function it'll serialize it

34:38it'll run it in the cloud they'll bring

34:40back a result or what's really cool

34:42about it and what we'd like to take

34:43advantage of is you can dynamically

34:45build yet further graphs from an already

34:47running graph so I mean in a nutshell is

34:49just a very easy way to run distributed

34:52computations frankly it's the best I've

34:55seen so we use it for our primary

34:59execution engine it's modular if

35:01somebody wants to write a an extension

35:03for something else they absolutely could

35:04you want to go back to celery by all

35:05means go right ahead but we're promoting

35:07desk as our as our primary executor yes

35:20do you expose the that you're running on

35:22tasks at all to the users thank do they

35:24get to see the task dashboard is there

35:26any exposure of that at all it's a great

35:28question so at the moment no but not

35:31because we've decided not to we just

35:33sort of haven't addressed that question

35:36so it's a good one we don't hide that

35:39we're running us in fact we would invite

35:41users who want to take advantage of the

35:42fact that we're running in DES to just

35:43open up a worker client and we're trying

35:45to figure out the right way to create a

35:47sort of first-class API where you could

35:49still run that locally effectively and

35:51do and do your debugging but also have

35:53it work properly in a distributed

35:54environment so we we don't hide it from

35:56users we also at the moment aren't

35:58explicitly promoting it to users it's

36:00just a it's an important act this

36:10all right well thank you all so much I

36:12appreciate your attention in your time

More from PyData

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.