Full transcript
0:00good afternoon my name is Jeremiah
0:01thanks so much for coming to my talk I
0:04hope everyone here knows this is a talk
0:06about the perils of workflow management
0:08and not one of the machine learning
0:09talks is taking place right now so if
0:11you want to take a moment reevaluate the
0:13room you're in that's totally fine for
0:15the rest of you I'm really happy to be
0:16here as one of the representatives of
0:18our DC Data community I have spent most
0:24of my career in various forms of data
0:26science usually in finance usually in
0:29risk management a few years ago I became
0:31a PMC of Apache airflow and this year I
0:34formed a software company called prefect
0:36here in DC I also want to mention I've
0:39had a long very friendly relationship
0:41with a company called quanto peein some
0:42of you may have attended James's talk
0:44this morning on using air flow in
0:46productions James is an SRE Encanto peon
0:49he's a really fantastic resource if you
0:51want to talk about deploying air flow in
0:54fact the first person who ever saw our
0:56prefect work was the head of data
0:57science at quant opions I think
0:59incredibly highly of their company
1:00however I do want to give some fair
1:03warning that this is going to be a very
1:04different type of talk than James and
1:07the and the inspiration for today's talk
1:09actually came from an error message task
1:13failed successfully the more I thought
1:17about this first I thought it was funny
1:18and then and then I I really liked it
1:20because this is sort of what I do I'm a
1:22risk manager and more recently I built
1:24workflow software so my job is to make
1:27sure that when failure happens it is
1:29handled correctly but even so what does
1:32it actually mean to fail
1:34successfully and and and maybe more
1:36importantly for all of you how do you
1:38turn that into actionable advice so I
1:40think we should start with the basics
1:42what is workflow management it's a good
1:45question and here's what I think in a
1:47very simple sense workflow management is
1:50when you have things you want done and
1:51you put a system in place to make sure
1:54they happen and that's it our team uses
1:56the terms positive and negative
1:58engineering to differentiate between
2:00these two types of work positive
2:02engineering is the work you actually
2:03want to do negative engineering is
2:06everything you have to do to make sure
2:07it gets done that's the error trapping
2:10and handling the type coercion the data
2:11serialization
2:12infrastructure all of that stuff the
2:15most data scientist data engineers want
2:16to be writing positive code and I
2:18believe it's the role of a good workflow
2:20system to support them by taking on the
2:22negative side in that sense negative
2:25engineering is a form of risk management
2:29now negative engineering requires you to
2:32anticipate and defend against all the
2:35possible ways your system could fail or
2:37in a more general sense behave some
2:40pessimists might say that's impossible
2:42and others missed the point by assuming
2:45it's easy they say well if you don't
2:47want errors just write better code and
2:49in my past life that would have been
2:51like saying oh you don't want to lose
2:52money just buy better stocks it's easy
2:55to say very very hard to do and it's
2:58frankly silly either way as a risk
2:59manager I can tell you that this is very
3:01doable in the right framework so to
3:04illustrate what negative engineering
3:05actually entails let's walk through an
3:08example
3:08so here's finally some Python this is a
3:10simple made-up task that I want to do
3:14and since it's what I want to do it's a
3:16form of positive engineering but now we
3:19are putting on our risk management hats
3:20and we start thinking about negative
3:21outcomes how do we defend against
3:24failure so the first thing that most
3:27people do is they trap the error right
3:29it's very easy it's very straightforward
3:31do this if it doesn't work raise a
3:32helpful error now a cynic might note
3:34that this just replaces one error with
3:37another and while that's true it could
3:39also be a case of better the error I
3:41know
3:41than the error I don't in either case we
3:44can do better and so most people will
3:46move on to retrial and so here we've
3:49added some logic to retry the task up to
3:51five times and it's gonna wait ten
3:54minutes between each attempt in case the
3:56system needs to cool down so we feel
3:58good about this we put it into
3:59production and immediately it starts
4:01failing something's wrong with the
4:03function arguments we don't have a way
4:05to replicate it locally so let's debug
4:08it in production and the way we're going
4:09to do this is with very specific errors
4:12when our arguments have the wrong type
4:13so have we actually solved a problem
4:16no but do we feel better about it maybe
4:20what we have done is taken a situation
4:23we already knew was bad
4:25and added an explicit hook to find out
4:27when it happens it may seem silly but
4:29that is a totally valid negative
4:31engineering exercise no one has ever
4:33complained about having too much
4:35information but now it turns out that
4:37our task is hogging resources and when
4:39it sleeps for 10 minutes it is still
4:41running so our full failure cycle takes
4:43almost an hour our team wants us to do
4:45better so we're gonna retry
4:47asynchronously and now in addition to
4:49writing error trapping code we're gonna
4:51write error trapping infrastructure
4:53which I'm obviously not gonna bother to
4:55write here but our function now returns
4:57these AP likes API like status messages
5:00so that some other system presumably can
5:02manage it and and retry it appropriately
5:04and it turns out we're not even done
5:06with that sometimes you might remember
5:09that this function which I don't even
5:10know where it is here anymore was doing
5:121 divided by a result and oh sorry
5:16and every now and then it turns out that
5:18that result can be 0
5:19so we need to introduce a special
5:21handler for that zero case lest that
5:24error take down our whole system so
5:26let's take a step back and see where we
5:27started and where we ended up this is
5:29what we wanted to do this is what we had
5:32to write to actually do it this is the
5:34positive engineering that's the negative
5:36engineering so yes my example is
5:39extremely contrived but if you think
5:41it's unrealistic then you need to talk
5:44to more data engineers for this example
5:46we worked with a single task to control
5:48issues that might arise in only its
5:50execution and that's great but even that
5:52doesn't justify a full workflow
5:54management system workflow managers
5:56become much more important when we start
5:58to deal with the dependencies between
5:59tasks and I would argue that what takes
6:03a series of function calls and graduates
6:06it to a workflow with a capital W is the
6:09fact that some process is monitoring and
6:11managing the full state of the system
6:13because while any task can in theory
6:16look out for its own state and its own
6:18failures we need something to ensure the
6:20integrity of the system as a whole and
6:23in order to do that we have to know how
6:25to work with failure the simple question
6:29is if a task has an error and no one is
6:32around to depend on that task and does
6:33anyone really care did it actually fail
6:37but if someone does depend on that task
6:39whether it's a person or another task
6:40then all of a sudden handling that
6:42failure becomes critical and so this
6:44idea of failing successfully is not just
6:47about reporting that an error happened
6:49it's about defining the behavior of the
6:51system once that error takes place and
6:53this is super important because task
6:55failure should not automatically mean
6:57system failure
6:59failing successfully means that failure
7:02is a part of our system it is a state
7:04just like success that we can work with
7:07that we can react to and that can drive
7:09business logic if you the designer want
7:13to apply meaning to that failure then
7:15that's totally fine maybe you want the
7:17workflow to stop as soon as it has a
7:19problem maybe you wanted to stop only
7:21after 10% of tasks fail
7:23maybe you expect failure and want to
7:25react to it every single time
7:27all of those are totally valid it's not
7:29for the workflow system to decide
7:31whether failure is necessarily good or
7:33bad it's job is to reveal it promote it
7:38and make it available for any business
7:40purpose and that is how you fail
7:43successfully now what happens when we
7:47don't embrace failure when we treat it
7:49as an outlier workflow systems are often
7:52victims of something that in my finance
7:53days we called wrong-way risk wrong-way
7:56risk is the idea that your exposure to a
7:59bad outcome is exaggerated when the bad
8:02outcome occurs in the classic example is
8:04that you buy insurance against some
8:06natural disaster and one day the
8:08disaster actually happens and you go to
8:10file a claim and you find out that your
8:12insurance company only sold policies to
8:14people in your neighborhood so
8:16everybody's making a claim and the
8:18insurance company is bankrupt and you
8:19wind up with nothing so you had
8:21wrong-way risk you thought you were
8:22insured but your insurance turned out to
8:24be negatively correlated with the thing
8:26that you are insuring against workflow
8:29management systems introduced wrong-way
8:31risk when they treat failure as an
8:33outlier or an afterthought by not taking
8:36it seriously or not providing
8:38first-class tools for working with it
8:39they themselves at exactly the
8:42moment that they should be most useful
8:43and the proof of that is simple if you
8:45could guarantee that your workflow would
8:47run perfectly you don't need a workflow
8:49manager at all
8:51just put in a script and run it whenever
8:52you need it it's only as we start to
8:54introduce some probability of failure
8:56the workflow system becomes really
8:58important it's how we're gonna deal with
8:59those failed states so therefore if the
9:02workflow system treats failure is
9:04unexpected it is least equipped to deal
9:07with the situation that we need it the
9:10most and that is wrong-way risk now a
9:13lot of people assume they can control
9:15this by building robust data pipelines
9:18but usually they were exclusively about
9:19the data and you talk about data lineage
9:21and data latency and data storage but
9:23you never hear about the workflow system
9:25itself collapsing and that's because of
9:27a rather massive assumption that we all
9:29make that our workflow systems are
9:30perfect that they handled themselves the
9:32right way and that they do not fail and
9:34I wish I could live in that world but I
9:36have touched almost every line of
9:37airflows code and I can assure you that
9:38there's nothing magic about a workflow
9:40system it has a strength it has a
9:42weakness and it doesn't care how much
9:44time you spent on your ETL checks it
9:45will fail if you don't use it the right
9:47way
9:47so you haven't set yourself up to fail
9:50successfully you haven't set yourself up
9:52to succeed in any way at all if you
9:54don't play to your workflow system
9:55strengths we can see how this will play
9:57out this is what a workflow system is
10:00supposed to do it's supposed to protect
10:02against problems that's the negative
10:04engineering aspect it's supposed to
10:06promise that work will get done that's
10:07the positive engineering side and of
10:09course it needs to provide us with the
10:10tools that we need for working with the
10:12results of those tasks good or bad but
10:14what really happens is quite different
10:17first workflow systems compromise in
10:21order to make good on their promise to
10:23protect against failure they limit what
10:25they let you do because you can't do
10:28very much they don't have to defend
10:29against very much and they minimize the
10:31number of things that can go wrong so
10:32the first sin is the compromise the
10:36second sin is the cheat the cheat is how
10:39we get around the compromise if your
10:41workflow system is compromised because
10:43it doesn't handle retries it's no
10:45problem just write a script like I did
10:46before that does the retry yourself at
10:49the end of the day it's all code and you
10:52can do whatever you want the only
10:53question is whether your code complies
10:55with the expectations of your workflow
10:57system so that might not seem like a
10:59problem but it can be really really
11:00dangerous you end up
11:02effectively running two systems neither
11:04aware of the other and without any
11:06common vocabulary for understanding what
11:08the other ones doing and that leads to
11:10the third sin corruption all workflow
11:14systems are designed with a philosophy
11:15of how workflow should behave if it's
11:17not written down then it's expressed in
11:19what kind of tasks they can run and what
11:20the contract is that those tasks fulfill
11:22the reason the previous sin was called
11:24cheating is because we were forcing the
11:26workflow manager to do something
11:28contrary to its philosophy contrary to
11:30what it was designed to do and that
11:32corrupts it how do you record a tasks
11:35retry attempt if you don't know what a
11:36retry is how do you decide if a task is
11:39ready to run based on data if you don't
11:40track data dependencies how do you
11:42resume from a failed state if your
11:44system was designed to exit on failure
11:46this is the most common type of airflow
11:49bug we see it happens when you try to do
11:51something that airflow doesn't really
11:52understand but lets you do anyway the
11:54system gets corrupted recovery can be
11:56impossible how would you like to see
11:59some examples first up something called
12:04X coms one of air flows compromises is
12:08that tasks communicate purely by a state
12:11in this case success or failure and
12:14therefore there is no way for them to
12:15exchange data or information people kept
12:18coming up with ways to get around this
12:19so a few years ago we cooked up
12:21something called X coms which stands for
12:23cross communication to a user X comes
12:26look like a first class way to move data
12:29between tasks but from airflows point of
12:31view air flows are excuse me ex-cons are
12:34a hack they're definitely a cheat under
12:38the hood and don't worry too much about
12:39this and the XCOM is just a way for
12:41tasks to impersonate an admin user and
12:44write arbitrary data to the airflow
12:45database tasks push or pull information
12:48from the database which creates strong
12:49data dependencies between tasks and
12:52these data dependencies can exist
12:53between two tasks between two dads or
12:56even between a single task and a past or
12:58future version of itself it is a very
13:01rigid and error-prone form of dependency
13:04and the kicker is that airflow the
13:06system we use to manage dependencies has
13:08no idea it exists it has no vocabulary
13:10to describe or enforce a database
13:13dependency
13:14and so the single most common airflow
13:16bug we see is users who have used X
13:18comms to hack strong data dependencies
13:21into a workflow system that doesn't
13:22track data dependencies what this means
13:25is that if a task ever fails and
13:27therefore doesn't push the XCOM the next
13:30task is also gonna fail an airflow has
13:32no idea why because we just we went
13:34outside the system it leads to a very
13:36frustrating chain of failures even worse
13:39X comms are serialized into the airflow
13:42database with no way to delete them and
13:44no exploration so if you're processing a
13:4610 gigabyte data frame and you pass it
13:48through 100 tasks then every single run
13:51is generating a terabyte of permanent
13:52storage which is just it's sort of
13:54insane and it all begs the question
13:56should I use X coms and the answer is
14:00probably not and the punchline
14:02is that that's me okay I wrote X coms
14:05about 3 years ago because I was
14:06frustrated by this problem and I own
14:09this cheat and I regret it and the truth
14:12is I don't think airflow is usable for
14:13data science without X comments I think
14:15you need to move data around data
14:17dependencies are just too important in
14:19the modern data stack but when I
14:21designed these X comes I was using them
14:23to pass around tiny little bits of
14:24transient information and I remember
14:26Maxime he's the original developer of
14:28airflow telling me that the system was
14:30prone to abuse but I didn't listen I
14:32just I thought this was too important to
14:34include in the system but sadly as I
14:37said it's become one of the largest
14:38source of largest sources of untraceable
14:41bugs in the air flow world so watch out
14:43for your ex comms
14:44the next thing I'd like to talk about is
14:47something called the branch operator
14:49just by truth the branch operator allows
14:52us to include branching logic which is
14:55pretty fantastic we need that a lot of
14:57the time and it's also pretty amazing
14:59because it's not possible in air flow
15:00the same compromised functionality that
15:03led me to make X comms says that there's
15:06no way for a single task to tell
15:08multiple downstream tasks that one of
15:10them should run and the other shouldn't
15:11all it can do is return success or
15:13failure but branching is important so so
15:15what do we do and of course we cheat we
15:20don't need to look too closely at the
15:21code but but if you do look in the code
15:23you'll see that the task is calling a
15:24method called skip
15:26and that seems kind of reasonable that
15:28it would call this method that lives on
15:30itself until you ask well why is this
15:31task skipping itself and the answer is
15:33it's not it's actually skipping the
15:36tasks that goes after it
15:37so the Skip method is another backdoor
15:40into the air flow metadata this task is
15:42manipulating the state of a task that
15:44hasn't been evaluated yet and from a
15:46data integrity standpoint that should
15:48horrify you right airflow the system
15:51that tracks dependencies and enforces
15:53contracts between tasks is letting a
15:55task impersonate an admin and mess
15:57around with other tasks it's a
15:59completely untraceable form of
16:01dependency and it's the most pure
16:02violation of air flows philosophy that
16:04I'm aware of but it's incredibly useful
16:08so we forgive it and in the interest of
16:11time I'm going to move through some of
16:12the some of the texts but these will be
16:14available later for you to read the last
16:16example that I want to run through in
16:18detail is about the intersection of air
16:19flow state and business logic so this
16:23isn't so much a compromise actually as
16:24it is a choice remember that a good work
16:27flow system is supposed to track state
16:28but be agnostic to how it's used for a
16:30business purpose to agnostic to what it
16:33signifies and that's what lets us fail
16:35successfully we elevate state to a place
16:37where you can make decisions based on it
16:39so this is an interesting situation this
16:41is a work flow that has two tasks one of
16:44them is very important it does something
16:45that you care about and the other one
16:47just cleans up if the first one has a
16:49problem so on the left this is a
16:51situation where the important task was
16:53successful and we didn't even bother
16:55running the cleanup and I think we all
16:57agree with air flow that that is a
16:58successful run of this workflow but on
17:00the right we have a situation where the
17:02important tasks failed and the cleanup
17:05task was successful now air flow thinks
17:07that this run was successful and to be
17:10honest I don't know if it was or wasn't
17:12that's up to the person who designed it
17:14maybe they only care that the workflow
17:16ran and so yes they would like to see
17:17this as a green light on their dashboard
17:19but I can easily imagine the situation
17:21where someone cares deeply about the
17:23important first task running in which
17:25case they'd like to see this as a
17:27failure but air flow does not give you a
17:30choice
17:31air flow has it hard-coded the task
17:34State excuse me dag State is always
17:36determined by the last tasks in the
17:39Dagon so this is a design failure this
17:42means that our workflow system is not
17:44capable of reflecting business logic and
17:47therefore we cannot fail successfully
17:50the system is corrupt again skip some
17:53texts in the interest of time and we'll
17:54just move very quickly if you're
17:56interested through a couple of other
17:57quick examples people use airflows
18:00global variables in place of parameters
18:02which Justin creates a very staple mess
18:05people abuse the fact that airflow
18:07reruns a dag every single time it runs a
18:09task to introduce tags that change while
18:12they're running which I don't even have
18:14to tell you how horrifying that should
18:15be off-schedule runs air flow used to
18:18not support off schedule runs we
18:20eventually sort of found a way to make
18:22it do it and we found that this this
18:24assumption was so baked into the system
18:26that both the scheduler and the UI broke
18:28as soon as we ran something off schedule
18:30in particular the schedule was working
18:32by adding an interval to the most recent
18:35run so if you had something ticking
18:36along every hour it was totally fine
18:38then all of a sudden if you ran
18:39something on the half-hour the scheduler
18:41would start adding the hour to the
18:42half-hour time and and keep going from
18:44there so again the idea is that these
18:47these are not impossible you can make
18:49the system do whatever you want but if
18:50you don't respect the assumptions of the
18:52system you're going to end up in an
18:53unexpected place and the last one I'll
18:56talk about is that air flows database
18:57has a unique constraint on a dag in its
18:59execution time which means that you
19:01cannot run a dag twice at the same time
19:04it doesn't support the concept of
19:06simultaneous runs of the same dag now if
19:08you live in air flows world that's
19:10totally fine that makes sense if you
19:11live in my world where I want to run
19:12workflows all the time no matter what
19:14for any circumstance that makes no sense
19:16at all so that's a that's a big problem
19:17to me I think it's extremely important
19:22for me to note that these problems did
19:23not start with air flow and they are not
19:25unique to air flow I just happen to know
19:26them very well in an air flow context
19:28from solving users problems I want to
19:31repeat that all work flow managers have
19:33strengths and weaknesses and identifying
19:35these weaknesses doesn't automatically
19:37mean it's bad it's just a benchmark for
19:39you to measure your appropriate use of
19:42the tool against and we can see how this
19:44has played out through history
19:45all right so in the beginning there was
19:47cron and in fact cron I was very
19:49surprised to learn they saw the way back
19:50to 90
19:5187 and I'll bet that every single person
19:53in this room has written a cron cheat
19:55because cron makes the compromise of
19:57course that there are no dependencies so
19:59if you want two things to run in
20:00sequence you schedule one for midnight
20:02you schedule one for 1:00 a.m. and you
20:04hope they work that's the crunchy and
20:06after cron we can move pretty pretty
20:08quickly this is Uzi which is cron in
20:10Java and XML
20:12this is Azkaban which is cron with
20:14dependencies this is Luigi with
20:16dependent which is cron with
20:17dependencies that are based on side
20:18effects SPARC is an interesting one to
20:21mention it doesn't feel like it fits in
20:24this pattern of data engineering
20:25workflow managers but if you have spent
20:27enough time in both the data engineering
20:28world and the data science world you'll
20:29quickly remember that spark and other
20:31tools like it including beam task even
20:33tensor flow these are just ways of
20:35describing computational graphs and
20:37moving data or some state through them
20:40so it just happens that SPARC passes
20:41data instead of States but it's the same
20:44idea and last but not least of course we
20:46have air flow which I would basically
20:47describe this cron with a database so
20:51the question is what have we learned
20:52from this history of work flow
20:54innovation which by the way I think
20:56stopped in 2015 this is a quote from the
21:00creator of Luigi said all workflow
21:02engines kind of suck
21:04including Luigi the question is well why
21:06why do they have to second and one
21:08answer is that they were originally
21:10created to solve a specific problem at a
21:13specific company which means that while
21:15we can today point back and say yeah
21:17they were compromised
21:18they weren't compromised in the moment
21:19they were appropriate to the task at
21:21hand it's only once they were open
21:22sourced and attempted to be widely
21:24applied that these compromises became
21:26evident and in the ahta in the absence
21:28of a modern alternative they've become
21:30weighed down by all of these cheats that
21:32I've been enumerated a common response
21:35from the authors of these platforms
21:37including myself is that users should
21:39just write idempotent tasks does
21:42everyone know what idempotency is it's
21:44it's it's basically the idea that if you
21:46run something multiple times it'll only
21:47cause an action once so this is sort of
21:51like the get-out-of-jail-free card of
21:53data engineering and data engineering
21:54you know workflow authors because of
21:56course if you have an idempotent task
21:58it'll work that you could use cron and
22:01just spam a server every two seconds
22:02with a bunch of item potent fast you
22:03don't need a workflow manager
22:05so it's like yeah this is true but does
22:07it really matter is it practical and the
22:09answers is no it's not so I'd like us
22:11you know you know to look to be clear
22:13idempotency is good but we need to live
22:15in the real world where it's unlikely
22:16and possibly impossible and perhaps
22:19impossible and so instead of pretending
22:21that people can write perfect code let's
22:24acknowledge that they write bad code and
22:26see what we can do to support them in
22:31today's talk I've shown that these
22:33workflow managers they commit these same
22:35basic sins over and over and because
22:37they implement compromise limited
22:39functionality they end up cheating to
22:41implement the features they want how can
22:42we overcome this we need to start with
22:44what users actually want to do and if
22:47you're the users in this room what you
22:49want to do is you want to write Python
22:51so this is what you know the the
22:54simplest ETL script I can imagine looks
22:56like in Python and this is an example of
22:58positive engineering this is what you
23:00want to do how would we design an
23:03idealized workflow system around this
23:05code well if you think about it almost
23:07everything we need is already here right
23:09we know what the what the tasks are we
23:11know what the workflow definition is we
23:13just need a way to communicate this to
23:15of this magic system that we're
23:17inventing and on and the way we need to
23:19do that is by creating entry points for
23:21our negative engineering hooks so here's
23:24what that might look like and if you
23:25don't see the difference I'll highlight
23:27it all we've done is introduce these
23:29decorators which will tell our system
23:32these are the tasks they're gonna be
23:34they're gonna run keep an eye on them
23:36manage their state manage their data
23:37dependencies and then down here we've
23:39created a workflow we're saying hey this
23:41is my computational graph it's any
23:43Python and I want you to pay attention
23:45to what happens this is going to be the
23:46workflow that we run in the future and
23:48that's it I don't have to explain
23:50anything else in theory it's all here so
23:52there's one last thing that we need and
23:55that's to import prefect so this is real
23:58this works today this is actually our
24:00hello world and this is the first time
24:01that we've ever shown it publicly which
24:03is kind of cool for us my team has
24:05worked incredibly hard to build this API
24:07not to mention the system that powers it
24:09and we're very excited that we're going
24:11to be open sourcing it in the very near
24:13future and I promise that it fails more
24:16successfully than anything you've
24:18scene in fact the most important thing
24:23that we did with prefect was that we
24:25allowed it to fail successfully prefect
24:27will track the state and result of every
24:29task in your workflow and let them
24:30depend on each other in a variety of
24:32ways we use these rich state objects the
24:34currency of the system and we allow
24:36users to apply semantics to those
24:38objects don't worry we have sensible
24:40defaults so you don't actually have to
24:41do anything more than what I showed a
24:43moment ago if you don't want to so a
24:45traditional data pipeline implemented in
24:47Prefect would ignore state and just pass
24:49data a more complex pipeline might use
24:51the state to make decisions maybe to
24:54kick off a cleanup job and oh really out
24:55there workflow could actually use could
24:58actually let failed tasks return data
25:01for example their own exception and use
25:04that to drive Diagnostics downstream as
25:06we think these abstractions become
25:07incredibly powerful and they are only
25:09possible because the system is agnostic
25:11to our applicants to its application it
25:13is failing successfully and I want to
25:16mention that under the hood we are using
25:18desc to drive pretty much all the
25:20execution we think that Basques is one
25:22of the best tools out there and was
25:23actually when I really got my hands on
25:25desk as a data scientist that I thought
25:27for the first time that something like
25:29prefect might even be possible and will
25:31have a lot more to say about that very
25:33soon so one last thing since I've been
25:36talking a lot about air flow today I
25:37spent years on air flow it's a very
25:39important tool to me it's a very
25:41important tool to the data engineering
25:42community I do think it has some flaws
25:44prefect does not share any code with air
25:47flow but it will run air flow for you
25:50and so we have some partners who are
25:51using prefect to modernize their air
25:53flow stack without losing any of the
25:55work that they put into building it so I
25:57want to thank you all for coming to my
25:59talk about failure if you'd rather talk
26:01about success we're working with a
26:03number of partners to develop prefect
26:05and if you're in this room today we
26:06would love to work with you you can
26:08reach us at this address which is hello
26:10at prefect IO or come find us outside
26:12we're the guys who have been handing out
26:13free compass coffee all day and since I
26:16think we have a little bit of time left
26:18um do we yeah then I'm happy to take
26:20some questions
26:21if there are any
26:26[Music]
26:28[Applause]
26:31or if there aren't yes one second hey
26:40Arno media capital one so the question I
26:43have is when we look at some of these
26:45workflows obviously there's all the
26:48problems that happen operationally and
26:49then there's all of the things than the
26:50challenges that we face what we're
26:52trying to develop them trying to figure
26:53it out debug it piece it through how can
26:56you compare and contrast what you're
26:57promising so far with Prefect
26:59in terms of as I am building a workflow
27:02what the debugging process looks like
27:05getting into the state being able to
27:06develop locally versus sort of that
27:08deployment stage so this will run on
27:12your local machine if you want to run it
27:15in a synchronous process so you can
27:16actually put breakpoints and inspect
27:18them that's totally fine we do that for
27:19our own testing if you want to run it in
27:21a local desk cluster or even using a
27:23desk synchronous executor that's totally
27:25fine and then you can also deploy this
27:26at you know into a cluster very easily
27:28but these are these are just functions
27:31there they're exposed in this form to
27:33you and they will do whatever it is that
27:34you want one of the one of the core
27:35design principles of prefect is that the
27:38tasks themselves are black boxes we we
27:40don't really care what's in them we're
27:42agnostic to it we actually don't even
27:43receive the code from our partners sort
27:46of an important privacy barrier that
27:47we've created so you can do whatever you
27:49want interact with your systems however
27:51you want we really don't care frankly so
27:55if you need to debug this locally go you
27:56just call it as a function you don't
27:57even have to run the flow it's it's not
27:59tied to anything unless you want it to
28:00be so this is actually we worked very
28:17hard on this so I like to show it off
28:18this is sort of our functional API it's
28:20not actually running the functions it's
28:21just building a computational graph
28:22similar to you know any deep learning
28:24tool that you might use today we have a
28:26fully class-based you know system behind
28:29it just like air flow if you want to
28:30call set upstream if you want to
28:31manipulate tasks and edges directly you
28:33you absolutely can if I had added a line
28:36here that was flowed out run
28:38you would receive back a dictionary with
28:40all of the state objects from all of the
28:42tasks in that run you could manipulate
28:45them and examine them and do whatever
28:46you want with them you can even crack
28:47open the hood and see how we're working
28:49with them so to be clear yeah we have a
28:54functional API we also have what we call
28:56our imperative API for more programmatic
28:58access but this is more fun could you
29:04talk more about integration with airflow
29:07I kind of made it you sounded like you
29:11made a promise if I keep working in air
29:12flow then you promise me something this
29:14is better later on I'm not like hosts in
29:17terms of like six months of development
29:18of like bunch of tags we we I challenged
29:23my team to did I do that I challenged my
29:27team to translate airflow into prefect
29:29which I thought would be rather doable
29:33based on my knowledge of prefect based
29:34on my knowledge of air flow we decided
29:37not to do that because there are some
29:38features of air flow that we don't
29:40support with a direct analog X comms
29:43being one of them we have what we think
29:44is a much stronger data dependency
29:47feature that we're that we're showing
29:48here and we didn't want to write
29:50translation logic that we would then
29:52have to maintain so we're working with
29:56some of our partners some of whom are
29:58actually other airflow PMC's who are
30:01using prefect to instrument their
30:04airflow deploys so you keep your airflow
30:06database running you keep air flow
30:07running we have flow in prefect which
30:12creates a copy a one four one copy of
30:14every task in your dag loads it into
30:16prefix so that you can use all of our
30:18instrument respectin tools you can use
30:20our API you can use our UI to examine it
30:22and run it but at the end of the day
30:24it's calling back to airflow to run your
30:26existing system so we've abandoned the
30:28idea of translating it because we just
30:30refuse to support some features
30:34maybe this is a weird thing to say but
30:36so I use a Z in my career before and I'm
30:38using air flow right now and you know
30:41I've never been a JavaScript programmer
30:43but you see the like the rise and death
30:44of like JavaScript frameworks and
30:47visualizations of that one of the
30:50reasons well we chose the Oracle flow
30:53because of the utility to us but we also
30:55were excited as someone on Google cloud
30:57platform that they kind of made it a
30:59first-class citizen with cloud composer
31:00do you see this world is like the coming
31:04and dying of the same way we have
31:05JavaScript frameworks like what's hot
31:07right now in terms of workflow
31:08orchestration and it's like next year
31:10like I'm gonna be applying for air flow
31:11jobs if you're like I don't care about
31:13that at all I in that time is it just
31:17like are there still like as you were
31:19saying people make these things and they
31:21open source and then everybody makes a
31:22ton of demands about what they should do
31:24and then they're like it wasn't designed
31:25to do that and they're like okay now on
31:27the next one I think that's the first
31:30time anyone's ever described the
31:31workflow manager space is hot to me I
31:33had I had to beg people to take me
31:36seriously when I said that this was even
31:37possible I mean we've had conversations
31:40with people who deny that it's possible
31:41until we show them how we how we did it
31:43air flow is close to four years old now
31:46and nothing has unseated it and so while
31:49it has gained momentum and to be clear I
31:50am an air flow PMC I am extremely
31:52excited to see that I think I have the
31:54third most commits of anyone in the
31:55world on air flow it's an amazing thing
31:58but it's not a hot space it's not a
32:01crowded space I'm frankly amazed that we
32:05see this emergence of an air flow
32:09developer like I'm not sure that's a
32:11good thing to be honest with you sort of
32:12my argument here is that the workflow
32:13system should fade right you shouldn't
32:16be an air flow developer you should be a
32:17developer who happens to use air flow
32:24any other questions yes we have we've
32:41designed our compute to be modular at
32:43the moment we preferred ask and we used
32:45ask running in kubernetes like how do
33:00you sorry I mean how do you had let's
33:04suppose suppose you have one step that's
33:06got a lot of different software than the
33:08next step like how would you address
33:10that a lot of different dependencies yes
33:12so software dependencies yeah no it's a
33:14good question today the the workflow is
33:16built as a container with a single set
33:17of dependencies you put them all in and
33:19we're working towards supporting an
33:21individual container per task yeah yeah
33:23okay okay is that a hint are we supposed
33:30to get out okay I'm happy to take
33:33questions as long as there are yes
33:42there you go thank you my apologies if
33:45you already answer this but I know many
33:47people have found frustrations with
33:51distributing airflow drops using celery
33:55just with the complications of a celery
33:58and things what's the story for prefect
34:01surround running a cluster where jobs
34:04are run across it sure so that's what we
34:08used ask for tasks is hands-down the
34:11best way to do that can you describe
34:13that a little bit further assuming
34:15people know I know about tasks and
34:18assuming others do can you describe how
34:20it yeah we're just best voice one of the
34:22easiest ways I've found to describe desk
34:24is actually a sort of a pure Python and
34:26friendly version of SPARC so desk is I
34:28think just an extremely easy way to
34:30distribute computation across any number
34:33of nodes whether it's your local laptop
34:35or something running in the cloud
34:36it'll take a function it'll serialize it
34:38it'll run it in the cloud they'll bring
34:40back a result or what's really cool
34:42about it and what we'd like to take
34:43advantage of is you can dynamically
34:45build yet further graphs from an already
34:47running graph so I mean in a nutshell is
34:49just a very easy way to run distributed
34:52computations frankly it's the best I've
34:55seen so we use it for our primary
34:59execution engine it's modular if
35:01somebody wants to write a an extension
35:03for something else they absolutely could
35:04you want to go back to celery by all
35:05means go right ahead but we're promoting
35:07desk as our as our primary executor yes
35:20do you expose the that you're running on
35:22tasks at all to the users thank do they
35:24get to see the task dashboard is there
35:26any exposure of that at all it's a great
35:28question so at the moment no but not
35:31because we've decided not to we just
35:33sort of haven't addressed that question
35:36so it's a good one we don't hide that
35:39we're running us in fact we would invite
35:41users who want to take advantage of the
35:42fact that we're running in DES to just
35:43open up a worker client and we're trying
35:45to figure out the right way to create a
35:47sort of first-class API where you could
35:49still run that locally effectively and
35:51do and do your debugging but also have
35:53it work properly in a distributed
35:54environment so we we don't hide it from
35:56users we also at the moment aren't
35:58explicitly promoting it to users it's
36:00just a it's an important act this
36:10all right well thank you all so much I
36:12appreciate your attention in your time