Full transcript
Trailer & Intro
0:00ML projects tend to fail in this common failure modes.
0:04One is what I call POC hell.
0:07Another failure mode is that it's easy to get maybe the first thing out of the door.
0:13But then it's kind of stuck together with duct tape and sticky gum.
0:18And it's hard to test or change.
0:21In my experience of working with data intensive ML initiatives or other
0:27technology specializations, I'd found that the ability to build the thing
0:32wasn't always the major success factor.
0:35Often it was understanding what the right thing to build was, making sure that
0:39there was alignment between technical teams and business teams and product
0:43teams on the right thing to build.
0:44You can't MLOps your problems away, just like how you can't
0:47DevOps your problems away.
0:49MLOps helps in a few ways, but MLOps is not going to write your tests for you.
0:54They are not going to talk to users and make sure you're
0:56writing the same right features or implementing the right features.
0:59It's useful to think about the essential characteristics of ML solutions.
1:05They'll be able to do things at speed and scale that no human could.
1:08But they'll make mistakes.
1:10We wouldn't consider an ML solution if we knew the right answer every time.
1:13We'd write some rules instead.
1:15One of the quotes, opening quotes in our book was from Edwards Deming, saying a bad
1:19system-would beat good person every time.
1:22The whole thesis of our book is how do we create those systems to help teams build
1:27the right thing, build the thing right, and in a way that's right for people.
1:45Hello, everyone.
1:45Welcome back to another new episode of the Tech Lead Journal podcast.
1:48Today, I have with me two Davids, very excited.
1:51One is David Tan.
1:52He is actually my ex-colleague in ThoughtWorks.
1:55And the other one is David Colls.
1:57So they are the co-authors, together with another one, Ada,
2:00who couldn't join us today.
2:02They wrote a book titled Effective Machine Learning Teams.
2:07Even though it's a machine learning as part of the book, but I think
2:09we are not going to cover solely just machine learning, because
2:12I'm also not an ML expert.
2:14But we'll discuss things like how we can build effective machine learning
2:17teams, what kind of practices that we should be aware of, and things like that.
2:22So welcome to the show.
2:23I'll maybe mention Dave as David Colls, right?
2:26And David for David Tan.
2:28So welcome to the show, guys.
2:30Yeah, thanks for having us Henry, happy to be here.
2:32Thanks, Henry.
2:33Nice to be here.
Career Turning Points
2:35Right.
2:35I always love to ask my guests maybe first to share a little bit of themselves by
2:39sharing your career turning points that you think we all can learn from that.
2:43So any one of you can start first.
2:45Cool.
2:46Um, so I'm currently engineering manager in the AI products group
2:50in Xero and my career into tech hasn't always been straight.
2:56Like I was actually non-technical when I graduated from university.
3:00So I was working in government for maybe two, three years.
3:04And I got, by accident, hand, foot and mouth disease and I had to
3:08like be quarantined for eight days.
3:10And from there I started like, you know, watching some YouTube videos
3:12about data analytics, about coding, and then tried some tutorials
3:18and, you know, it was really cool.
3:19So then decided to, you know, join a programming bootcamp, quit my job, did
3:24that, and then joined ThoughtWorks.
3:26And yeah, learned a lot about Agile, about test driven development, refactoring,
3:31CI/CD, got into ML engineering and so, you know, finally got into this
3:37exciting space of ML engineering, DevOps, building ML products.
3:41So yeah, I have to thank this 8 day quarantine back in the day
3:45for my career turning point.
3:47Yeah, maybe.
3:48How about you, Dave?
3:49If I'm to pick on an external theme like that, I'd probably choose deciding
3:54to play Ultimate Frisbee early in my career, where I met a number of
3:59people involved in the software scene.
4:01But, in particular, a ThoughtWorker who encouraged me
4:04to interview at ThoughtWorks.
4:06And when I think about turning points, I guess that kind of leads me into
4:10the fact that I started as a developer in technical, data intensive roles.
4:15I then, you know, looked for ways to make more impact in the leadership space.
4:20And when I joined ThoughtWorks, although I'd actually passed a coding
4:23interview, I joined as a project manager.
4:25Then I've worked for a number of years in Agile transformations and
4:29organizational transformation before then picking up the data practice.
4:33So I guess, my career has turned from deeply technical to leadership
4:40positions and organizational design a number of times over that period.
4:45Thank you so much for sharing your story.
4:46Very interesting, right?
4:47So the kind of serendipity that could happen throughout our career, right?
4:51And what led us to where we are at the moment.
Writing “Effective Machine Learning Teams”
4:54So today we are going to discuss topics from your book,
4:57Effective Machine Learning Teams.
4:58Maybe in the beginning, let's share what made you wrote this book.
5:03Yeah.
5:04This book started like as a blog post, actually.
5:06So back in 2019, we had wrapped up a project.
5:11We're involved a lot of refactoring of data science codebases.
5:15So we have data scientists wrote code in the Jupyter Notebook,
5:20no test, a lot of lines.
5:21You know, it was good, like it was solving a real business problem.
5:25And then we had to kind of productionize that.
5:27And so we started writing about how we did that.
5:30How do we encapsulate code into functions?
5:34How do we have automated tests?
5:35So we are, you know, clear about that, about the quality.
5:38How do we do CI/CD for those changes?
5:41And then we wrote a blog post and it kind of like exploded in, back then Twitter,
5:46someone from the R community shared it.
5:48And a lot of likes and it's like, okay, I think we're onto something here.
5:52There's this desire, I think, in the data science community or ML
5:55community to have more kind of engineering and robust practices.
6:00So yeah, you know, then we started writing more and more about...
6:04from the subsequent projects or work we did to kind of large
6:09scale ML training pipelines, large scale high volume ML products.
6:15And, you know, the practices that helped us build that reliably, build
6:19ML products that are safe, help you get fast feedback on your changes.
6:23And some of those things worked really well.
6:25So we wanted to share it more broadly with our readers.
6:29Uh, and that's just the engineering space.
6:31And I think Dave also has some other kind of, yeah, aspects that, you
6:35know, we started with this thread, calling engineering practices for ML.
6:39And then we discovered this whole kind of world of other
6:42practices that also help ML teams.
6:44Yeah.
6:45So, you know, in my experience of working with, I guess, data intensive
6:50ML initiatives or other technology specializations, I'd found that the
6:56ability to build the thing wasn't always the major success factor.
7:01Often it was understanding what the right thing to build was, making
7:04sure that there was alignment between technical teams and business teams
7:08and product teams on the right thing to build, which could be complex when
7:12people don't speak the same language.
7:14So, you know, what sort of processes and techniques can you use to build
7:18that alignment, especially when there's a lot of uncertainty both
7:22about the problem and the solution.
7:24And then also, because of the specialised nature of this work, it's often hard to
7:29build a team that can execute end to end.
7:32So then the factors around either building in a single team, how do you bring all of
7:39those different perspectives together, or when you have to rely on multiple teams to
7:43get a piece of work out the door, what's the best way to achieve fast flow and high
7:48quality in a way that's sustainable for the people working in that environment.
7:52So for me, those were the perspectives I wanted to bring to machine learning
7:57and product development, which is an essentially multidisciplinary activity.
8:02You know, how do we do that better as teams?
8:04Right.
8:05When I read your book, I must admit that I was thinking that I was going
8:09to learn a lot of the technical things about specific ML, you know, projects
8:13and ML products and things like that.
8:15But actually, you cover kind of like the holistic approach.
8:17not just the engineering, also like product, delivery, and things like that.
8:21And actually there are so many things that any software engineering
8:24team, I believe, could learn just by reading the book as well.
8:27I mean, putting aside some parts of the ML.
ML Engineering vs Other Types of Engineering
8:29But let's start maybe in the first place, because I'm sure many listeners
8:33here not coming from the ML background.
8:36What are the differences between, you know, ML engineering and the
8:40typical software engineering these days where people build websites,
8:43APIs, and things like that?
8:45Yeah, that's a great question.
8:47And I think ML systems fundamentally is orders of magnitude, I think,
8:52harder and more complex than software engineering in some ways.
8:57Like, of course, software engineering is hard, you know, you've got microservices,
9:01you've got distributed computing, you've got a lot of things, like in
9:05itself it's an art and whole discipline.
9:08ML itself, the difficulties manifest themselves in some common ways,
9:15like, you know, it's not just one service or one component.
9:19Usually, it's like a data pipeline, dependencies upstream, that, you
9:24know, it's complex in itself.
9:25You've got your scale or compute requirements for your own ML workload,
9:30which is also, can be done in many ways.
9:34You know, you can have a really large instance with really large
9:37memory, do everything there.
9:38Or, you know, at some point that breaks down, so how do you architect or make
9:43it simple to scale up large workloads?
9:46And then there's also the monitoring, like when you think about software
9:49products, typically your monitoring is kind of your four golden signals, right?
9:54Success rate, latency, things like that.
9:57You can do that for ML, and you can be all green and good, but that says
10:02nothing about like the quality of the predictions that models are producing.
10:07And then the correctness of those.
10:09So it adds another layer of challenges that I think the MLOps community
10:13have come to solve in recent years.
10:16And then further, even further down, like, yes, we've built this product, it's
10:19giving accurate predictions, to an extent.
10:23What is the kind of business impact or what levers or dials is it moving?
10:29So that's like kind of essentially the fundamental difference between the ML
10:35product and software engineering products.
10:36There's a small...
10:38There's this diagram from Google where like ML is just like a little box.
10:41And there's like infra, there's monitoring, there's
10:44data quality, data skew.
10:47So a lot of moving parts and that's why we wanted to bring a lot of practices
10:51that kind of help you tame the complexity in this different sub-components.
10:56And, yeah, when you look at how to manage the work as well, then, I guess, one of
11:00the characteristics of building an ML feature is there's a lot less certainty
11:05about how it's going to perform up front.
11:08And so often you have to integrate this exploratory data analysis or
11:13building prototypes or architecture search as it might be described into
11:18the product delivery process as well.
11:21And then, as you converge on a potential solution, then you're also dealing
11:26with another vector of change, which is the data and the world changing as
11:30well as David said around monitoring.
11:32Often that changes slowly, but sometimes we can see step changes.
11:36COVID was a classic example, with recommender systems moving from lowest
11:42price to highest, to best availability.
11:45And you know, ML models needing to catch up with that.
11:48So yeah, you both have less certainty and you have another vector of change to deal
11:51with when it comes to managing that work.
11:55Yeah, so thanks for sharing some of these complexities, right?
11:58So definitely for people who may not be able to relate.
12:01I think typically from what I can see, right, there are so many
12:04data pipelines when you build ML.
12:05You probably need more than just a small data, right?
12:08Something like a medium data or big data and the kind of compute that you need.
12:12Sometimes it's more like a distributed computing with
12:15large scale kind of pipelines.
12:17And I think typically what I can see as well, the feedback loop might be
12:20long in ML projects, simply because you have this training kind of thing.
ML and LLM
12:24In the first place, maybe let's clarify when you say ML, I think we are not
12:28just associating it with LLM, right?
12:29These days people are crazy about LLM and maybe they associate ML now as LLM.
12:34So maybe a little bit of insight here, like what do you mean by ML projects?
12:38Yeah, I guess, so yeah, great observation, Henry, and I was thinking, you know,
12:42as you talked about feedback cycles and big data and all of those elements.
12:46Yes, those are the things we traditionally associate with supervised
12:50machine learning, which I guess is our sort of primary focus in the book.
12:55Which might be anything from a binary classification problem, like is this
12:59transaction a fraudulent transaction?
13:01And you know, to solve that with a supervised paradigm, we get a lot of
13:04historical data about transactions, their amount, the age of the account.
13:10You know, maybe some information about the network around that account and
13:13then we label those transactions, whether they were fraudulent or not.
13:17That's actually a nice example of where the labels might be easy
13:20to get because we might get them from customer reports of fraud.
13:23But in general, that effort of labeling the data set can
13:25be really intensive as well.
13:27So that's the main paradigm we're looking at, but we think that, you
13:30know, looking through this lens of working with uncertainty, working with
13:34data, working with specializations.
13:37But there's a whole lot of other paradigms of ML or AI that might fit into that.
13:42So obviously working with a large language model in inference mode,
13:47that's still a challenging environment.
13:49The supervised paradigm might apply to fine tuning or even training from scratch
13:54your own large learning model as well.
13:57But then beyond that sort of paradigm, we're also looking at
14:01initiatives that might be, might use reinforcement learning, for instance.
14:05It's another approach to recommender systems.
14:08And this is where you might not need big data to start with.
14:11You might not have historical data to start with, but you might actually
14:14build or tune your data set as you go.
14:17We might also look further afield to techniques like simulation or operations
14:22research for optimisation, they share a lot of the same characteristics.
14:25Simulation is actually very similar to machine learning when
14:28it comes to productionising it.
14:30It's just that instead of a learned model of the world, you're working
14:33with an explicit model of the world.
14:35But you still need to go through feature engineering, ability to monitor
14:38and assess business impact and align to, um, stakeholder expectations too.
14:43So, while, yeah, while we're focused on supervised machine learning as a paradigm,
14:48yes, there's a big world out there.
14:50Anything that's decision making driven off data or models of the
14:53world or requires specialization might fit into this view that we've
14:57put forward in the book as well.
Why Many ML Projects Fail
14:59Thanks for the clarification.
15:01So maybe let's start going deeper into the book, right?
15:03I think one thing that piqued my interest when I read the book in the
15:06first few chapters of the book, right, I think you cover some statistics here.
15:10The first one is actually many ML projects doesn't make it to production, right?
15:15So and the second one, even though they reach production, they actually
15:19don't solve the real business problem or they don't bring any value, right?
15:22So maybe let's share a little bit why this is the case.
15:25What do you see out there as ML practitioners, right?
15:27So why is it so hard to actually bring ML projects to production?
15:32Yeah.
15:33Thanks, Henry.
15:34Like that's the crux of the book, in that ML projects tend to fail in this
15:40common like failure modes that we describe and I can go through them.
15:44And it's kind of preventable types of failure, like if we can learn from
15:51experience from history and that we would bring in the right tactics to
15:55make sure, like, as you mentioned, let's say a project that, you know,
15:59let's say we invest half a year, nine months into it and ship it.
16:03And we found that, oh, actually it's not solving the right problem.
16:06Or like users didn't really care about it.
16:08It's like, then what tactics can we bring in there with prototype
16:11testing or customer research?
16:14So that's kind of the whole kind of the main thesis of the book is like,
16:20how have we failed and what have we learned and how can we do better?
16:25So one common failure mode in ML projects is what I call POC hell.
16:31It's like a combination of factors, maybe business don't have enough of the
16:36appetite to release something to users.
16:39So we might, you know, take the first easy step to build a prototype, to
16:44test it, maybe internal demo of it.
16:47But the resources or risk to productionize that and put it in the
16:53hands of users is kind of too great.
16:55So with the lack of commitment there means that teams on the ground just kind
17:00of build POCs after POCs after POCs.
17:03So that's detrimental to, you know, the business in some ways, like
17:07you're missing out on opportunities.
17:08So detrimental to the morale of the team.
17:11You do a few and then, you know, after a while you say, what's the point?
17:14So, you move on and then you've got churn and you've got onboarding
17:17and you're losing talent, right?
17:20Another common, I think, failure mode is that it's easy to get maybe
17:25the first thing out of the door.
17:27You hustle, you get the data, you get the compute, train the
17:31model, you deploy the thing.
17:33But then it's kind of stuck together with duct tape and sticky gum.
17:39And it's hard to test or change after you evolve based on, you know, new
17:45customer feedback, things like-that.
17:47So then that comes to kind of continuous delivery, CI/CD, how confident can
17:51we be to test and deploy a change.
17:54And so this set of problems is solved by MLOps, by continuous
17:58delivery, from machine learning.
18:00There's probably a couple more, uh, I don't want to hog the mic.
18:04Dave, was there any other failure modes that came to mind or these challenges?
18:07Yeah, those are two big ones.
18:09It's, you know, is there any value in doing this at all?
18:13And that can be really challenging when the responsibility is divided across
18:17different teams and might even be in different parts of the organization.
18:21There might be a, I guess, the release of ChatGPT sparked a lot of curiosity
18:28in how do we best use LLMs in products.
18:31So, but we might have responsibility diffused across the organisation.
18:36We might have an ML team that's interested in looking at the
18:39use of large language models.
18:41We might have a product team that has their own roadmap, where they don't have
18:45these features anywhere on that roadmap.
18:48And then we might rely on data, especially in the case of Gen AI,
18:52that's unstructured data that's not well governed, that's hard to make available
18:57in a responsible way to these systems.
19:00And so while each group can do what's within their control to move a little
19:05bit towards a future state, it's not aligned or stitched up, you know,
19:09in a cohesive way that allows us to establish whether there is some value
19:14quickly with cheap and low risk tests.
19:16And then, as David said, once you've established there's some value, you
19:20know, if you have productionized something that's not supported by a lot
19:25of these good practices around MLOps and continuous delivery, which is often
19:29the case to understand if there is value there, then you start to understand how
19:34much value you're leaving on the table.
19:36And then this can be an opportunity to invest in those practices
19:40to be able to maximise the lifetime value of an ML product.
19:44But then that requires a certain, again, a certain degree of maturity and a
19:48certain ability to work across existing organisational boundaries or reshape them.
ML Success Modes
19:54Yeah.
19:55And there's also kind of the flip side of the failure mode is the success, right?
20:00So we've also worked with and seen teams, ML teams, successfully deliver,
20:04you know, really great ML products.
20:05And what tactics did they use to succeed?
20:08Like we could see things like customer testing, user testing, so before they
20:14build out a prototype or a MVP, which is a really expensive way to test
20:18something, you know, even just talking to users, understanding what is the pain
20:23point, what are their jobs to be done.
20:25Where could ML help?
20:26And then, you know, once they're bought in, they say, okay, this is the right ML
20:30product that will solve those problems.
20:33Then combating that failure mode of what Dave mentioned, like multiple teams
20:38kind of throwing across each other.
20:39If we remember the DevOps comic, where you have devs on one side, you have ops.
20:44So instead of throwing code between scientists and, you know, MLOps
20:50engineers, you know, we do the inverse Conway maneuver where we
20:54have a cross functional team shipping end to end for this initiative.
20:58And when we talked to the team members who worked on that project, they enjoyed that
21:03end to end ownership, the satisfaction of putting something, releasing to-customers.
21:08You get to touch different parts of the stack, learn about data
21:11science, you learn about Ops.
21:13You know, so they help them, you know, you can get fast feedback, you have
21:16the same standup, you ship things iteratively, you showcase it every,
21:20you know, two sprints, three sprints.
21:22So that feeling of flow and speed was like really amazing.
21:26And then, you know, when something is released, you know, you've
21:28got your continuous delivery, production monitoring, when
21:32things are not going well.
21:34Yeah, and you'll have safety for changes.
21:36If I need to make a refactoring, upgrade the library.
21:39If your CI checks are all green, your model monitoring dashboard is saying
21:44performance is good, then, you know, you merge that PR and you don't feel
21:48nervous or stressful about any releases.
21:51So yeah, there are also bright spots in how teams deliver ML products.
21:57Right.
21:57So I think when I heard you mentioned, you know, in the beginning about
22:00POC hell, you know, so many duct tapes when you bring production.
22:03I could relate to some other ML projects that I knew of before.
22:07So I think many, many, ML teams probably are in this mode, right?
22:10They keep building POC because maybe the expectation from stakeholders
22:14or from the business also kind of like with a lot of hype, right?
22:18Because these days people think AI can do a lot of magic, right?
22:21And especially with all the success of ChatGPT and all that, they
22:24even think that it is easy to do.
22:27But I think the reality may not be the case, right?
22:29And the other thing is about the practices, right?
22:32So many ML teams, first, either they don't have many good software engineers,
22:37so they are like data scientists or, you know, ML engineers who just have a
22:41lot of expertise in building the models.
22:43But actually they don't have all these other skills like refactoring, continuous
22:47delivery that you mentioned, right?
22:48Testing as well.
22:49Or the other flip side, which is a lot of software engineers
22:52being turned into ML engineers.
22:55So they are not necessarily an ML engineer, but they just come from,
22:58you know, like web application development and turned into ML engineer.
23:02So I think some of these things definitely I can relate.
23:04Yep, and just to add to that, there's also the flip side of engineers or ML engineers
23:09being excellent ML engineers, but not treating it as a data science problem
23:13where there is uncertainties, the need for experimentation to get fast feedback.
23:20So yeah, I think the, kind of emphasizes the need, I think, for
23:24cross functional collaboration from both, like data science, from MLOps,
23:29from, you know, SRE, things like that.
Ideal ML Engineering Team Composition
23:32So maybe let's go there, right?
23:33So the composition of the team, you mentioned a few times about cross
23:36functional teams and, you know, these silos between, for example, data
23:40science, maybe the ML Ops or whoever operates the system in the end.
23:45Maybe there are also other functions, like for example other teams, because
23:47ML typically also relies a lot of dependencies like data or maybe some
23:52kind of model, things like that.
23:53So maybe what are the best compositions here in your experience?
23:57Maybe you can advise us.
23:59I guess here that there's another question that is, and it's what we
24:03tackle towards the end of the book.
24:06We look at sort of multiple levels of team effectiveness.
24:10And one of the levels we look at is an individual in a team, the types of
24:14practices you can use, the expertise you can build as an individual.
24:18The next level we look at is within teams, again, how those practices around
24:22technology, delivery, and product.
24:24Augment teams, but also broader team dynamics.
24:28But then the third level we look at is between teams.
24:32So how can teams be effective between teams?
24:34And so the best makeup of a team depends a bit on the shape of the team
24:39and its interactions with other teams.
24:41And so this is where we use the team topologies model to identify the different
24:46types of ML teams that you might sit in.
24:49And so to run through it quickly and we can come back and dive into details.
24:53You know, you might have a stream aligned team, and this, we'd say,
24:57this is the basic unit to think of, first ML project, a stream aligned
25:01team that can deliver end to end.
25:02You're not going to be conducting groundbreaking ML research,
25:06you're going to be using tried and true techniques, in this case.
25:10Or at the other end where you have an established ML ecosystem internally,
25:14you can add another stream aligned team that draws on those existing
25:17services internally quite easily.
25:21Then as you scale beyond one team, there might be two routes that you take.
25:26And so you might go down the route of where you have a sort of low
25:30level, common concerns across different initiatives, that might be the opportunity
25:35for a technology platform to support ML at a lower level like compute
25:40and data and feature engineering.
25:43And so that would be a platform team in the team topologies team shape.
25:47And that would aim to provide as a service, and that would be composed
25:52differently from a stream aligned team.
25:54The other route you might take to scale is you might identify
25:56business clusters of ML needs.
25:58So there might be something around audience engagement or there might
26:02be something around asset valuation.
26:04Or there might be something around content moderation.
26:06And so those are a sort of specific set of business needs that also come with an
26:11associated set of ML paradigms as well, that don't necessarily have a common
26:16support in a technology platform at that right business level of abstraction.
26:20So then you might have what's called a complicated subsystem team.
26:24And again, the makeup of that might be a little bit different.
26:27And then as you, you know, as you have a bigger ecosystem, then there'll be
26:31a range of problems that come up of a similar nature, but like have a unique
26:36presentation each time, which might be around, say, privacy, or ethical use of
26:41data, or optimizing particular techniques.
26:45And this is where you might have an enabling ML team as well that sort
26:49of acts as a consultancy to other ML teams to make themselves redundant.
26:53And so that was a long way of saying, it depends on the ideal makeup of a team.
26:59But you might assume it's a stream aligned team and talk about the roles that sit
27:04in that and maybe I'll throw to David.
27:06I think team topologies is a really useful set of constructs,
27:10the ones that Dave went through.
27:12Because it helps teams scale through the team's API.
27:17So as Dave mentioned, you know, if we treat everything, let's say,
27:21as a stream aligned team, like in software engineering, we're very
27:24familiar with cross functional teams.
27:26Um, but then the failure mode there is that teams, like Conway's
27:30Law, right, we've got three teams, we're going to build three sets of,
27:33you know, basically architecture, rebuild certain tools that we need.
27:37Then that's where like maybe platform team comes in to abstract all of that.
27:41So then with that, what we've seen work really well in one particular
27:46case was this complicated subsystem team, like as you mentioned, ML
27:50is hard, a lot of moving parts.
27:52They took on the effort to build this ML product.
27:55Let's imagine it's, um, I can't say the specific example, but let's
27:59say it's a car valuations, right?
28:01That's kind of an ML product without a product.
28:04API to this team or this product is, I will give you the value of cars and now
28:11that is self serviceable by other teams.
28:13They could embed it on the mobile, mobile team could integrate with this
28:17API and expose the ML capability.
28:21A web team can do the same, or email marketing team can do the same, and
28:25then send personalized emails to say, you know, about car valuations.
28:29So then that was how that team scaled, rather than having to integrate with
28:33each particular thing, encapsulating that complicated subsystem as a kind of formal
28:38set of APIs, either through batch or through real time, that helped the team
28:42like achieve more and get more knowledge out of the ML product they built.
28:46Yes, yes, that's a great example that complicated subsystem team is probably
28:50going to have a bunch of specialists.
28:52It's going to look maybe most like what people imagine an ML team looks like.
28:57But as David said, to scale their impact in the organization, they do
29:01need that business domain expertise.
29:03They do need to deliver as a service instead of constantly
29:07collaborating with other teams.
29:09And you know, maybe some product thinking helps with
29:11that service definition as well.
29:13So, you know, that might be one team composition.
29:16Yeah, that's right.
29:18And if you replace that example there from car valuations, which is bit like,
29:22think not everybody can relate with that.
29:24Let's say it's like travel recommendations or product recommendations, then that one
29:29team that spent all the effort in building product recommendations now can impact
29:33multiple parts of the business by having the right team topology to, you know,
29:38have that fracture plane around, okay, my team is doing product recommendations.
29:42And yeah, you know, applying the kind of data product disciplines, exposing this
29:47as an API and encapsulating the details.
29:50Very exciting to hear about team topologies mentioned for
29:53ML product and ML teams, right?
29:55So I think it's kind of like back then, right, it was a revolutionary
29:58approach to how we kind of like create different teams.
30:01And I, I'm glad that you brought it up because still, I believe in many
30:04companies, they think, okay, we want to build ML product, ML project.
30:09They just hired a few data scientists or, you know, these ML experts, and they
30:13just asked them to kind of like build the model first without actually involving
30:17the other aspects of software engineers.
30:20It could be the stream aligned team from the product side.
30:22Or it could be, you know, the data, or could be anything, right?
30:25But I think that tends to kind of like have its challenges.
30:28So I think bringing the concepts such as team topologies to actually
30:31think holistically how we are going deliver the ML projects is something
30:35that's really, really important.
30:37Yeah, so a really great point, and I think one of the quotes, opening quotes in our
30:41book was from Edwards Deming, saying a bad system-would beat good person every time.
30:47So you can hire the smartest data scientists, put them in an environment
30:51where it's not, the right system, then, you know, we get what you described there.
30:57So the whole thesis of our book is how do we create those systems to
31:01help teams build the right thing.
31:04You know, solve the right problem, build the thing right.
31:07You know, engineering, ML engineering, data science.
31:10And then in a way that's right for people.
31:12I think what Dave puts in a really good way.
31:14Like it's not just shipping and shipping, but in a way that has the
31:19right team shape, right collaboration mode, right trust, psychological
31:22safety, and also the right processes to deliver, ship early and often.
31:27So yeah, I really was intrigued when you mentioned, you know, you can't
31:30just put a group of data scientists together or engineers together.
31:33You really got to create that system where they can, you know, ship to the
31:37right, you solve the right problems.
Building the Right ML Product
31:39Yeah, thanks for adding that.
31:41To come back to the theme of, you know, this product discipline, right,
31:44so I think we know that a lot of ML projects don't make into production,
31:48right, or they solve the wrong problem, or maybe don't bring value, right?
31:52So I think they are, like Dave mentioned in the beginning, there
31:54are a lot of uncertainties when you actually build ML model, ML product,
31:59right, because, I mean, the way it works also is kind of like prediction.
32:02It's kind of like there's some kind of ambiguities inside.
32:05LLM, there's a hallucination, right?
32:07So how can you actually come up with an approach such that, you know, when you
32:11first building the ML product, you can actually build something that is kind of
32:15like bringing the business value either to the users or to the organization?
32:20So I think this is probably one hard aspect as well.
32:22So typically how would you run this?
32:25Yeah, it is hard and it goes beyond the technical, although
32:29the technical informs it.
32:30I find it's useful to think about the essential characteristics of ML solutions.
32:36We're exploring them because we think they're going to be
32:40superhuman in some aspects.
32:43They're probably not going to beat the best experts in a field.
32:46That's maybe one myth that we should tackle straight up.
32:49But, you know, they'll be able to do things at speed
32:51and scale that no human could.
32:54But then we need to also consider that they'll make mistakes by the very nature.
32:59Again, we wouldn't consider an ML solution if we knew the right answer every time.
33:02We'd write some rules instead.
33:04So there'll be some percentage of mistakes, however small, in
33:07any ML solution, by design, as well as the unforeseen mistakes.
33:12And then when it comes to those mistakes by design, we really need to understand
33:16the cost sensitivity of what's, you know, what value do we get out of a bunch of
33:21right answers and what is the impact of maybe a very small number of wrong
33:25answers, but they could have a very huge impact across all sorts of dimensions, the
33:29financial, as well as security bias and fair treatment of all our stakeholders.
33:37So, yeah, we need to consider that fallible nature of them as well.
33:41And so, you know, we need to start with products that are
33:44designed to handle that failure.
33:45You know, they have some upside from when ML gets it right.
33:48But they're robust to the times when ML gets it wrong, because it will.
33:53So starting from that perspective, you know, we can then identify
33:56some experiments about how well does this need to work?
33:59We can start with very simple baselines.
34:01Sometimes, you know, even just predicting the majority class or
34:05random guessing and, you know, seeing how that works as a product.
34:09But, you know, ideally, to resolve this uncertainty, we're getting into
34:13some real data and understanding the predictive potential of the real data.
34:17And so this is, again, where it's like it's a challenging
34:20multidisciplinary exercise.
34:22We were trying to proceed on multiple fronts.
34:23Are we building the right thing?
34:25Can our solution support or a proposed solution support the
34:29performance that we expect?
34:31And so on.
34:33Yeah.
34:33And just to add to that, I think that emphasizes the importance of that cross
34:37functional nature of the work as well.
34:40Like if we frame this as a data science or ML problem, then we can try our level best
34:45to go from, let's say, 55% accuracy to 99.
34:50You'll never get to a hundred, and we will spend many, many weeks
34:55and months trying to get there.
34:57So it's not just the ML problem where we try to improve the model's
35:01accuracy or recall precision, but also how can we design for these
35:07failure modes of the ML model.
35:09So like for example, you know, displaying, designing a product in a way, right,
35:14to show users that, okay, this is not a confident prediction, or this is a
35:19prediction, but would you correct that?
35:22And also maybe even giving users options, like these are the top three, like,
35:26the model's top one prediction maybe, you know, not where we want it to be,
35:31but top three, okay, is much higher.
35:33So designing the product in a way that mitigates these
35:36failure modes of the ML model.
35:39And yeah, it's easy for teams if they don't have the right capabilities or
35:44right skillsets, to try to solve it.
35:47Like if you are a hammer, everything is a nail.
35:50And it's very costly to try to level up the accuracy.
35:53We may never get there, but yeah having that cross functional approach, like
35:57how do we design it in a-different way?
35:58How do we talk to users?
35:59Like do users find this okay?
36:01That's a kind of more holistic way to solve-the-problem.
36:04Yeah, that hammer and nail is really important as you start to shift to,
36:09okay, yeah, how do we make this viable?
36:12Or how do we, yeah, make it viable from the perspective of solving the
36:15problem effectively, but also from being economically sustainable to maintain it?
36:20And I guess there's another essential characteristic around that.
36:24The fact that ML solutions are narrow, the very training process is
36:28to optimize a loss function, which is a narrow definition of success.
36:32But they're composable.
36:33And, you know, this is, I think, where Gen AI can be really interesting.
36:37And I'm not the only one to take this perspective, but I've described like
36:41Gen AI as a stone soup for innovation.
36:44So the story of stone soup, if you haven't heard it, is that a weary traveler arrives
36:49at a village late at night, and all they have in their knapsack is a stone.
36:53So they go to the first house in the village and they ask the villager there,
36:57could I have an onion to make stone soup?
37:00You know, I've got the stone.
37:01All I need from you is an onion.
37:02That villager says, oh, great, yeah.
37:04I'll provide an onion.
37:05They go to the next house and repeat and ask for carrots and, you
37:08know, proceed around the village.
37:10By the end of that process, they cook up an amazing soup
37:13that feeds the whole village.
37:14And it was all cooked from a stone.
37:16Considering that, you know, you might have a spark of an idea, you might be able to
37:20prototype it easily, but it might actually be made up of many different components
37:24composed together to produce something that looks like it behaves intelligently.
37:29It needs to be factored into that process as well.
37:32And you know, one of those big components will be the
37:34differentiated data that you bring.
37:37And often like the step from prototyping something in a experimental environment
37:43to actually plumbing those data pipelines in a way that's sustainable for production
37:48use, you know, that can be a major step as well that needs to be considered upfront.
37:53And so, you know, when we've talked about, when we explored AI and ML
37:57initiatives, you know, we put a heavy weighting on that factor of where will
38:00the data come from and how will you deliver it to the product or application.
38:05You know, that's, again, you know, it's one of those laws where it always takes
38:09longer than you expect, even when you plan for it to take longer than you expected.
38:15Like any software engineering projects out-there, right?
38:17So I think the most important thing is ML product, ML project, right?
38:21So you need to still have the product thinking concept-in-the very beginning,
38:25right, so you, we've mentioned a little bit about, you know, user interviews,
38:28experiment, you know, building prototypes.
38:31Making sure the data that you feed into the ML training and all
38:35that is also appropriate, right?
38:37And I like the mentioning about failure modes, right?
38:40Because unlike other products out there, we kind of like know the input and output
38:44that we want the features to be, right?
38:46So ML product typically could fail in a, I don't know, unpredictable way, right?
38:51So we have so many things mentioned in the news, like for example, Google, you
38:55know, image classification, you know, the ChatGPT, and, you know, Gemini or Bard
39:00giving a wrong hallucinating answers.
39:03So like, how do you tackle that?
39:05Plus, I think the whole aspect of, you know, data security, bias, right?
39:09And also other associated aspects of fair data that you use in the training, right?
39:15I think it's also another thing that you should put your product thinking
39:18concept holistically so that you come up with a very useful and valuable product.
ML Engineering Best Practices
39:23So let's go to the other discipline, which is the engineering side, right?
39:27So I think what I could see in the past as well, like a lot of ML code is kinda like
39:33highly unstructured, I would say, right?
39:35So it's like procedural.
39:37There's no proper modeling.
39:38It's-just function calls over function calls, very complex with trace.
39:42So maybe in your view, and you mentioned in the beginning as well, you took a
39:46project to refactor ML project, right?
39:49So what are the disciplines that typically are lacking in the
39:52engineering aspect of ML product?
39:55And how we could do better?
39:57Yeah, I personally can resonate and relate with that experience.
40:01I had to get glasses recently, maybe because of old age, but I think mainly
40:06because looking at too much code all time and sometimes into late nights, cause
40:10of stress or, you know, which is other things, which we, as a healthy team, you
40:15know, if we do those right practices, with automated testing, with continuous
40:21delivery, automated deployment, then nobody should need work late-nights.
40:25So I think a couple of things that you mentioned there.
40:28Number one, I think, test automation is a big part.
40:31Like every ML engineer, data scientist we've worked with,
40:35they enjoy the automated tests that we introduced and added.
40:39And so it's, I think here at this point, it's a problem of information asymmetry.
40:45Like we've got pockets of teams of people who know how to do
40:48automated testing for ML systems.
40:50And every team that I've joined and worked with, they're just
40:53like, oh, I didn't know this.
40:54Oh, you could do that.
40:56So I think there's that desire and demand for, you know, more
40:59automated testing ML systems.
41:02One encapsulating story was, uh, we had a project that has kind
41:06of close to zero test coverage.
41:08It was an LLM system.
41:10So over time we added the test pyramid, like the simple things like unit tests,
41:15integration tests to touch our whole LLM application, check that it's okay.
41:20Those are still point based tests.
41:22Then we also have that kind of more like deeper model eval test,
41:27a suite of however many examples.
41:30We run it, five minutes later we know, okay, model accuracy
41:33is 75%, whatever number that is.
41:36So then we had this automated dependency manager, like on upgrade called
41:42Snyk, or Renovate, or Dependabot.
41:45So Snyk opened up a PR saying, you need to upgrade this.
41:50It automatically, the PR had all these green ticks, tests were passing.
41:54We automatically trigger model eval, we know, okay, performance is just as good.
41:59So in 15 minutes, we could merge the PR.
42:01No stress, no effort.
42:03So that's how we grew capacity of a team, just by having these
42:07automated testing, model eval.
42:09You touched a little bit on software design or code design as well.
42:14And, so yeah, that's, I think, any code base, not just ML, is susceptible to that.
42:20And so I think taking that next level of discipline to say, can I extract function?
42:27For this, can I have a readable variable name, not just df or x?
42:33All of those software hygiene practices, which we can link some in the show notes.
42:38Single responsibility principle.
42:41My favorite one is open close principle.
42:43Like can you design something that is open to extension, but you
42:47don't have to modify it every time.
42:50Probably I'm a bit going into too much detail there, but I think
42:54the design of it or lack of design sometimes stems from the lack of tests.
42:59As you know, if there's no test and nobody can refactor, refactoring is so
43:02scary, so risky, like nobody does that.
43:04We take the path of least resistance.
43:07So yeah, I think the teams that we've worked with, the moment we added
43:10that safety harness, or when you have that test on the path to production,
43:15then a lot of things can happen.
43:16One time we did a massive refactoring of like this variable that was
43:22linked in all different places.
43:24You know, it was, you use an IDE shortcut, replaced it in like, 150 places, and then,
43:30you know, test pass, commit, done, right?
43:32It was like, the effort reduced from maybe days of testing to, again, minutes.
43:38And then in 20 minutes, it was running in production.
43:41So yeah, it's like a lot of these practices that have been emerging.
43:45I think it's just about spreading it more and it's why we wrote the book
43:48so that our teams don't do late nights writing code like I did in one project.
43:54And yeah, just enjoy the flow.
43:56And work out work life balance.
43:59And yeah, zero stress the production deployments because, you know, tests are
44:03passing, you've got production monitoring.
44:05So a lot of these engineering practices will, you know, really help team feel
44:09the joy and flow of building ML products.
44:13Yeah, I think the call out around testing is really crucial and actually,
44:18allocating your effort effectively.
44:20So in a regular software product, we might use the test pyramid to direct effort.
44:26So that we have, you know, a large number of cheap, low level tests.
44:30You know, we have a medium number of integration level tests.
44:34And then, you know, we might have a small number of end to end tests.
44:38We can actually, when we're looking at ML applications and other data
44:41intensive applications, we can add a second dimension to that, so
44:44it's a grid instead of a pyramid.
44:47And that dimension is the data dimension.
44:49So, you know, we might have, you know, a lot of small cheap tests
44:53around individual data points.
44:55So canaries, I guess you might also describe those as.
44:59We might also then have samples, tests at a sample level.
45:03So these are going to give us a little bit more insight than a point data test.
45:08They're going to have some variability, so there's going to be some tuning
45:11there, but they offer us much faster feedback than the final level, which
45:17might be a kind of global, let's test on all the data that we have.
45:20And often people don't think hard enough about balancing the tests
45:24across that spectrum of data.
45:26So, you know, you might be able to do a very quick training run
45:30on a very small subset of data.
45:32If anything's misconfigured in the training run and the training doesn't work
45:34properly, for instance, you know that will break and you'll get that feedback really
45:38quickly rather than waiting hours for it.
45:40But you might also, you know, the training might pass, you might have a lower
45:44benchmark or threshold for acceptable performance on that low run, but at least
45:48you've tested end to end and got some feedback with a sample data set that
45:53it's likely to work at the large scale.
45:56So it's all about bringing that feedback back.
45:59And I think when we come to that uncertainty in the front end as well,
46:04also being thoughtful about how you test under those conditions of uncertainty
46:09when you don't even know what it is you're looking for in exploratory data analysis.
46:15And so there, it's kind of moving from the unknown-unknowns to known-unknowns
46:21to known-knowns through testing.
46:22You know, visualization is really key to be able to look at the data and understand
46:27what it's telling you, or use automated tools to find relationships in the data.
46:32And then when you sort of understand what the data is telling you.
46:35So that visualization, does it look right?
46:37Does it look like I expect?
46:39That can actually be a form of testing.
46:41I've called it visualization driven development at times as well.
46:44But then once you understand qualitatively what you're looking for,
46:49then, you know, there's a whole range of data science techniques that you
46:52can use to turn that into a binary expectation that can pass or fail.
46:57That, you know, might be useful in an exploratory environment, but then might
47:00also be something that you promote into a production integration pipeline
47:04as well as David was describing.
47:06So really, yeah, thinking hard about how you use testing to get
47:09fast feedback all the way through the life cycle is pretty crucial.
47:13And just to add to that, LLMs, as you mentioned, is all the
47:17rage for the past two years.
47:19And so, one of the teams we worked with, we had to innovate and
47:22think about how to shift left.
47:24So as Dave mentioned, full evaluation can be costly.
47:28It takes time.
47:30It can cost money, especially for LLMs.
47:33So what we ended up doing was to shift again that left.
47:36So before we kickstart the big eval or any further deployments,
47:41we had, uh, integration test.
47:43And one of the challenges that the team said was like, how can you test
47:47something that's non deterministic?
47:48You know, the answer is different every time.
47:50So we had to, you know, wrote a assertion function that asserts
47:55on intent rather than vocabulary.
47:58So we can still evaluate that this response, given these conditions,
48:02yes, is the intent of what we expect, what we had expected.
48:07We were using, um, PyHemcrest for that kind of matcher style extensions.
48:13But you could use, you know, other tools as well.
48:15But yeah, I think sometimes at the forefront of this new capability,
48:19we have to be a bit creative on how can we shift that left to get
48:23the feedback that Dave mentioned.
48:25Right.
48:25Thanks for mentioning some of these techniques.
48:28I'm always intrigued like you mentioned, right?
48:30So I'm always intrigued how can you test something that is non deterministic
48:33and you know, so many variables that could come in into play, right?
48:37So I think thanks for bringing also the importance of automated testing.
48:40I think I like what you mentioned when you explained that, right?
48:43There's a little bit of information asymmetry.
48:45Maybe there are some ML engineers who are never exposed to some of these techniques.
48:49And when they know it actually, they could actually follow the discipline
48:52and make sure that the products are getting better and better.
48:56And I like also the approach of, you know, slicing the data
48:59for different stages of tests.
49:01I think that's also key in making sure that the ML projects also kind of
49:04like still behaves as what we expect.
49:06Because like for example, you can tweak the model a little bit, you know,
49:09the output can change so much, right?
49:11So we don't want that happen in the production.
MLOps
49:14The other aspect of ML that people always talk about lately is about MLOps, you
49:19know, building platforms, you know, how-to actually deploy a model and operate it.
49:24So maybe a little bit here, what do you think about MLOps?
49:26Is it something that we all need for building ML product?
49:30And what problem does it solve?
49:32Yeah, I think it's yet another tool in our toolkit, which I very much welcome.
49:36You know, back in the day, we have to wrangle and think how to solve, you
49:40know, large scale distributed processing.
49:43But now there are these MLOps tools that lets you abstract away that concern.
49:48So in one case, one team, we have built an ML platform where now the data scientist,
49:53anyone who doesn't have, know anything about infrastructure or AWS or Kubernetes,
49:59they will just write plain Python.
50:01And say, I want to have this, you know, large vertical scaling this way.
50:05I want to fan out and all of that is-done in Python.
50:09Compute is one part of the MLOps stack.
50:11You know, there's experiment tracking, which has really helped us as well.
50:15Every pull request runs an experiment that reports some results that
50:19we can kind of check over time.
50:22If David creates a new PR that this is better than our champion model.
50:26So that MLOps practice was really, really welcome as well.
50:30The challenge here is like too many tools and it's hard to navigate.
50:35ThoughtWorks had this article called "A guide to evaluating ML platforms",
50:40which we could link in the show notes.
50:42That really helped like thinking about the capability.
50:45Like some platforms try to do everything, some are narrow.
50:48So how do you pick what's right?
50:50And how do you avoid like shotgun surgery, like vendor coupling, things like that.
50:55But yeah, end of the day, I think MLOps is about abstracting away
50:59complexity so that you can focus on solving the right problem, not having
51:03to deal with undifferentiated labor, you know, in your day to day work.
51:08And I think, yeah, coming back to the testing perspective as well, I think one
51:11of the things we highlight in the book, to get the most out of MLOps automation
51:16and abstraction, you also need to ensure that you're doing the right testing to
51:20give you confidence that when you're moving fast, you're doing so safely.
51:25Yep.
51:26And to add to that as well, like, that's a key point we make in the
51:29book in that MLOps, you can't MLOps your problems away, just like how
51:34you can't DevOps your problems away.
51:36Last week, you had on the show, um, or last episode, DX with Laura
51:40Tacho, then talking about DevEx.
51:42So I really like this diagram of the DevEx triangle.
51:46How do you get faster feedback loops?
51:48How do you manage cognitive load?
51:50How do you get in the flow state?
51:52MLOps helps in a few ways, but MLOps is not going to write your tests for you.
51:57They are not going to make, talk to users and make sure you're
51:59writing the same right features or implementing the right features.
52:02They're not going to make sure your code is nicely factored and readable
52:08so that you can stay in the flow.
52:10So yeah, I think it's another tool, but it needs to be coupled with these other
52:13disciplines that Dave and I mentioned in this podcast and in the book.
52:18Yeah, thanks for the plug for the developer experience as well, right?
52:22So don't forget any kind of ML product, essentially, it's also like
52:25a software engineering problem, right?
52:27So it's a socio technical, don't forget also the aspect of, you know, this
52:30feedback loop, psychological safety, also mentioned in the beginning, right?
52:34So all this, like I mentioned in the very beginning, right?
52:36It's not just ML technicals that you need to understand.
52:39But it's actually at the end, it's a software engineering thing that you
52:42have to handle really, really well.
Make Good Easy
52:44So we have talked a lot about the other things.
52:46As we move towards the end, is there anything that we haven't covered that
52:49you think should be mentioned as well?
52:52I think one key takeaway is how do we make good easy.
52:56As engineering leaders, as ML practitioners, we've talked a lot
53:01about a lot of different practices.
53:03If we can make good easy, then teams can kind of just by following the team
53:08practices, following exemplar, repos, then you get that for free in your CI/CD
53:15setup, your test strategy, even maybe hygiene checks of talking to users.
53:20Have you put a business case together before you start asking people
53:24to work on this for six months?
53:25So, yeah, it can get out of hand easily with so many moving parts.
53:30So I think as an engineering leader, how do we make good easy?
53:33Make teams on the ground, when they get the mission to do a certain
53:37piece of work, like it's kind of built into the way of working.
53:41Yeah, so I like it.
53:42Make good easy, right?
53:42So sometimes we all get excited about the technology, so many moving parts, so many
53:47technologies that we can play with, right?
53:48But we forgot the aspect to make it easy for people to adopt, make
53:52it easy to get the buy in as well.
53:54So I think thanks for mentioning that.
3 Tech Lead Wisdom
53:56So it's been an exciting conversation.
53:58I learned a lot about what it takes to actually build an ML project,
54:02which I find it really complicated.
54:04But as we reach the end of our conversation, I have one last
54:07question that I'd like to ask you.
54:08I call this the three technical leadership question.
54:11Just treat of it like an advice that you want to give to us as a listeners.
54:15So what will be the three technical leadership wisdom
54:17that you can share with us?
54:19Shall I go first?
54:20Yeah, if you like.
54:21I-think, um, again, in line with some of the philosophy of the book, being
54:26able to take different perspectives on technical problems is really
54:29key for your leadership growth.
54:31And so looking for opportunities to play different roles in projects, even
54:37for a short time, it gives you that understanding of what other stakeholders
54:41require and how to make them successful as you aim to be successful yourself.
54:47Yep.
54:47I thought of two.
54:49So one is, I call it focus, function, and fire.
54:54So this was an idea I got from Todd Henry in his book, I-think, [Herding] Tigers.
55:00So, you know, we are building things every day.
55:03Teams can get distracted with many things.
55:06So how can we, to set our team up for success, how can we give them that focus?
55:11That clear mission, the milestone, the why we're doing it for the
55:15customers to benefit for the business.
55:17The business case, like how will this benefit, you know,
55:20the metrics that we care about?
55:22Function, you know, the way of working.
55:24Instead of throwing work or communications over to teams, is
55:28have we got the right way of working.
55:30And Fire, you know, meaning the kind of implicit motivation,
55:34like why are we doing this?
55:35How is this helping people?
55:37How is this helping business?
55:39So that, that for me is one principle I take to, uh, my teams.
55:44The second one was really interesting and we covered this in the book as well
55:47about trust and psychological safety.
55:50So when it's absent from the room, then a simple conversation
55:56becomes a process of bureaucracy.
55:59Okay, I've got to write up this documentation, you've-got to
56:01pre-read it, and let's have a meeting in two weeks, discuss.
56:06And when it's also not there, then team members are afraid to voice
56:10concerns of how things might fail.
56:12So then we continue to track down the wrong direction.
56:15So yeah, as a kind of tech leaders, how do we make sure we have, you know, embody
56:19and encourage and ensure that we have that psychological safety in the team
56:23so that we can all do our best work?
56:26On David's final point, this is what I would say.
56:28That idea of trust is, is also really important in a multidisciplinary
56:32team and innovative initiatives under certain, under conditions of uncertainty.
56:38We want anyone to be able to speak out with good ideas or concerns
56:43that things might be broken.
56:45And so to be able to create those serendipitous moments, as well as the well
56:48understood moments that trust facilitates, is a pretty key focus for leaders.
56:54Yeah, I like the last aspect that you mentioned about psychological safety.
56:57So you also mentioned that if it's not there, right, things can tend to become
57:01like a bureaucratic kind of thing, right?
57:03So I-think that's really a good plug.
57:05So if people want to, you know, learn more about these exciting things
57:09that you mentioned in the book, or if they want to discuss with you,
57:12maybe is there place where can find any of you or both of you online?
57:16Yeah, you can find me on LinkedIn, David Colls.
57:20Yep.
57:20And I'm Davified and I can share a link in the show notes or we can
57:24share a link in the show notes.
57:26And our book, Effective Machine Learning Teams, is also the
57:29first chapter and preface.
57:31Actually, sorry, the preface is available for free on the link that
57:34we can share in the show notes.
57:36So that gives you an overview of everything we've talked about
57:39in this podcast in seven pages.
57:41And also the book itself, we can link it in the show note on where
57:44you can, yeah, read or listen to it.
57:47Thank you so much.
57:47So, uh, it's been a pleasure to have both of you Davids in the show.
57:51So, I hope people learn a lot about the aspects of machine learning
57:54model and software engineering good practices, anyway at the end.
57:58Thanks so much, Henry.
57:59This was a great chat.
58:00Thanks for having us, Henry.
58:01That was great.