Free YouTube Transcribe

Video transcript

Building Effective and Thriving Machine Learning Teams - David Tan & Dave Colls

Tech Lead Journal · 10,552 words · 48 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Trailer & Intro

0:00ML projects tend to fail in this common failure modes.

0:04One is what I call POC hell.

0:07Another failure mode is that it's easy to get maybe the first thing out of the door.

0:13But then it's kind of stuck together with duct tape and sticky gum.

0:18And it's hard to test or change.

0:21In my experience of working with data intensive ML initiatives or other

0:27technology specializations, I'd found that the ability to build the thing

0:32wasn't always the major success factor.

0:35Often it was understanding what the right thing to build was, making sure that

0:39there was alignment between technical teams and business teams and product

0:43teams on the right thing to build.

0:44You can't MLOps your problems away, just like how you can't

0:47DevOps your problems away.

0:49MLOps helps in a few ways, but MLOps is not going to write your tests for you.

0:54They are not going to talk to users and make sure you're

0:56writing the same right features or implementing the right features.

0:59It's useful to think about the essential characteristics of ML solutions.

1:05They'll be able to do things at speed and scale that no human could.

1:08But they'll make mistakes.

1:10We wouldn't consider an ML solution if we knew the right answer every time.

1:13We'd write some rules instead.

1:15One of the quotes, opening quotes in our book was from Edwards Deming, saying a bad

1:19system-would beat good person every time.

1:22The whole thesis of our book is how do we create those systems to help teams build

1:27the right thing, build the thing right, and in a way that's right for people.

1:45Hello, everyone.

1:45Welcome back to another new episode of the Tech Lead Journal podcast.

1:48Today, I have with me two Davids, very excited.

1:51One is David Tan.

1:52He is actually my ex-colleague in ThoughtWorks.

1:55And the other one is David Colls.

1:57So they are the co-authors, together with another one, Ada,

2:00who couldn't join us today.

2:02They wrote a book titled Effective Machine Learning Teams.

2:07Even though it's a machine learning as part of the book, but I think

2:09we are not going to cover solely just machine learning, because

2:12I'm also not an ML expert.

2:14But we'll discuss things like how we can build effective machine learning

2:17teams, what kind of practices that we should be aware of, and things like that.

2:22So welcome to the show.

2:23I'll maybe mention Dave as David Colls, right?

2:26And David for David Tan.

2:28So welcome to the show, guys.

2:30Yeah, thanks for having us Henry, happy to be here.

2:32Thanks, Henry.

2:33Nice to be here.

Career Turning Points

2:35Right.

2:35I always love to ask my guests maybe first to share a little bit of themselves by

2:39sharing your career turning points that you think we all can learn from that.

2:43So any one of you can start first.

2:45Cool.

2:46Um, so I'm currently engineering manager in the AI products group

2:50in Xero and my career into tech hasn't always been straight.

2:56Like I was actually non-technical when I graduated from university.

3:00So I was working in government for maybe two, three years.

3:04And I got, by accident, hand, foot and mouth disease and I had to

3:08like be quarantined for eight days.

3:10And from there I started like, you know, watching some YouTube videos

3:12about data analytics, about coding, and then tried some tutorials

3:18and, you know, it was really cool.

3:19So then decided to, you know, join a programming bootcamp, quit my job, did

3:24that, and then joined ThoughtWorks.

3:26And yeah, learned a lot about Agile, about test driven development, refactoring,

3:31CI/CD, got into ML engineering and so, you know, finally got into this

3:37exciting space of ML engineering, DevOps, building ML products.

3:41So yeah, I have to thank this 8 day quarantine back in the day

3:45for my career turning point.

3:47Yeah, maybe.

3:48How about you, Dave?

3:49If I'm to pick on an external theme like that, I'd probably choose deciding

3:54to play Ultimate Frisbee early in my career, where I met a number of

3:59people involved in the software scene.

4:01But, in particular, a ThoughtWorker who encouraged me

4:04to interview at ThoughtWorks.

4:06And when I think about turning points, I guess that kind of leads me into

4:10the fact that I started as a developer in technical, data intensive roles.

4:15I then, you know, looked for ways to make more impact in the leadership space.

4:20And when I joined ThoughtWorks, although I'd actually passed a coding

4:23interview, I joined as a project manager.

4:25Then I've worked for a number of years in Agile transformations and

4:29organizational transformation before then picking up the data practice.

4:33So I guess, my career has turned from deeply technical to leadership

4:40positions and organizational design a number of times over that period.

4:45Thank you so much for sharing your story.

4:46Very interesting, right?

4:47So the kind of serendipity that could happen throughout our career, right?

4:51And what led us to where we are at the moment.

Writing “Effective Machine Learning Teams”

4:54So today we are going to discuss topics from your book,

4:57Effective Machine Learning Teams.

4:58Maybe in the beginning, let's share what made you wrote this book.

5:03Yeah.

5:04This book started like as a blog post, actually.

5:06So back in 2019, we had wrapped up a project.

5:11We're involved a lot of refactoring of data science codebases.

5:15So we have data scientists wrote code in the Jupyter Notebook,

5:20no test, a lot of lines.

5:21You know, it was good, like it was solving a real business problem.

5:25And then we had to kind of productionize that.

5:27And so we started writing about how we did that.

5:30How do we encapsulate code into functions?

5:34How do we have automated tests?

5:35So we are, you know, clear about that, about the quality.

5:38How do we do CI/CD for those changes?

5:41And then we wrote a blog post and it kind of like exploded in, back then Twitter,

5:46someone from the R community shared it.

5:48And a lot of likes and it's like, okay, I think we're onto something here.

5:52There's this desire, I think, in the data science community or ML

5:55community to have more kind of engineering and robust practices.

6:00So yeah, you know, then we started writing more and more about...

6:04from the subsequent projects or work we did to kind of large

6:09scale ML training pipelines, large scale high volume ML products.

6:15And, you know, the practices that helped us build that reliably, build

6:19ML products that are safe, help you get fast feedback on your changes.

6:23And some of those things worked really well.

6:25So we wanted to share it more broadly with our readers.

6:29Uh, and that's just the engineering space.

6:31And I think Dave also has some other kind of, yeah, aspects that, you

6:35know, we started with this thread, calling engineering practices for ML.

6:39And then we discovered this whole kind of world of other

6:42practices that also help ML teams.

6:44Yeah.

6:45So, you know, in my experience of working with, I guess, data intensive

6:50ML initiatives or other technology specializations, I'd found that the

6:56ability to build the thing wasn't always the major success factor.

7:01Often it was understanding what the right thing to build was, making

7:04sure that there was alignment between technical teams and business teams

7:08and product teams on the right thing to build, which could be complex when

7:12people don't speak the same language.

7:14So, you know, what sort of processes and techniques can you use to build

7:18that alignment, especially when there's a lot of uncertainty both

7:22about the problem and the solution.

7:24And then also, because of the specialised nature of this work, it's often hard to

7:29build a team that can execute end to end.

7:32So then the factors around either building in a single team, how do you bring all of

7:39those different perspectives together, or when you have to rely on multiple teams to

7:43get a piece of work out the door, what's the best way to achieve fast flow and high

7:48quality in a way that's sustainable for the people working in that environment.

7:52So for me, those were the perspectives I wanted to bring to machine learning

7:57and product development, which is an essentially multidisciplinary activity.

8:02You know, how do we do that better as teams?

8:04Right.

8:05When I read your book, I must admit that I was thinking that I was going

8:09to learn a lot of the technical things about specific ML, you know, projects

8:13and ML products and things like that.

8:15But actually, you cover kind of like the holistic approach.

8:17not just the engineering, also like product, delivery, and things like that.

8:21And actually there are so many things that any software engineering

8:24team, I believe, could learn just by reading the book as well.

8:27I mean, putting aside some parts of the ML.

ML Engineering vs Other Types of Engineering

8:29But let's start maybe in the first place, because I'm sure many listeners

8:33here not coming from the ML background.

8:36What are the differences between, you know, ML engineering and the

8:40typical software engineering these days where people build websites,

8:43APIs, and things like that?

8:45Yeah, that's a great question.

8:47And I think ML systems fundamentally is orders of magnitude, I think,

8:52harder and more complex than software engineering in some ways.

8:57Like, of course, software engineering is hard, you know, you've got microservices,

9:01you've got distributed computing, you've got a lot of things, like in

9:05itself it's an art and whole discipline.

9:08ML itself, the difficulties manifest themselves in some common ways,

9:15like, you know, it's not just one service or one component.

9:19Usually, it's like a data pipeline, dependencies upstream, that, you

9:24know, it's complex in itself.

9:25You've got your scale or compute requirements for your own ML workload,

9:30which is also, can be done in many ways.

9:34You know, you can have a really large instance with really large

9:37memory, do everything there.

9:38Or, you know, at some point that breaks down, so how do you architect or make

9:43it simple to scale up large workloads?

9:46And then there's also the monitoring, like when you think about software

9:49products, typically your monitoring is kind of your four golden signals, right?

9:54Success rate, latency, things like that.

9:57You can do that for ML, and you can be all green and good, but that says

10:02nothing about like the quality of the predictions that models are producing.

10:07And then the correctness of those.

10:09So it adds another layer of challenges that I think the MLOps community

10:13have come to solve in recent years.

10:16And then further, even further down, like, yes, we've built this product, it's

10:19giving accurate predictions, to an extent.

10:23What is the kind of business impact or what levers or dials is it moving?

10:29So that's like kind of essentially the fundamental difference between the ML

10:35product and software engineering products.

10:36There's a small...

10:38There's this diagram from Google where like ML is just like a little box.

10:41And there's like infra, there's monitoring, there's

10:44data quality, data skew.

10:47So a lot of moving parts and that's why we wanted to bring a lot of practices

10:51that kind of help you tame the complexity in this different sub-components.

10:56And, yeah, when you look at how to manage the work as well, then, I guess, one of

11:00the characteristics of building an ML feature is there's a lot less certainty

11:05about how it's going to perform up front.

11:08And so often you have to integrate this exploratory data analysis or

11:13building prototypes or architecture search as it might be described into

11:18the product delivery process as well.

11:21And then, as you converge on a potential solution, then you're also dealing

11:26with another vector of change, which is the data and the world changing as

11:30well as David said around monitoring.

11:32Often that changes slowly, but sometimes we can see step changes.

11:36COVID was a classic example, with recommender systems moving from lowest

11:42price to highest, to best availability.

11:45And you know, ML models needing to catch up with that.

11:48So yeah, you both have less certainty and you have another vector of change to deal

11:51with when it comes to managing that work.

11:55Yeah, so thanks for sharing some of these complexities, right?

11:58So definitely for people who may not be able to relate.

12:01I think typically from what I can see, right, there are so many

12:04data pipelines when you build ML.

12:05You probably need more than just a small data, right?

12:08Something like a medium data or big data and the kind of compute that you need.

12:12Sometimes it's more like a distributed computing with

12:15large scale kind of pipelines.

12:17And I think typically what I can see as well, the feedback loop might be

12:20long in ML projects, simply because you have this training kind of thing.

ML and LLM

12:24In the first place, maybe let's clarify when you say ML, I think we are not

12:28just associating it with LLM, right?

12:29These days people are crazy about LLM and maybe they associate ML now as LLM.

12:34So maybe a little bit of insight here, like what do you mean by ML projects?

12:38Yeah, I guess, so yeah, great observation, Henry, and I was thinking, you know,

12:42as you talked about feedback cycles and big data and all of those elements.

12:46Yes, those are the things we traditionally associate with supervised

12:50machine learning, which I guess is our sort of primary focus in the book.

12:55Which might be anything from a binary classification problem, like is this

12:59transaction a fraudulent transaction?

13:01And you know, to solve that with a supervised paradigm, we get a lot of

13:04historical data about transactions, their amount, the age of the account.

13:10You know, maybe some information about the network around that account and

13:13then we label those transactions, whether they were fraudulent or not.

13:17That's actually a nice example of where the labels might be easy

13:20to get because we might get them from customer reports of fraud.

13:23But in general, that effort of labeling the data set can

13:25be really intensive as well.

13:27So that's the main paradigm we're looking at, but we think that, you

13:30know, looking through this lens of working with uncertainty, working with

13:34data, working with specializations.

13:37But there's a whole lot of other paradigms of ML or AI that might fit into that.

13:42So obviously working with a large language model in inference mode,

13:47that's still a challenging environment.

13:49The supervised paradigm might apply to fine tuning or even training from scratch

13:54your own large learning model as well.

13:57But then beyond that sort of paradigm, we're also looking at

14:01initiatives that might be, might use reinforcement learning, for instance.

14:05It's another approach to recommender systems.

14:08And this is where you might not need big data to start with.

14:11You might not have historical data to start with, but you might actually

14:14build or tune your data set as you go.

14:17We might also look further afield to techniques like simulation or operations

14:22research for optimisation, they share a lot of the same characteristics.

14:25Simulation is actually very similar to machine learning when

14:28it comes to productionising it.

14:30It's just that instead of a learned model of the world, you're working

14:33with an explicit model of the world.

14:35But you still need to go through feature engineering, ability to monitor

14:38and assess business impact and align to, um, stakeholder expectations too.

14:43So, while, yeah, while we're focused on supervised machine learning as a paradigm,

14:48yes, there's a big world out there.

14:50Anything that's decision making driven off data or models of the

14:53world or requires specialization might fit into this view that we've

14:57put forward in the book as well.

Why Many ML Projects Fail

14:59Thanks for the clarification.

15:01So maybe let's start going deeper into the book, right?

15:03I think one thing that piqued my interest when I read the book in the

15:06first few chapters of the book, right, I think you cover some statistics here.

15:10The first one is actually many ML projects doesn't make it to production, right?

15:15So and the second one, even though they reach production, they actually

15:19don't solve the real business problem or they don't bring any value, right?

15:22So maybe let's share a little bit why this is the case.

15:25What do you see out there as ML practitioners, right?

15:27So why is it so hard to actually bring ML projects to production?

15:32Yeah.

15:33Thanks, Henry.

15:34Like that's the crux of the book, in that ML projects tend to fail in this

15:40common like failure modes that we describe and I can go through them.

15:44And it's kind of preventable types of failure, like if we can learn from

15:51experience from history and that we would bring in the right tactics to

15:55make sure, like, as you mentioned, let's say a project that, you know,

15:59let's say we invest half a year, nine months into it and ship it.

16:03And we found that, oh, actually it's not solving the right problem.

16:06Or like users didn't really care about it.

16:08It's like, then what tactics can we bring in there with prototype

16:11testing or customer research?

16:14So that's kind of the whole kind of the main thesis of the book is like,

16:20how have we failed and what have we learned and how can we do better?

16:25So one common failure mode in ML projects is what I call POC hell.

16:31It's like a combination of factors, maybe business don't have enough of the

16:36appetite to release something to users.

16:39So we might, you know, take the first easy step to build a prototype, to

16:44test it, maybe internal demo of it.

16:47But the resources or risk to productionize that and put it in the

16:53hands of users is kind of too great.

16:55So with the lack of commitment there means that teams on the ground just kind

17:00of build POCs after POCs after POCs.

17:03So that's detrimental to, you know, the business in some ways, like

17:07you're missing out on opportunities.

17:08So detrimental to the morale of the team.

17:11You do a few and then, you know, after a while you say, what's the point?

17:14So, you move on and then you've got churn and you've got onboarding

17:17and you're losing talent, right?

17:20Another common, I think, failure mode is that it's easy to get maybe

17:25the first thing out of the door.

17:27You hustle, you get the data, you get the compute, train the

17:31model, you deploy the thing.

17:33But then it's kind of stuck together with duct tape and sticky gum.

17:39And it's hard to test or change after you evolve based on, you know, new

17:45customer feedback, things like-that.

17:47So then that comes to kind of continuous delivery, CI/CD, how confident can

17:51we be to test and deploy a change.

17:54And so this set of problems is solved by MLOps, by continuous

17:58delivery, from machine learning.

18:00There's probably a couple more, uh, I don't want to hog the mic.

18:04Dave, was there any other failure modes that came to mind or these challenges?

18:07Yeah, those are two big ones.

18:09It's, you know, is there any value in doing this at all?

18:13And that can be really challenging when the responsibility is divided across

18:17different teams and might even be in different parts of the organization.

18:21There might be a, I guess, the release of ChatGPT sparked a lot of curiosity

18:28in how do we best use LLMs in products.

18:31So, but we might have responsibility diffused across the organisation.

18:36We might have an ML team that's interested in looking at the

18:39use of large language models.

18:41We might have a product team that has their own roadmap, where they don't have

18:45these features anywhere on that roadmap.

18:48And then we might rely on data, especially in the case of Gen AI,

18:52that's unstructured data that's not well governed, that's hard to make available

18:57in a responsible way to these systems.

19:00And so while each group can do what's within their control to move a little

19:05bit towards a future state, it's not aligned or stitched up, you know,

19:09in a cohesive way that allows us to establish whether there is some value

19:14quickly with cheap and low risk tests.

19:16And then, as David said, once you've established there's some value, you

19:20know, if you have productionized something that's not supported by a lot

19:25of these good practices around MLOps and continuous delivery, which is often

19:29the case to understand if there is value there, then you start to understand how

19:34much value you're leaving on the table.

19:36And then this can be an opportunity to invest in those practices

19:40to be able to maximise the lifetime value of an ML product.

19:44But then that requires a certain, again, a certain degree of maturity and a

19:48certain ability to work across existing organisational boundaries or reshape them.

ML Success Modes

19:54Yeah.

19:55And there's also kind of the flip side of the failure mode is the success, right?

20:00So we've also worked with and seen teams, ML teams, successfully deliver,

20:04you know, really great ML products.

20:05And what tactics did they use to succeed?

20:08Like we could see things like customer testing, user testing, so before they

20:14build out a prototype or a MVP, which is a really expensive way to test

20:18something, you know, even just talking to users, understanding what is the pain

20:23point, what are their jobs to be done.

20:25Where could ML help?

20:26And then, you know, once they're bought in, they say, okay, this is the right ML

20:30product that will solve those problems.

20:33Then combating that failure mode of what Dave mentioned, like multiple teams

20:38kind of throwing across each other.

20:39If we remember the DevOps comic, where you have devs on one side, you have ops.

20:44So instead of throwing code between scientists and, you know, MLOps

20:50engineers, you know, we do the inverse Conway maneuver where we

20:54have a cross functional team shipping end to end for this initiative.

20:58And when we talked to the team members who worked on that project, they enjoyed that

21:03end to end ownership, the satisfaction of putting something, releasing to-customers.

21:08You get to touch different parts of the stack, learn about data

21:11science, you learn about Ops.

21:13You know, so they help them, you know, you can get fast feedback, you have

21:16the same standup, you ship things iteratively, you showcase it every,

21:20you know, two sprints, three sprints.

21:22So that feeling of flow and speed was like really amazing.

21:26And then, you know, when something is released, you know, you've

21:28got your continuous delivery, production monitoring, when

21:32things are not going well.

21:34Yeah, and you'll have safety for changes.

21:36If I need to make a refactoring, upgrade the library.

21:39If your CI checks are all green, your model monitoring dashboard is saying

21:44performance is good, then, you know, you merge that PR and you don't feel

21:48nervous or stressful about any releases.

21:51So yeah, there are also bright spots in how teams deliver ML products.

21:57Right.

21:57So I think when I heard you mentioned, you know, in the beginning about

22:00POC hell, you know, so many duct tapes when you bring production.

22:03I could relate to some other ML projects that I knew of before.

22:07So I think many, many, ML teams probably are in this mode, right?

22:10They keep building POC because maybe the expectation from stakeholders

22:14or from the business also kind of like with a lot of hype, right?

22:18Because these days people think AI can do a lot of magic, right?

22:21And especially with all the success of ChatGPT and all that, they

22:24even think that it is easy to do.

22:27But I think the reality may not be the case, right?

22:29And the other thing is about the practices, right?

22:32So many ML teams, first, either they don't have many good software engineers,

22:37so they are like data scientists or, you know, ML engineers who just have a

22:41lot of expertise in building the models.

22:43But actually they don't have all these other skills like refactoring, continuous

22:47delivery that you mentioned, right?

22:48Testing as well.

22:49Or the other flip side, which is a lot of software engineers

22:52being turned into ML engineers.

22:55So they are not necessarily an ML engineer, but they just come from,

22:58you know, like web application development and turned into ML engineer.

23:02So I think some of these things definitely I can relate.

23:04Yep, and just to add to that, there's also the flip side of engineers or ML engineers

23:09being excellent ML engineers, but not treating it as a data science problem

23:13where there is uncertainties, the need for experimentation to get fast feedback.

23:20So yeah, I think the, kind of emphasizes the need, I think, for

23:24cross functional collaboration from both, like data science, from MLOps,

23:29from, you know, SRE, things like that.

Ideal ML Engineering Team Composition

23:32So maybe let's go there, right?

23:33So the composition of the team, you mentioned a few times about cross

23:36functional teams and, you know, these silos between, for example, data

23:40science, maybe the ML Ops or whoever operates the system in the end.

23:45Maybe there are also other functions, like for example other teams, because

23:47ML typically also relies a lot of dependencies like data or maybe some

23:52kind of model, things like that.

23:53So maybe what are the best compositions here in your experience?

23:57Maybe you can advise us.

23:59I guess here that there's another question that is, and it's what we

24:03tackle towards the end of the book.

24:06We look at sort of multiple levels of team effectiveness.

24:10And one of the levels we look at is an individual in a team, the types of

24:14practices you can use, the expertise you can build as an individual.

24:18The next level we look at is within teams, again, how those practices around

24:22technology, delivery, and product.

24:24Augment teams, but also broader team dynamics.

24:28But then the third level we look at is between teams.

24:32So how can teams be effective between teams?

24:34And so the best makeup of a team depends a bit on the shape of the team

24:39and its interactions with other teams.

24:41And so this is where we use the team topologies model to identify the different

24:46types of ML teams that you might sit in.

24:49And so to run through it quickly and we can come back and dive into details.

24:53You know, you might have a stream aligned team, and this, we'd say,

24:57this is the basic unit to think of, first ML project, a stream aligned

25:01team that can deliver end to end.

25:02You're not going to be conducting groundbreaking ML research,

25:06you're going to be using tried and true techniques, in this case.

25:10Or at the other end where you have an established ML ecosystem internally,

25:14you can add another stream aligned team that draws on those existing

25:17services internally quite easily.

25:21Then as you scale beyond one team, there might be two routes that you take.

25:26And so you might go down the route of where you have a sort of low

25:30level, common concerns across different initiatives, that might be the opportunity

25:35for a technology platform to support ML at a lower level like compute

25:40and data and feature engineering.

25:43And so that would be a platform team in the team topologies team shape.

25:47And that would aim to provide as a service, and that would be composed

25:52differently from a stream aligned team.

25:54The other route you might take to scale is you might identify

25:56business clusters of ML needs.

25:58So there might be something around audience engagement or there might

26:02be something around asset valuation.

26:04Or there might be something around content moderation.

26:06And so those are a sort of specific set of business needs that also come with an

26:11associated set of ML paradigms as well, that don't necessarily have a common

26:16support in a technology platform at that right business level of abstraction.

26:20So then you might have what's called a complicated subsystem team.

26:24And again, the makeup of that might be a little bit different.

26:27And then as you, you know, as you have a bigger ecosystem, then there'll be

26:31a range of problems that come up of a similar nature, but like have a unique

26:36presentation each time, which might be around, say, privacy, or ethical use of

26:41data, or optimizing particular techniques.

26:45And this is where you might have an enabling ML team as well that sort

26:49of acts as a consultancy to other ML teams to make themselves redundant.

26:53And so that was a long way of saying, it depends on the ideal makeup of a team.

26:59But you might assume it's a stream aligned team and talk about the roles that sit

27:04in that and maybe I'll throw to David.

27:06I think team topologies is a really useful set of constructs,

27:10the ones that Dave went through.

27:12Because it helps teams scale through the team's API.

27:17So as Dave mentioned, you know, if we treat everything, let's say,

27:21as a stream aligned team, like in software engineering, we're very

27:24familiar with cross functional teams.

27:26Um, but then the failure mode there is that teams, like Conway's

27:30Law, right, we've got three teams, we're going to build three sets of,

27:33you know, basically architecture, rebuild certain tools that we need.

27:37Then that's where like maybe platform team comes in to abstract all of that.

27:41So then with that, what we've seen work really well in one particular

27:46case was this complicated subsystem team, like as you mentioned, ML

27:50is hard, a lot of moving parts.

27:52They took on the effort to build this ML product.

27:55Let's imagine it's, um, I can't say the specific example, but let's

27:59say it's a car valuations, right?

28:01That's kind of an ML product without a product.

28:04API to this team or this product is, I will give you the value of cars and now

28:11that is self serviceable by other teams.

28:13They could embed it on the mobile, mobile team could integrate with this

28:17API and expose the ML capability.

28:21A web team can do the same, or email marketing team can do the same, and

28:25then send personalized emails to say, you know, about car valuations.

28:29So then that was how that team scaled, rather than having to integrate with

28:33each particular thing, encapsulating that complicated subsystem as a kind of formal

28:38set of APIs, either through batch or through real time, that helped the team

28:42like achieve more and get more knowledge out of the ML product they built.

28:46Yes, yes, that's a great example that complicated subsystem team is probably

28:50going to have a bunch of specialists.

28:52It's going to look maybe most like what people imagine an ML team looks like.

28:57But as David said, to scale their impact in the organization, they do

29:01need that business domain expertise.

29:03They do need to deliver as a service instead of constantly

29:07collaborating with other teams.

29:09And you know, maybe some product thinking helps with

29:11that service definition as well.

29:13So, you know, that might be one team composition.

29:16Yeah, that's right.

29:18And if you replace that example there from car valuations, which is bit like,

29:22think not everybody can relate with that.

29:24Let's say it's like travel recommendations or product recommendations, then that one

29:29team that spent all the effort in building product recommendations now can impact

29:33multiple parts of the business by having the right team topology to, you know,

29:38have that fracture plane around, okay, my team is doing product recommendations.

29:42And yeah, you know, applying the kind of data product disciplines, exposing this

29:47as an API and encapsulating the details.

29:50Very exciting to hear about team topologies mentioned for

29:53ML product and ML teams, right?

29:55So I think it's kind of like back then, right, it was a revolutionary

29:58approach to how we kind of like create different teams.

30:01And I, I'm glad that you brought it up because still, I believe in many

30:04companies, they think, okay, we want to build ML product, ML project.

30:09They just hired a few data scientists or, you know, these ML experts, and they

30:13just asked them to kind of like build the model first without actually involving

30:17the other aspects of software engineers.

30:20It could be the stream aligned team from the product side.

30:22Or it could be, you know, the data, or could be anything, right?

30:25But I think that tends to kind of like have its challenges.

30:28So I think bringing the concepts such as team topologies to actually

30:31think holistically how we are going deliver the ML projects is something

30:35that's really, really important.

30:37Yeah, so a really great point, and I think one of the quotes, opening quotes in our

30:41book was from Edwards Deming, saying a bad system-would beat good person every time.

30:47So you can hire the smartest data scientists, put them in an environment

30:51where it's not, the right system, then, you know, we get what you described there.

30:57So the whole thesis of our book is how do we create those systems to

31:01help teams build the right thing.

31:04You know, solve the right problem, build the thing right.

31:07You know, engineering, ML engineering, data science.

31:10And then in a way that's right for people.

31:12I think what Dave puts in a really good way.

31:14Like it's not just shipping and shipping, but in a way that has the

31:19right team shape, right collaboration mode, right trust, psychological

31:22safety, and also the right processes to deliver, ship early and often.

31:27So yeah, I really was intrigued when you mentioned, you know, you can't

31:30just put a group of data scientists together or engineers together.

31:33You really got to create that system where they can, you know, ship to the

31:37right, you solve the right problems.

Building the Right ML Product

31:39Yeah, thanks for adding that.

31:41To come back to the theme of, you know, this product discipline, right,

31:44so I think we know that a lot of ML projects don't make into production,

31:48right, or they solve the wrong problem, or maybe don't bring value, right?

31:52So I think they are, like Dave mentioned in the beginning, there

31:54are a lot of uncertainties when you actually build ML model, ML product,

31:59right, because, I mean, the way it works also is kind of like prediction.

32:02It's kind of like there's some kind of ambiguities inside.

32:05LLM, there's a hallucination, right?

32:07So how can you actually come up with an approach such that, you know, when you

32:11first building the ML product, you can actually build something that is kind of

32:15like bringing the business value either to the users or to the organization?

32:20So I think this is probably one hard aspect as well.

32:22So typically how would you run this?

32:25Yeah, it is hard and it goes beyond the technical, although

32:29the technical informs it.

32:30I find it's useful to think about the essential characteristics of ML solutions.

32:36We're exploring them because we think they're going to be

32:40superhuman in some aspects.

32:43They're probably not going to beat the best experts in a field.

32:46That's maybe one myth that we should tackle straight up.

32:49But, you know, they'll be able to do things at speed

32:51and scale that no human could.

32:54But then we need to also consider that they'll make mistakes by the very nature.

32:59Again, we wouldn't consider an ML solution if we knew the right answer every time.

33:02We'd write some rules instead.

33:04So there'll be some percentage of mistakes, however small, in

33:07any ML solution, by design, as well as the unforeseen mistakes.

33:12And then when it comes to those mistakes by design, we really need to understand

33:16the cost sensitivity of what's, you know, what value do we get out of a bunch of

33:21right answers and what is the impact of maybe a very small number of wrong

33:25answers, but they could have a very huge impact across all sorts of dimensions, the

33:29financial, as well as security bias and fair treatment of all our stakeholders.

33:37So, yeah, we need to consider that fallible nature of them as well.

33:41And so, you know, we need to start with products that are

33:44designed to handle that failure.

33:45You know, they have some upside from when ML gets it right.

33:48But they're robust to the times when ML gets it wrong, because it will.

33:53So starting from that perspective, you know, we can then identify

33:56some experiments about how well does this need to work?

33:59We can start with very simple baselines.

34:01Sometimes, you know, even just predicting the majority class or

34:05random guessing and, you know, seeing how that works as a product.

34:09But, you know, ideally, to resolve this uncertainty, we're getting into

34:13some real data and understanding the predictive potential of the real data.

34:17And so this is, again, where it's like it's a challenging

34:20multidisciplinary exercise.

34:22We were trying to proceed on multiple fronts.

34:23Are we building the right thing?

34:25Can our solution support or a proposed solution support the

34:29performance that we expect?

34:31And so on.

34:33Yeah.

34:33And just to add to that, I think that emphasizes the importance of that cross

34:37functional nature of the work as well.

34:40Like if we frame this as a data science or ML problem, then we can try our level best

34:45to go from, let's say, 55% accuracy to 99.

34:50You'll never get to a hundred, and we will spend many, many weeks

34:55and months trying to get there.

34:57So it's not just the ML problem where we try to improve the model's

35:01accuracy or recall precision, but also how can we design for these

35:07failure modes of the ML model.

35:09So like for example, you know, displaying, designing a product in a way, right,

35:14to show users that, okay, this is not a confident prediction, or this is a

35:19prediction, but would you correct that?

35:22And also maybe even giving users options, like these are the top three, like,

35:26the model's top one prediction maybe, you know, not where we want it to be,

35:31but top three, okay, is much higher.

35:33So designing the product in a way that mitigates these

35:36failure modes of the ML model.

35:39And yeah, it's easy for teams if they don't have the right capabilities or

35:44right skillsets, to try to solve it.

35:47Like if you are a hammer, everything is a nail.

35:50And it's very costly to try to level up the accuracy.

35:53We may never get there, but yeah having that cross functional approach, like

35:57how do we design it in a-different way?

35:58How do we talk to users?

35:59Like do users find this okay?

36:01That's a kind of more holistic way to solve-the-problem.

36:04Yeah, that hammer and nail is really important as you start to shift to,

36:09okay, yeah, how do we make this viable?

36:12Or how do we, yeah, make it viable from the perspective of solving the

36:15problem effectively, but also from being economically sustainable to maintain it?

36:20And I guess there's another essential characteristic around that.

36:24The fact that ML solutions are narrow, the very training process is

36:28to optimize a loss function, which is a narrow definition of success.

36:32But they're composable.

36:33And, you know, this is, I think, where Gen AI can be really interesting.

36:37And I'm not the only one to take this perspective, but I've described like

36:41Gen AI as a stone soup for innovation.

36:44So the story of stone soup, if you haven't heard it, is that a weary traveler arrives

36:49at a village late at night, and all they have in their knapsack is a stone.

36:53So they go to the first house in the village and they ask the villager there,

36:57could I have an onion to make stone soup?

37:00You know, I've got the stone.

37:01All I need from you is an onion.

37:02That villager says, oh, great, yeah.

37:04I'll provide an onion.

37:05They go to the next house and repeat and ask for carrots and, you

37:08know, proceed around the village.

37:10By the end of that process, they cook up an amazing soup

37:13that feeds the whole village.

37:14And it was all cooked from a stone.

37:16Considering that, you know, you might have a spark of an idea, you might be able to

37:20prototype it easily, but it might actually be made up of many different components

37:24composed together to produce something that looks like it behaves intelligently.

37:29It needs to be factored into that process as well.

37:32And you know, one of those big components will be the

37:34differentiated data that you bring.

37:37And often like the step from prototyping something in a experimental environment

37:43to actually plumbing those data pipelines in a way that's sustainable for production

37:48use, you know, that can be a major step as well that needs to be considered upfront.

37:53And so, you know, when we've talked about, when we explored AI and ML

37:57initiatives, you know, we put a heavy weighting on that factor of where will

38:00the data come from and how will you deliver it to the product or application.

38:05You know, that's, again, you know, it's one of those laws where it always takes

38:09longer than you expect, even when you plan for it to take longer than you expected.

38:15Like any software engineering projects out-there, right?

38:17So I think the most important thing is ML product, ML project, right?

38:21So you need to still have the product thinking concept-in-the very beginning,

38:25right, so you, we've mentioned a little bit about, you know, user interviews,

38:28experiment, you know, building prototypes.

38:31Making sure the data that you feed into the ML training and all

38:35that is also appropriate, right?

38:37And I like the mentioning about failure modes, right?

38:40Because unlike other products out there, we kind of like know the input and output

38:44that we want the features to be, right?

38:46So ML product typically could fail in a, I don't know, unpredictable way, right?

38:51So we have so many things mentioned in the news, like for example, Google, you

38:55know, image classification, you know, the ChatGPT, and, you know, Gemini or Bard

39:00giving a wrong hallucinating answers.

39:03So like, how do you tackle that?

39:05Plus, I think the whole aspect of, you know, data security, bias, right?

39:09And also other associated aspects of fair data that you use in the training, right?

39:15I think it's also another thing that you should put your product thinking

39:18concept holistically so that you come up with a very useful and valuable product.

ML Engineering Best Practices

39:23So let's go to the other discipline, which is the engineering side, right?

39:27So I think what I could see in the past as well, like a lot of ML code is kinda like

39:33highly unstructured, I would say, right?

39:35So it's like procedural.

39:37There's no proper modeling.

39:38It's-just function calls over function calls, very complex with trace.

39:42So maybe in your view, and you mentioned in the beginning as well, you took a

39:46project to refactor ML project, right?

39:49So what are the disciplines that typically are lacking in the

39:52engineering aspect of ML product?

39:55And how we could do better?

39:57Yeah, I personally can resonate and relate with that experience.

40:01I had to get glasses recently, maybe because of old age, but I think mainly

40:06because looking at too much code all time and sometimes into late nights, cause

40:10of stress or, you know, which is other things, which we, as a healthy team, you

40:15know, if we do those right practices, with automated testing, with continuous

40:21delivery, automated deployment, then nobody should need work late-nights.

40:25So I think a couple of things that you mentioned there.

40:28Number one, I think, test automation is a big part.

40:31Like every ML engineer, data scientist we've worked with,

40:35they enjoy the automated tests that we introduced and added.

40:39And so it's, I think here at this point, it's a problem of information asymmetry.

40:45Like we've got pockets of teams of people who know how to do

40:48automated testing for ML systems.

40:50And every team that I've joined and worked with, they're just

40:53like, oh, I didn't know this.

40:54Oh, you could do that.

40:56So I think there's that desire and demand for, you know, more

40:59automated testing ML systems.

41:02One encapsulating story was, uh, we had a project that has kind

41:06of close to zero test coverage.

41:08It was an LLM system.

41:10So over time we added the test pyramid, like the simple things like unit tests,

41:15integration tests to touch our whole LLM application, check that it's okay.

41:20Those are still point based tests.

41:22Then we also have that kind of more like deeper model eval test,

41:27a suite of however many examples.

41:30We run it, five minutes later we know, okay, model accuracy

41:33is 75%, whatever number that is.

41:36So then we had this automated dependency manager, like on upgrade called

41:42Snyk, or Renovate, or Dependabot.

41:45So Snyk opened up a PR saying, you need to upgrade this.

41:50It automatically, the PR had all these green ticks, tests were passing.

41:54We automatically trigger model eval, we know, okay, performance is just as good.

41:59So in 15 minutes, we could merge the PR.

42:01No stress, no effort.

42:03So that's how we grew capacity of a team, just by having these

42:07automated testing, model eval.

42:09You touched a little bit on software design or code design as well.

42:14And, so yeah, that's, I think, any code base, not just ML, is susceptible to that.

42:20And so I think taking that next level of discipline to say, can I extract function?

42:27For this, can I have a readable variable name, not just df or x?

42:33All of those software hygiene practices, which we can link some in the show notes.

42:38Single responsibility principle.

42:41My favorite one is open close principle.

42:43Like can you design something that is open to extension, but you

42:47don't have to modify it every time.

42:50Probably I'm a bit going into too much detail there, but I think

42:54the design of it or lack of design sometimes stems from the lack of tests.

42:59As you know, if there's no test and nobody can refactor, refactoring is so

43:02scary, so risky, like nobody does that.

43:04We take the path of least resistance.

43:07So yeah, I think the teams that we've worked with, the moment we added

43:10that safety harness, or when you have that test on the path to production,

43:15then a lot of things can happen.

43:16One time we did a massive refactoring of like this variable that was

43:22linked in all different places.

43:24You know, it was, you use an IDE shortcut, replaced it in like, 150 places, and then,

43:30you know, test pass, commit, done, right?

43:32It was like, the effort reduced from maybe days of testing to, again, minutes.

43:38And then in 20 minutes, it was running in production.

43:41So yeah, it's like a lot of these practices that have been emerging.

43:45I think it's just about spreading it more and it's why we wrote the book

43:48so that our teams don't do late nights writing code like I did in one project.

43:54And yeah, just enjoy the flow.

43:56And work out work life balance.

43:59And yeah, zero stress the production deployments because, you know, tests are

44:03passing, you've got production monitoring.

44:05So a lot of these engineering practices will, you know, really help team feel

44:09the joy and flow of building ML products.

44:13Yeah, I think the call out around testing is really crucial and actually,

44:18allocating your effort effectively.

44:20So in a regular software product, we might use the test pyramid to direct effort.

44:26So that we have, you know, a large number of cheap, low level tests.

44:30You know, we have a medium number of integration level tests.

44:34And then, you know, we might have a small number of end to end tests.

44:38We can actually, when we're looking at ML applications and other data

44:41intensive applications, we can add a second dimension to that, so

44:44it's a grid instead of a pyramid.

44:47And that dimension is the data dimension.

44:49So, you know, we might have, you know, a lot of small cheap tests

44:53around individual data points.

44:55So canaries, I guess you might also describe those as.

44:59We might also then have samples, tests at a sample level.

45:03So these are going to give us a little bit more insight than a point data test.

45:08They're going to have some variability, so there's going to be some tuning

45:11there, but they offer us much faster feedback than the final level, which

45:17might be a kind of global, let's test on all the data that we have.

45:20And often people don't think hard enough about balancing the tests

45:24across that spectrum of data.

45:26So, you know, you might be able to do a very quick training run

45:30on a very small subset of data.

45:32If anything's misconfigured in the training run and the training doesn't work

45:34properly, for instance, you know that will break and you'll get that feedback really

45:38quickly rather than waiting hours for it.

45:40But you might also, you know, the training might pass, you might have a lower

45:44benchmark or threshold for acceptable performance on that low run, but at least

45:48you've tested end to end and got some feedback with a sample data set that

45:53it's likely to work at the large scale.

45:56So it's all about bringing that feedback back.

45:59And I think when we come to that uncertainty in the front end as well,

46:04also being thoughtful about how you test under those conditions of uncertainty

46:09when you don't even know what it is you're looking for in exploratory data analysis.

46:15And so there, it's kind of moving from the unknown-unknowns to known-unknowns

46:21to known-knowns through testing.

46:22You know, visualization is really key to be able to look at the data and understand

46:27what it's telling you, or use automated tools to find relationships in the data.

46:32And then when you sort of understand what the data is telling you.

46:35So that visualization, does it look right?

46:37Does it look like I expect?

46:39That can actually be a form of testing.

46:41I've called it visualization driven development at times as well.

46:44But then once you understand qualitatively what you're looking for,

46:49then, you know, there's a whole range of data science techniques that you

46:52can use to turn that into a binary expectation that can pass or fail.

46:57That, you know, might be useful in an exploratory environment, but then might

47:00also be something that you promote into a production integration pipeline

47:04as well as David was describing.

47:06So really, yeah, thinking hard about how you use testing to get

47:09fast feedback all the way through the life cycle is pretty crucial.

47:13And just to add to that, LLMs, as you mentioned, is all the

47:17rage for the past two years.

47:19And so, one of the teams we worked with, we had to innovate and

47:22think about how to shift left.

47:24So as Dave mentioned, full evaluation can be costly.

47:28It takes time.

47:30It can cost money, especially for LLMs.

47:33So what we ended up doing was to shift again that left.

47:36So before we kickstart the big eval or any further deployments,

47:41we had, uh, integration test.

47:43And one of the challenges that the team said was like, how can you test

47:47something that's non deterministic?

47:48You know, the answer is different every time.

47:50So we had to, you know, wrote a assertion function that asserts

47:55on intent rather than vocabulary.

47:58So we can still evaluate that this response, given these conditions,

48:02yes, is the intent of what we expect, what we had expected.

48:07We were using, um, PyHemcrest for that kind of matcher style extensions.

48:13But you could use, you know, other tools as well.

48:15But yeah, I think sometimes at the forefront of this new capability,

48:19we have to be a bit creative on how can we shift that left to get

48:23the feedback that Dave mentioned.

48:25Right.

48:25Thanks for mentioning some of these techniques.

48:28I'm always intrigued like you mentioned, right?

48:30So I'm always intrigued how can you test something that is non deterministic

48:33and you know, so many variables that could come in into play, right?

48:37So I think thanks for bringing also the importance of automated testing.

48:40I think I like what you mentioned when you explained that, right?

48:43There's a little bit of information asymmetry.

48:45Maybe there are some ML engineers who are never exposed to some of these techniques.

48:49And when they know it actually, they could actually follow the discipline

48:52and make sure that the products are getting better and better.

48:56And I like also the approach of, you know, slicing the data

48:59for different stages of tests.

49:01I think that's also key in making sure that the ML projects also kind of

49:04like still behaves as what we expect.

49:06Because like for example, you can tweak the model a little bit, you know,

49:09the output can change so much, right?

49:11So we don't want that happen in the production.

MLOps

49:14The other aspect of ML that people always talk about lately is about MLOps, you

49:19know, building platforms, you know, how-to actually deploy a model and operate it.

49:24So maybe a little bit here, what do you think about MLOps?

49:26Is it something that we all need for building ML product?

49:30And what problem does it solve?

49:32Yeah, I think it's yet another tool in our toolkit, which I very much welcome.

49:36You know, back in the day, we have to wrangle and think how to solve, you

49:40know, large scale distributed processing.

49:43But now there are these MLOps tools that lets you abstract away that concern.

49:48So in one case, one team, we have built an ML platform where now the data scientist,

49:53anyone who doesn't have, know anything about infrastructure or AWS or Kubernetes,

49:59they will just write plain Python.

50:01And say, I want to have this, you know, large vertical scaling this way.

50:05I want to fan out and all of that is-done in Python.

50:09Compute is one part of the MLOps stack.

50:11You know, there's experiment tracking, which has really helped us as well.

50:15Every pull request runs an experiment that reports some results that

50:19we can kind of check over time.

50:22If David creates a new PR that this is better than our champion model.

50:26So that MLOps practice was really, really welcome as well.

50:30The challenge here is like too many tools and it's hard to navigate.

50:35ThoughtWorks had this article called "A guide to evaluating ML platforms",

50:40which we could link in the show notes.

50:42That really helped like thinking about the capability.

50:45Like some platforms try to do everything, some are narrow.

50:48So how do you pick what's right?

50:50And how do you avoid like shotgun surgery, like vendor coupling, things like that.

50:55But yeah, end of the day, I think MLOps is about abstracting away

50:59complexity so that you can focus on solving the right problem, not having

51:03to deal with undifferentiated labor, you know, in your day to day work.

51:08And I think, yeah, coming back to the testing perspective as well, I think one

51:11of the things we highlight in the book, to get the most out of MLOps automation

51:16and abstraction, you also need to ensure that you're doing the right testing to

51:20give you confidence that when you're moving fast, you're doing so safely.

51:25Yep.

51:26And to add to that as well, like, that's a key point we make in the

51:29book in that MLOps, you can't MLOps your problems away, just like how

51:34you can't DevOps your problems away.

51:36Last week, you had on the show, um, or last episode, DX with Laura

51:40Tacho, then talking about DevEx.

51:42So I really like this diagram of the DevEx triangle.

51:46How do you get faster feedback loops?

51:48How do you manage cognitive load?

51:50How do you get in the flow state?

51:52MLOps helps in a few ways, but MLOps is not going to write your tests for you.

51:57They are not going to make, talk to users and make sure you're

51:59writing the same right features or implementing the right features.

52:02They're not going to make sure your code is nicely factored and readable

52:08so that you can stay in the flow.

52:10So yeah, I think it's another tool, but it needs to be coupled with these other

52:13disciplines that Dave and I mentioned in this podcast and in the book.

52:18Yeah, thanks for the plug for the developer experience as well, right?

52:22So don't forget any kind of ML product, essentially, it's also like

52:25a software engineering problem, right?

52:27So it's a socio technical, don't forget also the aspect of, you know, this

52:30feedback loop, psychological safety, also mentioned in the beginning, right?

52:34So all this, like I mentioned in the very beginning, right?

52:36It's not just ML technicals that you need to understand.

52:39But it's actually at the end, it's a software engineering thing that you

52:42have to handle really, really well.

Make Good Easy

52:44So we have talked a lot about the other things.

52:46As we move towards the end, is there anything that we haven't covered that

52:49you think should be mentioned as well?

52:52I think one key takeaway is how do we make good easy.

52:56As engineering leaders, as ML practitioners, we've talked a lot

53:01about a lot of different practices.

53:03If we can make good easy, then teams can kind of just by following the team

53:08practices, following exemplar, repos, then you get that for free in your CI/CD

53:15setup, your test strategy, even maybe hygiene checks of talking to users.

53:20Have you put a business case together before you start asking people

53:24to work on this for six months?

53:25So, yeah, it can get out of hand easily with so many moving parts.

53:30So I think as an engineering leader, how do we make good easy?

53:33Make teams on the ground, when they get the mission to do a certain

53:37piece of work, like it's kind of built into the way of working.

53:41Yeah, so I like it.

53:42Make good easy, right?

53:42So sometimes we all get excited about the technology, so many moving parts, so many

53:47technologies that we can play with, right?

53:48But we forgot the aspect to make it easy for people to adopt, make

53:52it easy to get the buy in as well.

53:54So I think thanks for mentioning that.

3 Tech Lead Wisdom

53:56So it's been an exciting conversation.

53:58I learned a lot about what it takes to actually build an ML project,

54:02which I find it really complicated.

54:04But as we reach the end of our conversation, I have one last

54:07question that I'd like to ask you.

54:08I call this the three technical leadership question.

54:11Just treat of it like an advice that you want to give to us as a listeners.

54:15So what will be the three technical leadership wisdom

54:17that you can share with us?

54:19Shall I go first?

54:20Yeah, if you like.

54:21I-think, um, again, in line with some of the philosophy of the book, being

54:26able to take different perspectives on technical problems is really

54:29key for your leadership growth.

54:31And so looking for opportunities to play different roles in projects, even

54:37for a short time, it gives you that understanding of what other stakeholders

54:41require and how to make them successful as you aim to be successful yourself.

54:47Yep.

54:47I thought of two.

54:49So one is, I call it focus, function, and fire.

54:54So this was an idea I got from Todd Henry in his book, I-think, [Herding] Tigers.

55:00So, you know, we are building things every day.

55:03Teams can get distracted with many things.

55:06So how can we, to set our team up for success, how can we give them that focus?

55:11That clear mission, the milestone, the why we're doing it for the

55:15customers to benefit for the business.

55:17The business case, like how will this benefit, you know,

55:20the metrics that we care about?

55:22Function, you know, the way of working.

55:24Instead of throwing work or communications over to teams, is

55:28have we got the right way of working.

55:30And Fire, you know, meaning the kind of implicit motivation,

55:34like why are we doing this?

55:35How is this helping people?

55:37How is this helping business?

55:39So that, that for me is one principle I take to, uh, my teams.

55:44The second one was really interesting and we covered this in the book as well

55:47about trust and psychological safety.

55:50So when it's absent from the room, then a simple conversation

55:56becomes a process of bureaucracy.

55:59Okay, I've got to write up this documentation, you've-got to

56:01pre-read it, and let's have a meeting in two weeks, discuss.

56:06And when it's also not there, then team members are afraid to voice

56:10concerns of how things might fail.

56:12So then we continue to track down the wrong direction.

56:15So yeah, as a kind of tech leaders, how do we make sure we have, you know, embody

56:19and encourage and ensure that we have that psychological safety in the team

56:23so that we can all do our best work?

56:26On David's final point, this is what I would say.

56:28That idea of trust is, is also really important in a multidisciplinary

56:32team and innovative initiatives under certain, under conditions of uncertainty.

56:38We want anyone to be able to speak out with good ideas or concerns

56:43that things might be broken.

56:45And so to be able to create those serendipitous moments, as well as the well

56:48understood moments that trust facilitates, is a pretty key focus for leaders.

56:54Yeah, I like the last aspect that you mentioned about psychological safety.

56:57So you also mentioned that if it's not there, right, things can tend to become

57:01like a bureaucratic kind of thing, right?

57:03So I-think that's really a good plug.

57:05So if people want to, you know, learn more about these exciting things

57:09that you mentioned in the book, or if they want to discuss with you,

57:12maybe is there place where can find any of you or both of you online?

57:16Yeah, you can find me on LinkedIn, David Colls.

57:20Yep.

57:20And I'm Davified and I can share a link in the show notes or we can

57:24share a link in the show notes.

57:26And our book, Effective Machine Learning Teams, is also the

57:29first chapter and preface.

57:31Actually, sorry, the preface is available for free on the link that

57:34we can share in the show notes.

57:36So that gives you an overview of everything we've talked about

57:39in this podcast in seven pages.

57:41And also the book itself, we can link it in the show note on where

57:44you can, yeah, read or listen to it.

57:47Thank you so much.

57:47So, uh, it's been a pleasure to have both of you Davids in the show.

57:51So, I hope people learn a lot about the aspects of machine learning

57:54model and software engineering good practices, anyway at the end.

57:58Thanks so much, Henry.

57:59This was a great chat.

58:00Thanks for having us, Henry.

58:01That was great.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.