Free YouTube Transcribe

Video transcript

From Chaos to Control: Why Data Contracts Matter

PyData · 6,659 words · 31 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:05Thank you. Thank you very much. First

0:08uh thank you for attending my

0:10presentation. I hope you will find it

0:12useful for you. Uh as you can see I'm

0:15going to talk about yeah data contracts.

0:18But first let me introduce myself.

0:21Uh recently two patients drive me uh

0:25building data platforms and pushing my

0:27limits in the water. One month ago I was

0:29capable to swim finally 5 kilometers in

0:31the sea in degrees and next month I hope

0:35to cross continents in Istanbul post

0:37forest trait and that also cool

0:40experience it's about cows in your head

0:41and how you control your emotions quite

0:44cool experience I would recommend you

0:46and let's switch back to my technical

0:48background in 2016 I joined Georgian

0:51startup Pulsar AI where I created first

0:53chatbot framework for Georgian linkage

0:55it was 2017 already and actually worked

0:59on different AI tools like social med

1:01media monitoring tools, sentiment

1:03analysis tools. Uh later I continued uh

1:06my career in Doppel freelance network.

1:09Probably you worry about it. It was

1:11quite long time and I worked on

1:13different scale projects from Fortune

1:15500 companies like media and to the

1:18smaller scale startup or hedge funds and

1:20since April 2024 I continued my career

1:24journey in indust windows is Lithuan US

1:26Lithuanian company uh door and windows

1:30manufacturing company and my first

1:32initial project was building the modern

1:35data infrastructure for them.

1:39Well, in indus we have different

1:42internal teams and different projects.

1:45uh production planning as name says like

1:48that software that helps uh ter people

1:51to plan these terminals and what's going

1:53inside in manufacturer uh internal HR

1:57software or we call it people hub it's

2:00for HR and drawing and calculating

2:03software um

2:06are different estimators or drafters

2:08they using it uh for estimate window

2:10size uh fix this parameters and do a lot

2:13of calculations because it kind of done

2:15for very big buildings sometimes for

2:17skyscrapers.

2:19uh this is just example of three

2:21projects but actually there's a much

2:23more why I'm talking about these

2:25projects because I built modern in data

2:28infrastructure I started collecting this

2:30data into data warehouse and after I run

2:33into very typical problem that happens

2:35in a lot of organizations

2:38software engineers start modifying the

2:41transactional databases schemas without

2:44letting us know and without thinking

2:47that it can break all downstreaming

2:50processes

2:52data pipelines was failing. DBT models

2:55also started failing or just not working

2:59and re reports they get delayed and in

3:02the end our end user managers were

3:05calling us and saying we need this

3:06report but it's getting delayed. We had

3:10this problem and fortunately for that

3:12moment uh my team in Lithuania they

3:15attended biggest data conference. It was

3:17also in Vnus and there was nice speech

3:19by Yenni from Zando and he was sharing

3:23um their experience how they fix this

3:25problem in Zando with data contracts

3:29and why we decided why not why not to

3:32try use this experience in our company

3:35and try maybe it will help us to

3:37coordinate uh working with different

3:39teams and finally fix this issue because

3:42before it was verbally agreement or

3:44different ways but it wasn't solution

3:47What is a data contract? Data contracts

3:49is machine machine readable document

3:51actually YAML file um that explicitly

3:54defines a structure, format, semantics,

3:57data quality, SLAs, term of use for

4:00exchanging data between producer and

4:02consumer producer that's these teams and

4:04we data team we are consuming it.

4:07Very key point in Yani's presentation

4:10called trust. But what what this says

4:13from official document instead of hoping

4:16now we have explicit agreement both

4:18sides having explicit agreement and

4:21continuously enforce it.

4:23Let's see further how does it look like

4:25actually it's bigger file I just split

4:27it into the chunks to go through it and

4:29show you per line what it means and how

4:32it works. First of that's fundamentals I

4:36call it this kind of passport of the

4:38data data contract. uh that's API

4:41version

4:43data contract not data contracts but

4:45software's API version that's going to

4:47process this because we had a case when

4:49you have YAML file if version changed it

4:52will not be able to process it and it

4:54fails like we have to follow this API

4:56version kind just explicitly defining

4:58the data the data contract ID unique

5:01identifier uh in this case it's orders

5:04because you can have thousands of data

5:06contract files uh name orders uh semi IC

5:10version uh this important and I will

5:12show you on the next slide how you can

5:14work with this and status of data

5:17contract in this case it's active

5:22semantic versioning in this case there's

5:24a three parameters that we could control

5:27first patch or you can see third

5:29parameter uh that most right one uh what

5:35it means no data impact sometimes in

5:38data YAML file You can just update

5:41definition of the column or just add

5:43some description. This has nothing to do

5:45with data. It's just descriptive

5:47information and actually will not break

5:49anything.

5:50Next one. Next parameter is minor and

5:53non-breaking change. Uh that's in middle

5:56number. This one

6:00um quite frequently even now during

6:02working process uh back end teams doing

6:04refactorings and adding new columns.

6:08uh table getting uh wider like bigger

6:10but it will not impact on us because the

6:13used columns is not changing only new

6:15columns is coming and that's not a

6:16problem and this is a queen or I don't

6:21know king [laughter]

6:23this what breaks everything like if back

6:26end team decides to refactor it and

6:29exactly deleting some columns or

6:31changing their names or changing formats

6:34it breaks everything

6:37and There's a teams they kind of keeping

6:38us updated that you know we are going to

6:40change but there's a teams that can just

6:43deploy it and say it works on our side

6:45and we little less care what's going on

6:47next and yeah this is major one

6:53okay next uh this is Kimma if previous

6:55slide I was calling it uh passport

7:00uh if previous slide was passport this

7:03is heart of our data contract as you can

7:06see like this uh physical type uh first

7:10of all it's name the name table name

7:13orders uh physical type is a table

7:16description

7:17information about table all the web shop

7:19orders actually it could be much longer

7:21I just cut it to uh fit this slide size

7:25and properties actually columns and here

7:28you can see like we have three type of

7:29information technical information uh

7:32semantic business semantic information

7:35and also So governments security related

7:38stuff. Okay. First column is order ID.

7:42Logical type it's string and that's a

7:44primary key. Customer ID business name

7:49customer identifier. Logical type

7:52string. Physical type var. We will come

7:55back to this part because logical type

7:58can be strict but physical type can be

7:59also other type. [snorts] It's required

8:03column. This is also important because

8:05we have to explicitly say don't delete

8:07it. This is a required column. Some

8:10example of data and classification. It

8:14can be internal, it can be public, it

8:16can be confidential data. Um in case in

8:19this case we are calling it internal

8:21because it can ina it can be used inside

8:24of company. In case of it could be

8:26public we could share it with external

8:28systems. For example, in my company, we

8:30have kind of data that even incre

8:35only concrete people can have access

8:37especially some financial data and so on

8:40and taking it as PII true this is

8:42sensitive data and further it could be

8:46masked or somehow changed for during

8:49further usage.

8:52Let's see another example. This order

8:54totals logical type integer physical

8:58type also integer description total

9:01order amount in sense. First let me

9:04mention this one is not directly

9:05connected to data contracts but quite

9:07interesting if you are working with

9:08financial data. Why not decimal? Why not

9:10float? Uh why in sense? Because there a

9:13floating point floating floating point

9:16error. And when you store data one data

9:19let's say 0.1 or 0.01 01 uh in decimal

9:23format um it cannot perfectly represent

9:26it as binary in the database. When you

9:29have it one you don't even realize it.

9:31But when you have let's say some tax

9:34records and there's millions of records

9:38it's

9:39aggregates and in the end you end up

9:41that number is not correct especially

9:43for sensitive financial data this

9:46getting a big problem. This is why it

9:48recommended to store it in integer. But

9:50what I wanted to mention also logical

9:52type options. This is kind of ch

9:56construction um constraint sorry uh and

10:00means that number cannot be negative. It

10:02only should be positive because money

10:04cannot be negative and this can be used

10:06for other cases as well.

10:10this uh remember two slides ago I said

10:13take attention but logical type can be

10:15string but physically it can be stored

10:17as a text the difference text is much

10:19bigger than vchart uh but sometimes we

10:21can have also UU ID uh logically it's

10:24still a string but it's stored as UU ID

10:28for optimization perspective and it's

10:29probably used as primary key for primary

10:31keys and example is shipped just text

10:38uh data quality

10:40in data quality from first glance you

10:43can recognize there's a two ways to do

10:46to run this data quality check one is

10:48table level or column level let's start

10:52with column level you can see the metric

10:54invalid values and pending shipped

10:58cancel it the statuses

11:00must be zero what this means like there

11:03cannot be other values there cannot be

11:05errors it should has only pending ship

11:08and cancel it. When this kind of error

11:11can happen, if you are working with user

11:13data, just aggregating after aggregation

11:15and group by you can see unknown,

11:18undefined or this kind of and with this

11:21way we are restricting. No, we will not

11:24um admit any other value except of this

11:27this kind of constraint that we can we

11:29can put on this column or just table

11:32level quality check. We are saying

11:34minimum there should be 100,000 rows. Uh

11:37not much actually. I remember for

11:39example when I was working with machine

11:40learning data we needed big amount of

11:43data like if there is a 2,000 you could

11:45not train your m classic classical

11:48machine learning models with this one

11:50like we could also put a constraint that

11:53we accept in only cases when it's more

11:55than 100,000

11:59u team

12:02under team we are just describing who is

12:05owner of the data

12:08uh and also support like it can be

12:10channels uh teams channel, discord

12:13channel, slack channel

12:16from first class uh for small companies

12:19this maybe doesn't look useful because

12:22everyone knows each other and it's not a

12:24problem but when you are working in big

12:26organization and there's a lot of people

12:29u for some moment you cannot know whom

12:32to address when there's a problem and

12:34this can be very useful

12:38Yeah.

12:40Terms of use,

12:43uh, purpose, limits and usage,

12:48purpose and usage, how this data can be

12:50used. Imagine software engineer decided

12:52to refactor table. But he should know

12:55who who is using it and for what

12:57reasons. And this can be very useful to

13:00identify like when you're going to make

13:02this change whom it could can impact and

13:07uh yeah used by analytics team for sales

13:10analysis.

13:12But also very interesting little

13:16hidden kind of issue that if you run

13:18query it will not fail but it can be an

13:21logical issue limitation contains only

13:24the last two years of data. Now imagine

13:26you have to create trend of revenues for

13:29last 5 years. You are building this

13:30dashboard. You are running SQL query for

13:33last 5 years. It will not fail. It will

13:35create for only two years.

13:37Technically it works. People can go and

13:39check. But what happens next? Like

13:42are we sure that there is no 3 years

13:45data because someone just deleted it or

13:49it's okay for business to have this last

13:51two years data? That's a big question

13:53and can confuse end user and I think uh

13:57having it explicitly defined that here

13:59it's only two years data could be very

14:01useful further in analytics

14:03customer properties sensitivity and

14:06secrets again like we are saying that

14:08this is secret data don't share or share

14:10with just concrete group of people

14:14uh SLAs

14:17in SLAs we are defining availability 90

14:19and 9% 99 and 9% And what it means

14:23definitely we have database and we can

14:25say it can be off for maintenance only

14:2933 minutes a month. Knowing this

14:31information can be useful when you

14:34create connection to database and will

14:36not end up with error. Like we

14:38specifying that there's a moment during

14:40the months or during time that it can be

14:42off for technical reason.

14:44Retention again we are going back to

14:46previous slide. Uh it's kind of

14:48limitation. We are saying we only have

14:50one year data. We are keeping only one

14:52year data.

14:55Support is business hours only during

14:57business uh support done only business

14:59hours. Now these three support,

15:01retention and availability they

15:03descriptive like we can just put it in

15:06the document to just describe the

15:08situation. But freshness it not just

15:11description it has it runs as a it has

15:14functional uh

15:18it's a kind of function that runs and

15:19can fail the test process because what

15:23it does actually it use order state

15:27current time stamp minus maximum order

15:29state and if it's going to be more than

15:3224 hours we have a lagging data this

15:35very frequently by the way happens in

15:37our in our company this happens because

15:39of sometimes failed ducks. Uh we are

15:41using DBT freshness check. It it was

15:43also mentioned like few presentations

15:45ago. Uh but again very useful in

15:48practice and you can take control that

15:51your data is fresh.

15:55Last one but a very important one that

15:58servers

15:59in this case is progress servers and

16:02defining all parameters. How we use it?

16:04Uh we have environment variables for

16:08production database and development

16:09database par with parameters and we can

16:12switch based on later we will see in the

16:14Python script how we doing it but yeah

16:17this can be used to connect to database

16:19to concrete database

16:24when you open data contract platform it

16:26has option that you can go and use

16:29interface to create your uh data

16:32contracts it's quite comfortable you

16:34just set and play with interface and

16:35generate this YAML file. But here's a

16:38trick like when we started using it, we

16:41had 300 tables to onboard. Now I think

16:44it's hard to imagine that I'm sitting

16:45and doing it for 300 tables here. It

16:48could be problem and this is time to

16:50mention AI finally that I could also

16:52mention in my presentation. I used Cors

16:56and the trick is was following in one

16:58folder I put few examples. I created few

17:01examples for concrete source in another

17:04folder. I just put details definitions

17:06of tables and wrote some prompt and

17:08saying here's examples here's a lot of

17:12tables generate data contracts for us

17:15and instead of like spending days and

17:17this manual routine work that's quite

17:19very boring I would say I even told my

17:21managers thanks AI it could be very

17:22boring for me playing with this

17:252 minutes 2 3 minutes and it's done

17:29there was a cases when AI little failed

17:32especially I think it's called medium

17:34minimum size, large size text like these

17:36kind of formats for some specific

17:38databases. But again, you are running

17:40tests, it fails, you're just checking

17:41manually. Instead of 300, maybe five or

17:4410 data contracts you should check.

17:48Okay, we have YAML file. We know what is

17:51data contract. What's Nate next? How to

17:53run it?

17:58Three different options to run data uh

18:00this test data contract. One is CLI

18:04opensource command line tool. As you can

18:06see it supports a lot of different uh

18:09databases uh platforms AWS glue,

18:13bigquery, iceberg.

18:16We now company using poss, MySQL, MSSQL

18:20that actually we use it but I would

18:22recom like list is very big. I could

18:24like continue for five minutes just

18:26saying what tools they use there. But

18:28you can go to this link and find if it's

18:30also appropriate for your company and

18:33yeah tool set you are using

18:37just example how output could look

18:41check that field line item present and

18:44test passed it has the type UU ID it's

18:48also passed sometimes they can just

18:50change UU ID to ID let's say so and it

18:53breaks because when ID is integer UU ID

18:56is other format and Especially when you

18:58are doing range sometime it can fail.

19:00This also important remaining is quite

19:03same but yeah has no missing values

19:07present and blah blah blah data. Uh this

19:10is the best part when you see data

19:13contract is valid. There's nothing to do

19:15with this.

19:38if you have big amount of data contracts

19:42and docker version um there is a docker

19:45image officially just need to pull and

19:47run also cool if uh there's a

19:50dependencies or something if you don't

19:52want to think about all of these

19:54problems just want to run I personally

19:56running with the when I testing locally

19:58in my laptop I'm using docker version

20:03okay we are the moment when we know how

20:06to run it we know this very cool tool

20:08exist nice but how we use it in indust

20:11because again we have a lot of projects

20:14and there was a lot of questions who is

20:16actually I will change the next slide

20:17who is going to responsible for like

20:19failed uh failed data contract it's

20:22going to be that software developer who

20:24is working in that team or it's going to

20:26be maybe analysts that should also

20:28update and push on the GitHub or it's

20:30going to be me like actually end up with

20:32me like [laughter] it's me who does it

20:34and yeah I am updating these data

20:37contracts

20:38Another question

20:40how we integrate it in our system. And

20:42they said 10 more than 10 projects.

20:44Should we refactor all these projects or

20:46how we will run it? Maybe on data

20:48engineering side I each time should run

20:50and check this databases if something

20:51changed

20:53>> [snorts]

20:53>> uh on schedule way. But after we made it

20:56event based which is more practical like

20:58as soon as this change happens it reacts

21:00instead of like running on schedule at

21:02daily basis let's say. So and this is

21:05also critical for me coordinating

21:07collaboration between developers

21:09engineers and analysts.

21:12three different teams and manager at

21:15different managers not the same manager

21:18we have to work together by the way they

21:20have their schedules their tasks and you

21:22know like one says I will do it after 2

21:24days and another one says I have more

21:26priority tasks and it it also cost cows

21:29we needed some solution for this one

21:31okay let's see how we created this data

21:35contract repository first actually you

21:37can see it's very easy this is

21:40screenshot of our repository just

21:42contracts folder with YAML files uh

21:46CI/CD YAML files GitLab CI/CD YAML files

21:49docker that's for Python with some

21:52dependencies and main pi uh main Python

21:56script that actually does all the job

22:03this is subfolders I wanted to show I

22:06wanted to show when I started my

22:08presentation I mentioned like we have

22:10This people hub let's say for internal

22:12HR software

22:14there is a mapping between this folder

22:17and the repository name like one folder

22:21name equal to one repository in the

22:23company and peoplehub contains YAML

22:26files separate YAML files for each

22:28table. There was idea maybe to keep all

22:31tables under one YAML file, but it could

22:34get very hard and mess because imagine

22:37having 20 tables under one YAML file. It

22:40could be hard. instead just having small

22:43but a lot of yaml files

22:46and mainpi mainpi script what it does it

22:50takes two arguments repository repo name

22:54and environment production or

22:55development because they're different

22:57databases actually depends on the

22:59environment.

23:00It runs data contracts tests and when it

23:04fails it sends email to ticketing system

23:07and automatically opens ticket

23:11This how it looks like. Data contract.

23:14This real scenario. This from real

23:16problem that I just made made a

23:18screenshot to show you. Zabix alert

23:22validation failed for stop project in

23:24the environment.

23:26Ticket owner first they uh assigned it

23:30to the to the teams. Um in this case

23:33Victoria we actually it assigns to

23:35engineer. I will mention it as well.

23:37Data quality validation. uh it failed

23:40and here's the reason why it fails, why

23:41it failed.

23:46Before I will continue, I want to say

23:48like when it fail

23:50this presentation I it will go even more

23:53detail in this

23:55how we integrated it with 10 projects

23:5810 and plus projects. This is some

24:01screenshot from CI Gitlab CI file. For

24:05each project, DevOps put these jobs

24:10based on the condition like CI commit

24:12branch should be equal to develop

24:14uh or this is for pro like main branch.

24:18It triggers

24:20another uh GitHub repository and runs

24:24pipeline in another repository. And here

24:27we have defined when the pipeline runs

24:29this job, it just triggers main pine

24:32main pine script.

24:34You see variable this is this sense uh

24:38pro repository name and this says

24:40environment like absolutely separate

24:43project it just run it it's not they not

24:46dependent each other loosely coupled

24:48system I would say like each working

24:50independently

24:52[snorts] and yeah it runs and checks

24:54okay for dev database for this project

24:56there's something changed run test

24:58failed or not

25:00and same actually for production job.

25:05Now important this coordination and how

25:08this done if it fails. First ticket uh

25:12if it fails on dev ticket assigned to

25:15their dev because changes actually in

25:18practice we found found there are two

25:20type of changes intentional they know

25:24what what they do and they really doing

25:26some refactoring or accidental temporary

25:29the developer was playing thinking okay

25:31this kind of schema could be even better

25:33than previous one wasn't necessary but

25:35he was playing in the end breaking

25:37everything like and when they received

25:40is they just roll back system and says

25:43okay we don't need to change

25:46even without disturbing data team

25:52if it's intentional and they know that

25:55they are changing something in the

25:56database then we are requesting from

25:59them to provide context

26:02why they in this case the context could

26:05be related for to to data engineer

26:08because I will know what's changing

26:10for data analysts sometimes they're

26:12changing generally logic or logic of the

26:14data and SQL queries DBT models they

26:18should completely refactor it because

26:20their logic is changing and they are

26:22providing this information under the

26:24ticket

26:26here's example creation date a real real

26:30case creation date was deprecated and

26:34back end was replacing it with created

26:36it

26:38do it to time zone standard

26:39standardization what's interesting

26:41imagine now they were storing data

26:44creation date in Lithuania time zone

26:48after they changing it to created at UTC

26:50time zone

26:52in aggregation everything going to be

26:54wrong numbers going to be wrong because

26:56now they are storing it different time

26:57zones and if not letting know analysts

27:02that this kind of change happened it

27:04could be disaster like for a lot of

27:06dashboard it's going to be like

27:07absolutely other numbers

27:09This why it can be also very useful.

27:13what should I do when receiving such a

27:16ticket and

27:18what we say if it's not big problem just

27:22continue deploying if it's or sometime

27:25we can request time uh to have meeting

27:27internally with analyst and engineers

27:29and decide how we will continue because

27:31it will it can require a lot of

27:33refactorings on our side before they

27:35will deploy it

27:37but some exceptions happens hot fixes

27:40some projects uh used by a lot of users

27:43like this estimators sitting in

27:46different offices doing this estimate

27:48for clients and if there's some big fail

27:51some problem happens they cannot sit and

27:52wait until analyst will discuss for a

27:54few days and do this um yeah

27:57brainstorming they immediately deploying

28:00this and we are fixing fixing it postf

28:02facto actually what was before data

28:05contracts quite frequently we are we

28:07were fixing it postfactum when

28:08everything started failing

28:12Yeah.

28:14Now the question, do we need to create

28:18data contracts for wall tables? No, it's

28:21not necessary. You can create it for uh

28:23most important ones. What we call

28:26important in my case in dbt there is a

28:28manifest JSON and if you get it and

28:31analyze you can find which tables

28:33actively used in the models and create

28:35statistics. This approach that I used

28:38and we got that some tables very

28:39actively used. By the way, even now I

28:41have one pending task they are changing

28:43that table and I said stop I'm on day

28:46off I'm in year one I will come back on

28:48Monday and I will fix it like this way

28:51we stopped it without that if they could

28:53change that projects table it could

28:56break almost all dashboard it's actively

28:58used this way you can on board but again

29:01with AI you can on board quite easily 50

29:05or 100 tables

29:08do we need to 35.

29:12Okay. And in the end, I want to express

29:15express my gratitude to my colleagues

29:17because I didn't work on this alone. Uh

29:19Yea, our team lead, and also Thomas

29:22DevOps. Uh thanks also them and we work

29:25together. I think it also Yeah. And

29:29thank you very much for attending.

29:31[applause]

29:34Miss Big Pleasure, I would like to

29:36answer on your questions.

29:38>> Thank you very much.

29:41Uh yeah have questions.

29:49>> Thanks Rolf. It's very close for me as a

29:52data engineer also. Uh my question is

29:55about the PII how you how you taking PII

29:58and who responsibility it is and what

30:02you do with the complex data types like

30:04JSON for example right from in some time

30:07to time uh PI data sensitive data using

30:10that kind of complex types right

30:12>> could you repeat first part API what

30:14>> uh about sensitive sensitive data who

30:17attack data whose responsibility is it

30:20>> interesting fact now we don't have such

30:22a like we have one sensitive data but

30:25actually devops state I will not give it

30:27to you right they're kind of cutting

30:29this table to three column like from

30:31five columns they're just cutting it and

30:34giving only three that's from financial

30:36data that could be very important but

30:39I'm not getting it and remaining data

30:41it's not sensitive and we don't have

30:43people who is actually doing it got it

30:45okay thank

30:55Thank you for very interesting talk. Uh

30:58the system of uh data contracts strikes

31:02me as uh being overlapping with uh the

31:06systems that are recently in fashion

31:08called semantic layers such as DBT

31:11semantic layer because it uh portrays

31:15also information about whatever this

31:17thing means for the business etc etc. uh

31:21do you see any overlap between these

31:23approaches and any synergy that might

31:26result in using both? Thank you.

31:29>> We are using both. Thank you first of

31:31all thank you for question and we

31:32already use semantic layer the our team

31:35working on it. uh what's difference um

31:39in our ca I showed you all the cases but

31:42actually we are not using all these

31:44quality checks because freshness check

31:46for example we implemented with dbt and

31:49data contracts mostly we use to specify

31:52column names and column types and just

31:55check if it exist or not whatever code

31:57could break it's kind of before dbt

32:00before data warehouse uh we check first

32:03and after run it what you mentioned at

32:05semantic It's kind of end of the story

32:07already when you have already and our

32:10analyst now defining these formulas and

32:14then yes it can it will be used in

32:15dashboards in calculations I can't

32:17provide more information because I'm not

32:19working on it but what I can say if you

32:21see the chain now here is a semantic

32:24layer and and here is a data contracts

32:27first.

32:30>> Yeah thank you. Um I have this like

32:33future looking the next step question uh

32:36and I would like your opinion on that.

32:38Do do you see a world where like the

32:42contracts are agents and like they

32:46negotiate between each other when the

32:49change happens and they try to negotiate

32:52and then sort it out before bringing

32:54human in a loop as as as much as

32:56possible. Like a mechanism could be for

33:00example

33:02uh there is a change in a contract and

33:03it broadcast the what the changes in the

33:06blast radius of that and they will other

33:09contractors subscribe to that and see if

33:11it's pertinent to them and then they

33:13will take account what changes has made

33:15and then incorporate that and there's

33:17like a peer-to-peer mechanism between

33:19the you know different departments and

33:21different contracts. I think my vision

33:23that I completely think that it could

33:25happen because I was checking stack

33:27overflow again we discussed it like just

33:29three or four year passed and stack

33:31overflow now seems like a museum you

33:33know it was dinar's era but just few

33:35years passed and with active development

33:38or of agents I think everything is

33:41possible and especially I have seen

33:43cases when the team integrated agents to

33:47analyze airflow logs and when something

33:50significant were happening I

33:52against quite quickly was reacting on it

33:55and yeah completely agree it's possible

33:58and in our company we also that was a

34:01question about DBT they have this agents

34:03and we already testing it uh what's the

34:06difference they fails for nuances like

34:09it it works good but when there's some

34:11nuances in filters or something it can

34:13fail not always but it can fail and we

34:15are exactly testing human versus against

34:19and data quality

34:24Thank you.

34:25>> So thank you. Uh I'm not a data

34:27engineer. So I'm uh I'm a I I feel I

34:32have to ask this question. You are uh

34:35constrating the uh software engineers

34:38not to alter the database by data

34:40contracts if I'm not mistaken.

34:44What what

34:45>> you are constraint constrainting the

34:49software engineers not to alter the

34:51database uh without authorization if I'm

34:55not mistaken.

34:56>> By the way, it's good question because

34:57we were discussing should we block

34:59someone or not block and like working

35:02process and this is why I say dev and

35:04pro mode depends like if it's in dev we

35:07just notifying them. Actually this kind

35:09of early notification system it's not

35:11it's not blocks anyone. Uh

35:13>> but we getting it earlier and can react

35:17instead of getting this panic everything

35:19fails quickly do it. It's a lot of

35:21stress on the team and also delays on

35:23reports. Instead we could get it earlier

35:26and can

35:27>> a followup question if I'm allowed

35:30follow

35:32>> a followup question.

35:33>> Yeah. Yeah.

35:34Do you think can you imagine uh a world

35:38that we have some uh contracts not just

35:42for the data but for the uh everything

35:45that AI generates

35:47including the code that we will say okay

35:50we will not read your code we will we

35:52don't care about what you are doing you

35:54have to be following not following but

35:57you have to meet this criteria you have

35:59to pass this unit test you have to have

36:01this kind of uh measurements And this is

36:04your contract. Follow this contract. Do

36:06whatever you want. Write in any language

36:09you want. Do you think it's possible? By

36:11the way, uh this uh data contracts is

36:14very universal because doesn't matter

36:16what language you're using and you can

36:18this YAML file you can do it with PHP

36:21you can do with Python like with Docker

36:23especially Docker version you just put

36:25in YAML file and it does for you like it

36:27doesn't matter

36:29environment like Python version

36:30officially supported but for other cases

36:33we were thinking maybe each team will

36:34integrate this test for them but after

36:37created this independent system just

36:39separate repository that runs

36:41Because before was the idea why if we

36:43put for each project as a part of like

36:46on the test stage uh this um data

36:49contracts with Python it was very

36:51problematic because some projects

36:53implemented in PHP but with this docker

36:55image again it's independent like you

36:58can just run it doesn't matter like what

37:00language it is if it related about

37:03independency of language but yeah with

37:04docker we reach this independency

37:07>> okay could do whatever they want

37:09>> against I I I think um based on the how

37:13the trends going I believe that in the

37:15future will do whatever they need.

37:17>> I hope it will we will have some delay

37:19because I don't want to lose my job.

37:22[laughter]

37:23I believe in good future but let's have

37:26a little delay.

37:31>> Yeah. Uh I wanted to ask if you're

37:33familiar with uh data observability

37:36tools like Monte Carlo and which do this

37:40proactive monitoring of changes and data

37:42changes and how you think this uh two

37:46approaches collaborate because this

37:48approach is preventing or notifying

37:50whenever change is happening once on the

37:52code level but the data observability

37:55tools like they're constantly

37:57monitoring. So if something happens you

37:59get notified or you get alert and then

38:01you get a sign of now also AI boat which

38:04helps you identify what changed. Do you

38:07use that as well or not?

38:09>> Uh not we are not using it but one you

38:11are from service titan maybe.

38:14>> Yeah because I had this conversation

38:15yesterday that

38:19>> with someone I forgot your I forgot your

38:21name but

38:22>> she was representing Yeah. She was from

38:25a service titan and I didn't know Monte.

38:28>> Yeah. Yeah. Yeah.

38:29Uh first of all I didn't know about that

38:31tool because it was where like Monte

38:33Carlo probably it's for statistic like

38:35simulation and I was little confused how

38:38Monte Carlo related to data like this

38:40way but what interesting why I don't

38:43know I was thinking okay this really

38:44exists why I didn't know about it the

38:47it's mostly accidentally when we had

38:50this data contract issue and when my lit

38:53uh was on this conference she bring this

38:56idea me at the moment when we had this

38:58issue I implemented it worked and we

39:00didn't continue like exploring other

39:02tools just accidentally these multiple

39:04factors happened together but what your

39:07colleague yesterday said Monte Carlo is

39:09not cheap tool it's expensive tool as

39:11far as I know

39:12>> depends how you use

39:13>> yeah but data contract is zero dollar

39:17>> yeah just follow up uh I trying to

39:20understand what is the best approach

39:22even it can be not multi there are open

39:25source versions and you can host it but

39:26the idea do you need after this do you

39:29need uh proactive monitoring of these

39:32changes or just what's in place is

39:35enough or you are able to uh catch

39:38everything with just this

39:40>> like currently what I see in company we

39:42don't need anything more especially this

39:43is for free we just use it

39:46>> forget about

39:47>> no no no I mean like this is okay this

39:49works for us we don't have any problem

39:51it really does what it should do like

39:53send us notification and we are reacting

39:56uh and also this manifesto that we sent

39:58across all teams that how we are going

40:00to react on it. is important because

40:02first when I integrated it a lot of

40:05developers were panicking okay it felt

40:07what should I do like no one know knew

40:10what they should do and they were

40:12texting me in panic okay I stuck but

40:14after preparing all this document pend

40:17actually little practice like failed and

40:19fixed everyone became little experienced

40:22and we are quite well managing these

40:24situations

40:25>> reactive not

40:27>> proactive yeah

40:28>> this one is productive

40:30>> this one not actively checks data. Yeah,

40:33we try to prevent and obser

40:40>> it's not no it this checks only the

40:43moment

40:44>> it happens after they did some change

40:46but maybe we can continue after

40:48>> yeah no we can

40:49>> I just I just want to show this one

40:51because

40:52this triggers only in two cases this

40:54kind of event based when they commit on

40:57develop or on pro main branch two cases

41:01it's not constantly It's not checks

41:02database like active

41:04>> if somebody if somebody goes manually

41:06changes drops the fields then you don't

41:08catch it right. Yes, we don't will not c

41:10but actually it's never like last two or

41:13three years it never happened. So good

41:15that you could disagree that

41:18>> like this very frequently happened what

41:20I mention it but that someone deletes

41:23actually I am quite actively saying guys

41:25if you want to work with table get read

41:26only access because it's good for you

41:28good for us and you know like if

41:30something happens you are calm that it's

41:32not because of you you have just read

41:34only access to the table

41:38that's actually

41:42we have like a

41:45It's nice actually. Yeah.

41:47>> The best part.

41:48>> Yeah. Thanks a lot for your insights. Uh

41:50I wanted to ask uh so whatever you

41:52presented relates uh to the tables. Uh

41:55do you think there are data contracts or

41:58something similar for other data

42:00modalities let's say uh images audio

42:04video and how they can look like if

42:07there are such things.

42:09I don't know again better to check I

42:11don't know if took took a picture of it

42:13but on their website there's a big table

42:15on GitHub mentioning everything and I

42:18personally not working with this kind of

42:20data but actually could be interesting

42:22because now we have AI team who is

42:25working especially with image processing

42:27yeah

42:28>> thank you

42:30question

42:34>> so my question is like during

42:37integration of this process

42:39Like you already mentioned that there

42:41was a kind of resistance or panic or

42:43people just

42:44>> don't want to do and what about like the

42:47current state of is there like uh like

42:51people who sabotaging like how you will

42:54deal is there just attempts to say okay

42:57let's don't do it it slows us down. So

42:59in terms of like human perspective, how

43:02difficult it to maintain? Okay, you have

43:04it now, but do you have like situations

43:06when they start to question like why we

43:08should keep this?

43:10>> Yeah, like it slows down. It was easier

43:12to tell our CEO like you know if we

43:15integrate it you will get a big benefit.

43:18If not like you will not see reports. It

43:20solved all the downstreaming problems.

43:23But in general we have clear plan like

43:25it fails ticket assigned to you and

43:28after discussion starting it's not about

43:30I wish or I not wish no we have strict

43:32uh requirement it fails ticket on you

43:35and we should work again we are showing

43:37that downstreaming issues like we had it

43:41a lot of times I hope I'm happy that

43:44actually engineers understanding but

43:46again presentation to co was mostly

43:50major part of this Any other questions?

43:56No. Only thing I learned today that you

43:59can easily spend a lot of money in Monte

44:02Carlo.

44:03>> That's right. There's a contract.

44:05>> Thank you very much.

44:06>> Thank you very much for attending. This

44:07game.

More from PyData

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.