Full transcript
0:05Thank you. Thank you very much. First
0:08uh thank you for attending my
0:10presentation. I hope you will find it
0:12useful for you. Uh as you can see I'm
0:15going to talk about yeah data contracts.
0:18But first let me introduce myself.
0:21Uh recently two patients drive me uh
0:25building data platforms and pushing my
0:27limits in the water. One month ago I was
0:29capable to swim finally 5 kilometers in
0:31the sea in degrees and next month I hope
0:35to cross continents in Istanbul post
0:37forest trait and that also cool
0:40experience it's about cows in your head
0:41and how you control your emotions quite
0:44cool experience I would recommend you
0:46and let's switch back to my technical
0:48background in 2016 I joined Georgian
0:51startup Pulsar AI where I created first
0:53chatbot framework for Georgian linkage
0:55it was 2017 already and actually worked
0:59on different AI tools like social med
1:01media monitoring tools, sentiment
1:03analysis tools. Uh later I continued uh
1:06my career in Doppel freelance network.
1:09Probably you worry about it. It was
1:11quite long time and I worked on
1:13different scale projects from Fortune
1:15500 companies like media and to the
1:18smaller scale startup or hedge funds and
1:20since April 2024 I continued my career
1:24journey in indust windows is Lithuan US
1:26Lithuanian company uh door and windows
1:30manufacturing company and my first
1:32initial project was building the modern
1:35data infrastructure for them.
1:39Well, in indus we have different
1:42internal teams and different projects.
1:45uh production planning as name says like
1:48that software that helps uh ter people
1:51to plan these terminals and what's going
1:53inside in manufacturer uh internal HR
1:57software or we call it people hub it's
2:00for HR and drawing and calculating
2:03software um
2:06are different estimators or drafters
2:08they using it uh for estimate window
2:10size uh fix this parameters and do a lot
2:13of calculations because it kind of done
2:15for very big buildings sometimes for
2:17skyscrapers.
2:19uh this is just example of three
2:21projects but actually there's a much
2:23more why I'm talking about these
2:25projects because I built modern in data
2:28infrastructure I started collecting this
2:30data into data warehouse and after I run
2:33into very typical problem that happens
2:35in a lot of organizations
2:38software engineers start modifying the
2:41transactional databases schemas without
2:44letting us know and without thinking
2:47that it can break all downstreaming
2:50processes
2:52data pipelines was failing. DBT models
2:55also started failing or just not working
2:59and re reports they get delayed and in
3:02the end our end user managers were
3:05calling us and saying we need this
3:06report but it's getting delayed. We had
3:10this problem and fortunately for that
3:12moment uh my team in Lithuania they
3:15attended biggest data conference. It was
3:17also in Vnus and there was nice speech
3:19by Yenni from Zando and he was sharing
3:23um their experience how they fix this
3:25problem in Zando with data contracts
3:29and why we decided why not why not to
3:32try use this experience in our company
3:35and try maybe it will help us to
3:37coordinate uh working with different
3:39teams and finally fix this issue because
3:42before it was verbally agreement or
3:44different ways but it wasn't solution
3:47What is a data contract? Data contracts
3:49is machine machine readable document
3:51actually YAML file um that explicitly
3:54defines a structure, format, semantics,
3:57data quality, SLAs, term of use for
4:00exchanging data between producer and
4:02consumer producer that's these teams and
4:04we data team we are consuming it.
4:07Very key point in Yani's presentation
4:10called trust. But what what this says
4:13from official document instead of hoping
4:16now we have explicit agreement both
4:18sides having explicit agreement and
4:21continuously enforce it.
4:23Let's see further how does it look like
4:25actually it's bigger file I just split
4:27it into the chunks to go through it and
4:29show you per line what it means and how
4:32it works. First of that's fundamentals I
4:36call it this kind of passport of the
4:38data data contract. uh that's API
4:41version
4:43data contract not data contracts but
4:45software's API version that's going to
4:47process this because we had a case when
4:49you have YAML file if version changed it
4:52will not be able to process it and it
4:54fails like we have to follow this API
4:56version kind just explicitly defining
4:58the data the data contract ID unique
5:01identifier uh in this case it's orders
5:04because you can have thousands of data
5:06contract files uh name orders uh semi IC
5:10version uh this important and I will
5:12show you on the next slide how you can
5:14work with this and status of data
5:17contract in this case it's active
5:22semantic versioning in this case there's
5:24a three parameters that we could control
5:27first patch or you can see third
5:29parameter uh that most right one uh what
5:35it means no data impact sometimes in
5:38data YAML file You can just update
5:41definition of the column or just add
5:43some description. This has nothing to do
5:45with data. It's just descriptive
5:47information and actually will not break
5:49anything.
5:50Next one. Next parameter is minor and
5:53non-breaking change. Uh that's in middle
5:56number. This one
6:00um quite frequently even now during
6:02working process uh back end teams doing
6:04refactorings and adding new columns.
6:08uh table getting uh wider like bigger
6:10but it will not impact on us because the
6:13used columns is not changing only new
6:15columns is coming and that's not a
6:16problem and this is a queen or I don't
6:21know king [laughter]
6:23this what breaks everything like if back
6:26end team decides to refactor it and
6:29exactly deleting some columns or
6:31changing their names or changing formats
6:34it breaks everything
6:37and There's a teams they kind of keeping
6:38us updated that you know we are going to
6:40change but there's a teams that can just
6:43deploy it and say it works on our side
6:45and we little less care what's going on
6:47next and yeah this is major one
6:53okay next uh this is Kimma if previous
6:55slide I was calling it uh passport
7:00uh if previous slide was passport this
7:03is heart of our data contract as you can
7:06see like this uh physical type uh first
7:10of all it's name the name table name
7:13orders uh physical type is a table
7:16description
7:17information about table all the web shop
7:19orders actually it could be much longer
7:21I just cut it to uh fit this slide size
7:25and properties actually columns and here
7:28you can see like we have three type of
7:29information technical information uh
7:32semantic business semantic information
7:35and also So governments security related
7:38stuff. Okay. First column is order ID.
7:42Logical type it's string and that's a
7:44primary key. Customer ID business name
7:49customer identifier. Logical type
7:52string. Physical type var. We will come
7:55back to this part because logical type
7:58can be strict but physical type can be
7:59also other type. [snorts] It's required
8:03column. This is also important because
8:05we have to explicitly say don't delete
8:07it. This is a required column. Some
8:10example of data and classification. It
8:14can be internal, it can be public, it
8:16can be confidential data. Um in case in
8:19this case we are calling it internal
8:21because it can ina it can be used inside
8:24of company. In case of it could be
8:26public we could share it with external
8:28systems. For example, in my company, we
8:30have kind of data that even incre
8:35only concrete people can have access
8:37especially some financial data and so on
8:40and taking it as PII true this is
8:42sensitive data and further it could be
8:46masked or somehow changed for during
8:49further usage.
8:52Let's see another example. This order
8:54totals logical type integer physical
8:58type also integer description total
9:01order amount in sense. First let me
9:04mention this one is not directly
9:05connected to data contracts but quite
9:07interesting if you are working with
9:08financial data. Why not decimal? Why not
9:10float? Uh why in sense? Because there a
9:13floating point floating floating point
9:16error. And when you store data one data
9:19let's say 0.1 or 0.01 01 uh in decimal
9:23format um it cannot perfectly represent
9:26it as binary in the database. When you
9:29have it one you don't even realize it.
9:31But when you have let's say some tax
9:34records and there's millions of records
9:38it's
9:39aggregates and in the end you end up
9:41that number is not correct especially
9:43for sensitive financial data this
9:46getting a big problem. This is why it
9:48recommended to store it in integer. But
9:50what I wanted to mention also logical
9:52type options. This is kind of ch
9:56construction um constraint sorry uh and
10:00means that number cannot be negative. It
10:02only should be positive because money
10:04cannot be negative and this can be used
10:06for other cases as well.
10:10this uh remember two slides ago I said
10:13take attention but logical type can be
10:15string but physically it can be stored
10:17as a text the difference text is much
10:19bigger than vchart uh but sometimes we
10:21can have also UU ID uh logically it's
10:24still a string but it's stored as UU ID
10:28for optimization perspective and it's
10:29probably used as primary key for primary
10:31keys and example is shipped just text
10:38uh data quality
10:40in data quality from first glance you
10:43can recognize there's a two ways to do
10:46to run this data quality check one is
10:48table level or column level let's start
10:52with column level you can see the metric
10:54invalid values and pending shipped
10:58cancel it the statuses
11:00must be zero what this means like there
11:03cannot be other values there cannot be
11:05errors it should has only pending ship
11:08and cancel it. When this kind of error
11:11can happen, if you are working with user
11:13data, just aggregating after aggregation
11:15and group by you can see unknown,
11:18undefined or this kind of and with this
11:21way we are restricting. No, we will not
11:24um admit any other value except of this
11:27this kind of constraint that we can we
11:29can put on this column or just table
11:32level quality check. We are saying
11:34minimum there should be 100,000 rows. Uh
11:37not much actually. I remember for
11:39example when I was working with machine
11:40learning data we needed big amount of
11:43data like if there is a 2,000 you could
11:45not train your m classic classical
11:48machine learning models with this one
11:50like we could also put a constraint that
11:53we accept in only cases when it's more
11:55than 100,000
11:59u team
12:02under team we are just describing who is
12:05owner of the data
12:08uh and also support like it can be
12:10channels uh teams channel, discord
12:13channel, slack channel
12:16from first class uh for small companies
12:19this maybe doesn't look useful because
12:22everyone knows each other and it's not a
12:24problem but when you are working in big
12:26organization and there's a lot of people
12:29u for some moment you cannot know whom
12:32to address when there's a problem and
12:34this can be very useful
12:38Yeah.
12:40Terms of use,
12:43uh, purpose, limits and usage,
12:48purpose and usage, how this data can be
12:50used. Imagine software engineer decided
12:52to refactor table. But he should know
12:55who who is using it and for what
12:57reasons. And this can be very useful to
13:00identify like when you're going to make
13:02this change whom it could can impact and
13:07uh yeah used by analytics team for sales
13:10analysis.
13:12But also very interesting little
13:16hidden kind of issue that if you run
13:18query it will not fail but it can be an
13:21logical issue limitation contains only
13:24the last two years of data. Now imagine
13:26you have to create trend of revenues for
13:29last 5 years. You are building this
13:30dashboard. You are running SQL query for
13:33last 5 years. It will not fail. It will
13:35create for only two years.
13:37Technically it works. People can go and
13:39check. But what happens next? Like
13:42are we sure that there is no 3 years
13:45data because someone just deleted it or
13:49it's okay for business to have this last
13:51two years data? That's a big question
13:53and can confuse end user and I think uh
13:57having it explicitly defined that here
13:59it's only two years data could be very
14:01useful further in analytics
14:03customer properties sensitivity and
14:06secrets again like we are saying that
14:08this is secret data don't share or share
14:10with just concrete group of people
14:14uh SLAs
14:17in SLAs we are defining availability 90
14:19and 9% 99 and 9% And what it means
14:23definitely we have database and we can
14:25say it can be off for maintenance only
14:2933 minutes a month. Knowing this
14:31information can be useful when you
14:34create connection to database and will
14:36not end up with error. Like we
14:38specifying that there's a moment during
14:40the months or during time that it can be
14:42off for technical reason.
14:44Retention again we are going back to
14:46previous slide. Uh it's kind of
14:48limitation. We are saying we only have
14:50one year data. We are keeping only one
14:52year data.
14:55Support is business hours only during
14:57business uh support done only business
14:59hours. Now these three support,
15:01retention and availability they
15:03descriptive like we can just put it in
15:06the document to just describe the
15:08situation. But freshness it not just
15:11description it has it runs as a it has
15:14functional uh
15:18it's a kind of function that runs and
15:19can fail the test process because what
15:23it does actually it use order state
15:27current time stamp minus maximum order
15:29state and if it's going to be more than
15:3224 hours we have a lagging data this
15:35very frequently by the way happens in
15:37our in our company this happens because
15:39of sometimes failed ducks. Uh we are
15:41using DBT freshness check. It it was
15:43also mentioned like few presentations
15:45ago. Uh but again very useful in
15:48practice and you can take control that
15:51your data is fresh.
15:55Last one but a very important one that
15:58servers
15:59in this case is progress servers and
16:02defining all parameters. How we use it?
16:04Uh we have environment variables for
16:08production database and development
16:09database par with parameters and we can
16:12switch based on later we will see in the
16:14Python script how we doing it but yeah
16:17this can be used to connect to database
16:19to concrete database
16:24when you open data contract platform it
16:26has option that you can go and use
16:29interface to create your uh data
16:32contracts it's quite comfortable you
16:34just set and play with interface and
16:35generate this YAML file. But here's a
16:38trick like when we started using it, we
16:41had 300 tables to onboard. Now I think
16:44it's hard to imagine that I'm sitting
16:45and doing it for 300 tables here. It
16:48could be problem and this is time to
16:50mention AI finally that I could also
16:52mention in my presentation. I used Cors
16:56and the trick is was following in one
16:58folder I put few examples. I created few
17:01examples for concrete source in another
17:04folder. I just put details definitions
17:06of tables and wrote some prompt and
17:08saying here's examples here's a lot of
17:12tables generate data contracts for us
17:15and instead of like spending days and
17:17this manual routine work that's quite
17:19very boring I would say I even told my
17:21managers thanks AI it could be very
17:22boring for me playing with this
17:252 minutes 2 3 minutes and it's done
17:29there was a cases when AI little failed
17:32especially I think it's called medium
17:34minimum size, large size text like these
17:36kind of formats for some specific
17:38databases. But again, you are running
17:40tests, it fails, you're just checking
17:41manually. Instead of 300, maybe five or
17:4410 data contracts you should check.
17:48Okay, we have YAML file. We know what is
17:51data contract. What's Nate next? How to
17:53run it?
17:58Three different options to run data uh
18:00this test data contract. One is CLI
18:04opensource command line tool. As you can
18:06see it supports a lot of different uh
18:09databases uh platforms AWS glue,
18:13bigquery, iceberg.
18:16We now company using poss, MySQL, MSSQL
18:20that actually we use it but I would
18:22recom like list is very big. I could
18:24like continue for five minutes just
18:26saying what tools they use there. But
18:28you can go to this link and find if it's
18:30also appropriate for your company and
18:33yeah tool set you are using
18:37just example how output could look
18:41check that field line item present and
18:44test passed it has the type UU ID it's
18:48also passed sometimes they can just
18:50change UU ID to ID let's say so and it
18:53breaks because when ID is integer UU ID
18:56is other format and Especially when you
18:58are doing range sometime it can fail.
19:00This also important remaining is quite
19:03same but yeah has no missing values
19:07present and blah blah blah data. Uh this
19:10is the best part when you see data
19:13contract is valid. There's nothing to do
19:15with this.
19:38if you have big amount of data contracts
19:42and docker version um there is a docker
19:45image officially just need to pull and
19:47run also cool if uh there's a
19:50dependencies or something if you don't
19:52want to think about all of these
19:54problems just want to run I personally
19:56running with the when I testing locally
19:58in my laptop I'm using docker version
20:03okay we are the moment when we know how
20:06to run it we know this very cool tool
20:08exist nice but how we use it in indust
20:11because again we have a lot of projects
20:14and there was a lot of questions who is
20:16actually I will change the next slide
20:17who is going to responsible for like
20:19failed uh failed data contract it's
20:22going to be that software developer who
20:24is working in that team or it's going to
20:26be maybe analysts that should also
20:28update and push on the GitHub or it's
20:30going to be me like actually end up with
20:32me like [laughter] it's me who does it
20:34and yeah I am updating these data
20:37contracts
20:38Another question
20:40how we integrate it in our system. And
20:42they said 10 more than 10 projects.
20:44Should we refactor all these projects or
20:46how we will run it? Maybe on data
20:48engineering side I each time should run
20:50and check this databases if something
20:51changed
20:53>> [snorts]
20:53>> uh on schedule way. But after we made it
20:56event based which is more practical like
20:58as soon as this change happens it reacts
21:00instead of like running on schedule at
21:02daily basis let's say. So and this is
21:05also critical for me coordinating
21:07collaboration between developers
21:09engineers and analysts.
21:12three different teams and manager at
21:15different managers not the same manager
21:18we have to work together by the way they
21:20have their schedules their tasks and you
21:22know like one says I will do it after 2
21:24days and another one says I have more
21:26priority tasks and it it also cost cows
21:29we needed some solution for this one
21:31okay let's see how we created this data
21:35contract repository first actually you
21:37can see it's very easy this is
21:40screenshot of our repository just
21:42contracts folder with YAML files uh
21:46CI/CD YAML files GitLab CI/CD YAML files
21:49docker that's for Python with some
21:52dependencies and main pi uh main Python
21:56script that actually does all the job
22:03this is subfolders I wanted to show I
22:06wanted to show when I started my
22:08presentation I mentioned like we have
22:10This people hub let's say for internal
22:12HR software
22:14there is a mapping between this folder
22:17and the repository name like one folder
22:21name equal to one repository in the
22:23company and peoplehub contains YAML
22:26files separate YAML files for each
22:28table. There was idea maybe to keep all
22:31tables under one YAML file, but it could
22:34get very hard and mess because imagine
22:37having 20 tables under one YAML file. It
22:40could be hard. instead just having small
22:43but a lot of yaml files
22:46and mainpi mainpi script what it does it
22:50takes two arguments repository repo name
22:54and environment production or
22:55development because they're different
22:57databases actually depends on the
22:59environment.
23:00It runs data contracts tests and when it
23:04fails it sends email to ticketing system
23:07and automatically opens ticket
23:11This how it looks like. Data contract.
23:14This real scenario. This from real
23:16problem that I just made made a
23:18screenshot to show you. Zabix alert
23:22validation failed for stop project in
23:24the environment.
23:26Ticket owner first they uh assigned it
23:30to the to the teams. Um in this case
23:33Victoria we actually it assigns to
23:35engineer. I will mention it as well.
23:37Data quality validation. uh it failed
23:40and here's the reason why it fails, why
23:41it failed.
23:46Before I will continue, I want to say
23:48like when it fail
23:50this presentation I it will go even more
23:53detail in this
23:55how we integrated it with 10 projects
23:5810 and plus projects. This is some
24:01screenshot from CI Gitlab CI file. For
24:05each project, DevOps put these jobs
24:10based on the condition like CI commit
24:12branch should be equal to develop
24:14uh or this is for pro like main branch.
24:18It triggers
24:20another uh GitHub repository and runs
24:24pipeline in another repository. And here
24:27we have defined when the pipeline runs
24:29this job, it just triggers main pine
24:32main pine script.
24:34You see variable this is this sense uh
24:38pro repository name and this says
24:40environment like absolutely separate
24:43project it just run it it's not they not
24:46dependent each other loosely coupled
24:48system I would say like each working
24:50independently
24:52[snorts] and yeah it runs and checks
24:54okay for dev database for this project
24:56there's something changed run test
24:58failed or not
25:00and same actually for production job.
25:05Now important this coordination and how
25:08this done if it fails. First ticket uh
25:12if it fails on dev ticket assigned to
25:15their dev because changes actually in
25:18practice we found found there are two
25:20type of changes intentional they know
25:24what what they do and they really doing
25:26some refactoring or accidental temporary
25:29the developer was playing thinking okay
25:31this kind of schema could be even better
25:33than previous one wasn't necessary but
25:35he was playing in the end breaking
25:37everything like and when they received
25:40is they just roll back system and says
25:43okay we don't need to change
25:46even without disturbing data team
25:52if it's intentional and they know that
25:55they are changing something in the
25:56database then we are requesting from
25:59them to provide context
26:02why they in this case the context could
26:05be related for to to data engineer
26:08because I will know what's changing
26:10for data analysts sometimes they're
26:12changing generally logic or logic of the
26:14data and SQL queries DBT models they
26:18should completely refactor it because
26:20their logic is changing and they are
26:22providing this information under the
26:24ticket
26:26here's example creation date a real real
26:30case creation date was deprecated and
26:34back end was replacing it with created
26:36it
26:38do it to time zone standard
26:39standardization what's interesting
26:41imagine now they were storing data
26:44creation date in Lithuania time zone
26:48after they changing it to created at UTC
26:50time zone
26:52in aggregation everything going to be
26:54wrong numbers going to be wrong because
26:56now they are storing it different time
26:57zones and if not letting know analysts
27:02that this kind of change happened it
27:04could be disaster like for a lot of
27:06dashboard it's going to be like
27:07absolutely other numbers
27:09This why it can be also very useful.
27:13what should I do when receiving such a
27:16ticket and
27:18what we say if it's not big problem just
27:22continue deploying if it's or sometime
27:25we can request time uh to have meeting
27:27internally with analyst and engineers
27:29and decide how we will continue because
27:31it will it can require a lot of
27:33refactorings on our side before they
27:35will deploy it
27:37but some exceptions happens hot fixes
27:40some projects uh used by a lot of users
27:43like this estimators sitting in
27:46different offices doing this estimate
27:48for clients and if there's some big fail
27:51some problem happens they cannot sit and
27:52wait until analyst will discuss for a
27:54few days and do this um yeah
27:57brainstorming they immediately deploying
28:00this and we are fixing fixing it postf
28:02facto actually what was before data
28:05contracts quite frequently we are we
28:07were fixing it postfactum when
28:08everything started failing
28:12Yeah.
28:14Now the question, do we need to create
28:18data contracts for wall tables? No, it's
28:21not necessary. You can create it for uh
28:23most important ones. What we call
28:26important in my case in dbt there is a
28:28manifest JSON and if you get it and
28:31analyze you can find which tables
28:33actively used in the models and create
28:35statistics. This approach that I used
28:38and we got that some tables very
28:39actively used. By the way, even now I
28:41have one pending task they are changing
28:43that table and I said stop I'm on day
28:46off I'm in year one I will come back on
28:48Monday and I will fix it like this way
28:51we stopped it without that if they could
28:53change that projects table it could
28:56break almost all dashboard it's actively
28:58used this way you can on board but again
29:01with AI you can on board quite easily 50
29:05or 100 tables
29:08do we need to 35.
29:12Okay. And in the end, I want to express
29:15express my gratitude to my colleagues
29:17because I didn't work on this alone. Uh
29:19Yea, our team lead, and also Thomas
29:22DevOps. Uh thanks also them and we work
29:25together. I think it also Yeah. And
29:29thank you very much for attending.
29:31[applause]
29:34Miss Big Pleasure, I would like to
29:36answer on your questions.
29:38>> Thank you very much.
29:41Uh yeah have questions.
29:49>> Thanks Rolf. It's very close for me as a
29:52data engineer also. Uh my question is
29:55about the PII how you how you taking PII
29:58and who responsibility it is and what
30:02you do with the complex data types like
30:04JSON for example right from in some time
30:07to time uh PI data sensitive data using
30:10that kind of complex types right
30:12>> could you repeat first part API what
30:14>> uh about sensitive sensitive data who
30:17attack data whose responsibility is it
30:20>> interesting fact now we don't have such
30:22a like we have one sensitive data but
30:25actually devops state I will not give it
30:27to you right they're kind of cutting
30:29this table to three column like from
30:31five columns they're just cutting it and
30:34giving only three that's from financial
30:36data that could be very important but
30:39I'm not getting it and remaining data
30:41it's not sensitive and we don't have
30:43people who is actually doing it got it
30:45okay thank
30:55Thank you for very interesting talk. Uh
30:58the system of uh data contracts strikes
31:02me as uh being overlapping with uh the
31:06systems that are recently in fashion
31:08called semantic layers such as DBT
31:11semantic layer because it uh portrays
31:15also information about whatever this
31:17thing means for the business etc etc. uh
31:21do you see any overlap between these
31:23approaches and any synergy that might
31:26result in using both? Thank you.
31:29>> We are using both. Thank you first of
31:31all thank you for question and we
31:32already use semantic layer the our team
31:35working on it. uh what's difference um
31:39in our ca I showed you all the cases but
31:42actually we are not using all these
31:44quality checks because freshness check
31:46for example we implemented with dbt and
31:49data contracts mostly we use to specify
31:52column names and column types and just
31:55check if it exist or not whatever code
31:57could break it's kind of before dbt
32:00before data warehouse uh we check first
32:03and after run it what you mentioned at
32:05semantic It's kind of end of the story
32:07already when you have already and our
32:10analyst now defining these formulas and
32:14then yes it can it will be used in
32:15dashboards in calculations I can't
32:17provide more information because I'm not
32:19working on it but what I can say if you
32:21see the chain now here is a semantic
32:24layer and and here is a data contracts
32:27first.
32:30>> Yeah thank you. Um I have this like
32:33future looking the next step question uh
32:36and I would like your opinion on that.
32:38Do do you see a world where like the
32:42contracts are agents and like they
32:46negotiate between each other when the
32:49change happens and they try to negotiate
32:52and then sort it out before bringing
32:54human in a loop as as as much as
32:56possible. Like a mechanism could be for
33:00example
33:02uh there is a change in a contract and
33:03it broadcast the what the changes in the
33:06blast radius of that and they will other
33:09contractors subscribe to that and see if
33:11it's pertinent to them and then they
33:13will take account what changes has made
33:15and then incorporate that and there's
33:17like a peer-to-peer mechanism between
33:19the you know different departments and
33:21different contracts. I think my vision
33:23that I completely think that it could
33:25happen because I was checking stack
33:27overflow again we discussed it like just
33:29three or four year passed and stack
33:31overflow now seems like a museum you
33:33know it was dinar's era but just few
33:35years passed and with active development
33:38or of agents I think everything is
33:41possible and especially I have seen
33:43cases when the team integrated agents to
33:47analyze airflow logs and when something
33:50significant were happening I
33:52against quite quickly was reacting on it
33:55and yeah completely agree it's possible
33:58and in our company we also that was a
34:01question about DBT they have this agents
34:03and we already testing it uh what's the
34:06difference they fails for nuances like
34:09it it works good but when there's some
34:11nuances in filters or something it can
34:13fail not always but it can fail and we
34:15are exactly testing human versus against
34:19and data quality
34:24Thank you.
34:25>> So thank you. Uh I'm not a data
34:27engineer. So I'm uh I'm a I I feel I
34:32have to ask this question. You are uh
34:35constrating the uh software engineers
34:38not to alter the database by data
34:40contracts if I'm not mistaken.
34:44What what
34:45>> you are constraint constrainting the
34:49software engineers not to alter the
34:51database uh without authorization if I'm
34:55not mistaken.
34:56>> By the way, it's good question because
34:57we were discussing should we block
34:59someone or not block and like working
35:02process and this is why I say dev and
35:04pro mode depends like if it's in dev we
35:07just notifying them. Actually this kind
35:09of early notification system it's not
35:11it's not blocks anyone. Uh
35:13>> but we getting it earlier and can react
35:17instead of getting this panic everything
35:19fails quickly do it. It's a lot of
35:21stress on the team and also delays on
35:23reports. Instead we could get it earlier
35:26and can
35:27>> a followup question if I'm allowed
35:30follow
35:32>> a followup question.
35:33>> Yeah. Yeah.
35:34Do you think can you imagine uh a world
35:38that we have some uh contracts not just
35:42for the data but for the uh everything
35:45that AI generates
35:47including the code that we will say okay
35:50we will not read your code we will we
35:52don't care about what you are doing you
35:54have to be following not following but
35:57you have to meet this criteria you have
35:59to pass this unit test you have to have
36:01this kind of uh measurements And this is
36:04your contract. Follow this contract. Do
36:06whatever you want. Write in any language
36:09you want. Do you think it's possible? By
36:11the way, uh this uh data contracts is
36:14very universal because doesn't matter
36:16what language you're using and you can
36:18this YAML file you can do it with PHP
36:21you can do with Python like with Docker
36:23especially Docker version you just put
36:25in YAML file and it does for you like it
36:27doesn't matter
36:29environment like Python version
36:30officially supported but for other cases
36:33we were thinking maybe each team will
36:34integrate this test for them but after
36:37created this independent system just
36:39separate repository that runs
36:41Because before was the idea why if we
36:43put for each project as a part of like
36:46on the test stage uh this um data
36:49contracts with Python it was very
36:51problematic because some projects
36:53implemented in PHP but with this docker
36:55image again it's independent like you
36:58can just run it doesn't matter like what
37:00language it is if it related about
37:03independency of language but yeah with
37:04docker we reach this independency
37:07>> okay could do whatever they want
37:09>> against I I I think um based on the how
37:13the trends going I believe that in the
37:15future will do whatever they need.
37:17>> I hope it will we will have some delay
37:19because I don't want to lose my job.
37:22[laughter]
37:23I believe in good future but let's have
37:26a little delay.
37:31>> Yeah. Uh I wanted to ask if you're
37:33familiar with uh data observability
37:36tools like Monte Carlo and which do this
37:40proactive monitoring of changes and data
37:42changes and how you think this uh two
37:46approaches collaborate because this
37:48approach is preventing or notifying
37:50whenever change is happening once on the
37:52code level but the data observability
37:55tools like they're constantly
37:57monitoring. So if something happens you
37:59get notified or you get alert and then
38:01you get a sign of now also AI boat which
38:04helps you identify what changed. Do you
38:07use that as well or not?
38:09>> Uh not we are not using it but one you
38:11are from service titan maybe.
38:14>> Yeah because I had this conversation
38:15yesterday that
38:19>> with someone I forgot your I forgot your
38:21name but
38:22>> she was representing Yeah. She was from
38:25a service titan and I didn't know Monte.
38:28>> Yeah. Yeah. Yeah.
38:29Uh first of all I didn't know about that
38:31tool because it was where like Monte
38:33Carlo probably it's for statistic like
38:35simulation and I was little confused how
38:38Monte Carlo related to data like this
38:40way but what interesting why I don't
38:43know I was thinking okay this really
38:44exists why I didn't know about it the
38:47it's mostly accidentally when we had
38:50this data contract issue and when my lit
38:53uh was on this conference she bring this
38:56idea me at the moment when we had this
38:58issue I implemented it worked and we
39:00didn't continue like exploring other
39:02tools just accidentally these multiple
39:04factors happened together but what your
39:07colleague yesterday said Monte Carlo is
39:09not cheap tool it's expensive tool as
39:11far as I know
39:12>> depends how you use
39:13>> yeah but data contract is zero dollar
39:17>> yeah just follow up uh I trying to
39:20understand what is the best approach
39:22even it can be not multi there are open
39:25source versions and you can host it but
39:26the idea do you need after this do you
39:29need uh proactive monitoring of these
39:32changes or just what's in place is
39:35enough or you are able to uh catch
39:38everything with just this
39:40>> like currently what I see in company we
39:42don't need anything more especially this
39:43is for free we just use it
39:46>> forget about
39:47>> no no no I mean like this is okay this
39:49works for us we don't have any problem
39:51it really does what it should do like
39:53send us notification and we are reacting
39:56uh and also this manifesto that we sent
39:58across all teams that how we are going
40:00to react on it. is important because
40:02first when I integrated it a lot of
40:05developers were panicking okay it felt
40:07what should I do like no one know knew
40:10what they should do and they were
40:12texting me in panic okay I stuck but
40:14after preparing all this document pend
40:17actually little practice like failed and
40:19fixed everyone became little experienced
40:22and we are quite well managing these
40:24situations
40:25>> reactive not
40:27>> proactive yeah
40:28>> this one is productive
40:30>> this one not actively checks data. Yeah,
40:33we try to prevent and obser
40:40>> it's not no it this checks only the
40:43moment
40:44>> it happens after they did some change
40:46but maybe we can continue after
40:48>> yeah no we can
40:49>> I just I just want to show this one
40:51because
40:52this triggers only in two cases this
40:54kind of event based when they commit on
40:57develop or on pro main branch two cases
41:01it's not constantly It's not checks
41:02database like active
41:04>> if somebody if somebody goes manually
41:06changes drops the fields then you don't
41:08catch it right. Yes, we don't will not c
41:10but actually it's never like last two or
41:13three years it never happened. So good
41:15that you could disagree that
41:18>> like this very frequently happened what
41:20I mention it but that someone deletes
41:23actually I am quite actively saying guys
41:25if you want to work with table get read
41:26only access because it's good for you
41:28good for us and you know like if
41:30something happens you are calm that it's
41:32not because of you you have just read
41:34only access to the table
41:38that's actually
41:42we have like a
41:45It's nice actually. Yeah.
41:47>> The best part.
41:48>> Yeah. Thanks a lot for your insights. Uh
41:50I wanted to ask uh so whatever you
41:52presented relates uh to the tables. Uh
41:55do you think there are data contracts or
41:58something similar for other data
42:00modalities let's say uh images audio
42:04video and how they can look like if
42:07there are such things.
42:09I don't know again better to check I
42:11don't know if took took a picture of it
42:13but on their website there's a big table
42:15on GitHub mentioning everything and I
42:18personally not working with this kind of
42:20data but actually could be interesting
42:22because now we have AI team who is
42:25working especially with image processing
42:27yeah
42:28>> thank you
42:30question
42:34>> so my question is like during
42:37integration of this process
42:39Like you already mentioned that there
42:41was a kind of resistance or panic or
42:43people just
42:44>> don't want to do and what about like the
42:47current state of is there like uh like
42:51people who sabotaging like how you will
42:54deal is there just attempts to say okay
42:57let's don't do it it slows us down. So
42:59in terms of like human perspective, how
43:02difficult it to maintain? Okay, you have
43:04it now, but do you have like situations
43:06when they start to question like why we
43:08should keep this?
43:10>> Yeah, like it slows down. It was easier
43:12to tell our CEO like you know if we
43:15integrate it you will get a big benefit.
43:18If not like you will not see reports. It
43:20solved all the downstreaming problems.
43:23But in general we have clear plan like
43:25it fails ticket assigned to you and
43:28after discussion starting it's not about
43:30I wish or I not wish no we have strict
43:32uh requirement it fails ticket on you
43:35and we should work again we are showing
43:37that downstreaming issues like we had it
43:41a lot of times I hope I'm happy that
43:44actually engineers understanding but
43:46again presentation to co was mostly
43:50major part of this Any other questions?
43:56No. Only thing I learned today that you
43:59can easily spend a lot of money in Monte
44:02Carlo.
44:03>> That's right. There's a contract.
44:05>> Thank you very much.
44:06>> Thank you very much for attending. This
44:07game.