Data Infrastructure Data Engineering Data Observability Data Data Engineer

What is Data observability, do I need it?

Jatin SolankiJan 23, 20236 min read

I started my career as a first-generation analyst focusing on writing SQL scripts, learning R, and publishing dashboards. As things progressed, I graduated into Data Science and Data Engineering where my focus shifted to managing the life-cycle of ML models and data pipelines. 2022 is my 16th year in the data industry and I am still learning new ways to be productive and impactful. Today, I am now the head of a data science & data engineering function in one of the unicorns and I would like to share my findings and where I am heading next.

When I look at the big picture, I realised that the problems most companies face are quite similar. Their vision towards being data-driven has turned into a BHAG — pronounced “bee hag” (Big Hairy Audacious Goal). We data folks like patterns, so here are my findings:

During 5 out of 10 review meets, I have witnessed people question the reliability of the data/report/dashboard. Additionally, HODs will also try to convince others that their data is the most accurate or reliable :)
A lot of times, HOD comes and says that the data is not updated. The data team is already working to fix the report/data table.
A new product got launched the week before, however, we are yet to figure out the performance. The data team is working on a query change and will soon update the CXO team.
Everyone has built expertise around writing complicated ML (machine-learned) models, however very few talk about or deploy inference monitoring. There is a high probability of model drift or performance drift in the coming weeks/months if not monitored or observed efficiently.
Very few companies deploy solutions or models to detect performance anomalies.

The list is long, I am sure you can relate or add more to this.

In a nutshell, I found that data reliability is a BIG challenge and there is a need for a solution that is easy to use, understand, deploy, and also not heavy on investment.

Hello, I am Jatin Solanki who is on a mission to build and develop a solution to make your data reliable.

What is needed to make your data more reliable?

Complexities around data infrastructure are surging as companies gear to get a competitive edge and out-of-the-box offerings.

Every company goes through a data maturity matrix. In order to reach a level where you deploy AI models or self-service models, you need to invest in a robust foundation.

In my opinion, the foundation begins with a reliable data source or defining source of truth. Your data models won’t be impactful if it’s ingested with bad data. You know it’s garbage in garbage out

On a high level, here are a few checks you can implement to ensure data reliability:

Volume: It ensures all the row/events are captured or ingested.
Freshness: Recency of the data. If your data gets updated every xx mins, this test will ensure its updated and raises an incident if not.
Schema Change: If there is a schema change or a new feature that was launched, your data team needs to be aware to update the scripts.
Distribution: All the events are in an acceptable range. e.g if a critical shouldn’t contain null values, then this test ensures to raise an alert for any null or missing values.
Lineage: This is a must-have module, however, we always underplay these ones. Lineage provides a handy info to the data team of the upstream and downstream.
Reconciliation: I would say recon or finding deltas between two given datasets. This could be used to understand the difference between stagingand production OR between source and destination . This could be effective in running some financial recon too, like payment gateway to the sales table.

What next? How do we implement this?

The most common question people face with:

Build versus Buy

I am a big fan of open source tech, however, in some critical modules, I prefer buying an out-of-the-box solution because it’s scalable and already tested in the market. Developing in-house might cost you around US2k per month and it includes a few hours of engineer’s time along with cloud cost.

If you are inclined toward buying an out-of-the-box solution, here are a few factors that should be part of your checklist.

Should be able to connect to popular sources which require minimal config.
Extract information automatically without the need for additional code.
No-code or CLI (I leave it to you)
Lineage and Catalog module.
Data Reconciliation along with scheduling feature.
Anomaly detection
Of course, Of course, all the tests we discussed earlier along with alerts should be in a position to tell where to debug.

A robust platform provides easy access to all the incidents and also evaluates the data health.

It should be in a position to automatically detect my critical data assets and apply hygiene checks.

Only platform to group alerts instead of pushing 100+ alerts.

At last, the solution should help you reduce data quality incidents and make your data more reliable.

So, do I need a data observability platform?

If your answer to any of the below questions or scenarios is “Yes”, then you should procure or deploy a data observability solution right away.

Dashboard not getting updated on regular basis?
Don’t know which report is accurate?
Business stakeholders are the first to learn about data incidents.
Questions during a meeting on the performance stats.
Have at least 2 members in the data team.
Deployed a business intelligence tool.

As software developers have leveraged on DataDog, Dynatrace, etc kind of solutions to ensure web/app uptime, data leaders should invest in data observability solutions to ensure data reliability.

Originally posted here

Similar Journal

ELT for the Data Consumer

You’ve likely heard about ELT — Extract Load and Transform… the Modern Data Stack’s evolution on ETL. This is a game changer by nature in that it enables organizations to ingest raw data into the data warehouse and transform it later. ELT gives end-users access to the entirety of the datasets they need by circumventing downstream issues of missing data that could prevent a specific business question from being answered.

jared parker8 min read

How All-in-One Tools Are Accelerating Data Democratization

A majority of business leaders believe data insights are key to the success of their business in a digital environment. However, many companies struggle to build a data-driven culture, with a key reason being the lack of a sound data democratization strategy.

jonas thordal7 min read

Beyond Observability for the Modern Data Stack

The term “observability” means many things to many people. A lot of energy has been spent—particularly among vendors offering an observability solution—in trying to define what the term means in one context or another.

Avadhoot Patwardhan8 min read

How to Make Better Decisions Together with Collaborative Analytics

Breaking down some of the problems I’ve seen in data collaboration and offering advice on how to make better, faster decisions with collaborative analytics.

Ryan Buick5 min read

+2 more

The Unbundling of SaaS Analytics

The modern data stack is on the rise. Many companies use raw data from their SaaS analytics tools as input for their data warehouse, but this introduces problems downstream. Are there better ways?

vincent hoogsteder5 min read

What Is Active Metadata, and Why Does It Matter?

Just like data mesh or the metrics layer, active metadata is the latest hot topic in the data world. As with every other new concept that gains popularity in the data stack, there’s been a sudden explosion of vendors rebranding to “active metadata”, ads following you everywhere and… confusion.

prukalpa ⚡ 9 min read

+1 more

What's the Difference Between Data Wrangling vs Data Cleansing vs Data Transformations

As the amount of data rapidly increases, so does the importance of data wrangling and data cleansing. Both processes play a key role in ensuring raw data can be used for operations, analytics, insights, and inform business decisions.

JD Prater6 min read

What is Data Onboarding? And 3 Ways It's Overwhelming Your Teams

Without a clear and quick process your dev, sales, and customer success teams can become overwhelmed by the amount of work required to delight new customers and ingest clean validated data.

JD Prater5 min read

The Modern Data Stack Ecosystem: Spring 2022 Edition

Without a clear and quick process your dev, sales, and customer success teams can become overwhelmed by the amount of work required to delight new customers and ingest clean validated data.

Jordan Volz25 min read

What is Data Observability?

Do you know the current status — quality, reliability, and uptime — of your data and data systems? Not last month or last week, but where they stand at this moment. As businesses grow, being able to confidently answer this question becomes more important. That’s because data needs to be clean, accurate, and up-to-date to be considered reliable for analysis and decision-making. This confidence comes through what’s known as data observability.

cody carmen4 min read

What is Data Reliability?

“There must be something wrong with Excel. I can't get these numbers to make sense.” For anyone who has had a similar experience of staring at a spreadsheet for far too long, we have news for you: Excel isn’t the problem; your data is.

cody carmen5 min read

Data Governance - A Thought Leader's Perspective

If you are a Data Leader in 2022, Data Governance is most definitely on your radar. Regardless of your organization's data maturity stage, chances are, you have already implemented or started implementing a Data Governance Strategy.

benedetta cittadin6 min read

Getting Started with Data Observability

In the past years, organizations have been investing heavily to convert themselves into data-driven organizations with the objective to personalize customer experiences, optimize business processes, drive strategic business decisions, etc. As a result, modern data environments are constantly evolving and becoming more and more complex. In general, more data means more business insights that can lead to better decision-making. However, more data also means more complex data infrastructure, which can cause decreased data quality, a higher chance of data breaking, and consequently erosion of data trust within organizations and risk of not being compliant with regulations. The data observability category — which has quickly been developing during the past couple of years — aims to solve these challenges by enabling organizations to trust their data at all times. Although the category is relatively young, there are already a wide variety of players with different offerings and applying various technologies to solve data quality problems.

benedetta cittadin13 min read

How to Create a Data Governance Team? 3 Essential Steps

Data governance is more than just having a strategy – it is about establishing a culture where quality data is achieved, maintained, valued, and used to drive the business. Modern-day businesses are supported by data and information in many ways and forms. In recent years, data has become the foundation for competition, productivity, growth, and innovation. We are seeing successful organizations shift their focus from producing data to consuming it, and data governance strategies becoming increasingly important to support their crucial business initiatives. Executives and shareholders are starting to realize that data is a strategic asset and data governance is a must if they want to get value from data.

tanmay sarkar15 min read

What Is Data Reliability And How Observability Can Help ?

Data matters more than ever – we all know that. But at a time when being a data-driven business is so critical, how much can we trust data and what it tells us? That’s the question behind data reliability, which focuses on having complete and accurate data that people can trust. This article will explore everything you know about data reliability and the important role of data observability along the way, including:

Eitan Chazbani7 min read

What Is Data Lineage?

The term “data lineage” has been thrown around a lot over the last few years. What started as an idea of connecting between datasets quickly became a very confusing term that now gets misused often. It’s time to put order to the chaos and dig deep into what it really is. Because the answer matters quite a lot. And getting it right matters even more to data organizations.

Eitan Chazbani12 min read