Reference data scattered across Excel, SharePoint, and multiple databases is a problem most organisations know well, and the usual fix is an expensive, standalone MDM platform.
Oleksii Zarembovskyi took a different approach: through custom software development, his team built MDM directly into the Databricks Medallion Architecture already in use. The result is a deployed solution with a controlled approval workflow, full audit trail, and governed reference data flowing straight into the Silver Layer no extra licensing platform required.
We talked to Oleksii about how the idea started, how it works, and how it could be adapted for clients running on Databricks.
Background & experience:
Over 6 years of experience in the data engineering field. Oleksii leads a team of 15+ Data Engineers, with a focus on organisational and technical debt coverage, performance improvements, and best practices implementation.
What inspired you to create your own MDM solution?
Oleksii: We were seeing a situation that is common across many organisations: reference data was distributed across multiple sources — Excel, SharePoint, relational databases — with no single, controlled source of truth.
But the challenge was not only where the data was stored. There was also no formal change management process: Who can propose a change? Who reviews and approves it? How do we track the history of changes? And how do we ensure data quality?
At the same time, specialised enterprise MDM platforms can involve significant licensing costs.
If Databricks is already the organisation's central data platform following a broader cloud migration, why not implement the required MDM and Reference Data workflow directly within that ecosystem?
That question became the starting point for the application — an architectural add-on to the existing Medallion Architecture.
How does the solution work in practice?
Oleksii: We built an application where both the UI and backend operate within the Databricks environment. This lets us reuse the platform's native authentication and authorisation capabilities instead of introducing another isolated system.
The workflow itself is built around clearly defined roles. Data Owners and Subject Matter Experts can propose changes, while Data Stewards review and approve them.
Once approved, changes are automatically pushed back to the Silver Layer. This gives projects and teams access to centralised, governed reference data instead of multiple competing versions of the same information.
Why is a controlled change management process so important?
Oleksii: Technically, storing a table is easy. Establishing trust in what that table contains is much harder. For every change, you need to know who proposed it, who reviewed it, who approved it, and exactly what was modified.
That is why the solution includes an audit trail, access controls, and mechanisms to identify and manage anomalies.
We also aligned the approach with DAMA principles. The goal was not simply to create a convenient interface for editing tables, but to establish a governed Data Management process.
What role does standardising business terminology play?
Oleksii: A major one. When different teams interpret the same entities differently, those inconsistencies eventually affect analytics and reporting. One principle behind the solution is using conformed data and consistent business terminology.
This helps reduce duplication and ambiguity while establishing a shared data narrative across projects.
Combined with a complete audit trail, it also makes results more reproducible and transparent for internal controls and audits.
What does the business case look like?
Oleksii: Cost efficiency was one reason I initiated the development.
Specialised enterprise MDM solutions can cost tens or even hundreds of thousands of dollars per year in licensing. If an organisation already uses Databricks, our approach can eliminate the need for an additional major licensing layer and instead leverage the existing platform, with costs primarily driven by the compute and storage actually consumed.
As a result, the potential difference in total cost of ownership can be substantial.
Of course, the exact economics depend on the scale, requirements, and architecture of each client, so the TCO should always be evaluated on a case-by-case basis.
Where does the solution stand today?
Oleksii: This is already a working solution. The application is deployed and in use.
The next step is to collect structured feedback from users and Data Stewards and use those insights to shape the next iteration of the product.
Another encouraging signal was the positive feedback the pilot received from the local Databricks team. We see clear potential to turn what started as an internal use case into a repeatable approach for a broader range of clients.
How scalable is the approach?
Oleksii: We deliberately avoided designing it around a single reference data set. Architecturally, the same approach can be applied to different types of tabular data, from reference data and Data Quality rules to anomalies and other scenarios that require controlled editing, approval workflows, and auditability.
The next step is to formalise the architecture, deployment guide, roles and processes, and the TCO model. This will help prepare the solution for potential packaging through the Databricks Brickbuilder Solutions program and create opportunities for future co-marketing and co-selling activities.
At the same time, the team is working on Unity Catalogue migration. How do these initiatives connect?
Oleksii: We see them as parts of a broader Data Governance story. Many organisations still operate with legacy approaches to metadata management and governance, which means moving to Unity Catalogue is not simply a technical migration. It requires proper pre-checks, mapping, quality testing, rollback planning, and alignment with compliance and security requirements.
We are already developing this approach through a real-world project. The next goal is to turn that experience into a standardised migration runbook and an end-to-end demonstration case we can reuse across future client engagements.
What is the bigger idea behind these initiatives?
Oleksii: For me, the goal is not simply to solve one isolated problem. It is to identify an approach that can be repeated and scaled. If we can take a real client challenge, build a working solution, validate it in practice, standardise the approach, and then apply it to other projects, we create something much more valuable than a one-off technical implementation.
That is how we see both the MDM application and our Unity Catalogue migration expertise: as practical experience that can evolve into a broader data strategy and repeatable Databricks solutions and services.
From an internal initiative to a scalable solution. This case demonstrates how an idea born from a real business challenge can evolve into a production-ready solution and become the foundation for a broader service offering.
The next chapter is about packaging that expertise, evolving the product, and exploring opportunities to scale it together with Databricks.
FAQs
Not exactly. Databricks is a unified data and AI platform built around the "lakehouse" architecture, which combines elements of data lakes and data warehouses. It does not operate like a conventional relational database management system such as MySQL or PostgreSQL. Rather, it offers a range of tools for data engineering, data science, machine learning, and analytics with respect to large-scale data stored in formats such as Delta Lake.
While Databricks does let you run SQL queries and manage structured data (and includes components like Unity Catalogue for data governance), its core purpose is broader than a DBMS. It's designed for processing, analysing, and building AI/ML models on massive datasets, often across distributed cloud storage rather than a single managed database engine.
Not exclusively, but it can serve that role. Databricks supports ETL/ELT workflows (via tools like Delta Live Tables, Spark jobs, and workflows), but it's really a broader data and AI platform. ETL is just one of many things you can do on it, alongside analytics, ML, and business intelligence.
Related insights
The breadth of knowledge and understanding that ELEKS has within its walls allows us to leverage that expertise to make superior deliverables for our customers. When you work with ELEKS, you are working with the top 1% of the aptitude and engineering excellence of the whole country.
Right from the start, we really liked ELEKS’ commitment and engagement. They came to us with their best people to try to understand our context, our business idea, and developed the first prototype with us. They were very professional and very customer oriented. I think, without ELEKS it probably would not have been possible to have such a successful product in such a short period of time.
ELEKS has been involved in the development of a number of our consumer-facing websites and mobile applications that allow our customers to easily track their shipments, get the information they need as well as stay in touch with us. We’ve appreciated the level of ELEKS’ expertise, responsiveness and attention to details.