Data Governance

Data Governance for a Financial Institution

Strategy, stewardship and cataloging to govern transactional data and distribute analytical data as a mesh.

BEREFERENCE

Overview

A financial institution operated with data spread across transactional systems and analytical environments, with no clear boundary between the two worlds.

Over a thousand tables supported business decisions and regulatory obligations, with no one formally accountable for any domain and no single place to discover what existed, where it came from, or who was allowed to access it.

The project established a data governance practice that began with strategy and ended in the teams’ daily routines, treating governance as distributed accountability rather than central control.

Challenge

The data environments had grown around each area’s needs, with no shared model underneath.

As a result, fundamental questions had no reliable answer:

  • Who is accountable for this data?
  • Did this number come from the transactional system or from an analytical transformation?
  • Who can access what, and on what grounds?
  • Where is this data produced, and which areas depend on it?
  • Does this metric mean the same thing here as it does in the source system?

The missing boundary between analytical and transactional workloads was the underlying problem: analytical queries reached operational data, and definitions diverged across reports with no one able to arbitrate which was correct.

The cost showed up at the front door. Getting one person ready to work with data took two days of access requests, informal explanations and discovery by conversation.

Approach

The work was split into four tracks, run in parallel with explicit dependencies between them:

  • Strategy: define what governance meant for the institution, which domains existed, and what obligations each one carried.
  • Stewardship: name owners per domain and turn that accountability into routine rather than a symbolic title.
  • Catalog and infrastructure: provide one place to discover, document and classify data, backed by infrastructure of its own.
  • Analytics engineering: establish the analytical transformation standard and the model for distributing data across domains.

Architecture

The first structural decision was to separate the transactional world from the analytical one explicitly.

Transactional systems became the governed source boundary, with access control of their own. The analytical environment began consuming that data through defined contracts instead of reaching into it directly, with Snowflake as the analytical platform and Trino as a federated query engine across the sources.

The domain became the unit of organization. Eight data domains were designed, each with its own owner, scope and obligations.

A data mesh model was built on that foundation: each domain publishes its data as a product, with a declared owner, documentation in OpenMetadata and associated access rules. Consumers find in the catalog what exists and under what conditions they may use it.

Implementation

Implementation began with the inventory: over a thousand tables mapped and classified by domain and sensitivity.

With that map in hand, the eight domains were redesigned and stewards appointed, taking part in decisions about their domains — who gets access, what each field means, what may be published.

OpenMetadata went live next, first as a record of what already existed and then as the entry point for anything new. Trino gave consumers a single query interface over sources that previously required separate access.

On the analytical side, dbt was adopted as the transformation standard, giving analytics engineers a versioned, testable and documented model. The resulting data products were then distributed across domains following the mesh model.

Culture work ran alongside every stage: governance holds only when teams absorb it into their routines, not when it depends on a central team to function.

Technology

  • Snowflake
  • Trino
  • dbt
  • OpenMetadata
  • Classification by domain and sensitivity
  • Domain-based access control
  • Data mesh as the distribution model

Impact

The most significant change was not technical — it was accountability.

Data stopped being a diffuse asset belonging to everyone and no one, and gained declared owners who answer for the quality, the meaning and the access to what they publish.

Business teams gained a place to discover what exists before asking for it, and technical teams gained criteria for deciding what to expose and to whom. Analytical workloads stopped competing with transactional ones.

The most visible effect was at the entrance: what used to take two days of informal coordination now takes an hour, because discovery, documentation and the access rule live in the same place. That made it practical to enable more than a hundred people to consume data on their own.

The practice scaled beyond its original scope because the domains themselves carry it out, rather than depending on a central governance team.

Results

2 days → 1 hourOnboarding a person to the data environment

  • 8 data domains redesigned, each with a declared owner, a defined scope and obligations of its own.

  • Over 1,000 tables mapped and classified by domain and sensitivity.

  • 100% of transactional data governed, with access control tied to the data’s domain and sensitivity rather than to inherited permissions.

  • More than 100 people enabled to discover and consume data on their own through the catalog.

  • Onboarding from 2 days to 1 hour, with discovery, documentation and the access rule in the same place.

  • A defined boundary between analytical and transactional, ending direct access from analytical workloads to source systems.

  • dbt adopted as the analytics engineering standard and mesh distribution in place, with each domain publishing its data as a product for the others.

All cases

Have a challenge
worth solving?
Let's talk.

Talk to BEREFERENCE