Data lineage - feeding third-party data catalogs to expose technical data lineage

{openAudit} lets you browse data lineage through its Mosaïc web portal, but the consolidated metadata can also be published to various data catalogs on the market, which often have significant limitations when it comes to data lineage. The benefit is having a single tool for optimal data governance.
The data lineage is reconstructed in the {openAudit} technical repository, then exposed to data catalogs via dedicated connectors.
This architecture makes it possible to decouple the data lineage collection and reconstruction mechanisms from the exposure layer used by governance teams.
Key point: the technical repository stays independent from the target catalog. The collection and lineage-reconstruction (and connection) mechanisms don’t depend on the data catalog in use.
Architecture

The collected metadata is consolidated in the oA.mdb columnar database after processing by the oA.Job engine.
The central repository notably contains:
- Field-level data lineage.
- Usage data.
- Associated governance information, where applicable.
The catalogs are then fed from this single repository via exposure connectors independent of the data catalog solutions in use.
Exposure connectors
The connectors handle:
- Automated publication of the lineage.
- Continuous synchronization.
- Restriction management.
- Support for the various Workspaces.
Data lineage exposure can be complete or partial depending on the display capabilities of the target catalog, keeping in mind that the data lineage produced upstream is itself available at field level.
Result: catalogs only consume the consolidated metadata, not the original technical sources.
Supported catalogs
The available connectors currently support the following data catalogs:
- DataGalaxy
- Zeenea
- Collibra
Usage management
When audit logs are available, observed usage is integrated into the technical repository and then exposed to the catalog.
This lets users distinguish:
- “Live” flows, consumed by tools or queried.
- Flows — and therefore objects — that see no consumption (often more than 80% of the total).
Typical use case: enrich the data catalog with usage information to identify objects that are candidates for rationalization or decommissioning.