A lakehouse with lineage and quality built in
Lakehouse architecture, real-time lineage, data quality and discovery for a data mesh platform serving many data products.
Challenge
A large data mesh platform needed one architecture for batch and real-time pipelines, and a way to trust and find the data products built on it.
Its users ranged from data engineers to people who do not write code.
What we did
- Designed and delivered a lakehouse on Apache Spark, Apache Iceberg, Apache Kafka, MinIO and Trino, unifying batch and streaming ETL.
- Built a real-time metadata and lineage system on OpenLineage.
- Built a hybrid search engine, full-text plus vector embeddings, on PostgreSQL for data discovery.
- Introduced data quality and validation pipelines with Great Expectations across batch and streaming workflows.
- Launched a self-service, no-code pipeline composer.
- Benchmarked large-scale processing frameworks to inform architecture decisions, and provided third-level support for critical components.
Outcome
- Batch and real-time pipelines run on one architecture.
- Every data product on the platform is traceable end to end.
- Data is easier to discover through combined keyword and semantic search.
- Non-technical users adopted the platform through the pipeline composer.
Technologies
Apache Spark, Apache Iceberg, Apache Kafka, Trino, MinIO, OpenLineage, Great Expectations, PostgreSQL