Skip to content

All case studies

A lakehouse with lineage and quality built in

Lakehouse architecture, real-time lineage, data quality and discovery for a data mesh platform serving many data products.

Client
Large-scale data mesh programme
Sector
Enterprise
Region
Middle East
Services
Data platform engineering

Challenge

A large data mesh platform needed one architecture for batch and real-time pipelines, and a way to trust and find the data products built on it.

Its users ranged from data engineers to people who do not write code.

What we did

  • Designed and delivered a lakehouse on Apache Spark, Apache Iceberg, Apache Kafka, MinIO and Trino, unifying batch and streaming ETL.
  • Built a real-time metadata and lineage system on OpenLineage.
  • Built a hybrid search engine, full-text plus vector embeddings, on PostgreSQL for data discovery.
  • Introduced data quality and validation pipelines with Great Expectations across batch and streaming workflows.
  • Launched a self-service, no-code pipeline composer.
  • Benchmarked large-scale processing frameworks to inform architecture decisions, and provided third-level support for critical components.

Outcome

  • Batch and real-time pipelines run on one architecture.
  • Every data product on the platform is traceable end to end.
  • Data is easier to discover through combined keyword and semantic search.
  • Non-technical users adopted the platform through the pipeline composer.

Technologies

Apache Spark, Apache Iceberg, Apache Kafka, Trino, MinIO, OpenLineage, Great Expectations, PostgreSQL

Facing something similar?

Tell us where your build is. We read every message and reply.