Skip to content

Technology radar

What we build with, and how firmly we recommend it. Every placement comes from running the technology in production on our engagements.

Adopt
Our default choice. Proven in production on our engagements.
Trial
Used on real work and worth choosing where it fits.
Assess
Promising. We evaluate it before committing a client to it.
Hold
We have run it in production and would not start new work on it.
AdoptTrialAssessHoldData platforms and lakehouseStreaming and orchestrationInfrastructure and DevOpsAI, ML and application engineering1234567891011121314151617181920212223242526272829303132333435363738394041424344454647

Data platforms and lakehouse

Adopt

  1. 1Apache IcebergOpen table format we build lakehouses on; engines can change without moving data.
  2. 2TrinoFederated SQL across sources, and our route to zero-copy data virtualisation.
  3. 3Apache SparkStill the workhorse for large batch and unified batch-streaming ETL.
  4. 4Data meshDomain ownership of data products, where the organisation is large enough to need it.
  5. 5Great ExpectationsData quality checks as code, in batch and streaming workflows.

Trial

  1. 6DatabricksStrong managed platform; we integrate it as one engine among several.
  2. 7Zero-copy data sharingOpen sharing protocols that serve tables to other platforms without duplication.
  3. 8OpenLineageA common lineage standard that gives end-to-end traceability across pipelines.
  4. 9Apache PolarisAn open Iceberg catalogue that keeps table metadata independent of any one engine.
  5. 10DuckDBIn-process analytics for local development, testing and small workloads.
  6. 11MeltanoOpen-source extract and load, defined as code and versioned with the project.

Assess

  1. 12DataHubCapable data catalogue; the operating cost needs weighing per organisation.
  2. 13DaftA distributed dataframe engine built for multimodal and AI workloads.
  3. 14Apache PinotReal-time OLAP for user-facing analytics at low latency.
  4. 15StarRocksA fast analytical engine that can query lakehouse tables directly.
  5. 16Apache OssieUnder evaluation.

Hold

  1. 17Apache HiveSuperseded by open table formats and modern query engines.
  2. 18Hadoop and Cloudera CDHWe migrate clusters like these; we do not start new platforms on them.

Streaming and orchestration

Adopt

  1. 19Apache KafkaThe backbone for event streaming and real-time synchronisation.
  2. 20Apache AirflowDependable orchestration for batch pipelines and data product generation.
  3. 21Spark Structured StreamingStream processing that shares code and skills with batch Spark.
  4. 22Event-driven architecturesSystems that react to events as they happen, decoupled through a log.

Trial

  1. 23Kafka ConnectGood for standard sources and sinks; custom connectors need care.

Assess

  1. 24Apache FlinkStateful stream processing where latency and event-time accuracy matter most.
  2. 25Apache FlussStreaming storage built for real-time analytics alongside Flink.

Hold

  1. 26StreamSetsPipelines are easier to version, test and review as code.

Infrastructure and DevOps

Adopt

  1. 27KubernetesOur default runtime for services and for data workloads.
  2. 28DockerOne packaging format from a laptop to production.
  3. 29GitLab CI/CDPipelines for build, test and regular production releases.
  4. 30Grafana and PrometheusMetrics and dashboards we put in place before a system goes live.
  5. 31Amazon Web ServicesThe cloud we know best, including managed Spark on EMR.
  6. 32Microsoft AzureA dependable choice, especially where the client estate is already there.

Trial

  1. 33HashiCorp VaultSecrets belong in a vault, not in application databases or configuration.
  2. 34GitHub ActionsSimple and close to the code; fine for most delivery pipelines.
  3. 35MinIOS3-compatible object storage for on-premise and sovereign lakehouses.
  4. 36SeaweedFSLightweight object storage for self-hosted deployments.

AI, ML and application engineering

Adopt

  1. 37FastAPI and PydanticTyped Python services and REST APIs with validation built in.
  2. 38PostgreSQLThe first database to reach for, including full-text and vector search.
  3. 39MLflowExperiment tracking and model lifecycle that teams can standardise on.
  4. 40AI-assisted developmentCoding assistants speed delivery when paired with rigorous review.
  5. 41Model Context ProtocolA standard way to give AI assistants governed access to tools and data.
  6. 42Vector databasesEmbedding search, combined with full-text where it helps, behind AI features and data discovery.
  7. 43Spring BootSolid, well-understood JVM services.
  8. 44React and TypeScriptFor the interfaces a data platform needs.

Trial

  1. 45StreamlitFast internal tools and data apps; not for customer-facing products.

Assess

  1. 46Knowledge graphsExplicit relationships between entities, as grounding for AI and for discovery.
  2. 47JevUnder evaluation.

Weighing one of these for your platform?

Tell us what you are deciding. We will say what we would choose and why.