Technology radar
What we build with, and how firmly we recommend it. Every placement comes from running the technology in production on our engagements.
- Adopt
- Our default choice. Proven in production on our engagements.
- Trial
- Used on real work and worth choosing where it fits.
- Assess
- Promising. We evaluate it before committing a client to it.
- Hold
- We have run it in production and would not start new work on it.
Data platforms and lakehouse
Adopt
- 1Apache IcebergOpen table format we build lakehouses on; engines can change without moving data.
- 2TrinoFederated SQL across sources, and our route to zero-copy data virtualisation.
- 3Apache SparkStill the workhorse for large batch and unified batch-streaming ETL.
- 4Data meshDomain ownership of data products, where the organisation is large enough to need it.
- 5Great ExpectationsData quality checks as code, in batch and streaming workflows.
Trial
- 6DatabricksStrong managed platform; we integrate it as one engine among several.
- 7Zero-copy data sharingOpen sharing protocols that serve tables to other platforms without duplication.
- 8OpenLineageA common lineage standard that gives end-to-end traceability across pipelines.
- 9Apache PolarisAn open Iceberg catalogue that keeps table metadata independent of any one engine.
- 10DuckDBIn-process analytics for local development, testing and small workloads.
- 11MeltanoOpen-source extract and load, defined as code and versioned with the project.
Assess
- 12DataHubCapable data catalogue; the operating cost needs weighing per organisation.
- 13DaftA distributed dataframe engine built for multimodal and AI workloads.
- 14Apache PinotReal-time OLAP for user-facing analytics at low latency.
- 15StarRocksA fast analytical engine that can query lakehouse tables directly.
- 16Apache OssieUnder evaluation.
Hold
- 17Apache HiveSuperseded by open table formats and modern query engines.
- 18Hadoop and Cloudera CDHWe migrate clusters like these; we do not start new platforms on them.
Streaming and orchestration
Adopt
- 19Apache KafkaThe backbone for event streaming and real-time synchronisation.
- 20Apache AirflowDependable orchestration for batch pipelines and data product generation.
- 21Spark Structured StreamingStream processing that shares code and skills with batch Spark.
- 22Event-driven architecturesSystems that react to events as they happen, decoupled through a log.
Trial
- 23Kafka ConnectGood for standard sources and sinks; custom connectors need care.
Assess
- 24Apache FlinkStateful stream processing where latency and event-time accuracy matter most.
- 25Apache FlussStreaming storage built for real-time analytics alongside Flink.
Hold
- 26StreamSetsPipelines are easier to version, test and review as code.
Infrastructure and DevOps
Adopt
- 27KubernetesOur default runtime for services and for data workloads.
- 28DockerOne packaging format from a laptop to production.
- 29GitLab CI/CDPipelines for build, test and regular production releases.
- 30Grafana and PrometheusMetrics and dashboards we put in place before a system goes live.
- 31Amazon Web ServicesThe cloud we know best, including managed Spark on EMR.
- 32Microsoft AzureA dependable choice, especially where the client estate is already there.
Trial
- 33HashiCorp VaultSecrets belong in a vault, not in application databases or configuration.
- 34GitHub ActionsSimple and close to the code; fine for most delivery pipelines.
- 35MinIOS3-compatible object storage for on-premise and sovereign lakehouses.
- 36SeaweedFSLightweight object storage for self-hosted deployments.
AI, ML and application engineering
Adopt
- 37FastAPI and PydanticTyped Python services and REST APIs with validation built in.
- 38PostgreSQLThe first database to reach for, including full-text and vector search.
- 39MLflowExperiment tracking and model lifecycle that teams can standardise on.
- 40AI-assisted developmentCoding assistants speed delivery when paired with rigorous review.
- 41Model Context ProtocolA standard way to give AI assistants governed access to tools and data.
- 42Vector databasesEmbedding search, combined with full-text where it helps, behind AI features and data discovery.
- 43Spring BootSolid, well-understood JVM services.
- 44React and TypeScriptFor the interfaces a data platform needs.
Trial
- 45StreamlitFast internal tools and data apps; not for customer-facing products.
Assess
- 46Knowledge graphsExplicit relationships between entities, as grounding for AI and for discovery.
- 47JevUnder evaluation.
Weighing one of these for your platform?
Tell us what you are deciding. We will say what we would choose and why.