Service · Software development

Custom data engineering platform development.

A production-ready, custom data engineering platform built around your sources, users and governance requirements. Ingestion, storage, transformation, quality and consumption as one coherent whole, rather than loose ETL scripts and tools nobody maintains any more. We build on managed components where we can, and write custom code where we must.

ETL pipelinesLakehouseStreamingData qualityOrchestration

Data engineering is the foundation, not just an integration.

A data engineering platform is something different from a few ETL scripts or a handful of API integrations. It is the complete foundation that carries your data from source to end user: reliable, repeatable and traceable. Where data integration consulting is about advice and building integrations between systems, and a real-time analytics platform focuses on streaming analysis, data engineering covers the pipelines, storage, orchestration and quality monitoring underneath.

We build data platforms for organisations that want to consolidate their source data into a data warehouse or lakehouse, replace their overnight batch jobs with a managed platform, or lay a foundation for analytics, machine learning and reverse ETL. No off-the-shelf package, no lock-in to a single vendor: a platform that fits your sources, scale and team.

The result is an environment in which new data sources land predictably, transformations are version-controlled, quality errors become visible before a dashboard displays them, and your analysts and data scientists work with data the business can trust.

In practice we see three main types of request: a scale-up without an existing warehouse that wants to start from scratch, an organisation with loose Python scripts and cron jobs that no longer scale, and an internal team that needs a reliable back-end dashboard on fresh data from multiple systems. For all three we build on managed components, choose the stack that suits your team size and budget, and make sure your people can ultimately manage it themselves.

We have worked for years with organisations that have outgrown Excel and standalone BI reports. For logistics companies, B2B SaaS businesses and operational teams in manufacturing and finance. Always with the same principles: choose managed where possible, write code where necessary, and make sure the platform stays readable for whoever has to build on it tomorrow.

Three types of data platform projects.

Depending on where your organisation stands: a platform built from scratch, a migration from loose scripts to a managed platform, or a focused back-end dashboard with fresh data from all your systems. In the first conversation we determine which project fits.

Greenfield platform · project spanning several sprints

Data engineering platform from scratch

For scale-ups and organisations without an existing data warehouse. We set up ingestion from your SaaS tools and databases, choose a lakehouse or warehouse suited to your scale, model the transformations, and put orchestration and monitoring in place. Delivered production-ready, with your team brought along for ongoing management. Typical scope: ingesting a handful of source systems, a staging and marts layer in dbt, one or two primary use cases live, and the governance foundations for further expansion.

Fivetran / AirbyteSnowflake / BigQuerydbtAirflow / Dagster
Migration · project spanning several sprints

Migrating from loose ETL scripts to a managed platform

For organisations with existing pipelines that no longer scale: cron jobs on standalone VMs, Python scripts without version control, unreliable overnight batches. We document the existing flows, rebuild them on a managed platform, and run both in parallel until we have confidence in the new pipeline. Only then do we switch over the downstream consumers, so dashboards and operational tools are never left without data.

Pipeline auditLift and shiftParallel runCutover
Back-end dashboard · compact project

Building a back-end dashboard on fresh data

An internal dashboard for your operations, finance or management team, fed by a lightweight data pipeline that brings together data from multiple sources and keeps it up to date. Not a Power BI report on an Excel export, but a small platform with its own ingestion, models and front end. Suited to a single focused question (margin per customer, stock rotation, OEE overview) for which BI tooling is overkill, but Excel has become unreliable.

Reverse ETLEmbedded analyticsCustom front endAuthentication and roles

What your platform contains at the end.

Not just pipelines that move data, but a complete data environment that your team can extend and manage itself.

  • Ingestion layer for your sourcesAPI pulls, change data capture (CDC) on transactional databases, file drops, and streaming ingestion via Kafka or Confluent where it fits.
  • Storage at the right layerData lake on S3 or ADLS for raw and semi-structured data, warehouse (Snowflake, BigQuery, Redshift, ClickHouse) or lakehouse such as Databricks Delta for the structured layer.
  • Transformation layer in dbt or SparkVersion-controlled SQL models or PySpark transformations, with tests, lineage and documentation. Custom Python or dataflow where the complexity calls for it.
  • Orchestration and schedulingAirflow, Dagster, Prefect or Mage, depending on your team and scale. Pipelines as code, with retries, alerts and clear dependencies between jobs.
  • Data quality and observabilityGreat Expectations, dbt tests, Soda or Monte Carlo for data quality testing. Anomalo or Datafold for monitoring drift and regressions. You'll know about failures before the business notices them.
  • Catalogue and governanceA data catalogue (Atlan, Unity Catalog, OpenLineage) so tables are discoverable, owners are known, and lineage from source to dashboard is visible. Including row- or column-level access policies where needed.
  • Consumption layerPower BI, Tableau or Looker for BI; embedded analytics in your own applications; reverse ETL via Hightouch or Census into operational tools; or a custom KPI dashboard.
  • Documentation and knowledge transferArchitecture overview, runbooks, a data dictionary and hands-on training for your analysts and data engineers. An optional maintenance contract is available for ongoing development.

When a data engineering platform is the right choice.

Four common situations in which clients come to us. If one of them sounds familiar, we'd be happy to talk further.

Greenfield

Scale-up without a warehouse

You're growing fast and the ad-hoc dashboard running on a production replica no longer holds up. You need a proper platform — ingestion, modelling, quality — before the whole organisation starts relying on the same Looker views and nobody remembers how the definitions came about.

Migration

Ad-hoc ETL scripts don't scale

Python scripts on standalone VMs, cron jobs without retries, nightly batches that sometimes fail to run. You want to move to a managed platform with versioning, alerts and lineage, without losing the existing flows or leaving the business without data in the meantime.

Multi-source

Ten-plus SaaS sources, one business view

HubSpot, Salesforce, Stripe, NetSuite, Zendesk, Mixpanel, production database — all with their own customer ID and their own definition of "active". You want a single view of the business, with clear definitions, governance and ownership for each metric.

ML & reverse ETL

Getting data back into operational tools

You want to write scores, segments or predictions from the warehouse back into HubSpot, Intercom or your own app. Or a feature store for ML models running in production, with versioning and monitoring so data drift is spotted in time.

Hybrid

Batch and streaming side by side

For most use cases batch is fine, but one or two processes need fresh data within seconds. You don't want to run two parallel platforms — we combine batch and streaming on a single stack, with clear patterns for when each path is used.

Backend dashboard

Operational dashboard on fresh data

A team — finance, operations, customer success — needs a dedicated dashboard that is always up to date and combines data from several systems. Excel exports and BI tools fall short; you want something that runs in a browser, with authentication and roles, fed by a lightweight pipeline.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

How a platform project runs with us.

1

Introductory meeting and source inventory

A conversation in which we establish which sources you have, which use cases the platform must serve, and what your team looks like. We review existing pipelines, BI tools, data ownership and the pain points that prompted you to start this project. It produces an initial direction: greenfield, migration or a dedicated backend dashboard.

2

Architecture and stack selection

A workshop with your data and IT team. At the end you have a reference architecture: which ingestion tools, which warehouse or lakehouse, which orchestration and quality framework, and which catalogue. Including a reasoned yes or no on alternatives, so you know not only what we choose but also why we don't choose the others. An initial plan and sprint roadmap are included.

3

Building in sprints

Each sprint delivers a working slice: one source fully ingested, one use case modelled end-to-end, or one quality check running live. You can test along the way, and your analysts see data flowing in. Over a multi-sprint engagement, the core platform takes shape with the first use cases, built in thin vertical slices rather than months of "ingestion phases" with no visible results.

4

Rollout and knowledge transfer

Production cutover, parallel runs where possible, documentation and hands-on sessions for your team. We write a runbook for the most common issues, a data dictionary for the marts layer, and an onboarding document for new analysts or engineers. Afterwards, optional ongoing management for security, monitoring and further development.

Frequently asked questions.

What clients usually want to know before we start.

Do you replace Snowflake or Databricks?
No. We build on managed platforms such as Snowflake, BigQuery, Redshift, ClickHouse or Databricks, and choose together which platform suits your scale, budget and team. We build the ingestion, models, orchestration, quality and consumption layers around it. The warehouse or lakehouse itself remains your managed service, so there's no lock-in to us. For heavier ML and data science workloads we more often look at Databricks; for classic analytics, Snowflake or BigQuery is often sufficient and cheaper to run.
Airflow or Dagster: what do you recommend?
Both are solid. Airflow is mature and has the broadest integration ecosystem; for teams already working with Airflow, it's often the obvious choice. Dagster and Prefect are more modern: asset-oriented, with a better developer experience and stronger lineage out of the box. For a new platform we more often choose Dagster; for migrations we more often rely on Airflow so existing DAGs can be reused. Mage is a lightweight alternative for smaller teams that don't want to run a full-blown scheduler. We choose based on your existing tooling, your team and the complexity of your dependencies, not on what's popular on LinkedIn.
Which tooling do you use for data quality testing?
In the transformation layer, dbt tests for uniqueness, not-null, referential integrity and custom business rules. For heavier checks on ingestion or staging tables, Great Expectations or Soda. For production monitoring of anomalies, freshness and drift, Monte Carlo, Anomalo or Datafold, depending on your budget and the metadata you already have. Data quality is a layer of the platform, not a separate tool bolted on afterwards. We explicitly model which tables are "trusted" and which are "raw" or "staging", so downstream consumers know what they can rely on.
Do you also do streaming with Kafka or Confluent?
Yes. For use cases that can't wait for a nightly batch, such as fraud detection, real-time personalisation or operational dashboards, we build on Kafka or Confluent Cloud, with Debezium for CDC and stream processing in Flink or Spark Structured Streaming. For pure analytics questions, batch is often cheaper and simpler; we choose per use case, not on ideology. A hybrid setup often makes sense: streaming for the few use cases where latency truly matters, batch for the rest. For pure real-time analysis, see also our page on a real-time analytics platform.
Is MLOps part of the platform?
Optional. We can place a feature store and model registry on top of the platform (Feast, MLflow) so data scientists work on the same data as the BI layer. Training and deploying models we handle in a separate engagement or in collaboration with your data science team. The data platform needs to be in place before MLOps makes sense: models built on unreliable data are a faster route to distrust than to value. Once the data foundation is in place, connecting the two is straightforward.
What determines the cost of a data platform engagement?
Mainly three factors: the number and complexity of the sources we need to ingest (a Salesforce instance with 200 custom fields is a different proposition from a standard Stripe connector), the complexity of the transformations and data models, and the level of maturity you want for governance, quality and observability. A first working platform with a few sources is very different from an enterprise platform with a catalogue, lineage and strict access policies. The licence costs of the chosen tooling also play a part, such as Snowflake or Databricks, Fivetran or Airbyte, Monte Carlo or dbt tests only, but those costs are billed directly to your account, not through us. We work on a fixed sprint budget and, after the introductory conversation, give you a realistic picture.
Do you work alongside our in-house data engineers?
Almost always. We prefer to work with your team rather than around it: pair programming, code review, joint sprint planning. Knowledge transfer happens in every sprint, not just at the end. For organisations that don't yet have an in-house data engineer, we can help with recruitment and onboarding once the platform is in place. See also our approach to enterprise software development.

Talk to us about your data engineering platform.

A thirty-minute introductory call, with no obligation. We listen to your sources, use cases and current data landscape, and give direction you can act on, even if you ultimately choose a different route.

Edit content