Service · Web development

Custom data warehouse development as the foundation for analytics and AI.

A central data layer on which your BI, dashboards, machine learning models and RAG applications can run. Built vendor-independently on Snowflake, BigQuery, Databricks or your own lakehouse: modular, with code ownership on your side, and GDPR-compliant from day one.

Snowflake / BigQueryLakehousedbt + AirflowELT with AirbyteAI-ready

A data warehouse is not a reporting database.

A data warehouse is the central storage and processing layer where you bring together all relevant data from your operational systems, transform it into reliable models, and make it available for analytics, BI, dashboarding and, increasingly, AI and machine learning applications. It is not a copy of your operational database; it is a purpose-built data model in which definitions are explicit, history is preserved and queries stay fast even across billions of rows.

We build data warehouses for mid-market and enterprise organisations that run multiple source systems (CRM, ERP, e-commerce, product analytics, marketing stacks) and want to consolidate their reporting and AI foundations into one reliable base. The conversation often starts with "our figures don't add up" or "every department calculates revenue differently" and ends with a data model in which business definitions are captured once and applied consistently through every dashboard, KPI report and LLM prompt.

A data warehouse rarely stands alone. It is usually the heart of a broader data engineering platform, with ingestion, orchestration, transformation and observability as surrounding layers. For the visualisation layer on top, one of two routes is usually sensible: a commercial BI tool such as Power BI or Tableau, or a custom BI layer that runs in your own brand style and can be embedded in your product. In the first conversation we advise which route suits, based on who will use it and which compliance requirements apply.

Four architectures for a data warehouse.

Traditional data warehouse (Inmon or Kimball): the classic approach using normalised data vault or dimensional star schemas. Suited to stable business processes where data models last for years and governance requirements are strict, such as financial reporting, annual accounts and regulatory data. Kimball, with fact and dimension tables, remains a strong choice for analytics-heavy workloads where speed and consistency outweigh flexibility. We build this on Postgres with Citus for on-premises deployments, or on a cloud warehouse such as Snowflake or Azure Synapse as data volumes grow.

Modern data stack (ELT and dbt): the standard of recent years. Raw data is loaded via Airbyte, Fivetran or Meltano into a cloud warehouse, transformations are built in dbt, orchestration runs through Airflow, Dagster or Prefect, and a BI layer sits on top. The ELT approach (extract, load, transform, rather than extract, transform, load) suits cloud warehouses that have the compute capacity to run transformations at query time. dbt has democratised the transformation layer: SQL with version control, tests, lineage and documentation as a by-product. For most mid-market organisations, this is the right starting architecture.

Lakehouse (Iceberg or Delta): the architecture that combines the best of a data lake and a data warehouse. It uses open file formats (Parquet) beneath a transactional table format layer (Apache Iceberg, Delta Lake, Apache Hudi) that delivers ACID guarantees without vendor lock-in. Suited to organisations with both structured and unstructured data that want several compute engines running on the same storage. Databricks delivers this with Delta Lake, Snowflake with Iceberg tables, or you can run your own open-source stack on S3 or GCS with Trino, Spark or DuckDB.

Data mesh (federated): for enterprise organisations with multiple business units or subsidiaries that want to keep managing their own data but need to collaborate across it. Each domain team owns its own data products, and a central team provides governance, the data catalogue and the federated query layer. It works well when the organisation is mature enough to carry decentralised ownership; for mid-sized companies it is usually overkill. We build mesh implementations on Snowflake Data Sharing, BigQuery Analytics Hub or a custom Trino layer.

Which route is right depends on your scale, use cases and organisational maturity. We deliberately don't always recommend the same thing, as one architecture for everyone is the wrong answer.

Three flavours of data warehouse.

Most warehouse projects fall into one of these three profiles. Which one fits depends on the number of source systems, the data volume, and how much you want to keep in-house. In the first conversation we'll discuss which route suits you and where the right cloud choice lies.

Compact project · fixed sprint budget

Single-cloud warehouse on BigQuery or Snowflake

A warehouse project in which we set up one cloud warehouse (BigQuery, Snowflake or Azure Synapse) with three to five source systems connected via Airbyte or Fivetran, dbt transformations with version control and tests, and Airflow or Dagster as orchestrator. Suited to mid-market organisations consolidating their data for the first time or building a first reporting layer. We deliver a central data model with core tables for your key entities (customers, orders, products, leads) plus the first mart layer for BI purposes. From there, your team can build further on it themselves, or you can have us continue development under a managed arrangement.

BigQuery / Snowflakedbt + AirflowAirbyte / FivetranPower BI / MetabaseGDPR-compliant
Mid-sized project · fixed sprint budget

Lakehouse with AI readiness

For organisations that have both structured and semi- or unstructured data (logs, JSON events, documents, images) and want their warehouse ready from the outset as a foundation for AI and LLM applications. We build a lakehouse on Databricks Delta Lake, Snowflake with Iceberg tables, or an open-source stack on S3 or Google Cloud Storage with Trino and DuckDB. On top come vector tables for RAG applications, feature stores for ML models, and a metadata catalogue that your AI development team can use without having to build data pipelines themselves. AI Act compliance layer included.

Databricks / IcebergDelta LakeVector storeFeature storeRAG-ready
Larger project · fixed sprint budget

Enterprise data mesh with federated governance

For enterprise organisations with multiple business units, subsidiaries or strict regulatory regimes (banks, insurers, healthcare groups, government). We build a federated architecture in which each domain team owns its own data product, with a central governance layer that manages the data catalogue, lineage, access control and quality monitoring. Components such as Snowflake Data Sharing, BigQuery Analytics Hub, or a Trino layer across multiple warehouses are combined with a data catalogue (DataHub, OpenMetadata, Collibra) and observability (Monte Carlo, Soda). Full GDPR and DPIA documentation, EU data residency, and a per-query audit trail.

Snowflake / TrinoDataHub / CollibraMonte Carlo / SodaEU residencyData mesh

What you get at the end.

A production-ready data warehouse that your team manages itself, plus everything around it so you can keep building without vendor lock-in. Code ownership is the starting point: all SQL, all DAGs, all infrastructure-as-code and all documentation sit with you in Git.

  • The warehouse itself, in productionProduction and staging environments in your cloud (GCP, AWS, Azure) or managed by us. Includes CI/CD, automated tests, monitoring, backups and a documented recovery procedure.
  • Source connectors and ingestion pipelinesConnected to your source systems: CRM (HubSpot, Salesforce, Pipedrive), ERP (SAP, NetSuite, Exact, AFAS), e-commerce (Shopify, Magento), product analytics (Mixpanel, Amplitude), marketing stacks (Meta, Google Ads), operational databases. Via Airbyte, Fivetran or Meltano, or with custom connectors where needed.
  • Transformation layer in dbtA complete dbt project with staging, intermediate and mart models, plus tests, exposures and documentation. Business definitions (revenue, margin, customer retention, CAC) are defined once in macros so they flow consistently through every report, dashboard and LLM prompt.
  • Orchestration and data qualityAirflow, Dagster or Prefect as orchestrator, with DAGs for all ELT jobs. Soda or Monte Carlo for data observability: freshness, volume checks, null rates, schema drift, and alerts to Slack or Teams if something breaks between source and mart.
  • Semantic layer and data catalogueA central place where business terms, column definitions and KPI formulas are documented, usable by your analysts, BI tool and any text-to-SQL or LLM layer. Integrated with DataHub, OpenMetadata or a lighter alternative, with lineage from source to dashboard.
  • BI and AI integrationsConnections to the BI tool of your choice (Power BI, Tableau, Looker, Metabase) or to a custom BI tool. For AI applications: feature store and vector tables ready for RAG, fine-tuning or classic ML models.
  • Compliance package and governancePII redaction in the staging layer, retention policies per table, EU data residency configuration where relevant, audit logging on every query, encryption in transit and at rest. GDPR documentation (DPIA), and for sector-specific regimes (NEN 7510 in healthcare, DORA in finance, the AI Act for ML models) the additional governance documents.
  • Complete codebase and documentationSource code, infrastructure as code (Terraform), dbt project, Airflow DAGs, architecture diagrams, ERDs, data dictionary, and an incident runbook. Everything lives in Git, with code ownership on your side, so another team can take over without vendor lock-in.
  • Training and ongoing managementSessions for your data engineers and analysts, a hands-on dbt workshop for the analytics team, and short videos for BI end users. Optionally, a management contract covering monitoring, security patches, cost optimisation and further development at a fixed monthly price.

When your own data warehouse is the right choice.

Six patterns in which organisations come to us for a warehouse project. If you recognise one of them, a conversation is usually worthwhile. We are honest about when a lighter alternative will do: sometimes a central dashboard is enough, and a full warehouse is over-investment.

Source system sprawl

Four or more systems that don't talk to each other

You have CRM, ERP, e-commerce, product analytics and marketing stacks, and each department exports separately to Excel. Definitions drift apart (sales calculates revenue differently from finance), reports differ by team, and nobody trusts the numbers any more. A central data layer in which everything comes together in one data model fundamentally solves this.

AI readiness

You want to apply LLMs or ML to your data

RAG applications, customer-specific chat assistants, predictive models for churn or fraud. All of these require a reliable data foundation on which feature stores and vector tables run. A warehouse is the typical prerequisite for serious AI development; without one there is no reproducibility, no lineage and no confidence in model outputs.

BI overlap

Dashboards for multiple departments

Operations, finance, marketing and management each want their own dashboard with their own KPIs. A central warehouse with a semantic layer is the natural foundation for a custom KPI dashboard per department, with consistently calculated metrics. No more conflicting definitions between reports.

Real-time challenge

Operational dashboards on live data

Logistics, energy, e-commerce, gaming: environments where a daily refresh is too slow and you need to steer on live status. A warehouse is often not enough in those cases; you need a real-time analytics platform alongside or on top of the warehouse, fed by event streams on Kafka or Materialize.

Compliance

GDPR, NEN 7510 or DORA requirements

Healthcare organisations handling patient data, financial institutions with DORA reporting, energy companies with regulatory logging. Governance, audit and data residency requirements become so stringent that ad hoc spreadsheets are no longer legally tenable. A warehouse with PII redaction, retention policies and EU data residency then becomes a necessity rather than a luxury.

Vendor independence

You want to escape a closed platform

You are locked into a SaaS reporting tool that does not give your data back, or into a legacy data warehouse whose licence costs rise faster each year than the value it delivers. A modular setup built on open formats (Parquet, Iceberg, dbt) gives you the option to switch compute engines in future without losing your data model.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

How a warehouse project works.

1

Introduction & use case discovery

A conversation to understand which decisions your organisation wants to support with data, which sources are involved, and which compliance requirements apply. We check straight away whether a full warehouse is the right investment; sometimes a lighter alternative is sufficient, and that is more honest advice than a heavier project.

2

Architecture & source inventory

A workshop with your team, plus interviews with the end users of the data (analysts, finance, operations). We draw up an architecture blueprint (warehouse choice, ELT pipeline, semantic layer, governance layer) and an initial data model for the key entities. At the end you have a defined scope, a technology stack choice and a realistic estimate of production cloud costs.

3

Foundation build in sprints

A working build every two weeks. We start with one source, one data model layer (staging, intermediate and the first mart) and basic orchestration. From sprint three onwards your team can already see first results in Power BI, Metabase or another BI tool. Subsequent sprints add sources and extend the data model.

4

Quality layer, observability and governance

dbt tests on all models, freshness and volume checks via Soda or Monte Carlo, setting up lineage and a data catalogue, and implementing PII redaction and retention policies. For AI applications, we also add the feature store and vector tables, with access controls that fit your position under the AI Act.

5

Compliance audit & penetration test

Before go-live: a penetration test on the cloud environment, a review of row-level security and encryption, an audit trail check, and data flow mapping for the DPIA. For sector-specific regimes we deliver the additional documentation that your data protection officer or regulator expects.

6

Rollout, training and ongoing development

Phased rollout to user groups, a hands-on dbt workshop for your analytics team, and ongoing management of security, performance and cost optimisation. Cloud warehouses can become significantly more expensive if nobody monitors query patterns, so we provide the tooling and runbook as standard to keep this under control.

The stack we use, and why we don't commit to a single vendor.

Our stack choices are pragmatic rather than religious. For ingestion, Airbyte and Fivetran are both good options, and the choice depends on your source portfolio and budget: Fivetran is more expensive but needs less maintenance, while Airbyte is open source and works excellently self-hosted. For custom-built connectors, we work with Meltano or our own Python layer. For orchestration, Airflow remains the de facto standard, but Dagster and Prefect have strengths in data asset tracking and developer experience respectively. We advise based on your team's skills and operational needs.

For transformation, dbt is the starting point: virtually all modern warehouse projects use dbt for SQL models with tests, lineage and documentation. For Python-based transformations (denoising, feature engineering, NLP) we use dbt-python, Polars or Spark, depending on scale. For the warehouse layer itself, we have extensive experience with Snowflake, BigQuery, Azure Synapse, Databricks, Redshift and ClickHouse, plus on-premises Postgres with Citus for situations where cloud is not an option.

For BI on top, we can work with any common tool (Power BI, Tableau, Looker, Metabase) or build a custom BI tool when embedded analytics or multi-tenancy is needed. For observability we use Monte Carlo or Soda, for the data catalogue DataHub or OpenMetadata. For AI integrations we work with LangChain, LlamaIndex or our own RAG orchestration, and we build vector tables on pgvector, Weaviate, Pinecone or Snowflake Cortex.

The thinking behind these choices is the same: you get a warehouse built on open standards (SQL, Parquet, Iceberg, dbt), so your team can switch vendors in future without losing your data model. Vendor independence is not a marketing slogan but an architectural principle. We build warehouses that you can maintain yourselves, or that another team of yours can take over.

Frequently asked questions.

What clients usually want to know before starting a warehouse project.

What exactly is a data warehouse?
A data warehouse is a central data system designed specifically for analytical workloads rather than operational transactions. You extract data from your operational systems (CRM, ERP, e-commerce, etc.) and transform it into a data model that makes analyses quick to answer, often in a star schema with fact and dimension tables, or in a data vault model. The difference from a data lake is that the data in a warehouse is explicitly modelled and transformed; the difference from your operational database is that the warehouse retains history and is optimised for analytical queries rather than point operations.
Which cloud warehouse do you usually recommend?
That depends on where your rest of the stack sits. For organisations already deeply invested in Google Cloud, BigQuery is usually the right choice: serverless, no capacity planning, native integration with the rest of GCP. For AWS organisations, Snowflake almost always falls into the same category, along with Redshift and Databricks. For Microsoft organisations, Azure Synapse or Databricks on Azure makes sense. For heavily operational workloads with high-cardinality dashboards (real-time analytics), ClickHouse is often the better choice. And for specific compliance requirements on data residency, we sometimes opt for on-premises Postgres with Citus.
What about GDPR and data residency?
Personal data (PII) in our setup receives a redaction layer at staging by default: the raw ingestion captures everything, but any field marked as PII is hashed, pseudonymised or fully redacted before it lands in the modelling layer. For EU data residency we configure the warehouse in an EU region (eu-west-3 for BigQuery, EU-Frankfurt for Snowflake, west-europe for Azure). Retention policies are set at column level, with automatic data deletion after the retention period. For a DPIA we deliver a data-flow mapping and a register of processing activities. For AI applications where PII ends up in features or vector stores, we extend this with an AI Act-compliant audit trail.
What is the difference with a data lake or a lakehouse?
A data lake is a raw storage layer for files (Parquet, JSON, CSV, logs) on object storage without a schema layer. You put everything in and work out what to do with it later. A data warehouse is the opposite: everything is explicitly modelled, transformed and optimised for analytical queries. A lakehouse combines both: open file formats on object storage (Parquet) with a transactional table-format layer (Iceberg, Delta Lake, Hudi) that provides ACID guarantees and warehouse-like query performance. For most mid-market organisations, a pure warehouse on BigQuery or Snowflake is the simplest choice. For organisations with large volumes of unstructured data or a strong focus on AI/ML, the lakehouse architecture becomes interesting.
How does a data warehouse compare to a data engineering platform?
A data warehouse is the central storage and compute layer. A data engineering platform is the broader context around it: ingestion tooling, orchestration, transformation frameworks, observability, governance, and the developer experience for your data team. In practice, the two overlap heavily: a serious warehouse project almost always includes a platform layer around it. Whether you tackle that as one project or as two consecutive phases depends on your starting position. If you already have a working warehouse but no proper ELT pipeline, we're more likely to talk about a platform project. If you're starting from scratch, a single combined project is usually more efficient.
How large does your organisation need to be to justify its own warehouse?
There's no fixed answer, but as a rough guide: an in-house warehouse becomes economically worthwhile when you have multiple source systems with overlapping entities, a data team or a strong desire to professionalise data work, and a use case that goes beyond purely ad-hoc reporting. Below that scale, a lighter solution, such as a central dashboard, a light ETL into an operational Postgres database, or a managed BI tool with direct source connectors, is often a better investment. We're honest about this: not every business needs a warehouse.
What about cloud costs?
Cloud warehouses usually work on a usage-based model: you pay for compute and for storage. Snowflake and BigQuery charge sharply for compute, so your query patterns strongly determine what you pay. The biggest pitfall is an unregulated team running queries that cause costs to explode. We therefore build cost monitoring in as standard (BigQuery slot reservations, Snowflake resource monitors, query-cost attribution per team) and run a monthly review of the most expensive queries. With every project we deliver a cost model for the first production year so there are no surprises.
Do you also build AI/ML models on the warehouse?
The warehouse itself is usually the right place to manage feature stores and vector tables. The models on top, whether classical ML (churn prediction, propensity scoring, fraud detection) or LLM applications (RAG, agentic flows, document classification), we build through our AI development practice. For classical ML we typically run Python frameworks (scikit-learn, XGBoost, LightGBM) on the warehouse data via dbt-python or Databricks. For LLM applications we connect the warehouse as a knowledge source to a RAG architecture with a vector store. Both require reliable data lineage and reproducibility, which is exactly what a good warehouse provides.
What if we already have a warehouse that no longer meets our needs?
Many of our projects are warehouse migrations rather than greenfield builds. Typical scenarios: a legacy on-premises warehouse on SQL Server or Oracle that has become too slow and too expensive to maintain, a Redshift cluster whose licensing costs and operational burden have grown heavy, or a first cloud warehouse set up without a modelling layer or governance that has since become a mess. We carry out an audit (where the pain lies, what the query patterns look like, what the current modelling layer is) and propose a migration path with parallel running, so you are never without reporting during the switchover. The migration often includes a data model clean-up at the same time, as the old setup has typically grown ad hoc over the years.
How long before we can go live?
For a first working warehouse on one cloud with three to five source systems, the foundation is in place after a few sprints, and your team can run the first dashboards in BI tools from that point on. A fully production-ready warehouse with a semantic layer, observability, governance, AI-readiness and multi-source integrations is a programme of several further sprints on top of that. We work in two-weekly sprints with a working build at the end of each, so you see early where things are heading and can steer along the way. For multi-tenant or mesh architectures the lead time is longer, as the governance and operating-model layer requires additional work.

Talk to us about your warehouse requirements.

A no-obligation introductory call of half an hour. We listen to who uses the data, which sources are involved and which compliance requirements apply, and we tell you honestly whether your own warehouse is the right route or whether a lighter alternative is enough. No sales funnel, just genuine sparring.

Fabian van Dijk
Business Developer · Appfront
fabian.vandijk@appfront.nl
Share LinkedIn Email

Edit content