Modern data stack: Layers, tools, and design guide

Contents
A modern data stack fails most often not from a bad tool choice, but from bolting cloud-native components onto a batch-era mental model. Teams swap Redshift for Snowflake, add dbt, and still wait days for trusted metrics because the semantic layer, observability, and governance pieces never made it into the design.
This guide breaks down what actually composes a modern data stack, which tools own which layer, and how to sequence adoption so you're not just modernizing infrastructure but rewiring how data moves and gets trusted across the org.
What is a modern data stack?
A modern data stack (MDS) swaps the single warehouse-plus-cron-script pattern for a modular set of cloud-native tools connected by ELT (extract, load, transform) pipelines instead of hand-coded ETL jobs. Ingestion pipelines land raw data first; the cloud data warehouse then becomes the compute engine that drives transformation, modeling, and analytics - not a passive storage dump.
The practical effect of a migration is easy to describe: raw data lands untouched, dbt models it inside the warehouse itself, and the brittle custom-script layer that broke on every schema change disappears.
IDC's Data Age 2025 study projected the global datasphere would reach 175 zettabytes by 2025, a growth curve legacy point-to-point pipelines were never built to absorb.
That scale is why organizations now assemble tools around people and use cases instead of one vendor's monolith, though tool sprawl carries its own total-cost-of-ownership problem worth weighing before the promise of faster insights turns into another dashboard nobody maintains.
The layers of a modern data stack
A modern data stack breaks into six layers: ingestion, storage, transform, semantic, activation, and observability. Each layer is a swappable part, not a monolith, which is the whole point of the MDS pattern versus a legacy warehouse build.
| Layer | Job | Common tools |
|---|---|---|
| Ingestion | Pull data from source systems, often via CDC | Fivetran, Airbyte |
| Storage | Hold raw and modeled data at scale | Snowflake, cloud data warehouse or lakehouse architecture |
| Transform | Model raw data into business logic | dbt (data build tool) |
| Semantic | Define metrics once so every tool agrees | dbt Semantic Layer, Cube |
| Activation | Push modeled data back into business tools | reverse ETL (Hightouch, Census) |
| Observability | Catch schema drift and pipeline failures early | Monte Carlo, Metaplane |
The semantic layer is the layer teams skip first and regret most. Without it, sales and finance build their own "revenue" logic in separate dashboards, and the numbers stop matching. A semantic layer sits between storage and every downstream tool, so people ask questions and get one consistent answer, not five.
Reverse ETL closes the loop the warehouse used to leave open: it pushes governed data back into Salesforce, HubSpot, or an ad platform, so operational teams act on the same numbers analytics sees.
More layers mean more contracts to maintain, and tool sprawl is now the main line item hiding inside MDS total cost of ownership.
Budgets are growing to absorb it: in dbt Labs' 2025 State of Analytics Engineering report, 30% of data teams reported budget growth, up from 9% the year before (dbt Labs via PR Newswire). Teams that skip the observability layer tend to find schema drift in a dashboard, weeks after it broke a pipeline, instead of in an alert.
A lakehouse architecture, layering warehouse-style modeling on top of open storage formats, is where several of these layers start to merge, which is the direction the stack follows next.
Which tools power each layer
Fivetran, Snowflake, and dbt cover three of the six layers in many modern data stack builds, and which tool serves each part matters less than the build-versus-buy call behind it.
Fivetran handles ingestion for teams that would rather pay per row synced than maintain custom connectors; Airbyte covers the same layer for teams with tighter budgets and a tolerance for self-hosting. Both push data into a cloud data warehouse rather than a self-managed cluster, the ELT (extract, load, transform) bet the whole MDS pattern now rests on.
Snowflake, alongside Databricks and BigQuery, dominates storage at scale. The three differ mainly on compute pricing and on how well they support lakehouse architecture for semi-structured data. Snowflake's dominance owes much to its three-layer architecture design, which separates storage, compute, and services for independent scaling.
For most analytics teams, dbt (data build tool) is the default transform layer. It turns SQL models into a version-controlled DAG, which is what makes schema drift catchable in CI rather than in a stakeholder's dashboard.
Higher up the stack, the semantic layer (Cube, dbt Semantic Layer) and reverse ETL (Census, Hightouch) handle activation, pulling modeled insights back out to the tools people actually use for sales and marketing follow-up.
Tool sprawl is the tax nobody budgets for. Six best-of-breed tools often cost more in integration and data observability overhead than one less elegant, more coupled platform.
Ingestion: Airbyte, Fivetran, CDC
Airbyte and Fivetran cover ingestion for most modern data stack builds, and the choice comes down to who owns the connector maintenance. Fivetran charges per row synced and handles schema drift automatically; Airbyte is open source, self-hosted, and cheaper at scale if your team can absorb the ops burden of running connectors against a growing warehouse.
Change data capture is the real latency lever underneath both tools.
Batch ingestion on a schedule works for reporting that tolerates a few hours of staleness; CDC streams row-level changes as they happen, which is what teams need when a cloud data warehouse feeds anything closer to real time.
Moving from scheduled batch loads to CDC can cut pipeline lag from hours to minutes without touching the transformation layer downstream. Match the ingestion pattern to how fresh the data actually needs to be, not to what's newest.
Warehouse vs lakehouse: Snowflake, BigQuery, Databricks
Snowflake, BigQuery, and Databricks now answer different questions about the same modern data stack problem: where do compute and storage decouple, and who pays for flexibility. Snowflake and BigQuery separate storage from compute cleanly, so a cloud data warehouse scales query power without re-architecting ingestion.
Databricks pushes further into lakehouse architecture, storing raw and curated data together and letting the same cluster serve SQL analytics and ML workloads. Teams weighing this tradeoff often bring in specialized Snowflake development services to design the ingestion and compute architecture correctly from the start.
The tradeoff shows up in total cost of ownership, not sticker price. Warehouses win on time-to-first-dashboard; teams get a working semantic layer over Snowflake in weeks. Lakehouse setups cost more engineering effort upfront but avoid duplicating storage for analytics and machine learning.
Both platforms are mature choices: Snowflake and Databricks have each been named Leaders in recent editions of Gartner's Magic Quadrant for Cloud Database Management Systems. If Snowflake is on your shortlist, start with what Snowflake is and how it's priced.
What is a semantic layer, and why it's the bottleneck
A semantic layer is the translation layer between raw warehouse tables and the metrics people actually argue about in a meeting: revenue, active users, churn. It sits between storage and the tools people query, so "active user" means the same thing whether someone opens Looker, pulls a spreadsheet, or triggers a reverse ETL sync into Salesforce.
The bottleneck isn't storage or ingestion anymore. Cloud data warehouse compute is cheap and elastic; the modern data stack solved that part years ago. What breaks at scale is metric-layer coupling: every BI tool wants to own its own definition of a metric, and each one drifts.
Looker's LookML was an early fix, but it tied metric logic to one BI tool. Teams that also needed that metric in a notebook, a reverse ETL push, or a data science model had to redefine it, and definitions diverged.
The typical symptom is three revenue numbers in finance, product, and sales dashboards, each technically correct and mutually contradictory. Centralizing metric logic in a semantic layer, decoupled from any single BI tool, is the pattern we recommend by default.
ETL vs ELT: When each still makes sense
ELT (extract, load, transform) is the default for most organizations building a modern data stack today. ETL still wins when data must be masked, filtered, or de-identified before it ever touches a cloud warehouse.
The split follows compliance, not preference.
A payments company is the clearest example. Tokenizing card numbers before they reach the warehouse keeps the warehouse out of PCI DSS scope, which is far cheaper than securing and auditing it to card-data standards.
Healthcare and European teams make similar calls for different reasons. A cloud warehouse can hold PHI under HIPAA if the vendor signs a business associate agreement (HHS), and GDPR allows personal data in a warehouse with a lawful basis. But data minimization, access-control simplicity, or a vendor contract that doesn't cover sensitive data often make it easier to mask or drop those fields before load.
ETL keeps that transformation step under tighter control, at the cost of slower iteration and more rigid workflows.
Everywhere else, ELT wins on scale and speed.
Tools like Fivetran, Airbyte, and Matillion push raw data into storage first; dbt handles transformation afterward, version-controlled and testable.
That ordering, plus a large open-source community building around it, is why ELT dominates analytics engineering now.
Our view: default to ELT, carve out ETL only where a named regulatory constraint (HIPAA, GDPR, PCI-DSS) forces masking before load. Mixing both without a clear rule, across mismatched technology and data stacks, is how tool sprawl and its total cost of ownership creep up on a data platform team.
How AI is changing the modern data stack
AI is changing modern data stack design by pushing structure earlier in the pipeline, not later. Agents built on top of the semantic layer only produce accurate insights if the underlying metrics are governed before an LLM ever queries them, at any scale.
That reverses the old assumption that business intelligence platforms clean up ambiguity downstream. Now dbt (data build tool) models double as data contracts: version-controlled, tested definitions that an AI agent references instead of re-deriving "revenue" from raw warehouse tables during processing.
This matters for prompt design too. When an agent's context window includes a governed metric definition rather than raw schema information, it stops guessing at business logic and starts citing a source of truth, which is a fundamentally different workflow than traditional BI.
Organizations that skip semantic governance get agents that hallucinate metrics with confidence, and confident hallucination is worse than a broken dashboard because no one questions the output during analysis.
Reverse ETL is following the same logic.
Instead of pushing transformed data back to Salesforce or HubSpot for people to read, it increasingly feeds context to AI agents working inside those same tools, closing the loop between data ingestion and action.
The same dbt Labs 2025 report found that 45% of respondents named AI tooling as a key investment priority for the year ahead, while poor data quality remained the most frequently reported challenge, cited by 56% (PR Newswire). That pairing is the whole argument: AI investment is rising faster than the data it depends on is improving.
Storage and ingestion technology matter less here than governance does.
The stacks that win won't be the ones with the newest tools. They'll be the ones where every agent, workflow, and dashboard draws from the same governed definitions.
Governance, security, and data observability
Governance in the modern data stack now means enforcing rules at the ingestion layer, not policing dashboards after the fact. A data catalog gives engineers and business people one place to check lineage, ownership, and freshness for every table moving through Fivetran, dbt, and the cloud data warehouse. Without it, organizations scale their tool count faster than their trust in the data those tools produce.
Data observability closes the gap a catalog alone can't. It watches volume, schema, and distribution patterns in near real time, catching a broken load before it reaches the semantic layer or a reverse ETL job pushes bad rows back into a CRM. In practice, the pattern that saves the most engineering time is upstream: catching schema drift in staging, before it corrupts a production model.
At current data volumes, manual governance is a losing proposition without automated cataloging and observability built into the stack from day one.
Security follows the same logic as storage and ingestion: push it into the pipeline definition, not a separate review step.
Row-level access policies defined once in Snowflake or the warehouse layer should propagate through every downstream tool, including whatever consumes insights next, rather than getting re-built per dashboard.
Teams that skip this step end up governing the same field five different ways across five different tools, which is its own quiet tax on total cost of ownership.
Orchestration: Airflow vs Dagster
Airflow suits task-based DAGs with mature scheduling; Dagster suits asset-based orchestration where data quality and lineage matter as much as the run itself. Airflow treats a pipeline as a sequence of tasks, it doesn't know or care what a task produces, just that it ran.
Dagster models the DAG around the assets themselves (a dbt model, a Snowflake table, a Fivetran sync), so failures surface as "this table is stale" rather than "this task exited non-zero."
Teams running a lean modern data stack with heavy dbt usage often move to Dagster once schema drift and partial loads start causing silent downstream breaks. The asset graph catches what a task-based DAG misses.
Airflow still wins on plugin count and hiring pool depth, which matters when the orchestration layer has to scale across a growing tool sprawl rather than one clean stack.
What comes after the modern data stack
The modern data stack is splitting into two directions: lakehouse architecture that collapses the warehouse-versus-data-lake tradeoff, and a semantic layer that decouples business logic from any single BI tool. Neither replaces MDS. Both extend it.
Lakehouse architecture (Databricks and Snowflake's Iceberg support are the clearest examples) lets teams store raw and structured data in open formats, then query it with warehouse-grade performance without duplicating storage. That matters when ingestion volume outgrows what a pure ELT-into-warehouse pattern was priced for. Once orchestration moves data reliably, the next question becomes where and how to store it.
If ingestion volume and architectural complexity outpace your team's capacity, partnering with a data engineering agency can help design and build the lakehouse layer correctly.
The semantic layer, meanwhile, answers a different problem: metric drift across dashboards. Tools like dbt's own semantic layer or Cube define revenue, churn, and active-user logic once, then serve it consistently to every BI tool and reverse ETL destination downstream.
Data mesh proposes something more radical, domain teams owning their own pipelines and infrastructure rather than a central platform team owning the stack.
It works at organizations with genuine domain complexity and tends to stall everywhere else, because it multiplies tool sprawl and total cost of ownership rather than reducing it.
If you lack that domain-aligned engineering capacity in-house, choosing the right data engineering company matters more than choosing the right tools.
Our view: a composable stack with a strong semantic layer gets most of data mesh's benefits without the organizational overhead, and it's where we'd start.
FAQ: Semantic layers, data mesh, ETL vs ELT, 2026 tools
What is a semantic layer and why does it matter?
Modern data stack vs data mesh: What's the difference?
When should you use ETL instead of ELT?
What are the best modern data stack tools in 2026?
Designing your modern data stack: Where to start
Start with an audit, not a rebuild: map dbt transformations and your Snowflake warehouse against the questions the business actually asks, then fix the weakest link, usually the semantic layer or observability, not storage.
A modern data stack scales when insights stay consistent across tools and teams, not when vendors change every year.
For a mid-market engineering organization with a small data team, the order of adoption matters more than the tool list. A sequence that keeps each step useful on its own:
- Ingestion and storage first. Connect the three to five source systems the business already argues about (usually the CRM, billing, and the product database) into one cloud warehouse. Managed connectors are worth paying for at this stage; a two-person team shouldn't be maintaining connector code.
- Transformation with tests from day one. Put every model in dbt under version control, with basic tests for uniqueness, nulls, and accepted values. Tests written later rarely get written.
- One governed metrics layer. Define the handful of metrics leadership reviews weekly (revenue, active users, churn) once, and point every dashboard at those definitions. This is the step that ends the "whose number is right" meetings.
- Observability once pipelines matter. Add freshness and volume alerts when a broken load would reach a customer-facing report or an AI feature, not before.
- Activation and orchestration last. Reverse ETL and an asset-aware orchestrator pay off once the models underneath are trusted. Before that, they just move untrusted numbers faster.
Two rules keep the stack from sprawling. Add a tool only when a named owner will maintain it. And retire something every time you add something, so the total count stays close to the number of layers you actually use.
Netguru's data engineering team runs this kind of audit, then builds or fixes the layers that are holding the stack back. If your metrics don't match across dashboards or your pipelines break on every schema change, talk to our team.
