Introduction
One of the most common questions I see on the Fabric community forums :
"Should I build this transformation in a Dataflow Gen2 or a Fabric Notebook?"
The official documentation says: "Use Dataflows for low-code transformations, use Notebooks for complex logic." That's not a decision framework. That's a slogan.
In practice, the wrong choice costs you CUs, debugging time, and maintainability. Here is a framework that actually answers the question built from running both tools on real workloads.
The 5 dimensions that matter
Before the decision tree: these are the factors that actually change the answer.
Dimension Why it matters
The decision tree
The CU cost reality
Dataflow Gen2 billing
Dataflow Gen2 bills per output query, at :
- 12 CU/second for the first 10 minutes of each query
- 1.5 CU/second beyond 10 minutes (8× cheaper)
Plus optional: High Scale Compute (Staging): 6 CU/s flat when enabled | Fast Copy: 1.5 CU/s when applicable.
Example : 10 output tables, each runs 3 minutes:
10 tables × 3 min × 60s × 12 CU/s = 21,600 CU consumed
Same 10 tables with query referencing :
1 source query (shared) + 10 small filter queries ≈ 80% reduction
Spark billing (Notebook / Spark Job Definition)
Spark jobs whether run via a Notebook or a Spark Job Definition bill per Spark session (not per query) :
- Starter pool: starts at 4 CU/s (8 vCores), scales up to a minimum of 8 CU/s after proactive scale-up. Custom Pool Small: 2 CU/s.
- Session is shared across all cells all your table loads in one run share the same session cost.
Example : 10 table loads in one Notebook :
1 session × 10 min × 60s × 4 CU/s = 2,400 CU consumed
Head-to-head comparison
Scenario Dataflow Gen2 vs Spark (Notebook)
Rule of thumb : For many small tables from the same source, Dataflow Gen2 with query referencing wins. For large volumes or complex logic, Notebooks win.
Scenario deep-dives
Scenario A : Loading a SharePoint Excel file into 15 Lakehouse tables
Use Dataflow Gen2. Create one base query that reads the Excel file, then reference it 15 times one per output table. SharePoint is read once; 15 tables are derived in-memory.
If you use a Notebook here, you open the Excel file 15 times (15 HTTP calls to SharePoint), parse it 15 times, and pay for 15 separate Spark reads of a non-native format.
Scenario B : Building a Gold layer fact table with SCD Type 2
Use a Notebook (Spark). This is not expressible in Power Query. Don't try.
This is not expressible in Power Query. Don't try.
Scenario C : Daily incremental load, 200k rows from SQL Server
Use Dataflow Gen2 with the watermark pattern.
- // Watermark helper query (Power Query M) LastLoadDate = List.Max(Table.Column(Lakehouse.Contents(...){[Name="fact_sales"]}[Data], "load_date"), #date(2020,1,1))
- // Main query — filter source before loading FilteredRows = Table.SelectRows(Source, each [created_at] > LastLoadDate)
A Spark job would cost 10–15× more per run for the same logic: the Spark session (4–8 CU/s) runs for its entire duration, even for a small incremental load.
Scenario D : 200M row fact table, complex window functions
Use a Notebook (Spark). Dataflow Gen2's Mashup engine runs on a single node at this scale you will hit memory limits or extreme runtimes. Spark distributes the workload across nodes.
When to use both together
The most powerful pattern is a pipeline that combines both :
Dataflow Gen2 → light extraction + normalization → Lakehouse (Silver) → Spark Notebook → complex transformation → Lakehouse (Gold)
Neither tool is asked to do something it's bad at.
Quick reference card
If you need...Use
Update May 2026
Thanks to Miles Cole (Principal PM @ Azure Data CAT) for technical corrections on Spark billing and pool sizing. Article updated accordingly.
References
- https://learn.microsoft.com/en-us/fabric/data-factory/pricing-dataflows-gen2
- https://learn.microsoft.com/en-us/fabric/data-engineering/how-to-use-notebook
- https://learn.microsoft.com/en-us/fabric/data-factory/tutorial-setup-incremental-refresh-with-dataflows-gen2
- https://learn.microsoft.com/en-us/fabric/data-engineering/lakehouse-and-delta-tables