Forum Discussion
Incremental refresh for dataflow gen2 and the BC connector
- 10 months ago
Hi SofBL
After going through your post, as per my thinking in your case, the key difference is that incremental refresh in a dataset and in a Dataflow Gen2 are enforced differently by the engine. In a Power BI dataset, the RangeStart and RangeEnd parameters only need to exist logically in the M query pipeline, and as long as the final query folds and the filter logic is applied somewhere upstream,even if the datetime column disappears later,the incremental refresh policy can still activate. However, Dataflow Gen2 enforces the rule at the final query stage: it requires the DateTime column to physically remain in the output table when configuring incremental refresh. Because your looping function expands company data and removes the date field before the final output, Dataflow Gen2 treats it as lacking a valid incremental column, even though function inputs use RangeStart and RangeEnd. That is why the same M logic works in a dataset but fails in Dataflow Gen2.
Additionally, Dataflow Gen2 currently does not allow append mode with incremental policies because the lakehouse engine expects to completely manage partitions for incremental loads. It needs replace mode so it can rewrite partition files; otherwise, append mode would lead to inconsistent partitions. This also explains the quarterly bucket size limitation, the engine needs predictable partition logic. The folding error you encountered further confirms that the Business Central connector still has folding limitations within Dataflow Gen2, even if it folds properly in a dataset, so the Fabric engine does not trust the query for incremental partition pruning.
I think by loading the full BC dataset once into Delta Lake, then using a second Dataflow Gen2 to append only rows with higher entry numbers,is a very practical pattern. It bypasses the datetime requirement, avoids folding constraints, and simulates incremental load efficiently by key-based tracking. This approach is commonly used in data engineering pipelines for immutable ledger-type tables, so you did not miss anything,in fact, you implemented a correct pattern where date-based incremental refresh isn’t suitable. Microsoft does not promote this pattern in official docs because incremental refresh in Fabric is positioned as a date-driven, fold-dependent mechanism, whereas key-based delta ingestion is an engineering workaround outside the standard UI-driven model. But in real-world Business Central workloads with multiple companies and immutable entry tables, your entry-no-based incremental ingest is arguably more efficient and scalable.
Hi SofBL
After going through your post, as per my thinking in your case, the key difference is that incremental refresh in a dataset and in a Dataflow Gen2 are enforced differently by the engine. In a Power BI dataset, the RangeStart and RangeEnd parameters only need to exist logically in the M query pipeline, and as long as the final query folds and the filter logic is applied somewhere upstream,even if the datetime column disappears later,the incremental refresh policy can still activate. However, Dataflow Gen2 enforces the rule at the final query stage: it requires the DateTime column to physically remain in the output table when configuring incremental refresh. Because your looping function expands company data and removes the date field before the final output, Dataflow Gen2 treats it as lacking a valid incremental column, even though function inputs use RangeStart and RangeEnd. That is why the same M logic works in a dataset but fails in Dataflow Gen2.
Additionally, Dataflow Gen2 currently does not allow append mode with incremental policies because the lakehouse engine expects to completely manage partitions for incremental loads. It needs replace mode so it can rewrite partition files; otherwise, append mode would lead to inconsistent partitions. This also explains the quarterly bucket size limitation, the engine needs predictable partition logic. The folding error you encountered further confirms that the Business Central connector still has folding limitations within Dataflow Gen2, even if it folds properly in a dataset, so the Fabric engine does not trust the query for incremental partition pruning.
I think by loading the full BC dataset once into Delta Lake, then using a second Dataflow Gen2 to append only rows with higher entry numbers,is a very practical pattern. It bypasses the datetime requirement, avoids folding constraints, and simulates incremental load efficiently by key-based tracking. This approach is commonly used in data engineering pipelines for immutable ledger-type tables, so you did not miss anything,in fact, you implemented a correct pattern where date-based incremental refresh isn’t suitable. Microsoft does not promote this pattern in official docs because incremental refresh in Fabric is positioned as a date-driven, fold-dependent mechanism, whereas key-based delta ingestion is an engineering workaround outside the standard UI-driven model. But in real-world Business Central workloads with multiple companies and immutable entry tables, your entry-no-based incremental ingest is arguably more efficient and scalable.