Forum Discussion
incremental refresh
Hi,
Incremental refresh can definitely help reduce the refresh time and manage a ~3 GB Import semantic model, but the fact that source data can be deleted is an important consideration.
By default, incremental refresh only refreshes the period defined in the incremental refresh policy. Therefore, if a record from an older/historical partition is deleted at the source, Power BI may not detect that deletion because that historical partition is no longer being refreshed.
Recommended approach
I would configure incremental refresh with a refresh window that covers the period in which updates/deletions can occur.
For example, if records can be modified or deleted within the last 30 days:
- Store/archive: 5 years
- Incrementally refresh: last 30 days
- The last 30 days will be reloaded on every refresh.
- Older partitions will remain unchanged.
This means a deletion that occurs within those 30 days will be reflected in Power BI during the next refresh.
If deletions can happen at any time, you need to consider a different approach, such as:
- Increase the incremental refresh window so that it covers the maximum period in which historical records can be deleted.
- If the source provides an audit/change tracking column, use the Detect data changes option where appropriate. Microsoft documents this as a way to identify periods where data has changed.
- If historical records can be deleted without any reliable way to identify those deletions, incremental refresh alone cannot guarantee that Power BI will discover them in old partitions. In that situation, consider periodically refreshing the historical partitions/full model or implementing a proper change-data-capture/ETL process upstream.
Important note about the first refresh
One thing to keep in mind is that incremental refresh does not immediately make the first service refresh fast.
After publishing the model, the initial refresh creates the historical and incremental partitions and loads the required historical data. Depending on the amount of data, this initial refresh can take considerable time.
Subsequent refreshes are generally much faster because only the partitions within the configured refresh window are refreshed.
For a 3 GB model, I would therefore test the initial refresh duration and memory/capacity requirements before moving it to production.
Another option: Fabric Lakehouse / Direct Lake
If you are already using Microsoft Fabric or are considering moving the data platform to Fabric, another architecture to evaluate is Lakehouse + Direct Lake.
With Direct Lake, the semantic model can consume Delta/Parquet data from OneLake without importing and duplicating the data into the semantic model. This can be particularly useful for large datasets and frequently changing source data.
So, broadly:
Current Import model → Incremental Refresh
Good option when you have a reliable date column and can define a safe refresh window for updates/deletions.
Large/frequently changing data + Fabric available → Lakehouse + Direct Lake
Worth considering if you are looking at a longer-term Fabric architecture.
References:
- Configure incremental refresh for Power BI semantic models
- Incremental refresh and real-time data overview
- Direct Lake overview
- Create a semantic model from a Fabric Lakehouse
Hope this helps.
Thanks!