Forum Discussion
Delayed data refresh in SQL Analytical Endpoint
- 2 years ago
So my info from Microsoft is that it should take a few seconds to a minute for the SQL Endpoint to discover the updated delta logs. They did have an issue which was fixed apparently.
If your issue persists it would be worth logging that with MS support and give them you workspace ID, lakehouse id, and time it was happening
I also load all my data using Synapse notebooks, using overwrite mode. When I query the table in the SQL endpoint in Fabric the data is updated within seconds from when the underlying delta table is updated most of the time - but sometime it takes up to a few minutes for the SQL endpoint to reflect the change in data.
This (somewhat short) delay is what messes with my API triggered dataset refreshes. I would assume that when my activity to load the delta lake table is done, its safe to run the dataset refresh. But this is not always the case. Whenever data has failed to appear in the dataset I always go right back to the SQL endpoint and query the table, at which point I can always see the new/todays data there
I have added a "5 minute wait" activity in my pipeline, and it now works "most of the time", but still sometimes fail to fetch new data, leading me to believe that 5 minutes is not enough in all cases. I will extend this wait activity to 10-15 minutes, to see if this will give the SQL endpoint sufficient time to do "whatever needs to be done" to refresh the shortcuts/links.
But as far as you know, the SQL endpoint should be always-up-to-date with the underlying data? There is no "reframing" activities going on behind the scenes, where it holds a cache of which parquet files makes up the latest verison of the Delta Lake table?
There certainly is reframing with power bi datasets in terms of updating metadata to understand which is the latest delta data. But for the SQL Endpoint I am not aware of any reframing process, it should read the latest transaction log and display the latest data. It's also the fact that it works "some of the time" in your scenario that's troubling me.
What's the volume of data?
- FelixL2 years agoAdvocate II
Yes, this is troubling for me to.
Most affected tables are rather small, around 100~ MB in total parquet size and 2~ Mil records.
I can easily reproduce the SQL endpoint delay issue by starting a notebook and performing a spark sql update against a delta lake table (update xxx set yyy = current_timestamp() , and in paralell checking the SQL endpoint with a query (select max(yyy) from xxx).
From the moment my spark update query returns "success" it takes anywhere from a few seconds up to a few minutes until i see my data change in the SQL endpoint (The old timestamp is returned for a number of T-SQL refreshes, until the new timestmap is finally returned).
I would expect to not ever be able to see the old timestamp if i start querying through the SQL endpoint after the spark sql query has finished - but this is not the case.
- AndyDDC2 years agoMost Valuable Professional
So my info from Microsoft is that it should take a few seconds to a minute for the SQL Endpoint to discover the updated delta logs. They did have an issue which was fixed apparently.
If your issue persists it would be worth logging that with MS support and give them you workspace ID, lakehouse id, and time it was happening
- FelixL2 years agoAdvocate II
That you for looking into this.
These few minutes delay would explain what I am seeing, and why my API triggered Power BI dataset refreshes fails to fetch new data.. I have had my Power BI refreshes running with a 10 minute delay for a few days now, and everything has been working so far. A 5 minute delay does not seem to be sufficient even now - I "sometime" miss data using only 5 minutes.
It's a real shame that there is "any" delay here though, since I cant be 100% sure "when" my data is available for pushing to Power BI via the endpoint. Its especially bad since this did not appear to be the case in Synapse Serverless.
I have some spark jobs loading data to some of my tables every 30 minutes (and running for 30~minutes), after which an incremental dataset load is triggered via the same pipeline running the notebooks. Having to add a 10 min wait actitvy in my pipeline here will in have a 30% negative effect on the loading times of my jobs (from job start, to when business users see data in Power BI).
I guess the long term solution here would be to use Direct Lake Semantic Models in Power BI, instead of SQL endpoint. This functionality is however still missing some ctirical features for me to be able to use right now (or I would have to heavily remodel my underlying Gold layer to handle the shortcomings of Direct Lake..)
For now though, I am satisfied in knowing why this is happening. Thanks.