Forum Discussion
Desktop refresh vs gateway refresh on service
Hi aj1973
I've not used dataflows until now. I'm trying to study up on that now. I know I'll be limited to some degree because I understand that merging tables in the dataflow requires a premium workspace, which we do not have.
Progress might be a little slow from here but I will work on your suggestion. Please forgive me if it takes a few days for me to report back on what I've tried and how it turned out.
Cheers
mmilegal
Hi mmilegal
No, you don't need premium capacity to use Dataflows. Premium capacity is for AI and ML capabilities.
To make it simple for you to understand, The Dataflow uses the Gateway to connect to your SQL server 2012, if it works then good. After you cennect to your source you will be using Power Query Online Platforme ( just like in your desktop) to ETL your data. After you prepare your data and refresh it you can then call it in your desktop and build the report. after that you just publish it and you all be good to go.
Let me know for more details
- mmilegal5 years agoFrequent Visitor
Hi aj1973
I'm embarrassed to keep missing the point - I suspect I am a bit like a kid who belongs in the beginner class who is trying to ask sensible questions in the advanced class! That's probably fairly accurate, I'm afraid! I've read a fair bit of content about dataflow vs dataset and watched a number of the ' guy in a cube' videos but I've ended up feeling like my only chance to grasp it was to try some practical implementation that relates to my real-world requirements. Unfortunately, I think my 'in at the deep end' approach has, this time, been more sink than swim.
Experiment 1
I tried a test dataflow with two of the tables from my old dataset. To try to test some of what I knew I'd need, I included the tables used in Example1 of my original post but decided to try only one of the three pivots. When I tried to edit the dataflow (in what I assume is the Power Query Online Platform - it looks near-identical to Power Query in Desktop) to merge the 2 tables, I got the message 'Computed tables require Premium to refresh. To enable refresh, upgrade this workspace to Premium capacity, or remove this table.'Experiment 2
At that point, I considered that I was only supposed to use the Dataflow to pull the source data so I saved the Dataflow without any transformation steps then opened a fresh pbix in the desktop and used the dataflows connector in the Get Data dialog in the desktop. In Desktop, I recreated the steps to merge, filter and pivot the data. When I applied the changes, the refresh began and was obviously going to take a while. I timed the refresh as processing around 950 rows per minute. The data being processed measures rows in the millions so I abandoned that.Experiment 3
My last throw of the dice was to replicate the merge/pivot steps in the online editor of the dataflow then create a new pbix in the Desktop and refresh the data in the Desktop. I kind of suspected that was a waste of time and I wasn't terribly surprised to see that the structure imported matched the output of the pivot that I had saved in the Dataflow but there was no data, presumably because the Dataflow was not refreshed for the reason above.Do those three experiments make sense?
Cheers
mmilegal
- aj19735 years agoCommunity Champion
It all makes sense, I like the way you described your 3 experiments.
From the first experiment I could tell that you don't have permission to use or store the dataflow in Azure Data Lake Storage https://docs.microsoft.com/en-us/power-query/dataflows/configuring-storage-and-compute-options-for-analytical-dataflows. That's why the dataflow didn't refresh at first. Bad news good news from your experiment is that your data source (SQL 2012 server) is accessible through your Gateway, that was the aim of all this experiment wasn't it?
So what is left for you to do is ask your Azure admin the permission to use the Data lake in order to save your dataflow. Once you have it then right after you build your dataflow in Power Query online and save the changes the service will ask you to refresh the dataflow, if it goes through then you will be good to go.
Let me know
- mmilegal5 years agoFrequent Visitor
Hi aj1973
I think I've managed to achieve the result but in a slightly different way - I might have dodged the issue in a manner that almost feels like cheating!
The bit about access to ADLS has kind of thrown me as I checked with the one person in our organisation likely to know about that and he doesn't believe we have a dedicated subscription for that.
However, it had occurred to me that everything else was working according to your suggestions and I could make the dataflow work so long as I didn't try to use merged or computed tables in the Power Query Online Platform. All the other fancy stuff like the pivots worked just fine. My 'cheat' solution - and I appreciate I'm fortunate that it was a solution available to me and not every user will be so lucky - is based on the fact that it was possible for me to use the SQL database from which I am pulling all my data. It was a very quick and simple process to create views in SQL to perform the limited actions that I might otherwise have done with a merge in PQOP. So, in case it helps anyone else with a similar challenge, the structure of my solution is something like this:
1. create views in SQL Server with the smallest amount of joins possible to allow filtering for further processing. That really was just limited to examples like pulling in a 'type' column from table B to apply to rows in table A to allow me to filter table A by type.
2. create a dataflow in a workspace to pull the data from SQL Server tables and views as per step 1. The most significant processing in PQOP is the pivots described as Example1 in my original post.
3. create a dataset in Desktop and use the dataflow connector in the Get Data dialog to pull data for the dataset from the dataflow created in step 2.
4. the most significant processing in the dataset is the DAX SUMMARIZE/FILTER processing outlined as Example2 in my original post.
I think I now have the entirety of the data required built into a single dataset that should be able to support the reporting requirements. The dataflow has refreshed in the cloud several times and, although it did fall over once, it seems to be refreshing in about 22 minutes. I've published the dataset and, scheduling a refresh for one hour after the scheduled refresh time of the dataflow, it refreshed last night in around 16 minutes. Size of the pbix is 199Mb. Not sure if all of those statistics necessarily are too meaningful without lots of other info not included here, but they are at least indicative of one user's experience.
I suspect much of my solution still bears the hallmarks of the enthusiastic amateur and is likely to fall short of best practice in some areas but I'm sure that will be the reality for many using Power BI in small organisations.
A very big thank you aj1973 - I know for certain I would not have worked out all of this on my own and I'm sure I've had a learning experience on the journey too. I very much appreciate both the knowledge you possess and your generosity with your time in sharing it.
Cheers
mmilegal