Forum Discussion

AGo's avatar
AGo
Icon for Post Patron rankPost Patron
6 years ago
Solved

Large amount of data refresh best practice

Hi everybody,

 

I've got Pro license. I've got a folder collecting CSVs from today to the past 12 month, 1 month is approx. 1,5GB. I could use directly folder source, but I'm using it through a dataflow. The refresh is very slow, in the next months I will may be over the 10GBs limit and I'm occupying Microsoft's and my bandwidth in a way that is nonsense, because this data could be subjected to incremental refresh day by day (production machines log). Premium capacity has no reason to be considered in relation to the smallness of the project and this only one necessity.

Does anyone have a best practice to extract and transfer data in csv in a compressed way? Do I need to make a program to inject these CSVs in a database that acts as a more compressed datasource?

Any advice?

 

Thanks very much in advance

  • AGo's avatar
    AGo
    6 years ago
    Thanks to you reply I accepted that I could upload the data on drive/sharepoint and process them directly in the cloud. I managed to open multiple csvs in multiple subfolders using Sharepoint folder connection method. Now the refresh takes a half hour to complete, maybe in the future it will be unsustainable with the data increasing and I'll have to pass to a datalake method, but for now it's working for this test phase. The develop is now made with a parameter limiting the number of rows to 1M to avoid bandwidth saturation and crashes when on desktop and then no limit when online. Thanks again.

7 Replies

  • Anonymous's avatar
    Anonymous
    Not applicable

    AGo You could spin up an Azure Analysis Services instance and build the model in there for cheaper. Then you could do a live connection from Power BI to that model without incurring the larger overhead. This would allow you to scale as well at a lower price point.

      • AGo's avatar
        AGo
        Icon for Post Patron rankPost Patron
        I read about this zip topic but it would be my last choice for a workaround because csvs increasing everyday must be automatically zipped in a multi-level hierarchical path in the single zip file.
        But the 10gb size limit refers to the pbix that is compressed, maybe there's no problem in my 10Tb limit on sharepoint/drive. I will test it because now these csvs are on a linux machine, and I have to set up a sync with drive.
        If it works I hope there is a method to open multiple csvs files on sharepoint/drive as for Folder connection.
        If only incremental refresh were available for Pro users...