Forum Discussion
Dealing with a huge csv file
- 1 year ago
I heard Dataflows is good for situations where one does not have permissions to access the database but another user doesalso called "breaking the security chain of custody" - I can not endorse that.
dataflow can be used as tool to reduce the semantic model size by trying to push much of the transformations away from the semantic model into the dataflowDataflows are glorified CSV files. Do these transforms at your own risk.
can a dataflow be used by a report in a different workspace?yes. Reusability is one of the positive features of dataflows. You will still need to have access to that "other" workspace.
The 800MB CV files you append to the semantic model, does it take you a lot of time? Are you able to do it without any issues using Power Bi desktop and selecting the TEXT/CSV connector? Do you use any special method other than the box standard import one would usually do using power bi desktop.
So when or under what scenario would dataflows be useful?
- lbendlin1 year ago
Super User
Doesn't take long at all. You can assume that CSV and Parquet are the best performing formats for ingesting data from Import Mode sources.
We use the SharePoint Folder connector exclusively as we want to combine all the CSVs (in this case 50 CSVs at 800 MB each) into the semantic model.
So when or under what scenario would dataflows be useful?1. dataflows can shield the developer (NOT the report user) from a slow data source.
2. dataflows can be re-used in multiple semantic models.
- mp3909881 year ago
Post Partisan
When you say dataflow shields the developer (NOT the report user) from a slow data source - doesn't a slow data source also affect a report user as well? If so, why does a dataflow only applicable to developer and not report user?
- lbendlin1 year ago
Super User
doesn't a slow data source also affect a report user as well?No, because dataflows are not even accessible to report users. The developer has to import the dataflow into a semantic model. The report users interact with the semantic model. The semantic model is stored in SSAS Tabular, in memory, with compression etc. That is usually faster than Direct Lake or Direct Query.
- mp3909881 year ago
Post Partisan
When you append additonal CSV files to the semantic model as you have mentioned, how does it know which CSV files are the new ones to append? Otherwise it will append the ones that have already been appended in the last refresh? Basically, how does Power BI identify the new CSV files?
- lbendlin1 year ago
Super User
Power BI has no memory. You need to do that partition management yourself.