Forum Discussion
LasseMr
4 years agoAdvocate II
Raw Dataflow-Computed Dataflow-Dataset Best Practice: Gen2 CPU Load
Dear Community, After having read countless pages on the subject and been in dialogue with MS Support I still fail to get a good answer on what is the "best" setup. We are running; Embedded...
lbendlin
4 years agoSuper User
Dataflows are good at raw storage, similar to Parquet blobls or Hadoop. They suck at anything requires computation, including incremental refresh. If your data source has an ok spooling performance there is no need for dataflows at all. (And don't get me started with Direct Query against dataflows).
If you must, feed your dataset from the raw dataflow and do the computations on the dataset side. Implement incremental refresh in your dataset, and do the partition management yourself.