Forum Discussion
Speed - Database vs Delta Files
- 9 months ago
Hi AJAJ, to answer you question about if pyspark would read a 50gb table into memory, no.
Pyspark uses lazy evaluation, which is a concept of transforming data without actually transforming data in memory until a pyspark action statement is called. So if you run the code spark.read("my_files"), pyspark is not actually reading the data at that moment. Also not when you run code like df.filter(...) or df.select(...), nothing is executed. Until you call an actionable function like df.show() or df.save(), then pyspark starts processing.
But it will not read 50gb into memory at once. Pyspark will read and process data in partitions and will spread the compute load over the available nodes in your cluster (or Fabric capacity). It's a method that is called distributed processing —for those who don't know yet.
A SQL Database is best at providing data at an instance when you invoke a select statement, as all the data is always hot —depending on the configuration of you database. This make a SQL ideal for application development, but because of this, it is also very expensive for storing data and compute. This is also why SPs instantly work, as they are part of the "always-on" environment of SQL.
In reporting environments we do not need the data to be "hot" all the times. Most of the time we query the data in batches at a specific time during the day once. Data lakes and data stored in Fabric then provide the cheapest storage option for your 50gb of data.
In a company you're not (always) there to code, you are there to make the best descisions for that company, which might be saving the company money or enhancing the business process with new functionality. Choosing between a SQL Server warehouse and Fabric lakehouse or warehouse might be one of those descisions. But, every decision can crown or kill you.
Hope this helps. If so, please give a Kudos 👍 and mark as Accepted Solution ✔️.
Hi AJAJ,
This would be a very interesting thing to benchmark.
The Fabric documentation does mention that a COPY statement is the most efficient at loading data into a warehouse if the file is already accessible, so what you're saying makes sense, but I would still love to see the numbers.
If you found this helpful, consider giving some Kudos. If I answered your question or solved your problem, mark this post as the solution.