Forum Discussion
Speed - Database vs Delta Files
- 9 months ago
Hi AJAJ, to answer you question about if pyspark would read a 50gb table into memory, no.
Pyspark uses lazy evaluation, which is a concept of transforming data without actually transforming data in memory until a pyspark action statement is called. So if you run the code spark.read("my_files"), pyspark is not actually reading the data at that moment. Also not when you run code like df.filter(...) or df.select(...), nothing is executed. Until you call an actionable function like df.show() or df.save(), then pyspark starts processing.
But it will not read 50gb into memory at once. Pyspark will read and process data in partitions and will spread the compute load over the available nodes in your cluster (or Fabric capacity). It's a method that is called distributed processing —for those who don't know yet.
A SQL Database is best at providing data at an instance when you invoke a select statement, as all the data is always hot —depending on the configuration of you database. This make a SQL ideal for application development, but because of this, it is also very expensive for storing data and compute. This is also why SPs instantly work, as they are part of the "always-on" environment of SQL.
In reporting environments we do not need the data to be "hot" all the times. Most of the time we query the data in batches at a specific time during the day once. Data lakes and data stored in Fabric then provide the cheapest storage option for your 50gb of data.
In a company you're not (always) there to code, you are there to make the best descisions for that company, which might be saving the company money or enhancing the business process with new functionality. Choosing between a SQL Server warehouse and Fabric lakehouse or warehouse might be one of those descisions. But, every decision can crown or kill you.
Hope this helps. If so, please give a Kudos 👍 and mark as Accepted Solution ✔️.
Hey AJAJ ,
You're not imagining it Spark notebooks and Delta Lake often feel slower than traditional SQL Server stored procedures, especially when you're coming from a decade of tuning high-performance OLTP/OLAP systems. But the reasons are architectural rather than inefficiencies, and the trade-offs depend on workload type.
Why SQL Server / Data Warehouse SPs Often Feel Faster
Traditional databases (SQL Server, Fabric DW, Synapse Dedicated, Snowflake, etc.) are built for tightly coupled compute + storage:
The engine knows the table structures, indexes, stats, partitions, data pages, etc.
Queries operate directly on local/attached storage.
The optimizer can make very aggressive, cost-based decisions.
Data rarely needs to be moved into a separate execution engine.
SPs benefit from plan caching and mature optimization.
Result: For single-node, I/O-optimized, relational workloads, SQL engines are extremely fast.
What Spark Is Doing Instead
Spark was designed for distributed processing of large, semi-structured/unstructured data (logs, events, files, ML workloads), not for low-latency SQL.
When Spark runs:
It loads data from object storage (ADLS/S3/Blob) into executors’ memory.
It executes operations in a distributed DAG (lazy evaluation).
It writes data back to object storage.
Yes — reading a 50GB Delta table means scanning some or all of those files.
But Spark uses:
predicate/file pruning
column pruning
caching
data skipping (Delta)
vectorized Parquet readers
This can make scans surprisingly fast, but it’s still a scan-based architecture, not index-based like SQL Server.
Is it “double work”?
Not exactly — Spark must read files because:
Object storage doesn’t support random access pages like SQL Server data files
There is no buffer pool, no indexes, no rowstore
Distributed compute requires data locality in memory/executors
So Spark doesn't have the same "just start reading pages off disk and execute" model that a database engine does.
For large-scale analytics, that model is actually a feature, not a flaw — but it’s different.
When Spark Is Slower
Spark will usually be slower for:
Small to medium datasets (<100GB)
Highly selective queries (which a SQL index would solve instantly)
Repeat/interactive workloads where DB caching beats Spark cluster startup
Single-record lookups or OLTP-style patterns
Complex joins on non-partitioned data
When Spark Is Faster
Spark shines when:
Data size exceeds what fits comfortably on a single DB engine
You need distributed compute for ML, ETL, or large transformations
You work with raw data in files (JSON, CSV, logs)
You benefit from cluster parallelism
🧠 The Key Point
Spark is a distributed compute engine designed for scale and flexibility.
SQL Server is a database engine designed for performance and efficiency.
Don’t treat Spark as a “slower SQL server.” It’s a different tool entirely.