Accelerate CSV Workloads with Native Execution Engine in Microsoft Fabric Spark
Comma-separated values (CSV) remain one of the most widely used file formats in data engineering. Whether ingesting log files, processing exports from legacy systems, or loading flat-file datasets, CSV workloads are a daily reality for Fabric Spark users.
The Native Execution Engine in Microsoft Fabric Spark now natively accelerates CSV file reads, delivering up to 2× faster performance on benchmark workloads with no code changes required.
New Native CSV Reader in Execution Engine
Previously, when the Native Execution Engine processed CSV files, reads fell back to Spark's default CSV reader. As a result, CSV workloads didn't benefit from the vectorized columnar execution that accelerates Parquet and Delta workloads.
The latest Native Execution Engine update includes a purpose-built CSV reader powered by the Velox engine with SIMD-optimized parsing. CSV file reads now execute directly within the vectorized pipeline, eliminating row-to-columnar conversion overhead and delivering significant performance improvements.
What Performance Can You Expect?
Benchmark Without NEE (Baseline) With NEE
| 20 GB CSV read/write workload | 56 seconds | 34–38 seconds (~1.5–1.65× faster) |
| TPC-DS benchmark suite (CSV) | Baseline | Up to 2× faster |
| TPC-H benchmark (SIMD optimized) | Baseline | ~35% improvement |
| TPC-DS end-to-end | Baseline | ~20% improvement |
Note: Performance gains vary depending on workload characteristics, schema complexity, and data distribution. CSV workloads with larger datasets and simpler schemas typically see the largest improvements.
Impact for Data Engineering Workflows
- Faster ETL Pipelines - CSV ingestion stages that previously bottlenecked pipelines can now complete more quickly, accelerating end-to-end data processing.
- Lower Compute Costs - Faster execution means workloads consume Fabric capacity for less time, potentially reducing compute costs.
- No Code Changes Required - If Native Execution Engine is enabled, CSV acceleration happens automatically. Existing notebooks, pipelines, and Spark SQL queries benefit immediately.
- Consistent Acceleration Across Formats - CSV joins involving Parquet, Delta, and other natively accelerated formats benefit from a unified performance model.
How It Works
Velox CSV Parser with SIMD Optimizations
The parser leverages CPU-level vector instructions to process multiple bytes simultaneously, accelerating:
- Field parsing
- Delimiter detection
- Type conversion
Mison-Based Structural Indexing
For compatible workloads, the engine uses structural indexing to identify field boundaries in a single pass before parsing. This reduces CPU utilization and improves performance on large files.
Together, these optimizations keep CSV data flowing through the vectorized columnar pipeline without falling back to Spark's row-based reader, preserving the performance characteristics of Native Execution Engine.
Supported CSV Options
The native CSV reader supports most commonly used Spark CSV options, including:
- Custom delimiters
- Custom quote characters
- Header inference
- Explicit schema specification
- Null value handling
- Multi-line records
- Encoding configuration
- Escape character configuration
For a complete list of supported options and known limitations, see the Native Execution Engine documentation.
Getting Started
CSV acceleration is available automatically whenever Native Execution Engine is enabled.
No additional configuration is required.
spark.conf.set("spark.native.enabled", "true")You can also enable Native Execution Engine at the environment level for all Spark sessions.
For setup instructions, see Enable Native Execution Engine.
Prerequisites
- Microsoft Fabric workspace with Spark enabled
- Fabric Runtime 3.5 or later
- Native Execution Engine enabled at the environment or session level
Next Steps
- Learn more about the Native Execution Engine
- Follow the step-by-step guide to Enable Native Execution Engine
- Explore supported data types and file formats in the Native Execution Engine documentation.
Feedback
- Share your experience in the Fabric Community Forums
- Ask questions on Microsoft Q&A