Blog Post

Fabric Updates Blog
3 MIN READ

Native Execution Engine now accelerates CSV workloads in Microsoft Fabric Spark

Santhosh_Ravin1's avatar
Santhosh_Ravin1
Icon for Microsoft Employee rankMicrosoft Employee
2 months ago

Accelerate CSV Workloads with Native Execution Engine in Microsoft Fabric Spark

Comma-separated values (CSV) remain one of the most widely used file formats in data engineering. Whether ingesting log files, processing exports from legacy systems, or loading flat-file datasets, CSV workloads are a daily reality for Fabric Spark users.

The Native Execution Engine in Microsoft Fabric Spark now natively accelerates CSV file reads, delivering up to 2× faster performance on benchmark workloads with no code changes required.

New Native CSV Reader in Execution Engine

Previously, when the Native Execution Engine processed CSV files, reads fell back to Spark's default CSV reader. As a result, CSV workloads didn't benefit from the vectorized columnar execution that accelerates Parquet and Delta workloads.

The latest Native Execution Engine update includes a purpose-built CSV reader powered by the Velox engine with SIMD-optimized parsing. CSV file reads now execute directly within the vectorized pipeline, eliminating row-to-columnar conversion overhead and delivering significant performance improvements.

What Performance Can You Expect?

Benchmark Without NEE (Baseline) With NEE

20 GB CSV read/write workload56 seconds34–38 seconds (~1.5–1.65× faster)
TPC-DS benchmark suite (CSV)BaselineUp to 2× faster
TPC-H benchmark (SIMD optimized)Baseline~35% improvement
TPC-DS end-to-endBaseline~20% improvement

Note: Performance gains vary depending on workload characteristics, schema complexity, and data distribution. CSV workloads with larger datasets and simpler schemas typically see the largest improvements.

Impact for Data Engineering Workflows

  • Faster ETL Pipelines - CSV ingestion stages that previously bottlenecked pipelines can now complete more quickly, accelerating end-to-end data processing.
  • Lower Compute Costs - Faster execution means workloads consume Fabric capacity for less time, potentially reducing compute costs.
  • No Code Changes Required - If Native Execution Engine is enabled, CSV acceleration happens automatically. Existing notebooks, pipelines, and Spark SQL queries benefit immediately.
  • Consistent Acceleration Across Formats - CSV joins involving Parquet, Delta, and other natively accelerated formats benefit from a unified performance model.

How It Works

Velox CSV Parser with SIMD Optimizations

The parser leverages CPU-level vector instructions to process multiple bytes simultaneously, accelerating:

  • Field parsing
  • Delimiter detection
  • Type conversion

Mison-Based Structural Indexing

For compatible workloads, the engine uses structural indexing to identify field boundaries in a single pass before parsing. This reduces CPU utilization and improves performance on large files.

Together, these optimizations keep CSV data flowing through the vectorized columnar pipeline without falling back to Spark's row-based reader, preserving the performance characteristics of Native Execution Engine.

Supported CSV Options

The native CSV reader supports most commonly used Spark CSV options, including:

  • Custom delimiters
  • Custom quote characters
  • Header inference
  • Explicit schema specification
  • Null value handling
  • Multi-line records
  • Encoding configuration
  • Escape character configuration

For a complete list of supported options and known limitations, see the Native Execution Engine documentation.

Getting Started

CSV acceleration is available automatically whenever Native Execution Engine is enabled.

No additional configuration is required.

spark.conf.set("spark.native.enabled", "true")

You can also enable Native Execution Engine at the environment level for all Spark sessions.

For setup instructions, see Enable Native Execution Engine.

Prerequisites

  • Microsoft Fabric workspace with Spark enabled
  • Fabric Runtime 3.5 or later
  • Native Execution Engine enabled at the environment or session level

Next Steps

Feedback

Updated 2 months ago
Version 1.0

1 Comment

  • Fantastic Feature. CSV Processing means quick dashboards with improved run time.