Bring more of your semi-structured data processing into the fast, vectorized execution path of Microsoft Fabric Spark.
JSON is one of the most common formats in modern data platforms. It carries application events, API payloads, operational telemetry, configuration data, and the metadata that coordinates data-driven processes. For many organizations, JSON is not an edge case. It is part of the critical path from ingestion through transformation and analytics.
The update of JSON support in the Microsoft Fabric Spark Native Execution Engine, now in preview, expands native acceleration to an important class of semi-structured workloads. Spark can now read and process JSON data through the Native Execution Engine's vectorized C++ path, helping more of the query remain columnar from the source through downstream transformations.
Why JSON performance matters
Analytics systems increasingly combine structured tables with semi-structured data. A pipeline might ingest JSON events from an application, use JSON control files to determine which tables to process, enrich the records with lakehouse data, and write curated Delta tables for reporting. JSON also appears behind the scenes in transaction metadata and other dependencies around Delta Lake processing.
These patterns make JSON parsing more than a file-read operation. It can influence the startup time, throughput, and end-to-end efficiency of an entire job. When a pipeline runs frequently or processes many files, even small costs in parsing and data conversion can accumulate across stages and workloads.
Common customer scenarios include:
- Ingesting application, device, web, and service telemetry.
- Processing nested records from APIs and partner data exchanges.
- Driving reusable pipelines with JSON configuration and control files.
- Reading schema, manifest, and metadata files during orchestration.
- Transforming semi-structured landing data into governed Delta tables.
How JSON fits into a lakehouse flow
A common lakehouse pattern begins with JSON arriving in the Files area of a lakehouse, through a OneLake shortcut, or from an upstream ingestion process. The records might represent customer activity, application operations, device measurements, or partner transactions. A Fabric notebook reads those files, applies a schema, selects the fields needed by the business, and prepares the data for additional processing.
The same job can then filter invalid or irrelevant events, flatten nested structures, derive business attributes, and combine the JSON records with trusted reference data. Aggregations create useful metrics, while the curated result is stored in Delta tables for downstream notebooks, pipelines, the SQL analytics endpoint, and Power BI. The JSON read is the entry point to this larger analytical flow, so accelerating it helps the job begin productive columnar processing sooner.
Metadata-driven frameworks amplify this effect. A reusable pipeline may read many small JSON documents that describe source locations, schemas, validation rules, transformation steps, and destinations. Those reads happen across multiple tables and recurring schedules. Keeping JSON parsing in the native path helps reduce repeated execution overhead and supports a more efficient foundation for standardized data engineering.
This matters because customers evaluate performance at the job and pipeline level, not only at an individual operator. A faster source reader is most valuable when its output can continue through filters, projections, joins, and aggregations without unnecessary transitions between execution models.
Keeping JSON processing in the native path
The Native Execution Engine accelerates supported Spark operations by offloading them from the JVM-based execution path to a vectorized native engine built on Velox and Apache Gluten (incubating). Columnar processing allows the engine to operate on batches of values instead of repeatedly materializing individual row objects. This design improves data locality, enables efficient use of modern processors, and reduces overhead across many analytical operations.
Before native JSON support, a query that encountered a JSON source used the Spark JVM path for JSON reading and parsing. Even when filters, projections, aggregations, or joins later in the plan were eligible for native acceleration, the data first passed through row-oriented processing and then transitioned into a representation suitable for the accelerated path. Those handoffs reduced the amount of work that could benefit from continuous columnar execution.
With this preview, JSON reading and parsing can run in the Velox-based native layer. Parsed values are produced as columnar batches that can flow directly into eligible native operators. By avoiding an early return to row-based JVM processing, Fabric Spark can reduce execution-path transitions and apply native acceleration across a larger portion of the job.
What this means for your workloads
The most important benefit is broader end-to-end acceleration. Customers can continue to use familiar Spark DataFrame and SQL patterns while the engine handles the execution-path improvements. There is no new JSON-specific programming model to learn and no need to rewrite existing transformations simply to access the native reader.
For ingestion workloads, native JSON processing can help increase throughput before data is standardized into Delta tables. For metadata-driven pipelines, faster reads of configuration and control data can reduce overhead that appears repeatedly across orchestrated jobs. For analytical workloads that query JSON directly, filters and projections can begin from a native columnar source rather than waiting for a JVM-based parsing stage.
The result is a more consistent performance model across common lakehouse formats. Teams can design pipelines around business requirements and data characteristics while Fabric expands the set of operations that remain on the accelerated path.
Use the Spark APIs you already know
Existing notebook code can continue to read JSON with standard Spark APIs. For example, a pipeline can load event data, select the fields needed for analysis, filter the records, and aggregate the results with the same DataFrame operations used today:
events = spark.read.json("Files/events/") daily_activity = ( events.filter("eventType IS NOT NULL") .groupBy("eventDate", "eventType") .count() ) When the plan uses supported operations, Fabric can execute the JSON read and downstream processing in the native columnar path. The optimization is delivered by the platform, so developers can focus on data quality, business logic, and the outputs their users need.
Build faster semi-structured data pipelines
JSON support is another step in expanding the performance coverage of the Native Execution Engine across real customer workloads. It brings acceleration closer to the point where semi-structured data enters the lakehouse and helps preserve columnar execution as that data is filtered, transformed, joined, and aggregated.
To learn how the engine works and how to use it with Fabric Spark, see Native execution engine for Fabric Data Engineering. You can also review Apache Spark runtime in Fabric and Lakehouse and Delta Tables for more information about the broader Fabric data engineering platform.
Get started by running a representative JSON workload in a Fabric notebook and comparing the end-to-end job experience. Review How to use notebooks for guidance and share your experience through the Microsoft Fabric Community. Your feedback helps us prioritize the next areas of acceleration.