Blog Post

Fabric Updates Blog
7 MIN READ

10 Things to know about GPU-powered Query Acceleration in Fabric Data Warehouse (Preview)

NevenaNikolic's avatar
NevenaNikolic
Icon for Microsoft Employee rankMicrosoft Employee
1 month ago

Query Acceleration for Microsoft Fabric Data Warehouse is a new GPU-powered capability that helps accelerate analytical workloads automatically. Query Acceleration identifies eligible SQL operations and offloads them to GPUs behind the scenes, helping improve query performance and throughput without requiring query rewrites, application changes, or specialized expertise.

Query Acceleration (Preview) can be enabled with a simple toggle. Sign up for preview access to explore faster analytics, lower query latency, and improved performance for eligible workloads.

1. Query Acceleration removes CPU execution bottlenecks in analytics workloads

Fabric Data Warehouse is optimized for analytical workloads on modern multi-core CPU infrastructure. However, as warehouse scale, concurrency, and query complexity increase, CPU resources increasingly become the limiting factor for specific classes of data-parallel operations. Large scans, joins, filters, and aggregations place substantial demands on compute, memory bandwidth, and data movement within the execution engine.

Query Acceleration extends the Fabric Data Warehouse execution engine with GPU-based query processing to address these bottlenecks. Rather than introducing a separate analytics service or requiring application changes, Query Acceleration is integrated directly into the query execution pipeline. During compilation and optimization, the engine identifies eligible portions of a query plan and selectively offloads supported operators to GPUs while the CPU continues to coordinate query orchestration and execution flow.

By accelerating the most computationally expensive stages of query execution, Query Acceleration increases available processing capacity, improves throughput under concurrent demand, and reduces execution time for eligible queries.

2. Query Acceleration gets you faster analytics without changing a single line of SQL

Query Acceleration in Fabric Data Warehouse is designed to be fully transparent to users: there is no need to rewrite queries, use special syntax, or manage hardware explicitly. The engine automatically identifies eligible portions of a SELECT query and offloads them to GPU acceleration when beneficial, based on factors such as query shape, operators, and supported functions. As a result, existing workloads can take advantage of acceleration immediately with no code changes, migration effort, or application modifications, requiring only a simple toggle to enable the feature.

Figure: Query Acceleration toggle among the workspace settings

 

 

 

 

 

 

 

 

3. Query Acceleration shines in high-concurrency analytics (higher throughput, lower tail latency)

Query Acceleration delivers the largest gains in high-concurrency analytical workloads, where many queries run simultaneously and compete for resources. By offloading eligible query fragments to GPUs, Fabric Data Warehouse significantly increases overall system throughput, allowing more queries to complete in parallel without overwhelming CPU capacity. This not only improves average query performance but also reduces tail latency (the slowest queries in a workload), which is critical for consistent user experiences in dashboards and interactive analytics. Because GPUs are optimized for highly parallel execution, they can efficiently absorb spikes in demand and process large volumes of data concurrently, helping ensure predictable performance even under heavy load.

4. Query Acceleration targets the expensive parts: scans, filters, joins, and aggregations

Query Acceleration helps you get insights faster by automatically accelerating the most performance-intensive parts of analytical workloads. Operations such as large table scans, filters, joins, aggregations, and other eligible query operators are ideal candidates for GPU-powered execution because they process large volumes of data in parallel and often account for the majority of query execution time. Rather than moving entire queries to the GPU, Fabric Data Warehouse intelligently accelerates only these high-impact execution fragments while the rest of the query continues to run on the CPU. This targeted approach delivers significant improvements in query latency and throughput with minimal overhead, allowing you to benefit from faster analytics without changing your existing T-SQL queries, DirectQuery semantic models, applications, AI-driven workloads, or other workflows.

5. Reliability is built in: Query Acceleration works seamlessly with your existing workloads

Query Acceleration is designed to be safe, adaptive, and non-disruptive by default. Whether a query is accelerated is determined dynamically by the optimizer and execution engine based on the query plan. Some queries are fully eligible for GPU execution, others are partially eligible, where only specific plan fragments are offloaded, and some are not eligible at all due to factors such as unsupported SQL functions, complex expressions, certain data types, or plan shapes that would require inefficient switching between CPU and GPU.

Crucially, acceleration is never “all or nothing.” The system continuously evaluates execution at runtime and applies acceleration only where it is both supported and beneficial. If an accelerated fragment encounters an issue, such as memory pressure, resource contention, or a runtime failure, the engine can gracefully fall back to CPU execution for that portion of the work. The query is then completed on the CPU without impacting correctness or stability.

6. Data formats and function coverage affect end-to-end speedups

Query Acceleration is built on Microsoft's Tensor Query Processor (TQP), which enables efficient execution of database operations on GPUs. Performance depends not only on GPU compute capabilities but also on data formats, encoding strategies, and execution characteristics, which is why speedups vary across workloads.

7. Query Acceleration scales with your data

Query Acceleration is built to accelerate large-scale analytical workloads, even when datasets exceed available GPU memory. Instead of requiring all data to fit on a GPU, it uses a partitioned execution model that processes data in smaller chunks on the GPU and combines the results transparently. If a partition cannot be processed on the GPU due to memory or compatibility constraints, it automatically falls back to CPU execution while the remaining partitions continue to benefit from acceleration.

Although GPUs typically have less memory capacity than CPUs, their higher memory bandwidth and parallel processing capabilities make them highly effective for scans, joins, and aggregations. Query Acceleration dynamically leverages both CPU and GPU resources, avoiding costly GPU disk spills and automatically redirecting unsupported operations to the CPU when needed. The query optimizer further expands acceleration opportunities by transforming certain operations into partition-friendly execution plans, enabling broader GPU utilization without requiring any changes to existing T-SQL queries.

8. Monitor Query Acceleration to maximize performance

When evaluating Query Acceleration, consider three key questions:

  • Was the query eligible for acceleration?
  • Which portions of execution were accelerated?
  • Where is the query spending most of its time?

Eligibility determines where acceleration can be applied, but execution characteristics ultimately determine the performance benefit. Many workloads see significant gains even with partial acceleration when their most expensive operations are accelerated. Understanding these factors can help you identify opportunities to maximize the value of Query Acceleration.

You can use SQL Server Management Studio (SSMS) and Query Insights to understand how Query Acceleration is impacting your workload. In SSMS, execution plans highlight where Query Acceleration was applied and which portions of a query were accelerated. Query Insights helps you track query performance over time, compare execution metrics, and see which queries and workloads benefit most from Query Acceleration.

Figure: An execution plan showing Query Acceleration operator

9. Query Acceleration delivers measurable performance improvements on benchmarks and production workloads

To quantify the impact of GPU acceleration, we ran industry benchmarks against three comparable cloud data warehouse providers and observed up to 7x faster performance across reporting, application, and AI-driven analytics scenarios. These are workloads where systems are typically pushed hardest, with multiple users and agents issuing queries simultaneously.

At a 100 GB data scale, what stands out is not just the speed of individual queries, but how the system behaves under load. As concurrency increases, most data warehouses slow down and become less predictable. In contrast, Fabric’s GPU-accelerated warehouse maintains stable performance, completing the full 22-query workload in approximately five seconds whether one user is running queries or 64 concurrent users.

Reported results show the largest gains in complex, compute-heavy queries. The largest improvements come from more complex queries (for example, multi-join and intrinsic-heavy patterns), while scan-heavy queries can be more sensitive to data transfer costs. Luckily, the architecture is designed to keep expanding coverage while managing the realities of data transfer cost.

10. Query Acceleration is grounded in published research: SIGMOD Best Industry Paper

Query Acceleration in Fabric Data Warehouse is built on Microsoft research and engineering in hardware-accelerated query processing. The work was published as “CoddSpeed: Hardware Accelerated Query Processing in Microsoft Fabric” in SIGMOD Companion ’26 and was selected as the Best Industry Paper. This recognition highlights the technical rigor and innovation behind the architecture powering Query Acceleration in Fabric Data Warehouse.

To learn more, refer to The 2026 ACM SIGMOD/PODS Conference: Bengaluru, India - SIGMOD Awards.

The paper goes beyond the productized Query Acceleration feature in Fabric Data Warehouse examining its underlying architecture, including GPU offload, efficient data movement, and reliable partitioned execution at scale. It also analyzes the impact of data formats, encoding, and networking on performance and presents results from benchmark and production workloads.

Summary

Query Acceleration brings GPU-powered analytics directly into Fabric Data Warehouse without requiring query rewrites or application changes. By accelerating eligible operations such as scans, joins, filters, and aggregations, it helps your organization process more data, support more users concurrently, and deliver insights faster. Built-in fallback mechanisms ensure reliability, while familiar tools such as SSMS and Query Insights make it easy to understand and monitor acceleration across workloads.

Whether you're building BI dashboards, powering AI experiences, or running large-scale analytical workloads, Query Acceleration helps you get more performance from the same SQL workflows you already use today.

Get started

GPU-Accelerated Fabric Data Warehouse is now available in preview across five regions. Sign up for preview access to explore faster analytics, lower query latency, and improved performance for eligible workloads.

Updated 1 month ago
Version 1.0
No CommentsBe the first to comment