At Microsoft Build, Fabric Data Warehouse announced GPU-accelerated query execution. I’ve been running it against TPC-H, and the numbers are worth sharing, along with the methodology, because I’d rather hand you enough detail to check my work, and let you it yourself.
This has been one of the most complex and deeply technical projects I’ve been involved in at Microsoft, reminding me of the days of developing Vertipaq more than 16 years ago. And it’s also been one of the most rewarding projects.
A few disclaimers up front: This is not an official performance test. This is me playing with freshly released tech and building an app over a weekend. I’ve almost certainly missed a few best practices — but then, so will most customers when they first turn this on. So, take these as “what a curious person gets out of the box,” not “what the engineering team can extract with perfect tuning.” That’s arguably the most honest number anyway.
Also, this feature is now in limited availability, primarily due to GPU availability, and not yet in all Azure regions. I really wanted to test it as a customer would, on a commercial installation (instead of an internal test lab), and finding the right region took a bit of effort.
Test environment setup (so you can trust the numbers)
First, it’s worth mentioning that this test — this technology — runs entirely on data in OneLake. All in parquet, no tricks or duplicate copies.
- Capacity: Fabric F64 Trial
- Dataset: TPC-H at 100 GB scale factor
- Topology: Client test machine and Fabric capacity in the same Azure region (no cross-region latency padding the results).
- Result set caching: Deliberately disabled. Cached result sets make any engine look brilliant. I wanted to measure actual query execution, not cache hits.
- Workload: The full TPC-H query set, 22 queries
- Turning it on: A single workspace toggle that starts in your Workspace settings → Data Warehouse → Query acceleration (Preview). One switch, no infrastructure to provision.
Figure: Workspace settings to enable query acceleration in Fabric.
Results for single users: Every query under 500ms
On a single connection, each of complete in under half a second on GPU. The slowest (Q21, Suppliers Who Kept Orders Waiting) averages 445ms. The fastest (Q6) was 140ms.
The order of execution for the queries is randomized, and while the query patterns follow TPC-H, the various scalar values used in query predicates are also randomized (to simulate a real dashboard pattern).
The CPU baseline for the same single-user run sits around 1.6s at the median — ry (or 4x at the mean), before any concurrency comes into play. In fact, excluding the one off at 458.5 ms, they all stay below 320 ms.
|
queryId |
queryName |
errors |
avgLatencyMs |
p95LatencyMs |
|
Q1 |
Pricing Summary Report |
0 |
248.5 |
255.1 |
|
Q2 |
Minimum Cost Supplier |
0 |
173.8 |
180.2 |
|
Q3 |
Shipping Priority |
0 |
217.2 |
221.7 |
|
Q4 |
Order Priority Checking |
0 |
190.1 |
196.8 |
|
Q5 |
Local Supplier Volume |
0 |
254.1 |
259 |
|
Q6 |
Forecasting Revenue Change |
0 |
140.3 |
144.2 |
|
Q7 |
Volume Shipping |
0 |
252.6 |
257.6 |
|
Q8 |
National Market Share |
0 |
273.3 |
282.5 |
|
Q9 |
Product Type Profit Measure |
0 |
315.8 |
320.2 |
|
Q10 |
Returned Item Reporting |
0 |
247.8 |
251 |
|
Q11 |
Important Stock Identification |
0 |
279.3 |
287.3 |
|
Q12 |
Shipping Modes and Order Priority |
0 |
196.7 |
203.8 |
|
Q13 |
Customer Distribution |
0 |
184.2 |
188 |
|
Q14 |
Promotion Effect |
0 |
153.6 |
157.4 |
|
Q15 |
Top Supplier |
0 |
171.5 |
176.1 |
|
Q16 |
Parts/Supplier Relationship |
0 |
155.9 |
162.2 |
|
Q17 |
Small-Quantity-Order Revenue |
0 |
230.9 |
240.1 |
|
Q18 |
Large Volume Customer |
0 |
306.9 |
315 |
|
Q19 |
Discounted Revenue |
0 |
197.2 |
202.9 |
|
Q20 |
Potential Part Promotion |
0 |
251 |
257.1 |
|
Q21 |
Suppliers Who Kept Orders Waiting |
0 |
444.8 |
458.5 |
|
Q22 |
Global Sales Opportunity |
0 |
171.3 |
175.9 |
It is worth noting that a sequential, standard TPCH suite execution (same tool, I just remove randomization of queries) is about 5.02 seconds in my experiment, which is impressive in itself!
Results for high concurrency
Single-user benchmarks are vanity. Real warehouses serve many analysts, dashboards, agents, and apps hammering the same capacity at once. This is where the GPU and CPU query execution engines stop resembling each other.
Throughput (queries/sec) as concurrency climbs:
|
Concurrent U |
CPU q/s |
GPU q/s |
|
1 |
0.65 |
4.56 |
|
64 |
3.70 |
111.4 |
|
128 |
7.19 |
104.8 |
|
256 |
5.44 |
60.3 |
|
512 |
10.89 |
115.9 |
|
1000 |
16.93 |
101.4 |
CPU throughput tops out around 17 q/s. GPU sustains roughly 100+ q/s across the range.
(Note of honesty: the 256-user GPU row is a noisy sample — clearly under its neighbors. I’m re-running it, but I’d rather show you the raw number than quietly delete the inconvenient one.)
Figure: Throughput vs concurrency — GPU sustains ~100+ q/s; CPU tops out near 17 q/s.
This is something I need to test with a larger capacity, perhaps F128 or F256.
At 1000 concurrent users:
|
CPU |
GPU |
|
|
Throughput |
16.9 q/s |
101.4 q/s |
|
p50 latency |
42.8 s |
8.1 s |
|
p95 latency |
97.7 s |
15.1 s |
Read the latency row carefully. Under heavy load, CPU p95 approaches 100 seconds — the kind of number where users assume the dashboard is broken and file a ticket. GPU holds p95 at 15 seconds with 6x the throughput. Fewer resources per query, more queries served, lower tail latency. That’s not just faster. At high concurrency, the overall workload can be cheaper because you're doing more work with the same capacity and getting more done for the same cost.
Figure: Latency vs concurrency — p90 (solid) and p50 (dotted), log scale. GPU stays roughly 6–7x below CPU.
So, is the GPU-acceleration really worth it?
The honest way to value isn't 'X times faster'; it's how much CPU capacity would I have to buy to match it?
At 1000 concurrent users on the same F64, GPU delivers 101 q/s versus CPU’s 17 q/s — almost exactly 6x the throughput from the same capacity, flipped on with one toggle.
Put differently: to match GPU throughput on CPU, you’d need roughly 6x the capacity — an F64 would have to grow to something like an F384. And that’s the generous reading, because it assumes CPU throughput scales linearly with capacity, which it doesn’t (look at the CPU column above — going from 64 to 1000 users barely moves it from 3.7 to 17 q/s). In practice, you’d likely need even more.
So, the value of the feature shows primarily at high concurrency, exactly what we expect from agents and dashboards operating directly over the lake.
The flip side, stated just as plainly: at a single user, GPU does 4.56 q/s vs CPU’s 0.65 — better, but a difference you may not even feel on one interactive query. The GPU-acceleration earns its money under load, not at the desk of one analyst running one query. If your warehouse is lightly used, plain CPU execution is the right, cheaper choice. If fifty dashboards hit it at 9am Monday, the math changes completely.
That’s the whole value story in one line: marginal at idle, decisive at concurrency.
The test harness: Built on Rayfin
The load-testing app (which, I’ll admit, looks better than most internal tools have any right to, and far better than anything I could’ve built manually) was built with Rayfin — Microsoft’s new open-source SDK and CLI for building Fabric apps, announced at Build 2026. You define the backend in code, run Rayfin up, and it deploys a fully managed, governed backend on Fabric with data landing in OneLake by default.
There’s a pleasing symmetry to it: I built the tool that stress-tests Fabric DW on Fabric, using the new app platform. Dogfooding all the way down.
Figure: Fabric DW TPC-H Concurrency analysis.
The tech, partnership and hardware powering this
The research behind these numbers is published in the CoddSpeed: Hardware Accelerated Processing in Microsoft Fabric paper, which won the Best Industry Paper award at SIGMOD 2026 — peer-reviewed by the database research community.
Full disclosure: I’m a co-author. I mention that not to take credit, but because you should know my bias when I tell you the research is solid.
NVIDIA accelerated computing is critical to these dramatic speedups. Custom CUDA kernels integrate NVIDIA accelerated computing directly into the Fabric Data Warehouse execution engine, enabling highly parallel processing. partnership with NVIDIA has been crucial for this release, and our teams worked hard together with our NVIDIA partners to get this product to market.
Ian Buck, Vice President of Hyperscale and HPC at NVIDIA shares with us his perspective:
“AI applications are redefining how a data warehouse needs to perform. As AI agents reason over enterprise data, analytics systems need low-latency performance for many simultaneous users. With NVIDIA accelerated computing and custom CUDA kernels built directly into Microsoft Fabric Data Warehouse, Microsoft is bringing the SQL workflows customers already use into the production AI era.”
The honest caveats
- I picked 100 GB because, while not very large by big data standards, it is a reasonable range for an enterprise dashboard. Also, it is a reasonable size for a weekend project and a Fabric Trial F64 capacity. More “official,” larger scale benchmarks executed by the Fabric DW team have been presented at Build and will be available.
- This is preview. Behavior will change, and your mileage will vary by workload.
- TPC-H is analytical and read-heavy — exactly where GPUs shine. If your workload is small, low-concurrency, or transactional, plain CPU execution is still sufficient, and you don’t need this.
- My benchmark noise (see the 256-user run) is a reminder that one run is an anecdote. Run your own workload before drawing conclusions.
But for the case GPUs are built for — high-concurrency analytical queries — the gap isn’t marginal. It’s the difference between a 15-second p95 and a 97-second one.
Intrigued? Learn more about it in our deep-dive blog post, A new analytics frontier: GPU-accelerated Fabric Data Warehouse, and and request access to the preview to try it with your own workload.