Forum Discussion
Diagnosing a slow Fabric Data Warehouse using the new skill
Hi arthurfr23,
Thanks for sharing this. I went through Microsoft's announcement and the current SQL DW operations documentation as well, and the time-correlation boundary you call out is one of the more important parts of this feature for production troubleshooting.
The skill is not really creating a new telemetry source. It is bringing existing signals from Capacity Metrics, Query Insights and SQL pool diagnostics into a bounded, read-only investigation workflow.
One detail I think is particularly important is Microsoft's documented distinction around capacity correlation.
The SQL DW operations documentation explicitly states that a Capacity Metrics operation ID is not the same identifier as Query Insights' distributed_statement_id. The skill therefore correlates the resolved Warehouse / SQL analytics endpoint and overlapping time window rather than joining those IDs.
It also treats Capacity Metrics CU seconds and Query Insights CPU milliseconds as different measurements. So an expensive workload overlapping a capacity spike is evidence of correlation, but it should not automatically be interpreted as exact query-level CU attribution.
There are a few operational boundaries around that correlation too: Query Insights retains 30 days of history and can lag by up to 15 minutes, while Capacity Metrics can expose fixed time windows rather than arbitrary timestamps.
I also noticed an interesting lifecycle-status inconsistency in the current Microsoft material.
The Fabric announcement explicitly announces the SQL DW operations skill as Generally Available, but the current Microsoft Learn page still carries an "Important" notice saying that the feature is in Preview.
So I would currently be cautious about assuming the lifecycle/support status purely from one of those pages until Microsoft aligns the documentation.
What I do like about the implementation is the emphasis on grounded diagnostics: the skill is designed to report evidence sources, treat zero-row results as valid findings, separate failures from cancellations, and avoid turning correlation into causation.
For me, that is probably the biggest practical improvement: not replacing Query Insights or Capacity Metrics, but reducing the manual correlation work while keeping the underlying evidence visible.
I'd be interested in whether the time-range behaviour you saw during testing was mainly the Capacity Metrics fixed-window behaviour, timestamp alignment, or something else.
AI-assisted drafting: AI was used to help structure and phrase this response. I reviewed and validated the technical content before posting.