Summary
Fabric Data Pipelines' ExecutePipeline activity, when run with waitOnCompletion: true, has a documented but unresolved platform behavior: after the invoked child pipeline actually finishes, the parent pipeline's next activity can sit in a "queued" state for 1-5 minutes before it resumes. This delay is not configurable and is not attributable to capacity, connections, or activity content — it is purely platform dispatch/finalize overhead. This is already discussed in Microsoft Q&A threads, where Microsoft support has acknowledged the symptom as "strange" without offering a resolution or workaround.
For any architecture that relies on nested ExecutePipeline calls with waitOnCompletion: true — which is the only way to build a fan-out/dynamic-dispatch orchestration pattern in Fabric today (see "Why this pattern exists" below) — this lag compounds directly with (a) the number of orchestration boundaries in a run and (b) the number of items processed within each boundary, and can end up being the majority of total wall-clock time even when the actual work is trivial.
Why this pattern exists (context Microsoft's engineering team may not have visibility into)
Fabric's ForEach activity does dynamic parallel iteration, but lane assignment is static: items are divided up front, so one lane can end up processing several slow items while other lanes sit idle once they've exhausted their assigned batch. There is no built-in dynamic work-stealing/pull-based queue primitive.
To work around this, we built a custom pattern: an Until loop with N parallel lanes, where each lane independently and atomically dequeues the next available work item from a SQL table (UPDATE TOP (1) ... WITH (UPDLOCK) ... OUTPUT ...) the instant it finishes its current item. This correctly solves the load-imbalance problem ForEach can't. But it requires every single work item to be dispatched as its own ExecutePipeline call with waitOnCompletion: true (so a lane can actually block on and then immediately re-loop after each item) — and in our specific chain, that call passes through two nested levels (an outer pipeline-switch/dispatch pipeline, then the actual work pipeline), so every item pays the finalize lag twice.
What we measured
A real production batch run (internal reference MasterPipelineRunID = b24c1690-...) where every work item hit a fast-path skip (no actual data movement — a delta pre-check determined nothing had changed) still took ~22 minutes wall-clock, broken down as:
| Execution Group | Entities | Actual work | Idle gap immediately after |
|---|---|---|---|
| 1 (Bronze) | 59 | 4m46s | 5m45s |
| 2 (Bronze SQL refresh) | 1 | 20s | 1m57s |
| 3 (Silver pre-load refresh) | 1 | 21s | 1m49s |
| 4 (Silver) | 27 | 2m51s | 3m46s |
| 5 (Silver post-load refresh) | 1 | 16s | — (last group) |
~8m54s of actual execution vs. ~13m17s of pure idle time between orchestration boundaries — roughly 60% of total runtime is platform overhead, not work, in the best case where nothing needed processing. The size of each gap scales with how many items were in the group that just finished (59- and 27-item groups produce multi-minute gaps; single-item groups still produce ~2-minute gaps), consistent with ForEach/Until performing internal aggregation/cleanup work on exit that scales with iteration count.
What we're asking
- Acknowledge and document the actual mechanism, not just the symptom. Is this a polling interval on the orchestration service side, a queue-drain step, or something else? Right now customers have no way to reason about or budget for this cost because its cause isn't documented anywhere we've found beyond community Q&A threads.
- A configurable or reduced finalize interval for ExecutePipeline, at least for pipelines invoked at high frequency/short duration (our per-item work often completes in under a minute, so a 1-5 minute finalize tax is a 100-500%+ overhead multiplier on top of real work).
- A first-class dynamic work-queue primitive (a native "pull next available item" pattern) so customers don't have to build this ourselves via nested ExecutePipeline calls in the first place — ForEach with genuinely dynamic (not pre-assigned) lane scheduling would remove the need for this pattern, and the finalize-lag cost with it, entirely.
We're happy to share the underlying pipeline JSON and the exact run telemetry above if it's useful for reproducing this on Microsoft's side.
Recent ideas
Increase or Make Configurable the Power BI Browser Tab Limit (Currently 20 Items)
Power BI users frequently work across multiple reports, dashboards, semantic models, notebooks, and Fabric artifacts simultaneously. The current browser limitation of 20 open Power BI/Fabric items re...sharadmahajan1 hour agoRegular VisitorNew233Views15likes1CommentEnable Fine-Grained, Least-Privilege Control Over Fabric Artifact Creation
Please introduce more granular permission controls that allow admins to determine which user groups can create specific Fabric artifacts. The current all-or-nothing model—where users can either creat...krish0077 hours agoNew MemberNew159Views2likes1CommentMS Fabric User Permission Segmentation by Artifacts
Currently, when we grant users permission to create Fabric artifacts, they can create anything available in the workspaces they have access to. This unrestricted access can lead to an unstructured en...Anonymous7 hours agoNot applicableNew327Views8likes2Comments