Forum Discussion
Issues with High Concurrency Mode in Notebook Pipelines
- 1 year ago
Hi djbc1986
Welcome to the Microsoft Fabric Community Forum.
To address performance issues in Microsoft Fabric pipelines, ensure all prerequisites for effective Spark session reuse are met. Confirm that all notebooks in the pipeline are configured to use the same Spark pool, as session reuse is only possible within a shared pool context. Each notebook must explicitly enable Spark session reuse by setting the same session_tag and selecting the option to reuse an existing Spark session in the notebook activity’s advanced settings. Notebooks should run sequentially to avoid session conflicts, with pipeline dependencies enforcing this order. Minimize idle times between executions to prevent session termination. Optionally, add an initialization notebook at the start to pre-warm the Spark environment. Monitor Fabric capacity utilization to ensure sufficient resources are available, reducing queue times and execution delays. These practices will enhance Spark efficiency and reduce delays.For reference , please go through the Microsoft official documents below:
Configure high concurrency mode for notebooks in pipelines - Microsoft Fabric | Microsoft LearnConfigure high concurrency mode for notebooks - Microsoft Fabric | Microsoft Learn
Concurrency limits and queueing in Apache Spark for Fabric - Microsoft Fabric | Microsoft Learn
If this response resolves your query, kindly mark it as Accepted Solution to help other community members. A Kudos is also appreciated if you found the response helpful.
Thank you for being part of Fabric Community Forum.
Regards,
Karpurapu D,
Microsoft Fabric Community Support Team. - 1 year ago
Hi djbc1986 ,
Thanks for sharing the detailed explanation and screenshots—they’re very helpful for diagnosing the issue.
What you’re experiencing is a common challenge when using high concurrency mode in notebook pipelines, especially in Fabric environments. Based on your screenshots:
- Each notebook step (1, 2, and 3) is taking a similar amount of time, even though only the second one is doing actual data processing.
- The Spark resource usage (Screenshot 3) shows very low efficiency (4.24%), and a significant part of the total duration is spent in queue (1m 32s out of 5m 18s).
- This suggests that even with high concurrency enabled and the same session_tag, the pipeline is not reusing the same Spark session as expected, resulting in each notebook having to wait for resources and session initialization.
A few suggestions to improve this:
Session Reuse:
Double-check that all notebooks in the pipeline are using the exact same session_tag and that session sharing is supported in your environment. In some cases, session reuse is only possible when the notebooks are set up in a very specific way, and minor differences in configuration (such as different libraries or environment settings) can prevent session reuse.Cluster Startup and Allocation:
The queue times are often related to cluster allocation or startup delays. If your workspace is running at or near Fabric capacity limits, there may not be enough resources available to start all notebooks at once—even with high concurrency. Check your Fabric capacity metrics and consider scaling up if you see frequent queuing.Pipeline Structure:
If your logging notebooks (1 and 3) are lightweight and don’t need to wait for the main processing, you could potentially restructure the pipeline so that these steps run in parallel rather than strictly in sequence, reducing overall wait time.Resource Release:
Ensure that notebooks are releasing resources properly at the end of each run (e.g., closing Spark sessions if not needed), so queued steps aren’t waiting for resource cleanup.Microsoft Docs & Support:
There are some known quirks with session_tag/session reuse in Fabric and Synapse environments. I recommend checking the latest Microsoft documentation and possibly raising a support ticket if you suspect a platform limitation.
Summary:
The main bottleneck here seems to be session or resource allocation rather than actual processing time. Focus on verifying session_tag consistency, monitoring Fabric capacity, and considering pipeline restructuring for lightweight steps.Let us know if adjusting these settings helps, or if you have any follow-up findings!
Good luck!
Hi djbc1986
Welcome to the Microsoft Fabric Community Forum.
To address performance issues in Microsoft Fabric pipelines, ensure all prerequisites for effective Spark session reuse are met. Confirm that all notebooks in the pipeline are configured to use the same Spark pool, as session reuse is only possible within a shared pool context. Each notebook must explicitly enable Spark session reuse by setting the same session_tag and selecting the option to reuse an existing Spark session in the notebook activity’s advanced settings. Notebooks should run sequentially to avoid session conflicts, with pipeline dependencies enforcing this order. Minimize idle times between executions to prevent session termination. Optionally, add an initialization notebook at the start to pre-warm the Spark environment. Monitor Fabric capacity utilization to ensure sufficient resources are available, reducing queue times and execution delays. These practices will enhance Spark efficiency and reduce delays.
For reference , please go through the Microsoft official documents below:
Configure high concurrency mode for notebooks in pipelines - Microsoft Fabric | Microsoft Learn
Configure high concurrency mode for notebooks - Microsoft Fabric | Microsoft Learn
Concurrency limits and queueing in Apache Spark for Fabric - Microsoft Fabric | Microsoft Learn
If this response resolves your query, kindly mark it as Accepted Solution to help other community members. A Kudos is also appreciated if you found the response helpful.
Thank you for being part of Fabric Community Forum.
Regards,
Karpurapu D,
Microsoft Fabric Community Support Team.