Forum Discussion

Perez0s0's avatar
Perez0s0
Frequent Visitor
8 months ago
Solved

pipeline fabric how to improve our performance?

Hi, 

Below our current pipeline thats reading seperate excel files from our onelake silver. There are 6 notebooks that are cleaning 6 seperate files and these are converted to delta table. After this it gets in our pipeline transferred to Gold warehouse. Resulting in a semantic model update. Currently this process is taking ~10 minutes. Mainly because each notebook takes about 1 minute. Is there a way to speed this up? Can these run parallel? our capacity is F4, so we ran into quiet some capacity struggles before.

We are both relatively new to Fabric and our data team is new as well so we're asking for some guidance 

  • Hi Perez0s0 ,

    If the notebooks are totally independent of each other, then you can start them all at once and next step to Gold would be coming from success of all 6 notebooks. This will make the Gold notebook/pipeline run ONLY after all 6 are successful.

    However, if running all at once takes a toll on your capacity, you can split them in pack of 3 something like below, where you are still achieving parallel processing to some degree.

     

     

    I do recommend adding a failure handling in this though (Either sending an email or running some error logging at each step failure)

    Also to add - agree with enabling high concurrency as suggested by AntoineW 

    this will help reduce the starting time on the spark as one notebook can share the session that was already started by other notebook.

6 Replies

  • Hi Perez0s0 ,

    If the notebooks are totally independent of each other, then you can start them all at once and next step to Gold would be coming from success of all 6 notebooks. This will make the Gold notebook/pipeline run ONLY after all 6 are successful.

    However, if running all at once takes a toll on your capacity, you can split them in pack of 3 something like below, where you are still achieving parallel processing to some degree.

     

     

    I do recommend adding a failure handling in this though (Either sending an email or running some error logging at each step failure)

    Also to add - agree with enabling high concurrency as suggested by AntoineW 

    this will help reduce the starting time on the spark as one notebook can share the session that was already started by other notebook.

  • I would focus on whichever activity is taking the longest in your pipeline first.

  • Hi Perez0s0

     

    In addition to the high concurrency mode that AntoineW mentioned, you may also want to look at combining notebooks so you have less notebooks overall to run. 


    If you found this helpful, consider giving some Kudos. If I answered your question or solved your problem, mark this post as the solution. 

  • Perez0s0's avatar
    Perez0s0
    Frequent Visitor

     after enabling high concurrency mode, adding a session tag didnt improve the process... 

    I will come back to this on monday.

  • Anonymous's avatar
    Anonymous
    Not applicable

    Hi Perez0s0 ,

    I would also take a moment to thank AntoineW , for actively participating in the community forum and for the solutions you’ve been sharing in the community forum. Your contributions make a real difference.
     

    I wanted to check if you had the opportunity to review the information provided. Please feel free to contact us if you have any further questions