Forum Discussion
NagaRK
Advocate I
1 year agoSpark in Notebook taking more time to process the data.
Hi all, I'm working on a diagnostic log ingestion engine built with PySpark and Delta Lake on Microsoft Fabric. My setup parses incoming ZIP logs from a server, transforms signal data per file in...
Srisakthi
Super User
1 year agoHi NagaRK ,
The performance is based on your spark setting,the capacity that you are using.
Make sure these settings are available in your spark environment and which is attached to your notebook
1. Native Execution Engine - uses vectorized engine and helps improves query performance
2. Apache Spark latest runtime
3. Leverage Autotune properties - helps in speed up workload execution and performance
4. Leverage Spark Resource Profiles - Microsoft fabric provides flexibility to utilise these resource profiles. By default all the workspaces is atatched to writeHeavy profile.
5. Make sure to enable v-order
You can refer this article on threadpool
Regards,
Srisakthi