Forum Discussion
User Data Functions and Spark
Use case: use Translytical Task Flows to update a file in lakehouse and refresh a delta table based on that file on which power BI report is built.
I wanted to refresh the delta table after updating the file using spark, but if I include any spark code I am getting the below invocation error.
{
"functionName": "test_spark_basic",
"invocationId": "xxxxxxx-b9bb-422e-8a80-xxxxxxxxx",
"status": "Failed",
"output": "",
"errors": [
{
"errorCode": "InternalError",
"message": "An internal execution error occured during function execution",
"properties": {
"error_type": "PySparkRuntimeError",
"error_message": "Java gateway process exited before sending its port number."
}
}
]
}
code I used for spark testing:
questions:
1. can we use spark inside User Data Functions? If yes, pls provide a guide. (I can see pyspark module in library section)
2. Is there any other way to refresh the delta table after modification of file from UDF itself?
Hi Anonymous,
Thank you for reaching out to Microsoft Fabric Community.
Spark is not supported inside User Data Functions, even though you may see pyspark in the library section. UDF’s run in a restricted python environment that does not include a Spark runtime, which is why you are getting the java gateway error.
- Use a Notebook step or Data Pipeline in the same Translytical Task Flow to refresh the Delta table. For example like below:
df = spark.read.format("parquet").load("Files/<lakehouse_name>/file_path")
df.write.format("delta").mode("overwrite").save("Tables/<lakehouse_name>/delta_table") - If your power bi report is connected to this table via Direct Lake mode, it will reflect the updates automatically no need of manual refresh.
If this post helps, then please consider Accepting as solution to help the other members find it more quickly, don't forget to give a "Kudos" – I’d truly appreciate it!
Thanks and regards,
Anjan Kumar Chippa
- Use a Notebook step or Data Pipeline in the same Translytical Task Flow to refresh the Delta table. For example like below:
1 Reply
- v-achippaCommunity Support
Hi Anonymous,
Thank you for reaching out to Microsoft Fabric Community.
Spark is not supported inside User Data Functions, even though you may see pyspark in the library section. UDF’s run in a restricted python environment that does not include a Spark runtime, which is why you are getting the java gateway error.
- Use a Notebook step or Data Pipeline in the same Translytical Task Flow to refresh the Delta table. For example like below:
df = spark.read.format("parquet").load("Files/<lakehouse_name>/file_path")
df.write.format("delta").mode("overwrite").save("Tables/<lakehouse_name>/delta_table") - If your power bi report is connected to this table via Direct Lake mode, it will reflect the updates automatically no need of manual refresh.
If this post helps, then please consider Accepting as solution to help the other members find it more quickly, don't forget to give a "Kudos" – I’d truly appreciate it!
Thanks and regards,
Anjan Kumar Chippa
- Use a Notebook step or Data Pipeline in the same Translytical Task Flow to refresh the Delta table. For example like below: