Forum Discussion
Any integration or tutorials for Spark Connect?
- 1 year ago
Hi dbeavon3 ,
Based on my understanding, since Spark connect requires remote connectivity, it needs a hostname which would be the IP address of the Spark Context. And since there is no authentication mechanism invovled with Spark-connect unless you manually setup a re-direction URL mechanism (authentication proxy), I don't believe Fabric will allow that level of configuration in their cloud system.
Using Managed Virtual networks with Fabric, you might get the URL of the Spark context and use it, but again this is just my assumption and as you said, there is no documentation, it is difficult to validate unless we do a PoC.
The following seems true after I read the description from MS site and Spark site.
Maybe someone copy/pasted from the OSS docs for Apache Spark.
Hi dbeavon3 ,
Based on my understanding, since Spark connect requires remote connectivity, it needs a hostname which would be the IP address of the Spark Context. And since there is no authentication mechanism invovled with Spark-connect unless you manually setup a re-direction URL mechanism (authentication proxy), I don't believe Fabric will allow that level of configuration in their cloud system.
Using Managed Virtual networks with Fabric, you might get the URL of the Spark context and use it, but again this is just my assumption and as you said, there is no documentation, it is difficult to validate unless we do a PoC.
The following seems true after I read the description from MS site and Spark site.
Maybe someone copy/pasted from the OSS docs for Apache Spark.
Hi thanks for the update. I was hoping to play with Spark Connect on Fabric, but I may have to revert to Databricks.
I actually use HDI in Azure more than anything else, but it is stuck on an earlier version of Spark. Hoping that changes in the near future. I love HDI, and wish Microsoft would give it a bit more TLC!
If I had time, I would work on a PoC. But even if it worked, it would probably be fragile, and would not be future-proof. The problem with Fabric Spark is that the cluster is not really a first-class member of the workspace. It is only created behind the scenes for the purpose of a notebook execution ("clients"). All the monetization is handled per notebook. So any of the "server" -oriented features of Spark are probably tucked away under the hood, and there is no surface-area for customers to interact with these features. It is unlikely that "Spark Connect" will be supported, until Microsoft finds a way to bill us for it, and make even more money than what they earn from notebooks. There is a tipping point where I'm guessing they would lose money, if customers are able to manage their own cluster, and successfully optimize the number of notebooks that are able to run on a given day.