Forum Discussion
QUESTION::NOTEBOOK::PYSPARK::CROSS WORKSPACE::LAKEHOUSE DETAIL ACCESS
Hi Element115
I don't find an out-of-box solution to get the desired outcome you want. Usually we would tend to use Power BI REST APIs or Fabric REST APIs to get the data. As you want to use PySpark, you can use some python libraries to call these REST APIs to get the data. Here are some of my ideas:
1. List all capacities and get their capacity IDs. Filter the result according to capacity type (sku) to remain only capacities that may have Fabric items. Capacities - List Capacities - REST API (Core) | Microsoft Learn
2. Iterate all above capacities to get the list of all workspaces in those capacities. Workspaces - List Workspaces - REST API (Admin) | Microsoft Learn
3. Iterate all of above workspaces to get the list of all lakehouses in these workspaces. Items - List Lakehouses - REST API (Lakehouse) | Microsoft Learn
4. Iterate all lakehouses to get the tables in each of them. Tables - List Tables - REST API (Lakehouse) | Microsoft Learn
5. To get the size of a delta table, currently there is no such API. You might refer to the solution in this thread QUESTION::LAKEHOUSE::SQL ENDPOINT::SYSTEM VIEWS - Microsoft Fabric Community
Hope this would be helpful.
Best Regards,
Jing
If this post helps, please Accept it as Solution to help other members find it. Appreciate your Kudos!
I have also updated the Spark runtime to use Spark 3.5 because the 3.5 API has this method:
addArtifacts(*path[, pyfile, archive, file]) Add artifact(s) to the client session.
but the runtime engine complains when I use it saying this method can only be used if you also use Spark Connect and I have no idea atm what is or how to use Spark Connect.
But this method looks like it can add lakehouses to the current SparkSession object that is running in a different workspace, and if so, then the metadata of these lakehouses should suddenly become available.
Anonymous Would you know how to use Spark Connect from a Fabric Notebook to be able to use SparkSession.addArtifacts() and add all the paths to other workspaces and lakehouses to the SparkSession?