Forum Discussion
Mounting NB on- the-fly
I must admit I am not sure what it means to Attach or Mount a lakehouse to a Notebook. Is Attach and Mount the same thing?
Anyway, I am unsure if the mssparkutils mount method is able to make a lakehouse the default lakehouse for the Notebook.
Okay, it can mount one or more lakehouses to the Notebook, which per my current understanding makes it possible to use relative directory paths (e.g. useful for os module and Pandas).
However, if you want to use Spark SQL, I guess we would need to be able to mount a lakehouse as the default Lakehouse of the Notebook? I am not sure if the mssparkutils mount method can do that.
It seems I am able to mount a Lakehouse as the default Lakehouse by using the solution to this forum thread: Solved: Re: How to set default Lakehouse in the notebook p... - Microsoft Fabric Community
However I need to use this code (need to add the -f argument):
%%configure -f
{
"defaultLakehouse": {
"name": "<lakehouseName>",
"id": "<lakehouseID>",
"workspaceId": "<workspaceID>"
}
}
It means the Livy session will need to restart. I don't know what the Livy session is, however it seems all variables will get lost by doing that.
I also didn't find out how to insert the lakehouseName, lakehouseID and workspaceID as variables.
So I had to hardcode those values in the JSON structure in the %%configure cell.
I guess it should be possible to pass variables into the JSON structure, however at the moment I don't know how to do that.
Other than that, it seemed to work, when I hardcoded the values.
This worked for me in a Notebook which has no default lakehouse. The default lakehouse gets attached programmatically in the %%configure code cell.
I hardcoded the values for lakehouseName, lakehouseID and workspaceID (as mentioned in previous comment).
%%configure -f
{
"defaultLakehouse": {
"name": "<lakehouseName>",
"id": "<lakehouseID>",
"workspaceId": "<workspaceID>"
}
}
%%sql
CREATE OR REPLACE TEMPORARY VIEW Dim_Product_temp_vw
AS
SELECT *, CURRENT_TIMESTAMP AS load_timestamp
FROM Dim_Product
df = spark.sql("SELECT * FROM Dim_Product_temp_vw")
df.write.mode("append").saveAsTable("Dim_Product_2")
However, if you want to switch to another default Lakehouse on-the-fly, then the temporary view will not be available after you switch to another default Lakehouse.
Because the variables get lost when running the %%configure -f code cell.
So I guess you would need to solve that using a workaround.
- frithjof_v2 years ago
Community Champion
smpa01
Could you explain more in detail what you want to do with the data?
Moving data from one workspace to another?There could be easier ways to accomplish this instead of attaching Lakehouse programmatically to Notebook.
- frithjof_v2 years ago
Community Champion
Here is a solution which doesn't involve mounting or attaching Lakehouses to the Notebook:
Perhaps you can use temporary views to be able to work with Spark SQL without having a default Lakehouse for your Notebook.
In this example I have a Notebook without any default Lakehouse and without any mounted/attached lakehouses.
I create a dataframe to connect to data from a lakehouse using the abfss path, and create a temporary view based on that dataframe.
df = spark.read.load("abfss://<workspaceID>@onelake.dfs.fabric.microsoft.com/<lakehouseID>/Tables/<tableName>") df.createOrReplaceTempView("MyView_Read")I use Spark SQL to do some modifications to the temporary view, and save it as another temporary view.
%%sql CREATE OR REPLACE TEMPORARY VIEW MyView_Write AS SELECT *, Current_Timestamp as LoadTime FROM MyView_Read;I create a new dataframe from the new temporary new.
I then write this new dataframe to another lakehouse table in another workspace.
df_write = spark.sql("SELECT * FROM MyView_Write") df_write.write.mode("append").save("<anotherWorkspaceID>@onelake.dfs.fabric.microsoft.com/<anotherLakehouseName>/Tables/<anotherTableName>")This seems to work fine for me.