frithjof_v's avatar
frithjof_v
Community Champion
1 year ago
Status:
New

Use SparkSQL without Default Lakehouse

Make it possible to write SparkSQL without having a Default Lakehouse.

 

With 3- or 4- part naming, i.e.

[workspace].[lakehouse].[schema].[table] 

 

there should be no need to attach a Lakehouse in order to use SparkSQL.

 

Needing to attach a Lakehouse is annoying and adds extra complexity.

5 Comments

  • Oh nice! this is a good one! Agreed it can be a pain, especially when merging from feature branch to main. Seems to work well with deployment pipelines though.
  • BHouston1's avatar
    BHouston1
    Regular Visitor
    Not a bad idea on the 4-part naming, but I could see an issue if a workspace is renamed and this breaks notebooks. I think that shortcuts are designed to solve for this, but then unfortunately you would need a default lakehouse before you can use them...
  • smpa01's avatar
    smpa01
    Community Champion

    Not exactly the way you are asking, but spark sql is capable of reading tables from unattached lakehouse by utilizing Azure Blob File System Secure path as following

     

    %%sql
    CREATE OR REPLACE TEMPORARY VIEW df
    USING delta
    OPTIONS (
      path 'abfss://'
    );
    
    SELECT * FROM df LIMIT 10;

     

  • BHouston1's avatar
    BHouston1
    Regular Visitor
    smpa01 Yes, that is available in out of the box Spark SQL, but I believe this idea is about being able to refer to tables using Fabric workspace/lakehouse conventions rather than abfss paths. Similar to Unity Catalog in Databricks. You therefore don't need a default lakehouse if you choose to always specify which lakehouse in each query.
  • GaryFfi's avatar
    GaryFfi
    Regular Visitor
    Would be very useful when writing common notebooks in a medallion architecture. For example, I have three notebooks to overwrite, append, and merge data from a Bronze lakehouse to a Silver warehouse. Right now, I am required to host it in each of the domain workspaces like finance sales and marketing just to attach/bind it to the proper persistent data store. The code is 100% same otherwise. This creates a maintenance nightmare.

Recent ideas