Forum Discussion

SamyAbdul's avatar
SamyAbdul
Frequent Visitor
23 hours ago

naming approach for hierarchy

Hi Experts , I have designed ADLS and Delta lakes in Microsoft Azure, and In Fabric Lakehouse design  should I follow the same file structure hierarchy as Source as

                  Source

                   year=2026

                      month =09

                        day =26

 As it would make various analytics engines to scan and retrieve the data quickly. I have seen some people just implementing  as 

Source

     2026

          09

             26

This might some times might cause contention as query engine and analytic engines might not able to identify 2026 as a year. Please suggest Thank you.                                                        

                       

2 Replies

  • ShivekMaharaj's avatar
    ShivekMaharaj
    Icon for Community Champion rankCommunity Champion

    Hi SamyAbdul​,

    I would keep the explicit Hive-style naming if you are actually partitioning the Delta table.

    Microsoft documents the standard pattern as:

    Source/
    year=2026/
    month=09/
    day=26/

    rather than:

    Source/
    2026/
    09/
    26/

    The reason is that year, month and day become explicit partition columns rather than just folder names.

    That allows query engines to understand the partition structure and perform partition pruning when filters use those columns.

    Microsoft documents this pattern in the Delta table partitioning guidance.

    For example, if you create the table with:

    PARTITIONED BY (Year, Month, Day)

    Fabric/Delta will generate the corresponding Hive-style folder layout automatically.

    I would therefore avoid manually creating numeric-only folders for managed Delta tables.

    One other point: I would not automatically partition every dataset down to day level.

    Microsoft recommends avoiding too many small partitions and suggests targeting roughly 1 GB or more per partition.

    So depending on data volume, something like:

    year + month
    
    may be better than:
    
    year + month + day

    For most Fabric Runtime 2.0 workloads, Microsoft also now recommends liquid clustering for general read-performance optimization, while traditional partitioning is especially useful when you need isolated concurrent writes.

    So my rule would be:

    Delta table
    -> define meaningful partition columns
    
    physical folders
    -> let Delta/Fabric create them
    
    naming
    -> use year=2026/month=09/day=26 rather than anonymous 2026/09/26

    AI-assisted drafting: AI was used to help structure and phrase this response. I reviewed and validated the technical content before posting.

  •  

    I agree with using Hive-style partition naming when you actually need partition columns to be recognized by the engine.

    One extra point is to avoid designing the folder hierarchy first and then forcing the Delta table to match it. In Fabric, I’d define the table partitioning strategy based on query patterns and data volume, then let Delta create the physical layout.

    For lower-volume datasets, year/month/day can create too many small partitions, so testing year/month or even no traditional partitioning can be better. The main goal should be efficient pruning without creating excessive small files.