Forum Discussion

Luigia-Costabil's avatar
Luigia-Costabil
Advocate I
2 months ago
Solved

Link to Fabric - Dataverse storage

Hi Fabric Community,

I'm trying to better understand the architecture behind Dataverse Link to Microsoft Fabric and how data is physically stored.

According to the documentation:

"Link to Fabric creates an optimized replica of your data in delta parquet format, using Dataverse storage."

and

"Tables added to OneLake consume Dataverse storage."

This raises a few questions for me.

My understanding was that Link to Fabric would replicate Dataverse data directly into OneLake/Fabric storage, similarly to how Fabric Mirroring works for other sources. However, the documentation seems to indicate that the analytics replica is actually stored in Dataverse-managed storage and exposed to Fabric.

Could someone clarify:

  1. Where is the replicated data physically stored when using Link to Fabric?

    • Dataverse storage?

    • OneLake storage?

    • Both?

  2. If the replica consumes Dataverse storage, what is actually stored inside the Fabric Lakehouse that gets created?

  3. Is Link to Fabric architecturally different from Fabric Mirroring?

    • If yes, what are the main differences?

  4. For large Dataverse environments, how do customers typically balance the trade-off between:

    • Dataverse Shortcuts

    • Link to Fabric

    • Copy Pipelines

    • Synapse Link

I'm particularly interested in understanding the storage and cost implications of each approach.

Thanks in advance!

  • Hi Luigia-Costabil ,

    Thankyou for reaching out to Microsoft Fabric Community.

    For the first question, the analytics replica created by Link to Fabric appears to be maintained using Dataverse-managed storage, which is why it consumes additional Dataverse capacity. Rather than creating a separate copy of the data in OneLake in the same way as Fabric Mirroring, Fabric establishes a direct and secure link to the Dataverse analytics replica, allowing you to work with the data from Fabric while the underlying storage remains associated with Dataverse.
    Regarding the Lakehouse that gets created in Fabric, it functions as an access and analytics layer over the Dataverse data rather than an independent storage location for a replicated copy of that data. This aligns with the documentation describing the use of shortcuts from Dataverse into OneLake.

    For the comparison with Fabric Mirroring, the architectures do appear to be different. With Link to Fabric, the analytics replica is managed through Dataverse and exposed to Fabric through the link. With Fabric Mirroring, the replicated data is maintained within Fabric-managed storage. Mirroring also provides a Fabric-native replication experience where Fabric manages the synchronization process after configuration.

    When choosing between the available options, the decision usually comes down to the level of control and data movement required. Dataverse Shortcuts are useful when you want lightweight access to Dataverse data from Fabric. Link to Fabric provides a streamlined analytics experience with minimal setup while Dataverse manages the replica. Copy Pipelines and Dataflows are typically preferred when you need full control over data movement, transformations, and storage within Fabric. Azure Synapse Link for Dataverse is often considered for advanced analytical scenarios that require additional flexibility over how exported data is consumed downstream.

    From a storage and cost perspective, the primary consideration is where the analytical copy of the data is maintained. If the analytics replica is managed through Dataverse, then Dataverse storage consumption becomes an important factor. Approaches that move or replicate data into Fabric shift more of the storage considerations toward Fabric capacity and storage usage. As a result, understanding where you want the data to live and which platform you want to optimize for is often one of the key factors when selecting the most appropriate integration pattern.

    For Reference:
    Analytics end-to-end with Microsoft Fabric

    Link your Dataverse environment to Microsoft Fabric and unlock deep insights 

    I hope this information helps. Please do let us know if you have any further queries

    Best Regards,
    Abdul Rafi



1 Reply

  • Hi Luigia-Costabil ,

    Thankyou for reaching out to Microsoft Fabric Community.

    For the first question, the analytics replica created by Link to Fabric appears to be maintained using Dataverse-managed storage, which is why it consumes additional Dataverse capacity. Rather than creating a separate copy of the data in OneLake in the same way as Fabric Mirroring, Fabric establishes a direct and secure link to the Dataverse analytics replica, allowing you to work with the data from Fabric while the underlying storage remains associated with Dataverse.
    Regarding the Lakehouse that gets created in Fabric, it functions as an access and analytics layer over the Dataverse data rather than an independent storage location for a replicated copy of that data. This aligns with the documentation describing the use of shortcuts from Dataverse into OneLake.

    For the comparison with Fabric Mirroring, the architectures do appear to be different. With Link to Fabric, the analytics replica is managed through Dataverse and exposed to Fabric through the link. With Fabric Mirroring, the replicated data is maintained within Fabric-managed storage. Mirroring also provides a Fabric-native replication experience where Fabric manages the synchronization process after configuration.

    When choosing between the available options, the decision usually comes down to the level of control and data movement required. Dataverse Shortcuts are useful when you want lightweight access to Dataverse data from Fabric. Link to Fabric provides a streamlined analytics experience with minimal setup while Dataverse manages the replica. Copy Pipelines and Dataflows are typically preferred when you need full control over data movement, transformations, and storage within Fabric. Azure Synapse Link for Dataverse is often considered for advanced analytical scenarios that require additional flexibility over how exported data is consumed downstream.

    From a storage and cost perspective, the primary consideration is where the analytical copy of the data is maintained. If the analytics replica is managed through Dataverse, then Dataverse storage consumption becomes an important factor. Approaches that move or replicate data into Fabric shift more of the storage considerations toward Fabric capacity and storage usage. As a result, understanding where you want the data to live and which platform you want to optimize for is often one of the key factors when selecting the most appropriate integration pattern.

    For Reference:
    Analytics end-to-end with Microsoft Fabric

    Link your Dataverse environment to Microsoft Fabric and unlock deep insights 

    I hope this information helps. Please do let us know if you have any further queries

    Best Regards,
    Abdul Rafi