data pipeline
61 TopicsHow are teams structuring Microsoft Fabric Warehouse for scalable ELT and reporting workloads?
I’m looking at architecture patterns for Microsoft Fabric Warehouse in environments where multiple data sources, transformation jobs, and Power BI workloads all depend on the same platform. A typical flow might look like: Source systems → ingestion → staging → transformations → Fabric Warehouse → semantic models → Power BI I’m particularly interested in how teams are handling: staging vs curated schemas ELT transformations inside the Warehouse incremental loading patterns data quality checks before publishing CI/CD across dev, test, and production workload separation for engineering vs reporting monitoring failed loads and long-running queries deciding when to use Warehouse vs Lakehouse for the same solution For teams running Fabric Warehouse at scale, what architecture patterns have worked best for keeping the platform maintainable and avoiding duplicated transformations?7Views0likes0CommentsConcurrency in Fabric Warehouse: how can multiple pipelines write to the same table?
Hi all, I'm designing ingestion for a Fabric Warehouse where several pipelines could write to the same tables. Examples include an hourly incremental load running at the same time as a backfill, and a few source systems landing in a shared fact table. My understanding is that Fabric Warehouse uses snapshot isolation, and that two transactions modifying the same table at the same time can conflict, with one failing. I'd like to design around that from the start instead of finding out in production. Questions: Is conflict detection at the table level, or finer-grained? Does it behave differently for INSERT-only loads vs UPDATE, DELETE or MERGE? What patterns are people using to avoid conflicts? Options I'm considering: Serializing writes per table through pipeline dependencies or a control table Landing each source in its own staging table, then running one consolidating MERGE Retry logic in pipelines for conflict errors Do long-running read queries (for example, Power BI refreshes) block or affect writes? Any monitoring tips? For example, which DMVs or Query Insights views show conflicts or long transactions? Thanks!Solved36Views1like3CommentsHow are teams using AI agents to automate modern data warehouse workflows?
I am exploring how AI agents can help improve data warehouse operations by automating repetitive tasks and assisting data teams with faster decision-making. Some areas I am interested in: Automating data pipeline monitoring and issue detection AI-assisted SQL generation and optimization Automatically identifying data quality issues Triggering workflows based on business events or data changes Generating documentation and insights from warehouse metadata With platforms like Microsoft Fabric Data Warehouse, how are teams approaching AI integration? Are you using: Fabric pipelines with AI-powered automation? Copilot or LLM-based assistants for warehouse development? Custom AI agents connected with data warehouse APIs? Would love to hear real-world architecture patterns and best practices from data engineers working with Fabric.43Views2likes4CommentsSwitch workspace from west US to east US
Wondering how to switch WS and all contents from west US to east US. Tried backup to Git repo and restore from there - Does not work synch to a new empty workspace. Original workspace was west US, (contains pipelines, notebook, warehouse) synched WS to git repo, disconnected WS Created new WS east US tried to synch git repo to new empty WS created on east US errored out on every level and nothing got synched. It just sucks with no import export option. Any other suggestions.Solved157Views2likes14CommentsPreparing Enterprise Data Warehouses for AI-Powered Applications
Hi everyone, I am exploring how organizations are preparing their enterprise data platforms for AI-powered applications. As more teams start building AI assistants and intelligent applications, having clean, structured, and accessible data becomes increasingly important. I would like to understand how the community is approaching this with Microsoft Fabric Data Warehouse. Some questions: - What data modeling approaches work best when preparing warehouse data for AI and analytics workloads? - How are teams balancing traditional BI reporting requirements with new AI use cases? - What strategies are you using for maintaining data quality and governance at scale? - Are there recommended patterns for connecting AI applications with enterprise warehouse data securely? Would love to hear practical experiences and lessons learned from teams working with Fabric Data Warehouse. Thanks!Solved81Views0likes5CommentsBest approach(es) to refresh Dev Warehouse from Prod Warehouse (Read-Write Copy) in Microsoft Fabric
We are currently evaluating the best practices for refreshing a Development Data Warehouse using data from a Production Data Warehouse within the same Fabric tenant. We have evaluated a few approaches based on the use case: Read-Only Access (Zero-Copy): If Dev only needs read access to Prod data, using OneLake Shortcuts is the clear choice. It provides instantaneous access with zero data duplication, no storage overhead, and minimal setup effort. Read-Write Access (Full Schema & Data Copy): If Dev requires a Read-Write copy (allowing developers to run DML operations like UPDATE, INSERT, or DELETE without impacting Production), our current options seem to be: Fabric Data Pipelines (Copy Activity): Scheduled or trigger-based pipelines to pull data from Prod to Dev. Cross-Database Querying / CTAS: Running T-SQL (CREATE TABLE AS SELECT ...) across workspaces. The Challenge: Both Read-Write options (Pipelines and CTAS) can become cumbersome to set up and maintain, especially as datasets grow. Additionally, duplicating large volumes of data incurs extra storage and CU (Capacity Unit) processing costs. Was wondering if there are better, more efficient options available in Microsoft Fabric to maintain or periodically refresh a Read-Write copy of a Warehouse in a Dev workspace? Is Table Cloning / Zero-Copy Clone supported or planned for cross-workspace scenarios? How are others managing periodic Dev environment refreshes from Prod without creating heavy data copy pipelines?323Views3likes11CommentsProblems in UK South
Anyone experiencing issues in UK South? Our issues this morning include: - Data engineering pipeline connections to Fabric warehouse SQL endpoints not working ('A task was cancelled') - Executing notebooks from engineering pipelines not working (timing out).Solved1.3KViews1like8CommentsFabric pipeline Deployment dev to prod
Hi there, I have been developing fabric data pipelines, warehouse, lakehouse in a fabric dev workspace. I noticed when deployment is done from WS_Dev to WS_Prod, pipelines are still pointing to Dev workspace. I can parameterize it. Where in Prod workspace do I need to configure that all objects should point to Prod workspace so code deployments from dev dont overwrite the variable in PROD. master pipeline with look up After lookup -- invoke pipeline child pipeline with copy activity.Solved2.6KViews0likes7CommentsStored Procedure to Manipulate Tables
I am migrating stored procedures from SSIS into a Fabric Warehouse stored procedure. When I try to run it from a script or stored procedure activity in a pipeline I get the error "queries referencing variables are not supported in distributed processing mode" similar to this article https://community.fabric.microsoft.com/t5/Data-Warehouse/SQL-statement-in-Warehouse/m-p/3777309 My stored procedure contains these things which I think is causing the error: Input variables Insert rows into an existing table Delete rows from an existing table What I've tried I tried to run through a Lookup activity similar to this article(https://community.fabric.microsoft.com/t5/Pipelines/How-to-Pass-Parameters-to-a-Stored-Procedure-in-Microsoft-Fabric/m-p/4270431) but am getting the same error code. Run through Stored Procedure activity - No luck Run through Notebook - Use Fabric tokens to connect to the warehouse directly as a database Seems to run?? I need to keep testing Is a Notebook really the only way to do this? Ideally I was hoping for a low code option as the eventual users prefer low codeSolved1.6KViews0likes5CommentsPipeline Duration Mismatch – Parent Pipeline Shows 3mins but Child Activities Total Only 1Min
I have two pipelines. The first pipeline contains two activities: Lookup and ForEach. Inside the ForEach activity, I’m invoking a second pipeline. The second (invoked) pipeline contains 6 activities, which connect to source SQL Servers via a gateway and load data into a Data Warehouse in Fabric. Here’s the issue: From the parent (driver) pipeline view, it shows that Pipeline 2 takes about 3 minutes to complete. However, when I open Pipeline 2 and sum the duration of the individual activities, the total comes to only 1 minute. Where is the remaining 2 minutes being spent?Solved1.5KViews0likes9Comments