data pipeline
630 TopicsFabric pipelines audit
Hi community, I am looking for a way to log pipelines details into an audit delta table. To be more specific, I want to log the start time, end time, status, error message if any etc. Having this by activity name would be even better. I was thinking to use a notebook that's being called at the end of each pipeline, so somehow to have only a notebook that can be used in all of my pipelines (because i have a lot of them). Is this something that you were able to build? Any guidance would be very much appreciated. thanks!26Views0likes2CommentsPower Automate Export to PDF Returns Blank Report for Report Connected to Shared Semantic Model
Hi Team, I am facing an issue with Power Automate's "Export To File for Power BI Reports" action. Scenario I have a Power BI Semantic Model (Dataset A). Report A is built directly on Dataset A. Report B is another report built using the same semantic model(shared dataset/thin report approach). Expected Behavior When Power Automate exports Report B to PDF, the report should contain the same data that is visible in Power BI Service. Actual Behavior Exporting Report A through Power Automate generates a PDF with data correctly displayed. Exporting Report B through Power Automate generates a PDF, but the visuals are blank and no data is shown. There are no export errors. Additional Findings Manual export from Power BI Service (File > Export > PDF) works correctly for both reports and the generated PDF contains data. The issue only occurs when exporting through Power Automate. Both reports use the same semantic model. The semantic model is accessible and contains data. If I export pages from Report A, data is visible in the PDF. If I export pages from Report B (connected report/thin report), the PDF is blank. Questions Does the Export To File for Power BI Reports action have any limitations with thin reports or reports connected to a shared semantic model? Are there any permission requirements (Build permission, RLS, semantic model access, etc.) that differ between manual export and Power Automate export? Has anyone experienced blank PDF exports when using a report connected to an existing semantic model while manual exports continue to work? Any guidance would be appreciated. Thank you.47Views1like1CommentHow to design a metadata-driven framework to run 600+ Fabric notebooks with parallel execution?
Hi Fabric Community, We are currently planning to migrate 600+ notebooks from Azure Databricks to Microsoft Fabric. Our current requirement is to build a metadata-driven orchestration framework using: Fabric Data Pipelines for orchestration SQL Database for metadata/control tables Fabric Notebooks for data processing OneLake/Lakehouse as the target storage Instead of creating separate pipelines for each notebook, we would like to have a generic metadata-driven pipeline that dynamically reads notebook information from SQL Database and executes the required notebooks. For example, our metadata table could contain: Notebook ID Notebook name/path Priority Dependency Parameters Data volume Expected runtime Compute/pool requirement Retry count Active/inactive flag Our main challenge We may need to execute approximately 40 notebooks in parallel. We would like to understand the recommended Fabric architecture for this scenario. Specifically: How should we design the metadata-driven pipeline to dynamically execute 600+ Fabric notebooks? How should we control parallelism when around 40 notebooks need to run simultaneously? Should we use a single Spark environment/pool, multiple environments/pools, or some other approach? How should we decide which Spark compute configuration should be used for each notebook based on: Data volume Memory requirement Processing time Shuffle-intensive workloads Small vs. large workloads? If 40 notebooks run simultaneously, how does Fabric manage the underlying Spark compute and capacity? Should we be concerned about resource contention or throttling? Is it recommended to classify notebooks into workload groups such as:and then control concurrency separately for each group? Small / Medium / Large / XLarge What is the recommended way to implement dependency management? For example:where Notebook B should only start after A succeeds. Notebook A → Notebook B → Notebook C What is the recommended approach for retry, failure handling, logging, and restartability for 600+ notebooks? Is there a recommended metadata-driven orchestration pattern/reference architecture in Microsoft Fabric for this scale? Are there any Fabric-specific limitations or best practices we should consider when running hundreds of Spark notebooks with high concurrency? Our main objective is to build a scalable framework where we can manage 600+ notebooks from metadata instead of maintaining hundreds of individual pipelines, while still controlling Spark compute and parallel execution efficiently. Any architectural recommendations, reference implementations, or real-world experience with similar workloads would be highly appreciated. Thanks!119Views1like11CommentsFabric Link Setup
Hello Community, I'm trying to Configure Fabric Link with FnO to sync My FnO data to Fabric. While creating Fabric Link for FnO I'm getting 2 Options as per below SS. One is F&O Entities (Caption 1) Second is F&O tables (Caption 2). I want to Understand How "FnO entities" and "FnO tables" different from each other, In which case I need to choose from second Option(F&O tables Caption 2) & in which case I need to choose from (F&O Entities Caption 1). Thank you43Views0likes2CommentsFail Message Not Working in Fail Activity
Good day I have a lookup activity that returns an output like this: { "count": 6, "value": [ { "SchemaTablename": "Schema.TableName", "QueryText": "Query goes here", "OnPremTableName": "[dbo].[TableName" }, After the Lookup, I have a ForEach Activity that contains 3 activities Lookup: OnPremColumnCount: returns the column count for a table on-premise. example output: { "firstRow": { "OnPremColumnCount": 30 } } Lookup: CloudColumnCount: returns the column count for a table in fabric warehouse. { "firstRow": { "OnPremColumnCount": 30 } } IF Condition: Checks to see if both tables have the same number of columns. If False, there is a fail activity. Now my issue is that I am trying to write an error message that tells me the number of columns on-prem, on cloud, and the on-premise and cloud table name. Here is what I have inside the Fail Message: @concat( 'Columns do not match. On-premise table column count: ', string(activity('OnPremColumnCount').output.firstRow),'.', ' ', 'Cloud table column count: ', string(activity('CloudColumnCount').output.firstRow),'.', string(activity('LookupGetTableList').output.OnPremTableName), string(activity('LookupGetTableList').output.SchemaTablename)) The Error Code is: ActivityFailed. I want it to show an error message like Columns do not match. On-premise table count: 20. Cloud table count: 15. [Schema].[MY_TABLE]. [Schema].[My_Table] I keep getting this error when I run the pipeline. I am not sure why this is happening. I suspect that I am not accessing the output of my previous activities correctly. More so I think it's an issue with accessing the output from the First LookUp Activity (LookupGetTableList). Here is a screenshot of my entire pipeline Your help is greatly appreciated. Thanks, ASolved97Views1like4CommentsHow are teams using AI agents to automate modern data warehouse workflows?
I am exploring how AI agents can help improve data warehouse operations by automating repetitive tasks and assisting data teams with faster decision-making. Some areas I am interested in: Automating data pipeline monitoring and issue detection AI-assisted SQL generation and optimization Automatically identifying data quality issues Triggering workflows based on business events or data changes Generating documentation and insights from warehouse metadata With platforms like Microsoft Fabric Data Warehouse, how are teams approaching AI integration? Are you using: Fabric pipelines with AI-powered automation? Copilot or LLM-based assistants for warehouse development? Custom AI agents connected with data warehouse APIs? Would love to hear real-world architecture patterns and best practices from data engineers working with Fabric.22Views1like1CommentHow are AI agents changing modern data engineering workflows?
I am exploring how AI agents can assist data engineering teams in managing complex data workflows and reducing repetitive operational tasks. Some areas I am interested in: AI agents for monitoring data pipelines and detecting failures Automated data quality checks and anomaly detection Intelligent assistance for ETL/ELT workflow optimization Generating documentation for datasets and transformations Using AI with Fabric pipelines, notebooks, and data workflows With the growth of AI-powered automation, I would like to understand how data engineers are approaching these patterns in real-world environments. What approaches, architectures, or best practices are teams using to combine Microsoft Fabric capabilities with AI agents?31Views0likes2CommentsPreparing Enterprise Data Warehouses for AI-Powered Applications
Hi everyone, I am exploring how organizations are preparing their enterprise data platforms for AI-powered applications. As more teams start building AI assistants and intelligent applications, having clean, structured, and accessible data becomes increasingly important. I would like to understand how the community is approaching this with Microsoft Fabric Data Warehouse. Some questions: - What data modeling approaches work best when preparing warehouse data for AI and analytics workloads? - How are teams balancing traditional BI reporting requirements with new AI use cases? - What strategies are you using for maintaining data quality and governance at scale? - Are there recommended patterns for connecting AI applications with enterprise warehouse data securely? Would love to hear practical experiences and lessons learned from teams working with Fabric Data Warehouse. Thanks!65Views0likes4CommentsFabric Copy Activity fails in Query mode but succeeds in Table mode through an on-premises gateway
Hi Fabric Community, I am using a Microsoft Fabric pipeline Copy Activity to copy data from a Fabric Lakehouse to an on-premises SQL Server database through a dedicated on-premises data gateway. I tested the Copy Activity using both of the following source connection types: Lakehouse connection Lakehouse SQL analytics endpoint connection With both connection types, the behavior is the same: Lookup Activity: Successful Script Activity: Successful Copy Activity using Table mode: Successful Copy Activity using Query mode: Fails When the Copy Activity source is configured with Use query = Query, it fails with the following error: ErrorCode=SqlFailedToConnect,'Type=Microsoft.DataTransfer.Common.Shared.HybridDeliveryException,Message=Cannot connect to SQL Database. Please contact SQL server team for further support. Server: 'xxxxxxxx.datawarehouse.fabric.microsoft.com', Database: 'LH_XYZ', User: ''. Check the connection configuration is correct, and make sure the SQL Database firewall allows the Data Factory runtime to access.,Source=Microsoft.DataTransfer.ClientLibrary,''Type=System.Data.SqlClient.SqlException,Message=A network-related or instance-specific error occurred while establishing a connection to SQL Server. The server was not found or was not accessible. Verify that the instance name is correct and that SQL Server is configured to allow remote connections. (provider: Named Pipes Provider, error: 40 - Could not open a connection to SQL Server),Source=.Net SqlClient Data Provider,SqlErrorNumber=53,Class=20,ErrorCode=-2146232060,State=0,Errors=[{Class=20,Number=53,State=0,Message=A network-related or instance-specific error occurred while establishing a connection to SQL Server. The server was not found or was not accessible. Verify that the instance name is correct and that SQL Server is configured to allow remote connections. (provider: Named Pipes Provider, error: 40 - Could not open a connection to SQL Server),},],''Type=System.ComponentModel.Win32Exception,Message=The network path was not found,Source=,' However, when I change the source configuration from Query to Table and select the source table directly, the Copy Activity completes successfully. Has anyone encountered this behavior? Does Query mode use a different runtime, connector path, or metadata-discovery process when the Copy Activity runs through an on-premises gateway? Are there any known limitations or additional gateway requirements for using a custom query as the source? Any guidance on how to diagnose or resolve this would be appreciated.90Views0likes9Comments