data engineering
3769 TopicsDiagnosing a slow Fabric Data Warehouse using the new skill
When a Fabric Data Warehouse slows down, the investigation usually means jumping between the Capacity Metrics app, Query Insights, and SQL pool diagnostics while manually correlating time ranges across three different tools. Microsoft just shipped the SQL DW operations skill as Generally Available, you can find the official announcement here: https://community.fabric.microsoft.com/t5/Fabric-Updates-Blog/Diagnose-Fabric-Data-Warehouse-workloads-with-the-SQL-DW/ba-p/5366102 I dug into what this actually changes in practice: how the skill consolidates capacity spikes, query regressions, and request failures into a single conversational interface, and where it still has blind spots worth knowing before you rely on it in a production incident. There is at least one behavior around time range correlation that surprised me during testing. To do that, I wrote a full article, and would love to hear if others are already using this in their warehouse troubleshooting flow. https://medium.com/@arthurfr23/diagnosing-fabric-data-warehouse-with-the-sql-dw-operations-skill-investigating-workloads-without-e4060fe8661e?sharedUserId=arthurfr2327Views2likes2CommentsFabric Link Setup
Hello Community, I'm trying to Configure Fabric Link with FnO to sync My FnO data to Fabric. While creating Fabric Link for FnO I'm getting 2 Options as per below SS. One is F&O Entities (Caption 1) Second is F&O tables (Caption 2). I want to Understand How "FnO entities" and "FnO tables" different from each other, In which case I need to choose from second Option(F&O tables Caption 2) & in which case I need to choose from (F&O Entities Caption 1). Thank you33Views0likes1CommentHow to design a metadata-driven framework to run 600+ Fabric notebooks with parallel execution?
Hi Fabric Community, We are currently planning to migrate 600+ notebooks from Azure Databricks to Microsoft Fabric. Our current requirement is to build a metadata-driven orchestration framework using: Fabric Data Pipelines for orchestration SQL Database for metadata/control tables Fabric Notebooks for data processing OneLake/Lakehouse as the target storage Instead of creating separate pipelines for each notebook, we would like to have a generic metadata-driven pipeline that dynamically reads notebook information from SQL Database and executes the required notebooks. For example, our metadata table could contain: Notebook ID Notebook name/path Priority Dependency Parameters Data volume Expected runtime Compute/pool requirement Retry count Active/inactive flag Our main challenge We may need to execute approximately 40 notebooks in parallel. We would like to understand the recommended Fabric architecture for this scenario. Specifically: How should we design the metadata-driven pipeline to dynamically execute 600+ Fabric notebooks? How should we control parallelism when around 40 notebooks need to run simultaneously? Should we use a single Spark environment/pool, multiple environments/pools, or some other approach? How should we decide which Spark compute configuration should be used for each notebook based on: Data volume Memory requirement Processing time Shuffle-intensive workloads Small vs. large workloads? If 40 notebooks run simultaneously, how does Fabric manage the underlying Spark compute and capacity? Should we be concerned about resource contention or throttling? Is it recommended to classify notebooks into workload groups such as:and then control concurrency separately for each group? Small / Medium / Large / XLarge What is the recommended way to implement dependency management? For example:where Notebook B should only start after A succeeds. Notebook A → Notebook B → Notebook C What is the recommended approach for retry, failure handling, logging, and restartability for 600+ notebooks? Is there a recommended metadata-driven orchestration pattern/reference architecture in Microsoft Fabric for this scale? Are there any Fabric-specific limitations or best practices we should consider when running hundreds of Spark notebooks with high concurrency? Our main objective is to build a scalable framework where we can manage 600+ notebooks from metadata instead of maintaining hundreds of individual pipelines, while still controlling Spark compute and parallel execution efficiently. Any architectural recommendations, reference implementations, or real-world experience with similar workloads would be highly appreciated. Thanks!76Views1like7CommentsFabric Copy Activity fails in Query mode but succeeds in Table mode through an on-premises gateway
Hi Fabric Community, I am using a Microsoft Fabric pipeline Copy Activity to copy data from a Fabric Lakehouse to an on-premises SQL Server database through a dedicated on-premises data gateway. I tested the Copy Activity using both of the following source connection types: Lakehouse connection Lakehouse SQL analytics endpoint connection With both connection types, the behavior is the same: Lookup Activity: Successful Script Activity: Successful Copy Activity using Table mode: Successful Copy Activity using Query mode: Fails When the Copy Activity source is configured with Use query = Query, it fails with the following error: ErrorCode=SqlFailedToConnect,'Type=Microsoft.DataTransfer.Common.Shared.HybridDeliveryException,Message=Cannot connect to SQL Database. Please contact SQL server team for further support. Server: 'xxxxxxxx.datawarehouse.fabric.microsoft.com', Database: 'LH_XYZ', User: ''. Check the connection configuration is correct, and make sure the SQL Database firewall allows the Data Factory runtime to access.,Source=Microsoft.DataTransfer.ClientLibrary,''Type=System.Data.SqlClient.SqlException,Message=A network-related or instance-specific error occurred while establishing a connection to SQL Server. The server was not found or was not accessible. Verify that the instance name is correct and that SQL Server is configured to allow remote connections. (provider: Named Pipes Provider, error: 40 - Could not open a connection to SQL Server),Source=.Net SqlClient Data Provider,SqlErrorNumber=53,Class=20,ErrorCode=-2146232060,State=0,Errors=[{Class=20,Number=53,State=0,Message=A network-related or instance-specific error occurred while establishing a connection to SQL Server. The server was not found or was not accessible. Verify that the instance name is correct and that SQL Server is configured to allow remote connections. (provider: Named Pipes Provider, error: 40 - Could not open a connection to SQL Server),},],''Type=System.ComponentModel.Win32Exception,Message=The network path was not found,Source=,' However, when I change the source configuration from Query to Table and select the source table directly, the Copy Activity completes successfully. Has anyone encountered this behavior? Does Query mode use a different runtime, connector path, or metadata-discovery process when the Copy Activity runs through an on-premises gateway? Are there any known limitations or additional gateway requirements for using a custom query as the source? Any guidance on how to diagnose or resolve this would be appreciated.87Views0likes9CommentsCopy Job - Amazon S3 connector fails with region signing error on generic endpoint
I'm trying to connect Fabric's Amazon S3 connector (Copy activity, pipelines) to a bucket in ap-southeast-2, using Access Key authentication. When I create a connection with the generic regional endpoint: https://s3.ap-southeast-2.amazonaws.com connecting fails with: Expression.Error: Invalid Url. The authorization header is malformed; the region 'us-east-1' is wrong; expecting 'ap-southeast-2' I've added s3:ListAllMyBuckets, s3:ListBucket, and s3:GetBucketLocation to the IAM user, and the error persists identically.72Views0likes7CommentsCross-tenant Fabric migration (no shared account): 6 things the docs don't tell you
Sharing this rather than asking — posting my notes in case they save someone else the time. I recently had to move semantic models, reports, notebooks, pipelines and a Lakehouse between two Fabric tenants where no single account had access to both — no guest access, no cross-tenant permissions, two entirely separate sign-ins. Deployment pipelines are same-tenant only. fabric-cicd assumes one tenant and a Git repo. So this became an export-to-file / import-from-file exercise against the REST API. Six things cost me real time. Posting them in case they save someone else the same days. A semantic model's connection lives in FOUR places in model.bim I started by rewriting the M expressions and the log said "0 expressions repointed" while the model kept pointing at the old source. The connection can sit in the M code (Sql.Database(...)), in a dataSources[] entry's connectionDetails.address, in Direct Lake entity partitions, and in M parameters with a meta annotation. Rewriting only one of them silently does nothing. I ended up operating on the raw model.bim JSON text so all four are covered. PBIR report binding changed shape between schema versions definition.pbir validates differently depending on its major version. Schema v2.0+ wants connectionString and nothing else — adding pbiServiceModelId gets you "the schema does not allow additional properties". Schema 1.x wants the full legacy set: pbiServiceModelId, pbiModelVirtualServerName ("sobe_wowvirtualserver"), pbiModelDatabaseName, name: "EntityDataSource", connectionType: "pbiServiceXmlaStyleLive". Omit them and you get "Cannot resolve neither report.json nor the PBIR report content in enhanced format". Preserve the original $schema and version from the source file and branch on the major version. Don't assume one shape. The Lakehouse SQL analytics endpoint is read-only over Delta CREATE VIEW, CREATE FUNCTION and CREATE PROCEDURE all work. CREATE TABLE does not — tables only come from Spark. If you're recreating a Lakehouse structure in a target tenant, that's two separate artefacts: a .sql script for the views and a PySpark notebook for the empty table structure. They also have to run in that order. Notebook writes need ?format=ipynb Creating or updating a notebook definition without it returns "PyToIpynbFailure: Convert py to ipynb failed" with no further explanation. Add the query parameter on both createItem and updateDefinition. Only notebooks need it. Delta rejects column names with spaces or punctuation AnalysisException [DELTA_INVALID_CHARACTERS_IN_COLUMN_NAMES] for anything containing space , ; { } ( ) newline tab = The obvious fix is to sanitise the names, but don't — your views and semantic models reference the original names. Enable column mapping on those tables instead: delta.columnMapping.mode = name, minReaderVersion = 2, minWriterVersion = 5. Original names preserved, Delta happy. Premium_ASWL_Error on refresh does not mean you need a gateway Most threads will tell you to set up an on-premises data gateway. In my case the model was set to "Default: Single Sign-On" and simply needed an explicit cloud connection created and bound to it. No gateway involved. Worth checking before you go down the gateway route. Bonus: pipelines fail with a bare "UnknownError" when they reference items that aren't in the target yet. The error never says which ones. I ended up scanning pipeline definitions for GUIDs and matching them against the source inventory to list the missing items by name before attempting the deploy, and deploying master pipelines after their children. Happy to go deeper on any of these if it's useful.33Views0likes1CommentDesigning AI Agent Workflows on Modern Data Platforms
Hi everyone, I have been exploring how AI agents can work with modern data engineering platforms to automate business processes. A common architecture I see is that data pipelines collect and prepare information, AI agents analyze the context and determine the next action, workflow services execute tasks through APIs or connected systems, and the results are stored for reporting or further processing. I am interested in how teams are approaching this with Microsoft Fabric. Are you using AI agents directly with Fabric workflows, or keeping the AI layer separate? What patterns work well for connecting AI agents with Lakehouse, Data Pipelines, or notebooks? How do you manage security and permissions when AI applications need access to enterprise data? Are there recommended approaches for combining Fabric workloads with external AI services? I would be interested to hear what architecture patterns others are using and what challenges you have encountered.50Views0likes3CommentsCapacity Metrics schema validation failed for all supported versions in FUAM
Hello, I am facing an issue while collecting data from the Microsoft Fabric Capacity Metrics App using the FUAM notebook code. The schema compatibility validation fails for every version checked by the notebook: INFO: Test for v53 failed INFO: Test for v47 failed INFO: Test for v40 failed INFO: Test for v37 failed The notebook then returns the following exception: ERROR: Capacity Metrics data structure is not compatible or connection to capacity metrics is not possible. Could someone please help clarify: Whether the Capacity Metrics App schema has recently changed. Whether the current FUAM release supports the latest Capacity Metrics App. Whether a newer validation query or updated FUAM notebook is available. Thank you.34Views0likes2CommentsFabric REST API endpoint for dataflow transactions should GET *sortable* results for a time-range
Dataflow refresh APIs seem to be the buggiest; the list of transactions on a GET for Fabric-dataflow, are not sorted by the latest. OData query options such as $top, $filter, $orderby do not work with this specific endpoint. Endpoint accepts the parameter but ignores it. Endpoint implements its own pagination mechanism. If GET ../transactions returns 10 rows, I can't assume it will have the latest refresh in those 10. This has been confirmed by running at different times of the day.6Views0likes0Comments