lakehouse
828 TopicsMirrored Metadata Catalog for Databricks - No Data Agent?
Has anyone tested Data Agents in Fabric (NL2SQL)? I'm connecting to a mirrored lakehouse that uses shortcuts to reach data in ADLS. The Data agents rely on the SQL endpoint, and related lakehouse tables. I found a blog that explicitly says this NL2SQL against a databricks catalog is possible. (also Google Gemini says it is possible too) Unlocking LLM-Powered through Data Agent from your Mirrored Databases in Microsoft Fabric | Microsoft Fabric Community However when I try to configure the data agent, and add the lakehouse as a data source, it gives a meaningless error: Couldn't add data source. Try adding the data source again. Has anyone tested metadata mirroring from databricks? These sorts of lakehouses are pretty important stragetic goal, but this experience makes be nervous. Are there any reasons why they should be a lot more buggy than a regular onelake lakehouse?29Views1like4CommentsFabric Copy Activity fails in Query mode but succeeds in Table mode through an on-premises gateway
Hi Fabric Community, I am using a Microsoft Fabric pipeline Copy Activity to copy data from a Fabric Lakehouse to an on-premises SQL Server database through a dedicated on-premises data gateway. I tested the Copy Activity using both of the following source connection types: Lakehouse connection Lakehouse SQL analytics endpoint connection With both connection types, the behavior is the same: Lookup Activity: Successful Script Activity: Successful Copy Activity using Table mode: Successful Copy Activity using Query mode: Fails When the Copy Activity source is configured with Use query = Query, it fails with the following error: ErrorCode=SqlFailedToConnect,'Type=Microsoft.DataTransfer.Common.Shared.HybridDeliveryException,Message=Cannot connect to SQL Database. Please contact SQL server team for further support. Server: 'xxxxxxxx.datawarehouse.fabric.microsoft.com', Database: 'LH_XYZ', User: ''. Check the connection configuration is correct, and make sure the SQL Database firewall allows the Data Factory runtime to access.,Source=Microsoft.DataTransfer.ClientLibrary,''Type=System.Data.SqlClient.SqlException,Message=A network-related or instance-specific error occurred while establishing a connection to SQL Server. The server was not found or was not accessible. Verify that the instance name is correct and that SQL Server is configured to allow remote connections. (provider: Named Pipes Provider, error: 40 - Could not open a connection to SQL Server),Source=.Net SqlClient Data Provider,SqlErrorNumber=53,Class=20,ErrorCode=-2146232060,State=0,Errors=[{Class=20,Number=53,State=0,Message=A network-related or instance-specific error occurred while establishing a connection to SQL Server. The server was not found or was not accessible. Verify that the instance name is correct and that SQL Server is configured to allow remote connections. (provider: Named Pipes Provider, error: 40 - Could not open a connection to SQL Server),},],''Type=System.ComponentModel.Win32Exception,Message=The network path was not found,Source=,' However, when I change the source configuration from Query to Table and select the source table directly, the Copy Activity completes successfully. Has anyone encountered this behavior? Does Query mode use a different runtime, connector path, or metadata-discovery process when the Copy Activity runs through an on-premises gateway? Are there any known limitations or additional gateway requirements for using a custom query as the source? Any guidance on how to diagnose or resolve this would be appreciated.105Views0likes10CommentsSharePoint Excel ingestion, Dataflow Gen2 vs shortcut + notebook CU efficiency
I need to ingest and transform multiple Excel files stored in SharePoint folders. I am comparing: Pure Dataflow Gen2 using SharePoint.Files and Power Query transformations. Lakehouse Files shortcut to the SharePoint folder, followed by all transformations in a Fabric notebook using PySpark. For the same files, transformations, output, schedule, and capacity: Is the shortcut + notebook approach expected to consume fewer Fabric Capacity Units than pure Dataflow Gen2? What is the recommended approach for SharePoint Excel-folder ingestion?126Views1like6CommentsHow are AI agents changing modern data engineering workflows?
I am exploring how AI agents can assist data engineering teams in managing complex data workflows and reducing repetitive operational tasks. Some areas I am interested in: AI agents for monitoring data pipelines and detecting failures Automated data quality checks and anomaly detection Intelligent assistance for ETL/ELT workflow optimization Generating documentation for datasets and transformations Using AI with Fabric pipelines, notebooks, and data workflows With the growth of AI-powered automation, I would like to understand how data engineers are approaching these patterns in real-world environments. What approaches, architectures, or best practices are teams using to combine Microsoft Fabric capabilities with AI agents?41Views0likes3CommentsFabric pipelines audit
Hi community, I am looking for a way to log pipelines details into an audit delta table. To be more specific, I want to log the start time, end time, status, error message if any etc. Having this by activity name would be even better. I was thinking to use a notebook that's being called at the end of each pipeline, so somehow to have only a notebook that can be used in all of my pipelines (because i have a lot of them). Is this something that you were able to build? Any guidance would be very much appreciated. thanks!45Views0likes4CommentsSalesforce data in Fabric with bronze/silver/gold: what would you change?
Hi community, I've recently built a Salesforce analytics setup on Microsoft Fabric with my team at datatobiz and I'd like to hear how you would do differently. Here is the setup: Ingestion: Dataflow Gen2 connects Salesforce to Fabric, with scheduled, incremental loads for objects such as Accounts, Opportunities, Orders, Products, Campaigns and Cases Bronze: raw Salesforce tables land in OneLake with the schema preserved, for auditability and lineage Silver: Fabric and Databricks notebooks handle duplicates, data types, business rules and object relationships Gold: subject-specific marts for Sales Performance, Customer 360 and Operations Pipelines: scheduled refreshes with monitoring and alerts Reporting: Power BI semantic models on star schemas, with DAX measures and role-based views Governance: Azure AD role-based access, Microsoft Purview lineage and Azure Key Vault I'd like your views on: Is Dataflow Gen2 a good fit for incremental Salesforce loads, or would you use Copy activity or notebooks instead? Has anyone used Fabric and Databricks notebooks together in a silver layer? How did it work for lineage and monitoring? What data quality checks do you automate between bronze, silver and gold? Would you change anything in this layering for Salesforce data? Thanks in advance!22Views0likes2CommentsFabric Link Setup
Hello Community, I'm trying to Configure Fabric Link with FnO to sync My FnO data to Fabric. While creating Fabric Link for FnO I'm getting 2 Options as per below SS. One is F&O Entities (Caption 1) Second is F&O tables (Caption 2). I want to Understand How "FnO entities" and "FnO tables" different from each other, In which case I need to choose from second Option(F&O tables Caption 2) & in which case I need to choose from (F&O Entities Caption 1). Thank you54Views0likes3CommentsHow to design a metadata-driven framework to run 600+ Fabric notebooks with parallel execution?
Hi Fabric Community, We are currently planning to migrate 600+ notebooks from Azure Databricks to Microsoft Fabric. Our current requirement is to build a metadata-driven orchestration framework using: Fabric Data Pipelines for orchestration SQL Database for metadata/control tables Fabric Notebooks for data processing OneLake/Lakehouse as the target storage Instead of creating separate pipelines for each notebook, we would like to have a generic metadata-driven pipeline that dynamically reads notebook information from SQL Database and executes the required notebooks. For example, our metadata table could contain: Notebook ID Notebook name/path Priority Dependency Parameters Data volume Expected runtime Compute/pool requirement Retry count Active/inactive flag Our main challenge We may need to execute approximately 40 notebooks in parallel. We would like to understand the recommended Fabric architecture for this scenario. Specifically: How should we design the metadata-driven pipeline to dynamically execute 600+ Fabric notebooks? How should we control parallelism when around 40 notebooks need to run simultaneously? Should we use a single Spark environment/pool, multiple environments/pools, or some other approach? How should we decide which Spark compute configuration should be used for each notebook based on: Data volume Memory requirement Processing time Shuffle-intensive workloads Small vs. large workloads? If 40 notebooks run simultaneously, how does Fabric manage the underlying Spark compute and capacity? Should we be concerned about resource contention or throttling? Is it recommended to classify notebooks into workload groups such as:and then control concurrency separately for each group? Small / Medium / Large / XLarge What is the recommended way to implement dependency management? For example:where Notebook B should only start after A succeeds. Notebook A → Notebook B → Notebook C What is the recommended approach for retry, failure handling, logging, and restartability for 600+ notebooks? Is there a recommended metadata-driven orchestration pattern/reference architecture in Microsoft Fabric for this scale? Are there any Fabric-specific limitations or best practices we should consider when running hundreds of Spark notebooks with high concurrency? Our main objective is to build a scalable framework where we can manage 600+ notebooks from metadata instead of maintaining hundreds of individual pipelines, while still controlling Spark compute and parallel execution efficiently. Any architectural recommendations, reference implementations, or real-world experience with similar workloads would be highly appreciated. Thanks!144Views1like12CommentsDesigning AI Agent Workflows on Modern Data Platforms
Hi everyone, I have been exploring how AI agents can work with modern data engineering platforms to automate business processes. A common architecture I see is that data pipelines collect and prepare information, AI agents analyze the context and determine the next action, workflow services execute tasks through APIs or connected systems, and the results are stored for reporting or further processing. I am interested in how teams are approaching this with Microsoft Fabric. Are you using AI agents directly with Fabric workflows, or keeping the AI layer separate? What patterns work well for connecting AI agents with Lakehouse, Data Pipelines, or notebooks? How do you manage security and permissions when AI applications need access to enterprise data? Are there recommended approaches for combining Fabric workloads with external AI services? I would be interested to hear what architecture patterns others are using and what challenges you have encountered.85Views1like5CommentsTitle: Microsoft Fabric – Impact of Enabling Case-Insensitive Collation Settings
Hi Fabric Community, We are evaluating the impact of enabling Case Insensitive (CI) settings in a Microsoft Fabric workspace. The workspace setting provides the following options: Case Sensitive: Latin1_General_100_BIN2_UTF8 Case Insensitive: Latin1_General_100_CI_AS_KS_WS_SC_UTF8 According to the documentation, these settings affect how SQL processes capitalization in object names and string data. Before enabling/using the Case Insensitive setting in our environment, we would appreciate the community's guidance on the following: Existing objects: If Case Insensitive is enabled at the workspace level, does it affect existing Warehouses or SQL Analytics Endpoints, or does it apply only to newly created objects? Existing data: Does changing the collation setting modify or transform the existing data stored in tables, or does it only change how SQL compares/processes string values? JOIN behavior: Could Case Insensitive collation change the results of existing SQL joins? For example, could the following values start matching when they previously did not? ABC vs abc Assets vs assets Project01 vs PROJECT01 WHERE conditions: Would a Case Insensitive collation change the results of existing filters such as:WHERE Account = 'ABC'when the stored value is abc? GROUP BY / DISTINCT: Could Case Insensitive collation change the results of GROUP BY, DISTINCT, or aggregation queries where values differ only by capitalization? Primary/unique keys: Are there any implications for duplicate detection or uniqueness when values such as ABC and abc exist in columns used as keys? Data pipelines / transformations: Could changing from Case Sensitive to Case Insensitive cause any unexpected behavior in Fabric Data Pipelines, Dataflows, Notebooks, or SQL transformations? Lakehouse vs Warehouse: Does this Case Sensitivity setting affect only Warehouses and SQL Analytics Endpoints, or can it also affect queries against Lakehouse tables? Power BI semantic models: Could changing the SQL collation impact Power BI semantic models, relationships, Direct Lake, Import, or DirectQuery reports? Performance: Is there any measurable performance difference between: Latin1_General_100_BIN2_UTF8 and Latin1_General_100_CI_AS_KS_WS_SC_UTF8 for large tables and joins? Recommended approach: For an enterprise Finance reporting solution where consistency of joins and mappings is critical, is it recommended to use Case Sensitive or Case Insensitive collation? Best practice for mixed requirements: If some business fields require case-insensitive comparisons while others require case-sensitive comparisons, what is the recommended approach in Fabric? Collation at column/query level: If the workspace/database is configured as Case Insensitive, is there a supported way to perform a case-sensitive comparison for specific columns or queries? Changing the setting later: If we initially create our Warehouse/SQL Analytics Endpoint using Case Sensitive collation, can we later change it to Case Insensitive, or would we need to create a new object? Production migration: If changing the collation requires recreating the Warehouse/SQL Analytics Endpoint, what is the recommended migration approach for existing production workloads? We are particularly interested in understanding whether enabling Case Insensitive collation can change query results, especially for existing SQL joins and mappings where source-system values may have differences in capitalization or leading/trailing spaces. Any guidance or real-world experience with Fabric enterprise implementations would be greatly appreciated. Thank you. AI-assisted drafting: AI was used to help structure and phrase this response. I reviewed and validated the technical content before posting.Solved52Views0likes6Comments