cicd
475 TopicsFabric Warehouse Git commit fails with “Maximum commit size exceeded” potentially due to XMLA.json
Based on the Fabric Warehouse Git integration documentation, the XMLA.json file should be excluded from Git integration workflows. Fabric should exclude this file from commits and updates to prevent the default semantic model metadata from being unintentionally stored in Git. [learn.microsoft.com] However, it appears that XMLA.json, or the associated semantic model metadata, may still be getting processed on the Fabric side while calculating or preparing the commit. This potentially causes the commit to exceed the maximum allowed size. Any One Face Same Issue, how to avoid above issue & prevent XMLA.json file from getting processes while git commit ?15Views0likes3CommentsGit Commit Issue
Hello Fabric Community, I have a workspace that is connected to my GitHub repo. I used a personal access token to integrate GitHub with Fabric. Initially, I created the token for only one day and committed my items to GitHub. When I connected my GitHub repo again with a new token, I was unable to commit the changes that had been made so far because I could not find any items pending to commit in the source control pane "Changes" tab. PFA. Please suggest a solution or let me know the proper way to integrate my changes into GitHub again. Thanks in advance for any insights or guidance!43Views0likes4CommentsGit exception should be more specific
I was unable to pull updates from git, with the below screenshot: But it's hard to verify which code paragraph is having issues, which made me hard to identify the issue. would that be possible to improve the error message to something like: Thanks!33Views1like3CommentsHow to implement CI/CD for Lakehouse Delta table schemas and SQL Analytics Endpoint
Hi Microsoft Fabric Community, We have set up an automated CI/CD pipeline using Azure DevOps Repos, Git integration, and Microsoft's fabric-cicd Python library. Our current workflow is: 1. Feature workspaces are connected to Git feature branches. 2. A Pull Request merges changes into the dev branch. 3. Our Azure DevOps pipeline deploys Fabric items across target workspaces (dev → test → prod). The current behavior is: Control-plane items such as Notebooks, Data Pipelines, and Copy Jobs deploy and parameterize successfully. The Lakehouse item itself and its OneLake shortcut definitions also deploy as expected. However, Delta table schemas under /Tables and SQL Analytics Endpoint objects such as Views, Stored Procedures, and TVFs are not captured or deployed. Git integration appears to track the Lakehouse metadata but not the underlying Delta table structures or SQL Analytics Endpoint DDL. We have the following questions for the product team and community: Question 1 - Lakehouse Delta Table Schemas How can we implement automated CI/CD for Delta table creation and schema evolution across environments without requiring manual execution? Question 2 - SQL Analytics Endpoint Objects Since DACPAC / SqlPackage does not support the read-only SQL Analytics Endpoint, what is the recommended approach for implementing CI/CD for Views, Stored Procedures, and TVFs using Azure DevOps? Question 3 - Fabric Roadmap Are there plans on the Microsoft Fabric roadmap for native Git integration to track and synchronize SQL Analytics Endpoint objects directly alongside other workspace items? Any reference architectures, repository templates, recommended tools, or best practices for implementing this type of CI/CD workflow would be greatly appreciated. Thank you!46Views0likes2CommentsUnable to Configure Deployment Rules for a Direct Lake Semantic Model in Fabric
Hi Fabric Community, I’m facing an issue with Fabric Deployment Pipelines and would appreciate some guidance. I have a Power BI semantic model using Direct Lake mode, and I’m deploying it between two Microsoft Fabric workspaces. I am an Admin in both the source and target workspaces. However, when I open the Semantic model deployment rules, both Data source rules and Parameter rules are disabled/grayed out, as shown in the screenshots. I also confirmed that the tables in the semantic model are using Direct Lake storage mode using the TABLETRaits() DAX query. My questions are: Why are Data source rules and Parameter rules disabled for this semantic model? Is there any additional configuration required to enable these deployment rules? If deployment rules are not supported for this scenario, what is the recommended approach for changing the data source between environments? Any advice or explanation would be greatly appreciated. Thanks in advance!52Views0likes4CommentsCDC vs Watermark: Choosing an Incremental Load Strategy in Microsoft Fabric
Why incremental loading matters If your source deletes rows or you need history, use CDC; if you only need inserts and updates and have a reliable "last modified" column, a watermark is simpler and works almost everywhere. That is the short answer. The rest of this post explains why. Reloading a whole table every night is fine at 50,000 rows. At 500 million it burns capacity units, hammers the source system, and stretches your refresh window. Incremental loading fixes this by moving only what changed since the last run. In Microsoft Fabric, the main tool for this is the Copy job in Data Factory. It supports both change detection approaches side by side, and can pick one per table. Knowing how each works helps you choose well and avoid silent data gaps. The watermark pattern A watermark load copies only rows whose "incremental column" value is higher than the highest value seen in the last successful run. The first run copies everything; each later run picks up where the last one stopped. The query behind it is simple: SELECT * FROM dbo.Orders WHERE ModifiedAt > @LastWatermark AND ModifiedAt <= @CurrentWatermark; In a classic pipeline you build this yourself: a Lookup for the old watermark, a Copy activity, and a step that saves the new value to a control table. Copy job removes that plumbing. It stores the watermark state for you, and a failed run resumes from the last successful one without data loss (Microsoft Learn). Copy job accepts these column types as a watermark: ROWVERSION – changes automatically on every insert or update; the most reliable choice on SQL Server-family sources. Datetime – columns such as ModifiedAt or LastUpdatedDatetime. Date – date-only columns; Copy job re-reads the last day to avoid gaps. String – values that can be read as datetimes. Integer – an increasing number. A column of another type can still work if you cast it to a supported type in the source query, as long as its values compare in order. File sources use the file's last-modified time as the watermark. The catch: a watermark only sees rows that still exist and whose column actually moved. Deleted rows are invisible, and rows with a NULL watermark are skipped on every incremental run after the first. Change data capture (CDC) CDC reads the source database's own change log instead of querying the table, so it captures inserts, updates and deletes exactly as they happened. You don't pick an incremental column; the database tells Fabric what changed. When a source table has CDC enabled, Copy job detects it, does an initial full load, and then replicates only changes on later runs (Microsoft Learn). Sources that support CDC in Copy job include Azure SQL Database, Azure SQL Managed Instance, on-premises SQL Server, Oracle, Snowflake, Google BigQuery, SAP Datasphere Outbound, and Fabric Lakehouse tables (via Delta Change Data Feed). You then choose how changes land in the destination: SCD Type 1 (Merge) – the default. The destination mirrors the current source: updates overwrite, deletes remove the row. SCD Type 2 – in preview. Copy job keeps every version, adding Valid_From, Valid_To and Is_Current columns. Deletes become soft deletes, so you keep a full audit trail with no custom code. Oracle sources don't support it yet. Copy job isn't the only CDC route in Fabric. Mirroring keeps a near real-time replica of an operational database in OneLake with almost no setup, and Eventstreams offers CDC connectors for streaming changes into Real-Time Intelligence. Microsoft's data movement decision guide compares the three. Side by side The deciding row is usually deletes: only CDC sees them. Consideration Watermark CDC Source prerequisite A column that always increases on insert or update CDC enabled on the source and supported by the connector Inserts Yes Yes Updates Only if the watermark column changes Yes Deletes No Yes History (SCD Type 2) No, unless you build it Built in (preview) Write methods Append or Merge Merge or SCD Type 2 Load on source Range scan on the watermark column each run Reads the change log; lighter on busy tables Setup effort Low; no DBA changes Needs CDC enabled, permissions, log retention Works with files Yes (last-modified time) No Source coverage Most database connectors A defined list of connectors Source: Incremental copy in Copy job, Microsoft Learn. Which one should you use? Three questions settle it for almost every table. If CDC is the right answer but you can't enable it on the source, fall back to a watermark and catch deletes with soft-delete flags or a periodic key comparison. Tips and pitfalls For watermark loads Prefer ROWVERSION over app timestamps. An application can forget to update ModifiedAt; the database never forgets a rowversion. Watch for NULLs. Rows inserted later with a NULL watermark are never picked up. Make the column NOT NULL with a default. Catch deletes another way. Use soft deletes (an IsDeleted flag that also bumps the watermark), or run a periodic key comparison against the source. Use Merge, not Append, for updates. Append writes a second copy of every updated row. For CDC loads Don't mix table types in one job. If a Copy job holds both CDC-enabled and non-CDC tables, it falls back to watermark for all of them. Split them into separate jobs. Size log retention. If the job is paused longer than the source keeps its change data, you will need a full reload. Know the current limits. Copy job captures net changes only (not every intermediate change) and supports only the default capture instance. For Lakehouse sources, enable Change Data Feed yourself; Copy job can't detect it (Microsoft Learn). For both Reset carefully. Resetting incremental state forces a full read on the next run but doesn't clear the destination. With Append, truncate the target first or you'll get duplicates. Plan for schema changes. New source columns aren't synced automatically, and a dropped column that's in your column mapping fails the run. Conclusion Watermark and CDC aren't rivals; most Fabric estates use both. Use a watermark for append-heavy fact tables, files, and sources you can't change. Use CDC wherever deletes or history matter, which is most dimension and master-data tables. Copy job makes the choice cheap to revisit: it detects CDC-enabled tables, handles state for you, and lets you reset per table. Start with the table that hurts most in your nightly full load, pick the method from the guide above, and measure the capacity you save. Sources Incremental copy in Copy job – Microsoft Learn Change data capture (CDC) in Copy job – Microsoft Learn Choose a data movement strategy – Microsoft Learn Copy job enhancements on incremental copy and CDC – Microsoft Fabric BlogDP-600 Voucher
a { text-decoration: none; color: #464feb; } tr th, tr td { border: 1px solid #e6e6e6; } tr th { background-color: #f5f5f5; } Hi everyone, I recently passed the DP-700 Microsoft Fabric Data Engineer Associate certification and am now planning to take DP-600 (Microsoft Certified: Fabric Analytics Engineer Associate) while the Fabric concepts are still fresh in my mind. Unfortunately, I wasn't able to use a previous DP-600 voucher when I first received it as I wasn't ready for the exam at the time. I was wondering if anyone is aware of any current voucher opportunities for DP-600, or has an unused voucher that they won't be using and is transferable (if permitted by the voucher terms). Any guidance would be greatly appreciated. Thanks in advance! Regards, Amanul56Views1like2CommentsMicrosoft Fabric Materialized Lake Views (MLVs) – Looking for Confirmation on a Few Points
Microsoft Fabric Materialized Lake Views (MLVs) – Looking for Confirmation on a Few Points Hi everyone, I'm currently working with Materialized Lake Views (MLVs) in Microsoft Fabric and trying to establish the right approach for implementing them in a proper Dev → UAT → Production setup. I have gone through the available documentation, but I would appreciate confirmation from people who are already using MLVs in real projects. MLV Refresh Failure and Monitoring My current understanding is that if an MLV fails during its initial execution or during a scheduled refresh: The MLV run will be marked as Failed. The failure can be investigated through the MLV run history/monitoring experience. The error code, error message and execution details can be used for troubleshooting. If an upstream MLV fails, dependent MLVs may be skipped. Questions: Is there a recommended/native way to send an email or Teams notification whenever an MLV refresh fails? What is the recommended production monitoring approach for MLV refresh failures? Is anything else required in addition to configuring the MLV schedule? ON MISMATCH FAIL vs ON MISMATCH DROP My current understanding is that ON MISMATCH is related to handling schema/structure mismatches encountered during MLV processing, rather than being a general data-quality validation mechanism. For example: ON MISMATCH FAIL would cause the MLV operation to fail when the applicable mismatch occurs, whereas: ON MISMATCH DROP would allow processing to continue while dropping the affected/incompatible data. Questions: Is this understanding correct? What exactly constitutes a "mismatch" for an MLV? What are the recommended scenarios for using FAIL vs DROP? Does ON MISMATCH FAIL affect the entire MLV refresh or only the affected data/operation? Dev → UAT → Production Deployment This is probably my biggest area of uncertainty. Suppose I create and test the following in DEV: DEV Workspace Lakehouse Source tables MLV 1 MLV 2 My expectation is that I should not manually recreate the MLVs in UAT and Production. Instead, I would use Fabric lifecycle/deployment capabilities to promote the artifacts: DEV → UAT → PROD Questions: Are MLV definitions currently supported for this type of Dev → UAT → Prod deployment? Does Deployment Pipeline automatically handle the corresponding Lakehouse/item dependencies? If the workspace names, Lakehouse names and item IDs are different between environments, how are those references handled? Does Fabric automatically rebind supported dependencies to the corresponding target-environment resources? Which MLV-related properties/configurations need to be manually configured or validated after deployment? For example: DEV Workspace: BI_DEV Lakehouse: Sales_LH_DEV ↓ UAT Workspace: BI_UAT Lakehouse: Sales_LH_UAT ↓ PROD Workspace: BI_PROD Lakehouse: Sales_LH_PROD How much of this environment-specific mapping is automatically handled by Fabric? MLV Refresh Schedules Across Environments Suppose my DEV MLV has a refresh schedule of every 1 hour. When I deploy it to UAT and Production: Is the schedule also deployed? If it is, can or should it be different between environments? What is the recommended approach for managing MLV schedules separately in DEV, UAT and PROD? Is schedule configuration considered part of the deployable MLV artifact or environment-specific configuration? MLV Dependencies / MLV on top of MLV Can we create an MLV that references another MLV? For example: Source Tables → MLV 1 – Silver → MLV 2 – Gold → MLV 3 – Aggregation Is this a supported and recommended architecture? If yes: Are there any limitations on the number or depth of MLV dependencies? How does refresh orchestration work? If MLV 1 fails, are MLV 2 and MLV 3 automatically skipped? How is the dependency and lineage handled during deployment from DEV → UAT → PROD? SQL View on top of MLV Can we create a regular SQL View on top of an MLV? For example: Source Tables → Materialized Lake View → SQL View → Semantic Model / Power BI If this is supported: Are there any performance implications? Does the SQL View remain virtual while the MLV remains materialized? Is this a recommended pattern for exposing MLV data to downstream consumers? View → MLV Conversely, can an MLV be created on top of a regular SQL View? For example: Source Tables → SQL View → MLV If supported, are there any restrictions around the types of views or queries that can be used as an MLV source? Recommended Production Architecture For a production implementation, would the following architecture be considered a reasonable pattern? Source Tables → MLV Layer → Silver MLVs → Gold MLVs → SQL Views (if required) → Semantic Model → Power BI Or is it generally better to consume MLVs directly from the semantic model where possible? General Best Practices For anyone using MLVs in production: What are the main limitations you have encountered? What should be considered when designing MLV dependencies? How do you handle monitoring and alerting? How do you manage Dev/UAT/Prod deployment? Are there any MLV-specific CI/CD considerations that are easy to miss? What would you recommend avoiding when designing an MLV architecture? I'd particularly appreciate confirmation from anyone who has implemented MLVs across multiple Fabric environments such as DEV, UAT and PROD. Thanks in advance!Solved104Views2likes5CommentsFabric Data Pipeline – Does Office 365 Outlook Activity Support Service Principal Authentication?
Hi everyone, I’m looking for confirmation from anyone who has successfully implemented **Service Principal authentication with the Office 365 Outlook activity in a Microsoft Fabric Data Pipeline**. We currently use the **Office 365 Outlook activity**, authenticated with a personal user's OAuth login. This causes problems with CI/CD because the connection does not reliably carry over between workspaces/environments. We’re therefore looking at moving to **Service Principal authentication**, which now appears as an option in the connector. We have created an App Registration with: * Microsoft Graph * `Mail.Send` – Application permission Before granting admin consent, I’d like to clarify a few things. **1. Does Service Principal authentication actually work end-to-end with the Office 365 Outlook activity?** Has anyone successfully configured this in Fabric and used it across multiple workspaces/environments through CI/CD? There are previous discussions suggesting that although Service Principal authentication appears in the UI, it may not actually work with the Outlook activity. https://community.fabric.microsoft.com/discussions/ac_dataengineering/how-to-setup-service-principal-for-microsoft-outlook-365-connector/5189483 If you have this working, it would be great to know: * What configuration/permissions you used * Whether admin consent was required * Whether the connection survives deployment between workspaces * Whether the Outlook activity successfully sends emails using the Service Principal **2. Can the `Mail.Send` permission be restricted?** `Mail.Send` as an Application permission is quite broad. Is there a supported way to restrict the Service Principal to specific mailbox(es), for example using **Exchange Online RBAC for Applications**? **3. Is Microsoft Graph via an HTTP/Web activity a better option?** If Service Principal authentication with the built-in Outlook activity is unreliable, would calling Microsoft Graph `sendMail` through an HTTP/Web activity be a more reliable and least-privilege approach for CI/CD? If anyone has this working in **production or across multiple Fabric environments**, I’d really appreciate details of the configuration and any limitations you encountered.Thanks in advance.Solved91Views2likes6Comments