dataflow
4637 TopicsOne of my Dataflow Gen 1 stopped working
Hello everyone, I have a few DataFlow Gen1, all of them are working except one. "TauxHorairePanierMoyen" worked perfectly for months, and this morning I can't update the file. It shows this kind of error : Error: Request ID: 82fd0734-bf12-4ff2-b85d-79e2981cb408 Activity ID: 0b7f17c0-b8c4-4f96-beef-896adb689b04 When I go to Power Query it loads perfectly, but when I update, it doesn't work. --> It's a combined Excel Files. Requested on Nom du flux de données État d'actualisation du flux de données Nom de la table Nom de partition Actualiser l’état Heure de début Heure de fin Durée Lignes traitées Octets traités (Ko) Validation maximale (Ko) Temps processeur Temps d’attente Moteur de calcul Erreur 30/09/2026 11:29 TauxHorairePanierMoyen Échec TauxHorairePanierMoyen Non disponible Échec 30/09/2026 11:29 30/09/2026 11:35 00:06:00.7300 Non disponible Non disponible Non disponible Non disponible Non disponible Non disponible Error: Request ID: 82fd0734-bf12-4ff2-b85d-79e2981cb408 Activity ID: 0b7f17c0-b8c4-4f96-beef-896adb689b04 My Power Query code : Source = SharePoint.Files("XXX", [ApiVersion = 15]), #"Lignes filtrées" = Table.SelectRows(Source, each Text.Contains([Name], "TauxHorairePanierMoyen")), #"Fichiers masqués filtrés" = Table.SelectRows(#"Lignes filtrées", each [Attributes]?[Hidden]? <> true), #"Appeler une fonction personnalisée" = Table.AddColumn(#"Fichiers masqués filtrés", "Transformer le fichier", each #"Transformer le fichier"([Content])), #"Colonnes renommées" = Table.RenameColumns(#"Appeler une fonction personnalisée", {{"Name", "Source.Name"}}), #"Autres colonnes supprimées" = Table.SelectColumns(#"Colonnes renommées", {"Source.Name", "Transformer le fichier"}), #"Colonne de table développée" = Table.ExpandTableColumn(#"Autres colonnes supprimées", "Transformer le fichier", Table.ColumnNames(#"Transformer le fichier"(#"Exemple de fichier"))), #"Type de colonne changé" = Table.TransformColumnTypes(#"Colonne de table développée", {{"Nom Affaire", type text}, {"ORV", type text}, {"Nature ORV", type text}, {"Code équipe", type text}, {"Libellé équipe", type text}, {"N° client", type text}, {"Nom client", type text}, {"Nature (PGC)", type text}, {"Type ligne (PR/MO/DIV)", type text}, {"Date facture", type date}, {"No facture", Int64.Type}, {"Libellé réceptionnaire", type text}, {"Référence", type text}, {"Désignation", type text}, {"Quantité servie", type number}, {"Montant net", type number}, {"Montant brut", type number}, {"Montant remise", type number}, {"Montant remise CBQ", Int64.Type}, {"Remise client (Code)", type text}, {"Libellé remise CBQ", type text}, {"Famille comptable Reconstitué (Code)", type text}}) in #"Type de colonne changé" File Exemple : let Source = SharePoint.Files("XXX", [ApiVersion = 15]), #"Lignes filtrées" = Table.SelectRows(Source, each Text.Contains([Name], "TauxHorairePanierMoyen")), #"Fichiers masqués filtrés" = Table.SelectRows(#"Lignes filtrées", each [Attributes]?[Hidden]? <> true), Navigation = #"Fichiers masqués filtrés"{0}[Content] in Navigation Parameter let Source = Excel.Workbook(Paramètre, null, true), Navigation = Source{[Item = "TauxHoraire&PanierMoyen", Kind = "Sheet"]}[Data], #"En-têtes promus" = Table.PromoteHeaders(Navigation, [PromoteAllScalars = true]) in #"En-têtes promus" Transform File Exemple : let Paramètre = #"Exemple de fichier" meta [IsParameterQuery = true, IsParameterQueryRequired = false, Type = type binary, BinaryIdentifier = #"Exemple de fichier"] in Paramètre Thanks for your help. Kinds regards, Aude-Marie33Views0likes2CommentsPower BI Report Usage Metrics – Workspace-Level and Organization-Level Usage
Power BI Report Usage Metrics – Workspace-Level and Organization-Level Usage Hi everyone, I am working on a Power BI governance requirement around report usage metrics, and I am trying to understand the recommended approach at both the workspace level and organization/tenant level. I have some initial understanding of the available options, but I would like to get confirmation from the community and understand if there are better/native approaches. Workspace-Level Report Usage If I have multiple reports within a single workspace, I would like to see the usage metrics for ALL reports in that workspace in one place. For example: Workspace A Report 1 → Views / Unique Viewers / Last Viewed Report 2 → Views / Unique Viewers / Last Viewed Report 3 → Views / Unique Viewers / Last Viewed Instead of opening Usage Metrics separately for each report. My understanding is that the Usage Metrics semantic model can potentially be used to create a custom usage report for the workspace. However, I would like to confirm: How exactly can we get usage metrics for all reports within a single workspace? Is there a supported way to remove/modify the report-level filtering so that the Usage Metrics semantic model covers all reports in that workspace? Is this the recommended approach for workspace-level usage reporting? Are there any limitations with this approach? Organization-Level Report Usage The next requirement is to go beyond a single workspace. I would like to build a centralized view such as: Organization → Workspace → Report → Views → Unique Viewers → Last Viewed → Other Usage Metrics For example: Organization Workspace A Report 1 Report 2 Report 3 Workspace B Report 4 Report 5 Workspace C Report 6 Report 7 All of this should ideally be available in one centralized Power BI governance report. For this requirement, I know there are some possible approaches, such as: Using Power BI/Fabric Admin Portal capabilities with the required admin access Using Power BI Admin APIs Using Power BI Activity Log APIs Building a custom data collection process and storing the information in a Fabric Lakehouse/Warehouse/SQL database But my questions are: Are these the only practical/supported ways to achieve organization-wide report usage? Is there any native Power BI/Fabric feature that already provides tenant-level usage metrics across all workspaces? Can the standard Usage Metrics semantic model be consolidated across multiple workspaces? Is there another Microsoft-supported API or Fabric capability that is better suited for this requirement? What permissions/admin roles are actually required for the different approaches? Workspace-Level vs Organization-Level I am particularly interested in understanding whether the recommended architecture is different for these two scenarios: Scenario A: Single Workspace → All Reports → Usage Metrics Scenario B: Entire Organization → All Workspaces → All Reports → Usage Metrics Would we use the Usage Metrics semantic model for Scenario A and Admin APIs/Activity Logs or another centralized source for Scenario B? Or is there a common approach that can handle both? Historical Usage Another concern is data retention. My understanding is that the current Usage Metrics experience has a limited historical window, with the new Usage Metrics experience retaining 30 days of data. If we need to retain usage information for: 6 months 1 year Multiple years what is the recommended approach? Would this be the appropriate architecture? Power BI / Fabric Usage Data → API / Activity Log / Other Source → Fabric Lakehouse / Warehouse → Historical Usage Table → Centralized Governance Dashboard Or is there a Microsoft-native approach that can retain this historical information without building our own storage layer? Final Architecture Question Ultimately, I am trying to determine the recommended approach for: Organization ↓ Workspace ↓ Report ↓ Usage Metrics ↓ Historical Usage I would appreciate guidance on: How to achieve workspace-level usage for all reports How to achieve organization-wide usage across all workspaces Whether Admin Portal/API/Activity Log approaches are the only options Whether there is a native Fabric/Power BI capability I may be missing How to handle the 30-day Usage Metrics retention limitation Recommended architecture for long-term Power BI governance and usage tracking Thanks in advance to anyone who can share their experience or recommended approach.80Views1like4CommentsBest Practices for Handling Incremental Data Loads in Dataflow Gen2
Hi everyone, I am exploring different approaches for handling incremental data loads with Dataflow Gen2 in Microsoft Fabric. For larger datasets, refreshing the entire dataset every time can become inefficient, so I am interested in understanding how others are designing their dataflows to process only new or changed records. A few questions: How are you identifying and filtering changed records between dataflow runs? Is it better to manage incremental logic directly inside Dataflow Gen2, or use a Fabric pipeline to control the process? How do you handle failed runs or partially processed data without creating duplicate records? Are there any recommended patterns for maintaining good performance as the volume of historical data grows? I would appreciate hearing about approaches that have worked well in real Fabric environments.56Views1like6CommentsFabric Data Agent Access issue
I have created a fabric data agent using semantic model of a published report, It is working fine but I'm not able to share it with others I published the agent in M365 copilot and share with my teams of 100 people only 4 or 5 people can go to M365 copilot and search the agent and added the agent to their agent tab others are not even seeing the agent but if I share the agent link from M365 copilot then they see the agent but agent is not answering the questions but it's perfectly answering for me and other 4 to 5 people who see the agent when I surfed it said it said there are 2 level of access to get answer from the agent 1. agent level access 2. underlying data level access I have workspace contributor access for workspace, read, write, reshare permission for the agent. most of them share the same access but there are not able to see the only difference I found is I'm having premium per user license and most of other have pro license so I doubted it but internet says license is not an issue can anyone pls help me with this issue21Views0likes0CommentsF64 Import Model Refresh Fails Due to Memory: Can Data Duplication Be Avoided
Hello, We are experiencing refresh failures for a large Power BI Import semantic model on an F64 Microsoft Fabric capacity. The refresh fails because the operation exceeds the available memory. Our main objective is to resolve the refresh failure. We are also trying to understand how the model is compressed and which objects consume the most storage, because reducing the model size may help reduce the refresh-time memory requirement. Current Environment The semantic model contains three fact tables and multiple dimension tables. The data source is an Azure Synapse Analytics dedicated SQL pool. The semantic model and reports run on an F64 Microsoft Fabric capacity. The semantic model uses Import storage mode. The semantic model size is approximately 14.8 GB. We import only the columns required for reporting, relationships, calculations, and security. One fact table contains approximately 3.8 billion rows. This large fact table is divided into partitions of approximately 100 million rows using XMLA. We refresh only the partitions that contain changed data. The other two fact tables have been denormalized by adding only the required descriptive attributes that were previously obtained from dimension tables. We denormalized those two fact tables because some table visuals were failing with the following error when using the normalized design: Query exceeded available resources. The denormalized design improved the table visual behaviour, but it increased the semantic model size. Main issue: Refresh failure due to memory usage Even though we refresh only the affected partitions of the 3.8-billion-row fact table, the semantic model refresh is failing because of the memory constraint on the F64 capacity. We would like to understand how memory is used during an Import semantic model refresh. Our understanding is that Power BI needs to preserve the currently queryable version of the model or partition while it processes a new version. This could temporarily require memory for both existing and newly processed data, along with additional memory for compression, dictionaries, relationships, indexes, and transaction commit operations. Could you please clarify the exact refresh behaviour? Specifically: Does Power BI temporarily maintain two copies of the complete semantic model, or only two versions of the partition being refreshed? Which model structures may be duplicated or rebuilt during a partition refresh? How much additional working memory is normally required beyond the stored semantic model size? Could a 14.8 GB semantic model exceed the F64 memory limit during refresh because the existing model, refreshed partition, dictionaries, relationship indexes, and processing workspace are held in memory at the same time? Does refreshing one partition at a time materially reduce peak memory, or can model-level structures still cause high memory consumption? Can refresh-time model duplication be disabled? Is there any supported way to prevent or reduce the apparent duplication of data during refresh? For example, can an Import refresh directly modify or update the existing data in a partition instead of building a separate replacement version and switching to it after processing completes? We would like to know whether any of the following is possible: Update existing compressed data in place. Append new rows directly to an existing processed partition. Delete or modify individual rows without recreating the affected partition. Disable the refresh transaction or shadow-copy behaviour. Disable the retention of the old partition version during processing. Commit the refresh in smaller stages to reduce peak memory. Release the old partition from memory before loading the replacement. Process only model metadata, relationships, or calculations without creating another data copy. Configure a lower-memory refresh mode for large Import semantic models. Control the number of rows or segments processed in each refresh transaction. If in-place updates or disabling refresh-time duplication are not supported, what is the recommended approach for refreshing a model of this size on F64? We understand that transactional refresh behaviour may be required to keep the existing semantic model available to report users and to allow rollback if processing fails. However, we would like confirmation of whether this behaviour can be changed or optimized. VertiPaq dictionary and compression question We are also investigating which columns contribute most to the semantic model size. Consider the following simplified example: LargeFact[EntityIdentifier] SecondFact[EntityIdentifier] EntityDimension[EntityIdentifier] All three columns use the same data type and contain many of the same identifier values. Does VertiPaq create: one shared dictionary for the identifier across the complete semantic model, one dictionary per table, one independent dictionary per physical column, or separate dictionaries or encoding structures at the partition or segment level? Does the relationship between the fact and dimension tables allow the key dictionaries to be reused, or does every physical column maintain its own dictionary and encoded value storage? For example, suppose the fact table containing 3.8 billion rows has a foreign-key column referencing another table. If that foreign-key column contains substantially fewer distinct values than the total row count, is its storage broadly composed of: a dictionary containing the distinct values, encoded references for the 3.8 billion rows, column segments, relationship indexes, hierarchy structures, and other internal storage objects? We would like to understand whether the storage is primarily caused by: the dictionary, the encoded values across 3.8 billion rows, internal relationship structures, partition-related structures, or a combination of these. Storage query used I used the following DAX query to identify the columns and internal storage objects consuming the most space: // Check largest storage entity EVALUATE VAR SegmentSizes = GROUPBY ( INFO.STORAGETABLECOLUMNSEGMENTS(), [TABLE_ID], [COLUMN_ID], "Records", SUMX ( CURRENTGROUP(), [RECORDS_COUNT] ), "SegmentUsedBytes", SUMX ( CURRENTGROUP(), [USED_SIZE] ), "SegmentAllocatedBytes", SUMX ( CURRENTGROUP(), [ALLOCATED_SIZE] ) ) VAR ColumnDetails = SELECTCOLUMNS ( INFO.STORAGETABLECOLUMNS(), "TABLE_ID", [TABLE_ID], "COLUMN_ID", [COLUMN_ID], "Table", [DIMENSION_NAME], "Column", [ATTRIBUTE_NAME], "ColumnType", [COLUMN_TYPE], "DataType", [DATATYPE], "Encoding", [COLUMN_ENCODING], "DictionaryBytes", [DICTIONARY_SIZE] ) VAR Combined = NATURALLEFTOUTERJOIN ( SegmentSizes, ColumnDetails ) RETURN SELECTCOLUMNS ( Combined, "Table", [Table], "Column or object", [Column], "Object type", [ColumnType], "Data type", [DataType], "Encoding", [Encoding], "Records", [Records], "Segment MB", DIVIDE ( [SegmentUsedBytes], 1024 * 1024 ), "Dictionary MB", DIVIDE ( [DictionaryBytes], 1024 * 1024 ), "Approximate total MB", DIVIDE ( [SegmentUsedBytes] + COALESCE ( [DictionaryBytes], 0 ), 1024 * 1024 ), "Allocated MB", DIVIDE ( [SegmentAllocatedBytes], 1024 * 1024 ) ) ORDER BY [Approximate total MB] DESC The query showed identifier columns from multiple tables among the largest storage objects. The most significant result was a foreign-key column in the 3.8-billion-row fact table. The query reported that this column, or its related internal storage object, was consuming up to approximately 13 GB. Could you please confirm whether this query correctly estimates storage usage by column or internal storage object? Main questions Our primary questions are: How can we resolve the refresh failure caused by the F64 memory constraint? Does Import refresh temporarily duplicate the complete model, or only the affected partition and related structures? Is there any supported way to update an existing Import partition in place? Can refresh-time duplication or transactional processing be disabled? Is there a supported configuration that reduces peak refresh memory? Is the storage query above correct, particularly the approximately 13 GB reported for a foreign-key object? Are VertiPaq dictionaries maintained per column, per table, per partition, or per model? What architecture would Microsoft recommend for an Import semantic model containing a 3.8-billion-row fact table on F64? Our immediate objective is to complete the refresh successfully. Model-size optimization is important mainly because it may reduce the peak memory required during processing. At the same time, we need to avoid reintroducing the 'Query exceeded available resources.' error in the report table visuals. Thank you.38Views2likes2CommentsFrom Chaos to Clarity: How the Manufacturing Dashboard Helps Everyone Understand the Business
Let's walk through what each page actually shows, and why it matters. Page 1: Customer Insights — "Who are we selling to, and are they happy?" Every business lives or dies by its customers, but it's surprisingly hard to keep track of hundreds of relationships in your head. This page acts like a customer relationship "scoreboard." It shows things like: * How many active customers the business currently has * Which customers are considered healthy (buying regularly, paying on time) versus at risk (going quiet, showing warning signs) * How revenue is spread across different customers and regions — are we relying too heavily on just a few big accounts? Think of it like a doctor's checkup, but for relationships instead of health. If ten customers suddenly go from "healthy" to "at risk," that's a signal to act before they walk away for good — not after the revenue has already dropped and everyone's asking why. Page 2: Production & Supply — "Are we making enough, and can we deliver it?" This is the page for anyone who cares about what's actually happening on the factory floor and getting product out the door. It's less about money and more about operations. Key questions it answers: * How much did we produce this month, and is that on target? * Are shipments going out on time, or are we falling behind on delivery promises? * Is one particular plant underperforming compared to the others? If you're a plant manager, this is likely the first thing you check every morning — before coffee, even. It tells you at a glance whether today is a "business as usual" day or a "we need to fix something now" day. Page 3: Overview — "How's the whole business doing, in 30 seconds?" This is the page built for someone with almost no time to spare — a CEO, an investor, or anyone who just needs the headline numbers without digging through details. At the top sit eight simple cards, each showing one important number and whether it went up or down compared to last month: * Total Revenue — how much money came in * Total Production — how much product was made * Capacity Utilization — how much of the factories' potential is actually being used * Gross Profit — how much money is left after production costs * Supply Fulfillment — how reliably orders are being delivered * Inventory Value — how much stock is sitting in the warehouse * Active Customers — how many customers are currently buying * Overall Operational Efficiency — a single score summarizing how smoothly everything is running Below that, charts break revenue and production down by month, by plant, and by region — so if a number drops, you don't just see that it dropped, you can immediately see where. Was it one plant having a bad month, or a slowdown across an entire region? There's also a simple "Alerts & Insights" section that puts the numbers into plain words — things like "Supply on track: fulfillment is 94% and improving" — so nobody has to guess what a chart is trying to tell them. Page 4: Inventory & Working Capital — "Do we have too much stock, or too little?" This page tackles a balancing act every manufacturing business faces. Keep too much inventory sitting around, and you're tying up cash and warehouse space that could be used elsewhere. Keep too little, and you risk running out of product right when a customer needs it — losing sales and trust. This page shows: * Total Inventory Value — how much money is currently tied up in stock * Inventory Turnover — how quickly that stock is being sold and replaced (a higher number generally means things are moving efficiently) * Days Inventory Outstanding — roughly how many days' worth of stock is sitting around unused * Raw Material vs. Finished Goods Stock — how much is still waiting to be turned into product versus ready to ship * Slow-Moving Inventory — stock that isn't selling and may need attention * Stockout Risk — a warning flag for items at risk of running out * Inventory Accuracy — how well the recorded stock counts match what's physically in the warehouse There's even a detailed table listing specific materials by name, plant, and status (like "Slow-Moving" or "Healthy"), so instead of a vague warning, someone gets a precise to-do list of exactly what needs a closer look. Why This Approach Works for Everyone The real magic of this dashboard isn't any single chart — it's the consistency. Every page follows the same basic structure: 1. Filters at the top (date range, region, plant, product category) so anyone can narrow the view to exactly what matters to them 2. A handful of key numbers, shown as simple cards with an up or down arrow — no complicated formulas to interpret 3. A plain-language "Alerts & Insights" section that explains, in a sentence or two, what changed and why it matters That means a machine operator, a plant manager, a customer success rep, and a CEO can all open the same dashboard and immediately find what's relevant to them — no translation needed, no waiting for someone else to "run the numbers." In a business with as many moving parts as manufacturing, that kind of shared clarity isn't a luxury. It's what keeps everyone — from the shop floor to the boardroom — pointed in the same direction.31Views0likes0CommentsCheck the incremental refresh
Hello, How can I verify that the incremental refresh has been successfully applied after a dataset refresh? Are there any logs, indicators, or recommended methods to confirm that the refresh was processed incrementally rather than as a full refresh? Thank you.Solved51Views0likes3CommentsLakehouse vs. Warehouse in Microsoft Fabric: Which One Should You Actually Use?
If you're new to Microsoft Fabric, you've probably hit this wall already: you go to create an item, and Fabric hands you two very similar-sounding options - Lakehouse and Warehouse. Both store tabular data in Delta format in OneLake and provide SQL access, but they offer different development and transactional experiences. Both let you query with SQL.Fabric presents both as analytical data-store options, and their capabilities overlap enough that the choice is not always obvious. So which one do you pick? Short answer: it depends on who's writing the queries and what shape your data is in. Long answer: keep reading — and try the quick self-check below before you scroll to the recommendation. 60-Second Self-Check Answer these three questions honestly before you create your next Fabric item: 1. Who will write the transformation logic? - A) Python/Spark notebooks, data engineers comfortable with PySpark - B) SQL analysts and BI developers who live in T-SQL 2. What does your source data look like? - A) Mixed — JSON, Parquet, CSV, streaming events, semi-structured - B) Mostly clean, structured, relational-shaped data 3. What's the end consumption pattern? - A) A mix of ML, notebooks, ad-hoc exploration, and reporting - B) Primarily Power BI reports and governed semantic models Mostly A's → Lean Lakehouse. Mostly B's → Lean Warehouse. Mixed bag? → You're not alone — see the "Can I use both?" section below. Lakehouse: The Flexible One A Lakehouse stores data as files (Delta Parquet) in OneLake, and gives you two ways to work with it: - Notebooks (PySpark, Spark SQL) for engineers who want full control - A SQL analytics endpoint that auto-generates so SQL folks can still query the same tables Pick Lakehouse when: - Your data arrives messy, semi-structured, or in large volumes that benefit from Spark's distributed processing - Your team already thinks in notebooks and data science workflows - You want schema flexibility - Delta tables support schema enforcement and controlled schema evolution, offering flexibility while preserving data reliability. - You're building a medallion architecture (bronze → silver → gold) and need engineering muscle at the bronze/silver layers Watch out for: - The SQL endpoint is read-only - The SQL analytics endpoint is read-only for Lakehouse table data, so it does not support INSERT, UPDATE, or DELETE against those Delta tables. Modify or load the data through Spark or another supported ingestion and transformation experience. You can still create supported SQL objects such as views, functions, and stored procedures in the endpoint. - Fabric provides automatic Delta table optimizations, but advanced workloads may still benefit from deliberate file sizing, data layout, partitioning, or optimization strategies. (OPTIMIZE, V-Order, partitioning) python #Typical Lakehouse bronze-to-silver pattern in a notebook df = spark.read.format("json").load("Files/raw/events/") df_clean = df.dropDuplicates().withColumn("load_date", current_date()) df_clean.write.format("delta").mode("overwrite").saveAsTable("silver_events") Python from pyspark.sql.functions import current_date() df = spark.read.format("json").load("Files/raw/events/") df_clean = df.dropDuplicates().withColumn("load_date", current_date()) df_clean.write.format("delta").mode("overwrite") saveAsTable("silver_events") Warehouse: The Familiar One Fabric Warehouse provides a rich T-SQL-first data warehousing experience. — think of it as a cloud data warehouse that happens to store data in OneLake under the hood. Pick Warehouse when: - Your team's primary skill is T-SQL, not Spark/Python - You need full DML (`INSERT`, `UPDATE`, `DELETE`, `MERGE`) with transactional guarantees - You're modeling a governed, relational structure — star schemas, stored procedures, views - You want a more traditional data-warehouse development experience (cross-database queries, You want a familiar relational warehouse development experience with T-SQL, views, stored procedures, and support for compatible SQL tools.) Watch out for: - Less flexible with wildly semi-structured or streaming-first data — you'll typically land that in a Lakehouse first, then move it in Warehouse is designed primarily for T-SQL rather than Spark-native development. If Spark is central to your transformation logic, Lakehouse is usually the more natural starting point. sql -- Typical Warehouse transformation pattern MERGE INTO dbo.DimCustomer AS target USING staging.CustomerUpdates AS source ON target.CustomerID = source.CustomerID WHEN MATCHED THEN UPDATE SET target.Email = source.Email WHEN NOT MATCHED THEN INSERT (CustomerID, Email) VALUES (source.CustomerID, source.Email); Can I Use Both? Yes - and honestly, A supported and commonly discussed architecture is to use a Lakehouse for engineering-oriented layers and a Warehouse for a curated relational serving layer. do exactly this: - Lakehouse for bronze/silver ingestion and heavy transformation (Spark does the messy work) - Warehouse for the polished gold layer that analysts and Power BI consume with familiar T-SQL Both use OneLake, and Fabric supports patterns such as shortcuts and cross-database queries that can reduce or avoid unnecessary data duplication. If you physically load curated data into separate Warehouse tables, however, that creates another stored representation. Try It Yourself Before your next project kickoff, run this checklist with your team: [ ] Who owns the transformation code - engineers or SQL analysts? [ ] Does the source data need Spark-level flexibility, or is it already relational? [ ] Do you need multi-table transactions and T-SQL DML, or can table changes be implemented through Spark and Delta operations? [ ] Could a hybrid (Lakehouse → Warehouse) actually be the real answer? No data to test with yet? A quick way to practice both patterns above is to grab a free sample dataset (or a ready-made dashboard layout to reverse-engineer) from a site like Docynx. Over to you I'd love to hear how your team decided: Did you go Lakehouse, Warehouse, or both? What tipped the decision — team skill set, data shape, or something else entirely? Drop your setup in the comments — especially if you've got a "we picked wrong and had to migrate" story, those are always the most useful ones.73Views0likes0CommentsBackward Deployment
Hello ! I have a deployment pipeline in this order : A (Development) -> B (Preproduction) -> C (Production). This is the standard order of the workspace, but there were some reports that have been created directly in B, so is it possible to deploy from B to A? without unassigning the worskpace from the actual pipeline. Thank you.102Views0likes6Comments