dataflow
43 TopicsFrom Chaos to Clarity: How the Manufacturing Dashboard Helps Everyone Understand the Business
Let's walk through what each page actually shows, and why it matters. Page 1: Customer Insights — "Who are we selling to, and are they happy?" Every business lives or dies by its customers, but it's surprisingly hard to keep track of hundreds of relationships in your head. This page acts like a customer relationship "scoreboard." It shows things like: * How many active customers the business currently has * Which customers are considered healthy (buying regularly, paying on time) versus at risk (going quiet, showing warning signs) * How revenue is spread across different customers and regions — are we relying too heavily on just a few big accounts? Think of it like a doctor's checkup, but for relationships instead of health. If ten customers suddenly go from "healthy" to "at risk," that's a signal to act before they walk away for good — not after the revenue has already dropped and everyone's asking why. Page 2: Production & Supply — "Are we making enough, and can we deliver it?" This is the page for anyone who cares about what's actually happening on the factory floor and getting product out the door. It's less about money and more about operations. Key questions it answers: * How much did we produce this month, and is that on target? * Are shipments going out on time, or are we falling behind on delivery promises? * Is one particular plant underperforming compared to the others? If you're a plant manager, this is likely the first thing you check every morning — before coffee, even. It tells you at a glance whether today is a "business as usual" day or a "we need to fix something now" day. Page 3: Overview — "How's the whole business doing, in 30 seconds?" This is the page built for someone with almost no time to spare — a CEO, an investor, or anyone who just needs the headline numbers without digging through details. At the top sit eight simple cards, each showing one important number and whether it went up or down compared to last month: * Total Revenue — how much money came in * Total Production — how much product was made * Capacity Utilization — how much of the factories' potential is actually being used * Gross Profit — how much money is left after production costs * Supply Fulfillment — how reliably orders are being delivered * Inventory Value — how much stock is sitting in the warehouse * Active Customers — how many customers are currently buying * Overall Operational Efficiency — a single score summarizing how smoothly everything is running Below that, charts break revenue and production down by month, by plant, and by region — so if a number drops, you don't just see that it dropped, you can immediately see where. Was it one plant having a bad month, or a slowdown across an entire region? There's also a simple "Alerts & Insights" section that puts the numbers into plain words — things like "Supply on track: fulfillment is 94% and improving" — so nobody has to guess what a chart is trying to tell them. Page 4: Inventory & Working Capital — "Do we have too much stock, or too little?" This page tackles a balancing act every manufacturing business faces. Keep too much inventory sitting around, and you're tying up cash and warehouse space that could be used elsewhere. Keep too little, and you risk running out of product right when a customer needs it — losing sales and trust. This page shows: * Total Inventory Value — how much money is currently tied up in stock * Inventory Turnover — how quickly that stock is being sold and replaced (a higher number generally means things are moving efficiently) * Days Inventory Outstanding — roughly how many days' worth of stock is sitting around unused * Raw Material vs. Finished Goods Stock — how much is still waiting to be turned into product versus ready to ship * Slow-Moving Inventory — stock that isn't selling and may need attention * Stockout Risk — a warning flag for items at risk of running out * Inventory Accuracy — how well the recorded stock counts match what's physically in the warehouse There's even a detailed table listing specific materials by name, plant, and status (like "Slow-Moving" or "Healthy"), so instead of a vague warning, someone gets a precise to-do list of exactly what needs a closer look. Why This Approach Works for Everyone The real magic of this dashboard isn't any single chart — it's the consistency. Every page follows the same basic structure: 1. Filters at the top (date range, region, plant, product category) so anyone can narrow the view to exactly what matters to them 2. A handful of key numbers, shown as simple cards with an up or down arrow — no complicated formulas to interpret 3. A plain-language "Alerts & Insights" section that explains, in a sentence or two, what changed and why it matters That means a machine operator, a plant manager, a customer success rep, and a CEO can all open the same dashboard and immediately find what's relevant to them — no translation needed, no waiting for someone else to "run the numbers." In a business with as many moving parts as manufacturing, that kind of shared clarity isn't a luxury. It's what keeps everyone — from the shop floor to the boardroom — pointed in the same direction.36Views0likes0CommentsLakehouse vs. Warehouse in Microsoft Fabric: Which One Should You Actually Use?
If you're new to Microsoft Fabric, you've probably hit this wall already: you go to create an item, and Fabric hands you two very similar-sounding options - Lakehouse and Warehouse. Both store tabular data in Delta format in OneLake and provide SQL access, but they offer different development and transactional experiences. Both let you query with SQL.Fabric presents both as analytical data-store options, and their capabilities overlap enough that the choice is not always obvious. So which one do you pick? Short answer: it depends on who's writing the queries and what shape your data is in. Long answer: keep reading — and try the quick self-check below before you scroll to the recommendation. 60-Second Self-Check Answer these three questions honestly before you create your next Fabric item: 1. Who will write the transformation logic? - A) Python/Spark notebooks, data engineers comfortable with PySpark - B) SQL analysts and BI developers who live in T-SQL 2. What does your source data look like? - A) Mixed — JSON, Parquet, CSV, streaming events, semi-structured - B) Mostly clean, structured, relational-shaped data 3. What's the end consumption pattern? - A) A mix of ML, notebooks, ad-hoc exploration, and reporting - B) Primarily Power BI reports and governed semantic models Mostly A's → Lean Lakehouse. Mostly B's → Lean Warehouse. Mixed bag? → You're not alone — see the "Can I use both?" section below. Lakehouse: The Flexible One A Lakehouse stores data as files (Delta Parquet) in OneLake, and gives you two ways to work with it: - Notebooks (PySpark, Spark SQL) for engineers who want full control - A SQL analytics endpoint that auto-generates so SQL folks can still query the same tables Pick Lakehouse when: - Your data arrives messy, semi-structured, or in large volumes that benefit from Spark's distributed processing - Your team already thinks in notebooks and data science workflows - You want schema flexibility - Delta tables support schema enforcement and controlled schema evolution, offering flexibility while preserving data reliability. - You're building a medallion architecture (bronze → silver → gold) and need engineering muscle at the bronze/silver layers Watch out for: - The SQL endpoint is read-only - The SQL analytics endpoint is read-only for Lakehouse table data, so it does not support INSERT, UPDATE, or DELETE against those Delta tables. Modify or load the data through Spark or another supported ingestion and transformation experience. You can still create supported SQL objects such as views, functions, and stored procedures in the endpoint. - Fabric provides automatic Delta table optimizations, but advanced workloads may still benefit from deliberate file sizing, data layout, partitioning, or optimization strategies. (OPTIMIZE, V-Order, partitioning) python #Typical Lakehouse bronze-to-silver pattern in a notebook df = spark.read.format("json").load("Files/raw/events/") df_clean = df.dropDuplicates().withColumn("load_date", current_date()) df_clean.write.format("delta").mode("overwrite").saveAsTable("silver_events") Python from pyspark.sql.functions import current_date() df = spark.read.format("json").load("Files/raw/events/") df_clean = df.dropDuplicates().withColumn("load_date", current_date()) df_clean.write.format("delta").mode("overwrite") saveAsTable("silver_events") Warehouse: The Familiar One Fabric Warehouse provides a rich T-SQL-first data warehousing experience. — think of it as a cloud data warehouse that happens to store data in OneLake under the hood. Pick Warehouse when: - Your team's primary skill is T-SQL, not Spark/Python - You need full DML (`INSERT`, `UPDATE`, `DELETE`, `MERGE`) with transactional guarantees - You're modeling a governed, relational structure — star schemas, stored procedures, views - You want a more traditional data-warehouse development experience (cross-database queries, You want a familiar relational warehouse development experience with T-SQL, views, stored procedures, and support for compatible SQL tools.) Watch out for: - Less flexible with wildly semi-structured or streaming-first data — you'll typically land that in a Lakehouse first, then move it in Warehouse is designed primarily for T-SQL rather than Spark-native development. If Spark is central to your transformation logic, Lakehouse is usually the more natural starting point. sql -- Typical Warehouse transformation pattern MERGE INTO dbo.DimCustomer AS target USING staging.CustomerUpdates AS source ON target.CustomerID = source.CustomerID WHEN MATCHED THEN UPDATE SET target.Email = source.Email WHEN NOT MATCHED THEN INSERT (CustomerID, Email) VALUES (source.CustomerID, source.Email); Can I Use Both? Yes - and honestly, A supported and commonly discussed architecture is to use a Lakehouse for engineering-oriented layers and a Warehouse for a curated relational serving layer. do exactly this: - Lakehouse for bronze/silver ingestion and heavy transformation (Spark does the messy work) - Warehouse for the polished gold layer that analysts and Power BI consume with familiar T-SQL Both use OneLake, and Fabric supports patterns such as shortcuts and cross-database queries that can reduce or avoid unnecessary data duplication. If you physically load curated data into separate Warehouse tables, however, that creates another stored representation. Try It Yourself Before your next project kickoff, run this checklist with your team: [ ] Who owns the transformation code - engineers or SQL analysts? [ ] Does the source data need Spark-level flexibility, or is it already relational? [ ] Do you need multi-table transactions and T-SQL DML, or can table changes be implemented through Spark and Delta operations? [ ] Could a hybrid (Lakehouse → Warehouse) actually be the real answer? No data to test with yet? A quick way to practice both patterns above is to grab a free sample dataset (or a ready-made dashboard layout to reverse-engineer) from a site like Docynx. Over to you I'd love to hear how your team decided: Did you go Lakehouse, Warehouse, or both? What tipped the decision — team skill set, data shape, or something else entirely? Drop your setup in the comments — especially if you've got a "we picked wrong and had to migrate" story, those are always the most useful ones.82Views0likes0CommentsDataflow Gen2 vs Copy Job vs Pipeline: Choosing by Workload, Not by Habit
Every Fabric team has a default tool. People who came from Azure Data Factory build a pipeline for everything. Power BI people open Dataflow Gen2 for everything. Newer teams might put every table into a Copy Job because the wizard is quick. Each tool does its own job very well. Problems start when it gets used for another tool's job: pipelines full of hand-built watermark logic, dataflows used only to copy tables, and Copy Jobs expected to handle transformations they were never built for. This post gives you a simple way to pick the right tool based on what the workload needs. One sentence each Copy Job moves data from A to B, including incremental loads, with as little setup as possible. Dataflow Gen2 shapes data. It cleans, merges, reshapes and applies business rules using Power Query. Pipeline coordinates work. It runs steps in order, handles dependencies and failures, and ties everything together. If you remember only one thing, remember this: Copy Job moves, Dataflow shapes, Pipeline orchestrates. Start with three questions about the workload Ask these before you open any editor. Am I moving data or changing it? If the data arrives at the destination looking mostly like the source, with only column mapping or type changes, you are moving data. If you are joining, deduplicating, deriving columns or applying business logic, you are transforming it. How does the source change? A one-time or full reload is a different problem from an ongoing incremental sync. Incremental loads based on change data capture (CDC) are different again. How many steps depend on each other? A single load on a schedule is one thing. "Load these, then validate, then transform, then run a stored procedure, then alert someone if it fails" is a workflow. Your answers usually point clearly to one tool. Copy Job: when the job is data movement Copy Job is the newest of the three and the one most often overlooked by teams stuck in old habits. Microsoft's decision guide lists its main scenarios as incremental copy and replication (both watermark-based and native CDC), data lake and storage migration, medallion ingestion, and out-of-the-box multi-table copy. Microsoft Learn The incremental support is the main reason to use it. In a pipeline, the Copy activity handles incremental copy through pipeline expressions and control tables, and only with watermarks. That means you build and maintain the control table, the lookup, the parameterized query and the watermark update yourself. Copy Job does this for you. Microsoft Learn Choose Copy Job when: You need to ingest many tables from a database into a lakehouse or warehouse. You want an initial full load followed by incremental updates, and you don't want to build watermark logic. The source supports CDC and you want inserts and updates merged into the destination automatically. The destination is outside Fabric. Copy Job supports 40+ destination connectors. Microsoft Learn Microsoft's own example fits this well: an analyst who needs multi-table selection across regional SQL Server instances, a bulk initial load, and then CDC-based incremental merges picks Copy Job because it supports both watermark-based and native CDC incremental copying through a wizard, and automatically detects CDC-enabled tables. Microsoft Learn Don't choose Copy Job when: You need real transformation. Its transformation support is rated low, and that is intentional. Microsoft Learn The load is one step in a larger workflow with conditions, retries and downstream dependencies. That belongs in a pipeline. Dataflow Gen2: when the job is shaping data Dataflow Gen2 is Power Query running at Fabric scale. It is the right tool when the value lies in the transformation logic. It offers 170+ built-in connectors, 300+ transformation functions in a visual interface, and data profiling tools for checking data quality. Microsoft Learn A common complaint is that dataflows are slow for large volumes. That used to be a fair criticism, but Fabric has added several performance features aimed at specific workloads: Fast Copy is for direct, high-throughput copies from a supported source with no transformations. It uses the same backend as the pipeline Copy activity. microsoftmicrosoft Modern Evaluator helps when you are shaping data from connectors that don't fold, or only partly fold. microsoft Partitioned Compute is for large, partitioned or multi-file datasets that can be processed in parallel. microsoft Staging lets you land raw data first and transform it afterwards (ELT), so ingestion and transformation don't compete in one pass. microsoft Choose Dataflow Gen2 when: Business logic is the main work: cleansing, standardizing codes, merging sources, calculated columns. The people who own the logic know Power Query and would struggle to maintain Spark or SQL. You are combining files, APIs, SharePoint lists and databases into one clean dataset. You want visual data profiling while you build. Don't choose Dataflow Gen2 when: You are only copying tables. Fast Copy makes this workable, but Copy Job is simpler, gives you incremental loads without extra work, and has less to maintain. Your destination isn't supported. Dataflow Gen2 lists around 7+ destination connectors, compared with 40+ for the copy tools. Microsoft Learn The transformations are very complex or code-heavy. At that point a notebook is usually the better choice. Pipeline: when the job is coordination A pipeline is not mainly a data movement tool. It is the orchestrator. Microsoft describes it as low-code orchestration that groups several activities together to complete a task. The Copy activity inside it is powerful, and it remains a strong option for very large migrations. The guide describes it as the best low-code choice for moving petabytes of data into lakehouses and warehouses, either ad hoc or on a schedule. Microsoft LearnMicrosoft Learn The real reason to use a pipeline is the control flow: dependencies, branching, retries, failure handling, parameters, and calling other items. Choose a pipeline when: Several steps must run in a set order, where step B only runs if step A succeeds. You need logic such as If/Else, ForEach over a metadata list, or waiting on an external event. You are combining different item types, such as a Copy Job, a dataflow, a notebook, a stored procedure and a web call. You need error handling and notifications around the whole process. You are doing a large, custom migration where you want detailed control over the Copy activity. Microsoft's scenario for pipelines describes a workflow that runs stored procedures, calls web APIs, moves files and executes other pipelines. That is orchestration, not simple ingestion. Microsoft Learn Don't choose a pipeline when: It would contain a single activity that just runs one dataflow on a schedule. Dataflows can be scheduled on their own. You would be rebuilding incremental loading by hand when Copy Job already supports your source. Side-by-side Copy Job Dataflow Gen2 Pipeline Main job Move and replicate Transform and shape Orchestrate Incremental loads Built in (watermark and CDC) Possible, but not its strength Manual (watermark and control tables) Transformation Low High None itself (calls other items) Destinations 40+ ~7+ Depends on activities Authoring Wizard Power Query Visual canvas and expressions Best owner Data integrator, analyst Analyst, data engineer Data engineer Warning sign you picked wrong Adding transformation workarounds Dataflow has no transformation steps Pipeline has one activity Common mistakes that come from habit The "everything is a pipeline" team. Every source table gets a Lookup, a ForEach, a parameterized Copy activity and a stored procedure to update the watermark. It works, but you now maintain a small custom framework that Copy Job gives you ready-made. Keep the pipeline and let it call the simpler pieces. The "everything is a dataflow" team. Dataflows with forty queries that just select a table and load it. Move the raw ingestion to Copy Job and keep dataflows for the layer where logic actually happens. The "Copy Job does it all" team. Trying to handle business rules with column mappings, then adding SQL views downstream to fix what should have been transformed properly. Once logic appears, add a Dataflow Gen2 or a notebook. The single-activity pipeline. A pipeline whose only purpose is to run one item on a schedule adds a layer to monitor without adding control. Use the item's own schedule until you really need dependencies. The pattern that usually wins: use all three, each for its own job For a typical medallion architecture, the tools fit together naturally: Bronze with Copy Job. Ingest source tables with built-in incremental or CDC loads. No custom watermark logic. Silver and Gold with Dataflow Gen2 (or notebooks for heavy, code-first logic). Apply cleansing, conformance and business rules where they are visible and easy to maintain. Pipeline around everything. Run ingestion, then transformation only if ingestion succeeded, then refresh or post-processing steps, and send an alert if anything fails. Each tool does what it is best at, and each is simpler because it isn't doing another tool's job. A 30-second decision checklist No transformation, and data must stay in sync over time? → Copy Job One-off or very large custom migration needing fine control? → Copy activity in a pipeline Transformation logic is the main work, and owners know Power Query? → Dataflow Gen2 Complex, code-first transformation at scale? → Notebook Multiple steps, dependencies, branching or error handling? → Pipeline, calling the tools above Closing thought The right question isn't "which tool do we use?" It is "what does this workload need?" Movement, shaping and coordination are three different problems, and Fabric gives you a dedicated tool for each. Teams that choose by workload build less custom plumbing, find problems faster, and hand solutions over more easily, because each piece does one clear job. Next time you start a new load, answer the three questions first, then pick the tool.73Views1like0CommentsOrchestration Tool Selection in Microsoft Fabric Data Factory
Why Tool Choice Is an Engineering Decision Picking the wrong orchestration primitive is a debt that compounds silently. A team that builds everything in notebooks because 'that's what we know' will eventually own fifty interdependent notebooks with hard-coded paths, bespoke retry loops, and monitoring gaps — maintained by two people. The failure mode is not technical; it is organisational. When the business analyst can't touch the transformation and the ops team can't centralise the logs, delivery velocity collapses. Microsoft Fabric Data Factory ships three distinct orchestration primitives — Pipelines, Dataflow Gen2, and Notebooks — because no single tool is optimal across all workloads. This article helps you map each tool to the scenario it was built for, with concrete code and configuration examples. Tool Selection Matrix Use the matrix below as your starting point. Real solutions usually combine more than one tool; treat each row as guidance for an individual task within a broader workflow. Decision Flowchart Step through the questions below in order, stopping at the first 'yes'. Most real pipelines combine several tools, so run the flowchart once per task — not once per project. Pipelines — Orchestration and Control Flow Pipelines are the scheduler and coordinator. They do not transform data; they determine when, in what order, and under what error conditions other activities run. When to choose a Pipeline Scheduling: Trigger-based execution (tumbling window, schedule, storage event). Fan-out: ForEach iterates over a dynamic list — e.g., 10 source schemas loaded in parallel. Error handling: Until activity with configurable back-off; conditional branching on activity outcome. Centralised monitoring: Run history, duration, and failure reason are surfaced in the Fabric monitoring hub — no custom logging code required. Dataflow Gen2 — No-Code Transformations Dataflow Gen2 exposes the Power Query engine behind a visual interface. Any engineer who has used Excel Power Query or Power BI can build and maintain transformations without writing a single line of Python or SQL. When to choose Dataflow Gen2 Analyst ownership: Transformations that business analysts or BI developers will modify post-deployment. Moderate complexity: Filters, joins, type coercions, aggregations, pivots — anything expressible in the Power Query formula language (M). Connector breadth: 300+ connectors out of the box; no custom connector code needed. Small-to-medium data: Power Query engine handles millions of rows comfortably; hand off to Spark Notebooks for billions. Notebooks — Programmatic Transformations Spark Notebooks are the right tool when the transformation logic exceeds what a visual canvas can express, when you need ML libraries, or when data volume demands distributed processing. When to choose a Notebook Complex logic: Nested conditionals, custom scoring functions, recursive lookups. Machine learning: scikit-learn, MLflow, Spark MLlib — all available in the Fabric Spark runtime. Big data: Billions of rows — Spark partitions the work across the cluster automatically. Developer ownership: Engineers maintain the code; version it in Git like any other source artefact. Common Integration Patterns Production-grade data platforms rarely use just one tool. The patterns below represent the most common compositions and the scenarios they address. Pattern 1 — Schedule → Dataflow → Warehouse A schedule trigger fires a Pipeline that runs a Dataflow Gen2 activity to clean and join source data, then writes the result directly to a Warehouse table. Operations monitors via Pipeline run history; analysts modify the Dataflow independently. Use case: Daily sales consolidation — analysts own the transformation, engineers own the schedule. Pattern 2 — Pipeline → Notebook → Email Notification A Pipeline runs a Notebook activity that executes complex PySpark logic (e.g., anomaly detection), then passes the output path to a Web activity that calls a Logic Apps endpoint to send an alert email. Use case: Nightly data quality checks that trigger ops alerts when anomaly thresholds are breached. Pattern 3 — Dataflow seeds Lakehouse, Notebook reads for ML A Dataflow cleans raw customer records and writes to Delta tables. A separate Notebook reads those tables to train a churn prediction model, logging the run to MLflow. The two artefacts are decoupled — analysts iterate on data prep, data scientists iterate on the model. Use case: Productionising an ML workflow without coupling data engineering to data science release cycles. Pattern 4 — End-to-End ETL Pipeline A single Pipeline chains four activities: (1) Copy Data extracts files from an external SFTP; (2) a Dataflow Gen2 activity standardises column names and types; (3) a Notebook applies business rules and ML scoring; (4) a second Copy Data loads results to the Warehouse; (5) a Web activity posts a completion notification to Teams. Performance Considerations Different tools have different performance envelopes. Matching data volume to the right engine avoids both over-engineering and under-provisioning. Pipeline — Copy Data activity The Copy Data activity is optimised for high-throughput bulk transfers. Increase Data Integration Units to scale throughput horizontally. Enable staging through Azure Blob Storage for transfers between incompatible source and destination connection types. This activity performs best on large file copies and full database extracts where no row-level transformation is required. Dataflow Gen2 The Power Query engine handles datasets comfortably in the millions-of-rows range. Enable staging in the Dataflow settings pane to off-load compute from the on-premises data gateway and execute transformations in the cloud. For datasets larger than this threshold, the latency and memory characteristics of the Power Query engine make a Spark Notebook the better choice. Notebooks — Spark Spark Notebooks distribute computation across a cluster and are the appropriate engine for datasets in the billions-of-rows range. Tune executor core and memory configuration in the Spark pool settings. Apply Delta Lake Z-ORDER clustering on frequently filtered columns to reduce the volume of data scanned per query. Use broadcast joins for small dimension tables to avoid expensive shuffle operations across the cluster. Key Takeaways for Data Engineers Getting tool selection right from the start avoids costly rewrites and knowledge silos later. Three principles guide sound architectural decisions in Data Factory: 1. Match the tool to the person who will maintain it — not just the person who builds it. A Dataflow maintained by an analyst and a Notebook maintained by an engineer are both valid choices; a Notebook that only two engineers understand is a liability. 2. Pipelines orchestrate; they do not transform — resist embedding transformation logic in pipeline expressions or Copy activity mapping columns. Keep orchestration and transformation concerns in separate artefacts. 3. Composition beats monoliths — a Pipeline that orchestrates a Dataflow and a Notebook is easier to debug, test, hand over, and evolve than a 500-line PySpark Notebook that handles extraction, transformation, loading, and notification in a single script. The architectural principle to internalise: Pipeline orchestrates, Dataflow transforms visually, Notebook transforms programmatically. Maintain this separation of concerns consistently and the architecture will scale with your team and your data volumes.93Views0likes0CommentsBuilding Your First Pipeline in Microsoft Fabric
Data rarely lives where you need it. More often than not, the first real challenge of any analytics project is simply getting information from one place to another, reliably and on a schedule. This is exactly where Microsoft Fabric pipelines shine. If you have worked with Azure Data Factory before, much of this will feel familiar, but Fabric brings everything together inside a single, unified workspace and adds a few welcome conveniences along the way. In this article, we will walk through creating your very first Fabric pipeline from the ground up. We will start by setting up the pipeline and exploring the activities available to us, then move on to scheduling and monitoring. Finally, we will build a complete, end-to-end data movement that copies a file from an Azure Data Lake Storage (ADLS) account into a lakehouse. By the end, you will have a working pipeline and a solid mental model of how the pieces fit together. Creating the Pipeline Begin by opening your Newsletter_Pipelines workspace. To create a new pipeline, click on New Item. Give your pipeline a name — in this case, we will call it Ingestion Data. Once you click OK, you are taken to the pipeline authoring canvas shown below. From here, we can start building the pipeline from scratch. Click on the activity you want to add — we will choose Copy Data. Next, click on Activities. Exploring Activities At the top of the canvas you will find the activities bar. Clicking the Activities tab reveals the most commonly used activities right away. If you need something beyond those, the three dots open up a much longer list of options. From here you can, for instance, connect to Databricks or trigger a notification. One of the most useful additions in Fabric is the Outlook activity, which lets you send an email directly from the pipeline. You can even post a notification to Microsoft Teams. These options were not available in Azure Data Factory, so for data engineers they are a genuinely handy way to keep stakeholders informed. Scheduling and Running the Pipeline Adding activities is only half the story — at some point you will want the pipeline to run automatically. To handle that, click the Run button at the top. From there you can run the pipeline immediately, or set up a schedule. Clicking the Schedule button opens the scheduling options, where you can define how and when the pipeline runs. If you would rather use a different kind of trigger — a storage-based trigger, for example — you can configure that here as well. And whenever you want to check on past executions, View Run History gives you a complete record of every run. On the right-hand side, a slider lets you zoom in and out of the canvas as needed. A quick orientation tip for anyone new to Data Factory and pipelines: clicking anywhere outside the activity box surfaces the properties for the overall pipeline configuration in the bottom panel, while clicking inside the box shows the settings specific to that individual activity. Your First End-to-End Data Movement Now it is time for a complete Fabric data movement. Our goal is to build a pipeline that moves data from an Azure Data Lake Storage account into a lakehouse. To see what we are working with, head back to the ADLS account and open the landing container. Inside, you will find the source files. Of the three files here, suppose we only want to move orders.csv from ADLS to the lakehouse. To do that, we will use the Data Factory pipeline we just started building. We already have a pipeline with a single Copy activity in place. The first thing worth doing is renaming that activity to something meaningful — for example, Transfer data from ADLS to a Lakehouse. Naming the activity this way is not mandatory, but it is good practice. A descriptive name makes the pipeline far easier for other developers to understand at a glance; they can tell what the activity does without having to open it and inspect the source and destination. Configuring the Source To move the data, click on the Source tab and set up a connection to the source. Selecting Browse all reveals the wide range of data sources Fabric supports. We want to connect to an Azure Data Lake Storage account, but if you choose View more you will see that Fabric can also connect to SharePoint, Salesforce, Oracle, FTP, SFTP, and many others. Since we only need Data Lake Storage, type Data Lake at the top, find Azure Data Lake Storage, and click it to open the connector. With the connector open, we need to supply the URL, and Fabric will then ask for authentication. The credentials can be an organizational account, a SAS token, or an account key. In this example we will use a SAS token to connect to the Azure Data Lake Storage account. To generate one, go back to the Data Lake Storage account and search for SAS, then select Shared access signature. On the SAS configuration screen, allow all resource types, grant all permissions, and set the expiry far enough out so that you can still access it later. Then click Generate SAS and connection string. You will receive both a connection string and a SAS token. For now, copy the connection string and return to Fabric. When Fabric asks whether you are creating a new connection, choose yes, and paste in the URL. Paste the URL carefully. The example shows a format like https://<storage-account>.dfs.core.windows.net. When you copy and paste, however, the value often comes through as a blob type rather than dfs. Simply change blob to dfs manually and it will work correctly. We want the path to point as far as the landing container, so type landing here. Finally, give the connection a name — we will call it ADLS connection. Since this is neither a private nor an on-premises network, no data gateway is required, so leave that blank. Under authentication type, choose Shared access signature. Now return to the Azure portal, copy the SAS token, come back to Fabric, and paste it in. To recap: we supplied the URL, changed blob to dfs, named the connection, and provided the SAS token. With that done, click Connect. The connection takes a moment to establish and will eventually succeed. To confirm everything is working, click Test connection — you should see that the connection is successful. Next, choose which file or folder to bring in using the file path. The connection is made, but now we browse to a specific file — in this case, orders.csv. Click Browse, open the landing container, select orders.csv, and click OK. Because the file is a CSV, set the file format to DelimitedText. If you need finer control, the Settings tab lets you adjust details such as the column delimiter (useful when the file is not comma-separated) and whether the first line is a header. For now, click OK and move on to the destination. Before moving on, you can confirm the data looks right using the Preview tab. Clicking Preview Data shows the incoming records, and in this case everything reads perfectly. Configuring the Destination Now switch to the Destination tab. As before, there is no existing connection to our lakehouse, so click Browse All. We want to connect to the lakehouse, so type lakehouse. Rather than selecting the New Fabric item option, we will use the OneLake Catalog. Click View More, and under it you will find the e-commerce catalog — exactly where we want our data to land. Select it to make the connection to Orders_lakehouse. With the connection in place, we are almost ready to run. Fabric asks whether to copy the data into the Tables folder or the Files folder; we will save to Files for now. When prompted for a file path, you can browse the folders available in the lakehouse. We will create a new path called copied via pipeline and copy the data there. You can also set the output file format by clicking through the options — we will keep the data in Parquet format. Select Parquet, and we are done. All that remains is to execute the pipeline so it pulls the data from ADLS and pushes it into the lakehouse. Running and Verifying the Pipeline Click the Run button, then choose Save The pipeline now starts running, and full execution takes a little while. The Output tab shows live details. Initially the run sits in a queued state, waiting for resources. Once resources are assigned, it moves into the in-progress state and then completes. You can watch it progress from queued, to in progress, to succeeded. Our pipeline has executed successfully. Clicking into the run details, we can see it copied 3,000 records from the source file into the lakehouse, which now also holds 3,000 records. You will notice that the data read size and data written size differ. That is because Parquet is a compressed format, so the data footprint shrinks once written. To confirm the data actually landed, navigate to the lakehouse, where you will see the new copied via pipeline folder. Open it, and the Parquet files are right there. Wrapping Up And that is how you copy data from an ADLS account into a lakehouse using Fabric pipelines. In just a handful of steps, we created a pipeline, explored the activities Fabric offers, set up scheduling and monitoring, and built a complete data movement from source to destination — verifying along the way that every record arrived safely. The real takeaway is how approachable this process has become. What once required stitching together separate tools now happens inside a single, cohesive workspace, complete with conveniences like Outlook and Teams notifications that simply were not available before. With this foundation in place, you are well positioned to build richer pipelines: chaining multiple activities, adding transformations, and orchestrating sophisticated, scheduled workflows. Your first pipeline is rarely your last, but it is the one that makes everything that follows feel possible.466Views4likes0CommentsSite to Insight: Fabric’s New SharePoint Picker
One of the greatest things about attending FABCON in Atlanta recently wasn’t just the incredible community or the deep-dive sessions – it was those “aha!” moments during the keynote announcements. Among the many massive reveals, there was one specific update that caught my attention as a massive “quality of life” win: the SharePoint Site Picker (Preview). For anyone who works between SharePoint and Fabric daily, this was easily one of my favourite announcements of the conference. Here’s why it’s a gamechanger.639Views1like1CommentCreating Azure Data Lake Storage (ADLS) in Azure: A Step-by-Step Guide
In modern data platforms, building efficient and reliable data pipelines is at the heart of every data engineering workflow. This is where Microsoft Fabric Data Factory comes into play. It provides a powerful, intuitive interface to design, orchestrate, and automate data movement across different systems. With its visual pipeline designer, rich set of activities, and built-in scheduling and monitoring capabilities, Data Factory enables data engineers to create scalable workflows—from simple data ingestion to complex transformations—without heavy coding. As we move from understanding the interface to building real solutions, the next essential step is preparing a robust data source. In this article, we will start by creating an Azure Data Lake Storage (ADLS) account in Azure, upload sample data, and lay the foundation for our upcoming data pipelines. We’ll begin by setting up an Azure Storage account, where we’ll upload some sample data that will later be used in our data pipeline. I’ll assume you already have access to a Microsoft Azure account. Start by opening a new browser tab and navigating to the Azure portal (portal.azure.com). From there, we’ll proceed to create the storage account within your Azure environment. To set up the storage account, simply search for “Storage” in the Azure portal, where you’ll find the Create option—go ahead and click it. You’ll then be prompted to select your subscription and define a new resource group. Next, provide a unique name for your storage account. You can leave the region and other settings as default if they suit your needs and proceed by clicking Next. To configure it as an Azure Data Lake Storage account, make sure to enable the Hierarchical Namespace option. After that, continue clicking Next through the remaining steps. In the final stage, you’ll see a summary page displaying all your selected configurations. Take a moment to review everything, and once you’re satisfied, click Create to proceed. The deployment may take a few moments to complete. Once the Azure Storage account is successfully created, navigate to the resource to continue. Once the storage account is ready, the next step is to create a container. Navigate to Data Lake Storage, add a new container, give it a name like Building Container, and then click Create. Within this container, we can now upload our sample data. Simply select the required files—such as orders, products, and customers—and click Upload to add them. That’s it—your data has been successfully uploaded to the Azure Data Lake Storage account. In the next article, we’ll build a pipeline to read this data from ADLS and load it into a Fabric Lakehouse.1.2KViews6likes1CommentDesigning a Reusable Power BI Semantic Model for Multi-Fact Analysis
We are going to design a Power BI semantic model that serves multiple reports built on top of a Snowflake / SQL warehouse using a typical Gold layer (fact and dimension tables) for a hybrid DeFi & TradFi analytics platform. The original question came from a real analytical query that joins two fact tables (FACT_TRADE_EXECUTION and FACT_WALLET_ACCOUNT) with several conformed dimensions (DIM_PROTOCOL, DIM_INVESTOR_TIER, DIM_ASSET_CLASS, DIM_INSTRUMENT_TYPE) and later mixes ledger-statement logic with window functions (ROW_NUMBER, SUM() OVER (...)). The design decision was: one semantic model per fact table, one model per business subject area, or a single wide custom query that pre-computes everything? This article summarizes the conclusions and aligns them with current Microsoft and community guidance.2.5KViews2likes0CommentsFrom Power BI to Microsoft Fabric What Data Professionals Should Know
Microsoft Fabric is changing how we think about analytics, data engineering, and BI. If you already work with Power BI, SQL, or Azure tools, this blog will help you understand what Fabric really is, how its components fit together, and where to start without getting overwhelmed.