events
6 TopicsI Let Fabric's New Data Engineering Agent Run My ETL: Here's What Happened
Last week I handed Microsoft Fabric's new Data Engineering Agent a messy ETL job that took me 11 hours to build by hand last quarter. It had a working, validated pipeline in front of me in under an hour. I still changed 3 of its 9 steps before I'd let it near anything important. That is the short version. The long version is more interesting, because the places where it went wrong tell you exactly how to use it well. Some context first. At FabCon Europe 2026 in Barcelona (September 28 to October 1), Microsoft announced the Data Engineering Agent in preview. It can plan and run migrations, ETL and performance work under guardrails that engineers define. It builds on technology from Osmos, the agentic data engineering startup Microsoft acquired in January 2026, and was demoed on stage as "Project Osmos". Copilot has been helping us write code for a while now. This is a different idea. You describe the outcome, and the agent does long-running, multi-step work on its own while you supervise. I wanted to know whether that holds up on a real, slightly ugly dataset rather than a keynote demo. So I tried it. What the Data Engineering Agent actually is The agent is an autonomous worker inside your Fabric workspace, not a chat assistant that suggests code. You give it a goal, and it does the multi-step work itself. Its roots are in Osmos, a Seattle startup whose "AI Data Engineer" already generated production-grade PySpark notebooks inside Fabric. Microsoft bought the company in January 2026 and folded the team into Fabric engineering. Based on Microsoft's FabCon session on Project Osmos, a typical run looks like this: Read the request. You describe the outcome, even loosely ("land the orders CSVs in a clean silver table"). Inspect the workspace. It looks at your existing code, Lakehouses and data layout before writing anything. Plan. It breaks the job into steps and can hand parts to sub-agents for long-running work. Build. It creates and updates notebooks and other Fabric items. Validate. It checks its own output so you can trust what was built. Hand back to you. A human reviews and approves. Microsoft is explicit that supervision stays part of the process. The headline use cases are migrations, ETL and performance tuning. I chose ETL, because that is where most of us spend our week. The test: a real pipeline, not a demo I picked a job I had already built by hand, so I could judge the agent's work line by line against my own. The data is sales orders from a mid-sized retail client . Every day their order system drops a CSV export into a Lakehouse folder. The business wants a clean orders table for analysts and a daily sales summary for a Power BI report. Item My setup Source data 14 daily CSV exports, 18.4 million rows, about 2.1 GB, in Files/raw/orders/ Target silver_orders (cleaned Delta table) and gold_daily_sales (aggregate by day, store and category) Known traps Duplicate order IDs from system retries, two date formats, a currency column with nulls, cancelled orders flagged as negative quantities Capacity F64 trial capacity, sandbox workspace Hand-built baseline About 11 hours across two days, including testing The exact prompt I gave it: Ingest the daily order CSVs in Files/raw/orders into a cleaned silver Delta table called silver_orders. Remove duplicates, standardise dates and currency, and exclude cancelled orders. Then build gold_daily_sales with total revenue, order count and average basket size by day, store and category. The pipeline must run incrementally each day. I deliberately kept the prompt a little vague. I didn't say which date formats exist, how cancellations are encoded, or which column identifies a duplicate. In real life, requirements rarely arrive as a perfect spec. I wanted to see whether the agent would investigate, ask, or just guess. What happened, step by step 1. It read the workspace before writing anything The first thing I noticed: it didn't start coding. For the first four minutes it profiled the data. It sampled each CSV, inferred the schema and looked at my existing notebooks in the workspace. It came back with a short summary of what it found. Two findings impressed me: It spotted both date formats (2026-09-14 and 14/09/2026) and flagged that files from 9 September onwards switched format. It noticed my existing notebooks used a nb_<layer>_<entity> naming pattern and said it would follow it. It missed one thing, which matters later: it reported quantity as "integer, some negative values" without asking what negative quantities meant. 2. The plan it proposed Next it produced a 9-step plan and stopped for my approval: Create a staging notebook to read all CSVs with an explicit schema Normalise both date formats to a single order_date column Fill null currency values from the store's default currency Deduplicate on order_id, keeping the latest record Remove cancelled orders Write silver_orders as a Delta table, partitioned by order_date Build gold_daily_sales with the three requested measures Add incremental logic based on the source file date Run validation checks and produce a summary report This is a sensible plan. It is roughly what I wrote on a whiteboard last quarter. But reading it closely, I changed three things before approving: Step 4: it planned to overwrite silver_orders on each run. I changed it to a MERGE on order_id, because late-arriving corrections are common in this data. Step 5: it planned to remove cancellations using a status column. That column exists, but this client never fills it. Cancellations show up as negative quantities. I told it so. Step 8: it wanted to detect new data by file modified time. I asked it to use the date in the file name instead, which is more reliable when files are re-uploaded. It accepted all three changes and rewrote the plan in a few seconds. 3. The build Once approved, it worked for 38 minutes without needing me. It created three notebooks (nb_stg_orders, nb_silver_orders, nb_gold_daily_sales) and a small pipeline to run them in order. The code was clean and readable. Here is its date handling, which is close to what I would write myself: from pyspark.sql import functions as F df = df.withColumn( "order_date", F.coalesce( F.to_date("order_date_raw", "yyyy-MM-dd"), F.to_date("order_date_raw", "dd/MM/yyyy") ) ) And the merge into silver, after my change in the plan: from delta.tables import DeltaTable silver = DeltaTable.forName(spark, "silver_orders") (silver.alias("t") .merge(updates.alias("s"), "t.order_id = s.order_id") .whenMatchedUpdateAll() .whenNotMatchedInsertAll() .execute()) It asked me one question mid-way: whether average basket size should be revenue per order or items per order. That was exactly the right question, and one I had to ask the client myself last time. 4. Self-validation When the build finished, it ran its own checks and produced a short report: Row counts reconciled from raw to staging to silver, with each dropped row explained (duplicates, cancellations) No nulls left in order_date, currency or store_id No duplicate order_id values in silver_orders Daily totals in gold_daily_sales matched a direct aggregate of silver_orders One check failed on the first pass. 41,200 rows had a null order_date after parsing. It traced the cause itself: a third date format (14-Sep-2026) appeared in two of the older files, which its initial sampling had missed. It added a third pattern to the coalesce, re-ran, and the check passed. That moment sold me on the validation step. A human could easily miss 0.2% of rows quietly disappearing. 5. My review against my own version I compared its output with the pipeline I built by hand last quarter, running both on the same 14 files. Measure Agent Me, by hand Time to working pipeline 52 minutes (incl. my review) About 11 hours Plan steps I changed 3 of 9 Not applicable Data issues found 4 of 4, after one self-correction 4 of 4 Full run time on 18.4M rows 6 min 40 s 5 min 55 s Silver row count 17,962,118 17,962,118 The row counts matched exactly. My hand-built version ran about 45 seconds faster, mostly because I had tuned partitioning. More on that below. Would I merge its code as-is? After my three plan changes, yes, with one small performance fix. Guardrails: the part you shouldn't skip An agent that can create and change items in your workspace needs limits before it needs a prompt. Microsoft designed this agent around guardrails that engineers define, and setting them up was the best 15 minutes of the whole test. What I set up before the run: A sandbox workspace. It's a preview feature, so it ran on a copy of the data, nowhere near production. Scoped data. It could read Files/raw/orders/ and write only to new tables. My existing tables were off limits. Approval points. It had to stop for my approval after the plan, and again before the first write to any Delta table. Git in the loop. The workspace was connected to a Git repo, so every notebook it created or changed landed as a diff I could review line by line. In practice, it never tried to step outside these limits. The approval after the plan turned out to be the most valuable one: that is where I caught all three of my changes. If you set only one guardrail, make it that one. Where it struggled No agent review is credible without the failures, so here are the three that mattered. It assumed instead of asking about business rules. It saw negative quantities and planned to use an empty status column for cancellations. Had I approved the plan without reading it, about 310,000 cancelled orders would have flowed into the gold table and inflated revenue. It asked a good question about basket size, but not this one. Its sampling missed a rare date format. The third format appeared in only two files. Its validation caught the problem, but only after the first run. On a much larger dataset, that first run costs real capacity. Correct, but not tuned. It partitioned silver_orders by order_date, which produced many small files for this data volume. Switching to monthly partitions and running OPTIMIZE closed most of the 45-second gap with my version. When I asked it to "make the silver write faster", it suggested the same fix, but it didn't do it unprompted. Three caveats worth knowing before you try it: It's a preview. Features, behaviour and pricing can change before general availability. Don't build production processes on it yet. Watch your capacity. Long-running agent work consumes capacity units on top of the Spark jobs themselves. Check the Capacity Metrics app after your first run so you know what a typical job costs. It doesn't know your business. It can read your schema, not your meeting notes. Every rule I wrote into the prompt was handled correctly. Every rule I left out was a guess. The verdict The Data Engineering Agent is like a fast, tireless junior engineer: excellent at the mechanics, careful about checking its own work, and still in need of a senior reviewing the plan. What I'd tell my team after this test: Spend your time on the prompt and the plan, not the code. Every problem I found was visible in the plan before a single line ran. Write business rules down explicitly. Cancellations, duplicates, currency defaults. If it's tribal knowledge, the agent will guess. Trust the validation step, but read its report. It caught a data loss bug I might have missed myself. Tune afterwards. Treat its first version as correct, then ask it for performance improvements. Try it now if you have a sandbox workspace, a backlog of repetitive ingestion or migration jobs, and someone who can review PySpark. Wait for GA if you need it in production, your pipelines are full of undocumented business logic, or your capacity budget is tight. The bigger shift is in our role. Going from 11 hours to under an hour didn't remove me from the work. It moved my effort from typing code to writing clear requirements, setting guardrails and reviewing output. That is a skill worth building now, preview or not. Have you tried the Data Engineering Agent yet? Tell me in the comments what you threw at it, and where it surprised you. Sources FabCon Europe 2026 announcements (Kanerika) From Prompt to Finished Project: Introducing Project Osmos (FabCon Europe 2026 session) Microsoft Acquires Osmos (Redmond Magazine) Fabric January 2026 Feature Summary (Microsoft Fabric Blog)FabCon & SQLCon 2026: The Ultimate Roundup of Announcements
While I wasn’t able to catch the action live in Barcelona last week for FabCon and SQLCon 2026, keeping tabs on the announcements has been a wild ride. Microsoft Fabric has officially cleared the 40,000-customer milestone and the event heavily emphasized Frontier Transformation – moving past standard AI productivity hacks and embedding autonomous agents directly into corporate workflows. Yet, among all the flashy Copilot integrations and agentic apps, one specific announcement completely stole the show for me: the introduction of the F0 capacity tier with on-demand billing. For anyone who has ever wrestled with budget approvals just to test out a new data platform, this is a game-changer. Let’s dive into why F0 is my favorite takeaway, alongside a full roundup of everything else that dropped in Barcelona. 1. My #1 Favorite: The F0 Capacity SKU & Pay-As-You-Go Billing Historically, dipping your toes into Microsoft Fabric meant committing upfront to a reserved compute tier (like an F2 or higher). If you just wanted to spin up a quick test environment, build a proof of concept, or play with OneLake shortcuts, you still had to navigate a monthly capacity bill. What’s new: Microsoft introduced F0, a zero-provisioned capacity SKU featuring true on-demand billing. Why it’s a massive deal: It acts as the “serverless” equivalent for Fabric. You can now evaluate features, run lightweight development tasks and test out workloads without locking into an upfront monthly reservation. Compute costs finally align strictly to what actually runs, shattering the barrier to entry for smaller teams and developer sandboxes. 2. Grounding Microsoft Copilot with Fabric IQ AI tools often hallucinate because they lack proper context. Microsoft is tackling this by wiring business context directly into Microsoft Copilot. What’s new: Fabric IQ is generally available as a shared intelligence bridge, connecting OneLake data, Power BI semantic models and operational metrics directly into Copilot Chat and Cowork – with zero additional AI token costs. Why it matters: Instead of guessing, your organization’s Copilot can answer business questions using the exact semantic data definitions your teams already trust and govern. Read more: Check out Arun Ulag’s keynote overview on the official Microsoft Azure Blog. 3. Power BI’s Evolution: Agentic App Creation Power BI is stepping past standard dashboards into operational application building. What’s new: You can use natural language inside Power BI Desktop to build, preview and publish purpose-built data applications straight from a trusted semantic model. These apps can handle inputs, write-back data and power workflows. Why it matters: Power BI Pro and Premium Per User (PPU) users will get access to Fabric Apps and Fabric Database capabilities (up to 1 GB per app) at no additional cost. Read more: Dive into Mohammad Ali’s deep dive on Power BI’s next chapter: The evolution of business intelligence. 4. Fast-Tracking Apps to Production with Fabric Apps & Rayfin Prototyping an app with AI is easy, but getting it to a secure production environment used to mean heavy lifting. Fabric Apps solves this. What’s new: Using the open-source Rayfin SDK and CLI, developers can build app backends and deploy them directly to Fabric with built-in TypeScript functions, a Secret Store, PostgreSQL support and private-by-default security. Why it matters: Apps plug straight into existing Lakehouses and Warehouses without data duplication. Read more: Catch Sachin Patney’s post, From prompt to production: What’s new in Fabric Apps. 5. Expanding the OneLake Ecosystem & IQ Sharing What’s new: Microsoft rolled out IQ sharing (in preview), allowing organizations to securely share governed data, markdown agent instructions and ontologies across different teams, partners and external ecosystems. Read more: Read Dipti Borkar’s breakdown on What’s new in Microsoft OneLake and its rapidly growing ecosystem. 6. Autonomous Data Engineering Agents What’s new: Driven by technology from the Osmos acquisition, Microsoft introduced a preview of the Fabric data engineering agent. Rather than just auto-completing code snippets, this agent handles complex, long-running engineering tasks like migrations, ETL pipeline creation and modernizations based on human guardrails. Read more: Check out Bogdan Crivat’s blog on Bringing governed analytics into the flow of work: Fabric Analytics at FabCon Europe 2026. 7. SQLCon: Scale, Intelligence, and Serverless Pauses Running parallel to FabCon, SQLCon delivered major updates for database administrators balancing traditional relational databases with modern demands. What’s new: A new Database Hub (Public Preview) serves as a single pane of glass to monitor health and security across SQL Server, Azure SQL and PostgreSQL. Plus, DiskANN vector indexes are now generally available to supercharge native semantic search using T-SQL. Read more: Read Shireesh Thota’s update on SQLCon Barcelona 2026: Advancing SQL with greater control, scale and intelligence. Final Thoughts: Looking Ahead Reflecting on everything announced in Barcelona, it’s clear that Microsoft is moving faster than ever to bridge the gap between heavy-duty data engineering and practical, AI-driven applications. While agentic workflows and Copilot enhancements grab the headlines, structural updates like the F0 tier are what truly make this ecosystem accessible to everyone from lone developers to enterprise architects. Honestly, I am still trying to digest all these announcements made and also can’t wait to try them soon! We are officially stepping into an era where building, scaling and operationalizing data isn’t just about managing tables – it’s about powering the intelligence layer that runs the business. Are you as thrilled about the F0 capacity tier as I am? What feature from FabCon are you testing first? Let me know in the comments below!31Views2likes0CommentsTired of Hacking DAX Just to Make a Chart Look Right? Fabric Apps Let You Build the Exact Dashboard
Ever needed a dashboard that looks a very specific way- a big trend number with a mini sparkline, a real-time ticking counter, a layout Power BI's report canvas just won't give you- and ended up stuffing extra measures and workarounds into your semantic model just to fake it? Fabric Apps let you build that exact look as a small custom web app that plugs into your existing model, and an AI coding agent can build most of it for you. Here's what it actually is, in plain terms, with one simple example.Introducing the Fabric Extensibility Toolkit Contest 🎉
We’re thrilled to announce our newest community challenge: the Fabric Extensibility Toolkit Contest - your chance to build innovative workload items, inspire the ecosystem, and help ignite the momentum behind Fabric extensibility.20KViews11likes0Comments