capacities
30 TopicsI Let Fabric's New Data Engineering Agent Run My ETL: Here's What Happened
Last week I handed Microsoft Fabric's new Data Engineering Agent a messy ETL job that took me 11 hours to build by hand last quarter. It had a working, validated pipeline in front of me in under an hour. I still changed 3 of its 9 steps before I'd let it near anything important. That is the short version. The long version is more interesting, because the places where it went wrong tell you exactly how to use it well. Some context first. At FabCon Europe 2026 in Barcelona (September 28 to October 1), Microsoft announced the Data Engineering Agent in preview. It can plan and run migrations, ETL and performance work under guardrails that engineers define. It builds on technology from Osmos, the agentic data engineering startup Microsoft acquired in January 2026, and was demoed on stage as "Project Osmos". Copilot has been helping us write code for a while now. This is a different idea. You describe the outcome, and the agent does long-running, multi-step work on its own while you supervise. I wanted to know whether that holds up on a real, slightly ugly dataset rather than a keynote demo. So I tried it. What the Data Engineering Agent actually is The agent is an autonomous worker inside your Fabric workspace, not a chat assistant that suggests code. You give it a goal, and it does the multi-step work itself. Its roots are in Osmos, a Seattle startup whose "AI Data Engineer" already generated production-grade PySpark notebooks inside Fabric. Microsoft bought the company in January 2026 and folded the team into Fabric engineering. Based on Microsoft's FabCon session on Project Osmos, a typical run looks like this: Read the request. You describe the outcome, even loosely ("land the orders CSVs in a clean silver table"). Inspect the workspace. It looks at your existing code, Lakehouses and data layout before writing anything. Plan. It breaks the job into steps and can hand parts to sub-agents for long-running work. Build. It creates and updates notebooks and other Fabric items. Validate. It checks its own output so you can trust what was built. Hand back to you. A human reviews and approves. Microsoft is explicit that supervision stays part of the process. The headline use cases are migrations, ETL and performance tuning. I chose ETL, because that is where most of us spend our week. The test: a real pipeline, not a demo I picked a job I had already built by hand, so I could judge the agent's work line by line against my own. The data is sales orders from a mid-sized retail client . Every day their order system drops a CSV export into a Lakehouse folder. The business wants a clean orders table for analysts and a daily sales summary for a Power BI report. Item My setup Source data 14 daily CSV exports, 18.4 million rows, about 2.1 GB, in Files/raw/orders/ Target silver_orders (cleaned Delta table) and gold_daily_sales (aggregate by day, store and category) Known traps Duplicate order IDs from system retries, two date formats, a currency column with nulls, cancelled orders flagged as negative quantities Capacity F64 trial capacity, sandbox workspace Hand-built baseline About 11 hours across two days, including testing The exact prompt I gave it: Ingest the daily order CSVs in Files/raw/orders into a cleaned silver Delta table called silver_orders. Remove duplicates, standardise dates and currency, and exclude cancelled orders. Then build gold_daily_sales with total revenue, order count and average basket size by day, store and category. The pipeline must run incrementally each day. I deliberately kept the prompt a little vague. I didn't say which date formats exist, how cancellations are encoded, or which column identifies a duplicate. In real life, requirements rarely arrive as a perfect spec. I wanted to see whether the agent would investigate, ask, or just guess. What happened, step by step 1. It read the workspace before writing anything The first thing I noticed: it didn't start coding. For the first four minutes it profiled the data. It sampled each CSV, inferred the schema and looked at my existing notebooks in the workspace. It came back with a short summary of what it found. Two findings impressed me: It spotted both date formats (2026-09-14 and 14/09/2026) and flagged that files from 9 September onwards switched format. It noticed my existing notebooks used a nb_<layer>_<entity> naming pattern and said it would follow it. It missed one thing, which matters later: it reported quantity as "integer, some negative values" without asking what negative quantities meant. 2. The plan it proposed Next it produced a 9-step plan and stopped for my approval: Create a staging notebook to read all CSVs with an explicit schema Normalise both date formats to a single order_date column Fill null currency values from the store's default currency Deduplicate on order_id, keeping the latest record Remove cancelled orders Write silver_orders as a Delta table, partitioned by order_date Build gold_daily_sales with the three requested measures Add incremental logic based on the source file date Run validation checks and produce a summary report This is a sensible plan. It is roughly what I wrote on a whiteboard last quarter. But reading it closely, I changed three things before approving: Step 4: it planned to overwrite silver_orders on each run. I changed it to a MERGE on order_id, because late-arriving corrections are common in this data. Step 5: it planned to remove cancellations using a status column. That column exists, but this client never fills it. Cancellations show up as negative quantities. I told it so. Step 8: it wanted to detect new data by file modified time. I asked it to use the date in the file name instead, which is more reliable when files are re-uploaded. It accepted all three changes and rewrote the plan in a few seconds. 3. The build Once approved, it worked for 38 minutes without needing me. It created three notebooks (nb_stg_orders, nb_silver_orders, nb_gold_daily_sales) and a small pipeline to run them in order. The code was clean and readable. Here is its date handling, which is close to what I would write myself: from pyspark.sql import functions as F df = df.withColumn( "order_date", F.coalesce( F.to_date("order_date_raw", "yyyy-MM-dd"), F.to_date("order_date_raw", "dd/MM/yyyy") ) ) And the merge into silver, after my change in the plan: from delta.tables import DeltaTable silver = DeltaTable.forName(spark, "silver_orders") (silver.alias("t") .merge(updates.alias("s"), "t.order_id = s.order_id") .whenMatchedUpdateAll() .whenNotMatchedInsertAll() .execute()) It asked me one question mid-way: whether average basket size should be revenue per order or items per order. That was exactly the right question, and one I had to ask the client myself last time. 4. Self-validation When the build finished, it ran its own checks and produced a short report: Row counts reconciled from raw to staging to silver, with each dropped row explained (duplicates, cancellations) No nulls left in order_date, currency or store_id No duplicate order_id values in silver_orders Daily totals in gold_daily_sales matched a direct aggregate of silver_orders One check failed on the first pass. 41,200 rows had a null order_date after parsing. It traced the cause itself: a third date format (14-Sep-2026) appeared in two of the older files, which its initial sampling had missed. It added a third pattern to the coalesce, re-ran, and the check passed. That moment sold me on the validation step. A human could easily miss 0.2% of rows quietly disappearing. 5. My review against my own version I compared its output with the pipeline I built by hand last quarter, running both on the same 14 files. Measure Agent Me, by hand Time to working pipeline 52 minutes (incl. my review) About 11 hours Plan steps I changed 3 of 9 Not applicable Data issues found 4 of 4, after one self-correction 4 of 4 Full run time on 18.4M rows 6 min 40 s 5 min 55 s Silver row count 17,962,118 17,962,118 The row counts matched exactly. My hand-built version ran about 45 seconds faster, mostly because I had tuned partitioning. More on that below. Would I merge its code as-is? After my three plan changes, yes, with one small performance fix. Guardrails: the part you shouldn't skip An agent that can create and change items in your workspace needs limits before it needs a prompt. Microsoft designed this agent around guardrails that engineers define, and setting them up was the best 15 minutes of the whole test. What I set up before the run: A sandbox workspace. It's a preview feature, so it ran on a copy of the data, nowhere near production. Scoped data. It could read Files/raw/orders/ and write only to new tables. My existing tables were off limits. Approval points. It had to stop for my approval after the plan, and again before the first write to any Delta table. Git in the loop. The workspace was connected to a Git repo, so every notebook it created or changed landed as a diff I could review line by line. In practice, it never tried to step outside these limits. The approval after the plan turned out to be the most valuable one: that is where I caught all three of my changes. If you set only one guardrail, make it that one. Where it struggled No agent review is credible without the failures, so here are the three that mattered. It assumed instead of asking about business rules. It saw negative quantities and planned to use an empty status column for cancellations. Had I approved the plan without reading it, about 310,000 cancelled orders would have flowed into the gold table and inflated revenue. It asked a good question about basket size, but not this one. Its sampling missed a rare date format. The third format appeared in only two files. Its validation caught the problem, but only after the first run. On a much larger dataset, that first run costs real capacity. Correct, but not tuned. It partitioned silver_orders by order_date, which produced many small files for this data volume. Switching to monthly partitions and running OPTIMIZE closed most of the 45-second gap with my version. When I asked it to "make the silver write faster", it suggested the same fix, but it didn't do it unprompted. Three caveats worth knowing before you try it: It's a preview. Features, behaviour and pricing can change before general availability. Don't build production processes on it yet. Watch your capacity. Long-running agent work consumes capacity units on top of the Spark jobs themselves. Check the Capacity Metrics app after your first run so you know what a typical job costs. It doesn't know your business. It can read your schema, not your meeting notes. Every rule I wrote into the prompt was handled correctly. Every rule I left out was a guess. The verdict The Data Engineering Agent is like a fast, tireless junior engineer: excellent at the mechanics, careful about checking its own work, and still in need of a senior reviewing the plan. What I'd tell my team after this test: Spend your time on the prompt and the plan, not the code. Every problem I found was visible in the plan before a single line ran. Write business rules down explicitly. Cancellations, duplicates, currency defaults. If it's tribal knowledge, the agent will guess. Trust the validation step, but read its report. It caught a data loss bug I might have missed myself. Tune afterwards. Treat its first version as correct, then ask it for performance improvements. Try it now if you have a sandbox workspace, a backlog of repetitive ingestion or migration jobs, and someone who can review PySpark. Wait for GA if you need it in production, your pipelines are full of undocumented business logic, or your capacity budget is tight. The bigger shift is in our role. Going from 11 hours to under an hour didn't remove me from the work. It moved my effort from typing code to writing clear requirements, setting guardrails and reviewing output. That is a skill worth building now, preview or not. Have you tried the Data Engineering Agent yet? Tell me in the comments what you threw at it, and where it surprised you. Sources FabCon Europe 2026 announcements (Kanerika) From Prompt to Finished Project: Introducing Project Osmos (FabCon Europe 2026 session) Microsoft Acquires Osmos (Redmond Magazine) Fabric January 2026 Feature Summary (Microsoft Fabric Blog)FabCon & SQLCon 2026: The Ultimate Roundup of Announcements
While I wasn’t able to catch the action live in Barcelona last week for FabCon and SQLCon 2026, keeping tabs on the announcements has been a wild ride. Microsoft Fabric has officially cleared the 40,000-customer milestone and the event heavily emphasized Frontier Transformation – moving past standard AI productivity hacks and embedding autonomous agents directly into corporate workflows. Yet, among all the flashy Copilot integrations and agentic apps, one specific announcement completely stole the show for me: the introduction of the F0 capacity tier with on-demand billing. For anyone who has ever wrestled with budget approvals just to test out a new data platform, this is a game-changer. Let’s dive into why F0 is my favorite takeaway, alongside a full roundup of everything else that dropped in Barcelona. 1. My #1 Favorite: The F0 Capacity SKU & Pay-As-You-Go Billing Historically, dipping your toes into Microsoft Fabric meant committing upfront to a reserved compute tier (like an F2 or higher). If you just wanted to spin up a quick test environment, build a proof of concept, or play with OneLake shortcuts, you still had to navigate a monthly capacity bill. What’s new: Microsoft introduced F0, a zero-provisioned capacity SKU featuring true on-demand billing. Why it’s a massive deal: It acts as the “serverless” equivalent for Fabric. You can now evaluate features, run lightweight development tasks and test out workloads without locking into an upfront monthly reservation. Compute costs finally align strictly to what actually runs, shattering the barrier to entry for smaller teams and developer sandboxes. 2. Grounding Microsoft Copilot with Fabric IQ AI tools often hallucinate because they lack proper context. Microsoft is tackling this by wiring business context directly into Microsoft Copilot. What’s new: Fabric IQ is generally available as a shared intelligence bridge, connecting OneLake data, Power BI semantic models and operational metrics directly into Copilot Chat and Cowork – with zero additional AI token costs. Why it matters: Instead of guessing, your organization’s Copilot can answer business questions using the exact semantic data definitions your teams already trust and govern. Read more: Check out Arun Ulag’s keynote overview on the official Microsoft Azure Blog. 3. Power BI’s Evolution: Agentic App Creation Power BI is stepping past standard dashboards into operational application building. What’s new: You can use natural language inside Power BI Desktop to build, preview and publish purpose-built data applications straight from a trusted semantic model. These apps can handle inputs, write-back data and power workflows. Why it matters: Power BI Pro and Premium Per User (PPU) users will get access to Fabric Apps and Fabric Database capabilities (up to 1 GB per app) at no additional cost. Read more: Dive into Mohammad Ali’s deep dive on Power BI’s next chapter: The evolution of business intelligence. 4. Fast-Tracking Apps to Production with Fabric Apps & Rayfin Prototyping an app with AI is easy, but getting it to a secure production environment used to mean heavy lifting. Fabric Apps solves this. What’s new: Using the open-source Rayfin SDK and CLI, developers can build app backends and deploy them directly to Fabric with built-in TypeScript functions, a Secret Store, PostgreSQL support and private-by-default security. Why it matters: Apps plug straight into existing Lakehouses and Warehouses without data duplication. Read more: Catch Sachin Patney’s post, From prompt to production: What’s new in Fabric Apps. 5. Expanding the OneLake Ecosystem & IQ Sharing What’s new: Microsoft rolled out IQ sharing (in preview), allowing organizations to securely share governed data, markdown agent instructions and ontologies across different teams, partners and external ecosystems. Read more: Read Dipti Borkar’s breakdown on What’s new in Microsoft OneLake and its rapidly growing ecosystem. 6. Autonomous Data Engineering Agents What’s new: Driven by technology from the Osmos acquisition, Microsoft introduced a preview of the Fabric data engineering agent. Rather than just auto-completing code snippets, this agent handles complex, long-running engineering tasks like migrations, ETL pipeline creation and modernizations based on human guardrails. Read more: Check out Bogdan Crivat’s blog on Bringing governed analytics into the flow of work: Fabric Analytics at FabCon Europe 2026. 7. SQLCon: Scale, Intelligence, and Serverless Pauses Running parallel to FabCon, SQLCon delivered major updates for database administrators balancing traditional relational databases with modern demands. What’s new: A new Database Hub (Public Preview) serves as a single pane of glass to monitor health and security across SQL Server, Azure SQL and PostgreSQL. Plus, DiskANN vector indexes are now generally available to supercharge native semantic search using T-SQL. Read more: Read Shireesh Thota’s update on SQLCon Barcelona 2026: Advancing SQL with greater control, scale and intelligence. Final Thoughts: Looking Ahead Reflecting on everything announced in Barcelona, it’s clear that Microsoft is moving faster than ever to bridge the gap between heavy-duty data engineering and practical, AI-driven applications. While agentic workflows and Copilot enhancements grab the headlines, structural updates like the F0 tier are what truly make this ecosystem accessible to everyone from lone developers to enterprise architects. Honestly, I am still trying to digest all these announcements made and also can’t wait to try them soon! We are officially stepping into an era where building, scaling and operationalizing data isn’t just about managing tables – it’s about powering the intelligence layer that runs the business. Are you as thrilled about the F0 capacity tier as I am? What feature from FabCon are you testing first? Let me know in the comments below!30Views2likes0CommentsTesting Fabric Runtime 2.0 Before It Becomes the Default
Why you should test this week, not next month Runtime 2.0 has been GA since August, but it is still opt-in. According to the Runtime 2.0 docs, Microsoft plans to make it the default for new workspaces and new Environment items in late September 2026. That is right now. At the same time, Runtime 1.3 (Spark 3.5) is listed with an end-of-support date of 30 September 2026. It moves into Long Term Support from 1 October for six months, through March 2027 (lifecycle page). So you have a runway, but it is not long. In my trainings, the question I get most is: "Will my notebooks break when this flips?" My honest answer: most will run fine, a few will fail loudly, and a small number will return different results without failing. That last group is the reason to test. This post walks through the exact method I use to test in isolation, without touching production. What you need: a Fabric capacity (F2 or trial is fine), Contributor or higher on a workspace, and 30–45 minutes. What actually changes between 1.3 and 2.0 Every core component moves a major or minor version. Versions below are from the official Runtime 1.3 and Runtime 2.0 pages. Component Runtime 1.3 Runtime 2.0 What to watch Apache Spark 3.5 4.1 ANSI SQL mode is on by default in Spark 4 Delta Lake 3.2 4.2 New 4.x table features can break reads from other Fabric engines Python 3.11 3.13 Pinned pip packages may have no 3.13 wheel Java 11 21 Custom JARs built for old JDKs Scala 2.12.17 2.13 Any Scala JAR compiled for 2.12 will not load OS Azure Linux 2.0 Azure Linux 3.0 Native libraries in custom packages R 4.4.1 4.5.2 SparkR is deprecated in Spark 4.x The four areas that cause real trouble in my experience: ANSI mode. In Spark 3.5, CAST('abc' AS INT) quietly returns NULL. In Spark 4.x it throws an error. The same applies to integer overflow and invalid dates. Pipelines that relied on silent nulls now fail, which is actually good, but you need to know before a 2 a.m. run does. Scala 2.13 / Java 21. If your Environment has custom JARs, they must be rebuilt. There is no workaround. Python library pins. A requirements.txt or Environment library list pinned for Python 3.11 is the most common failure at publish time. Delta 4.x features. The docs are explicit: Delta Lake 4.x-specific features are experimental and only work in Spark experiences. If the SQL analytics endpoint, Power BI Direct Lake or the Warehouse read the same table, do not enable them. Check Delta Lake interoperability first. The good news: Runtime 2.0 ships with the Native Execution Engine, which Microsoft benchmarks at up to 6x faster than open-source Spark on TPC-DS, with no code change. The safe test method: a side-by-side Environment The trick is simple. Never change the runtime at workspace level first. An Environment item pinned to 2.0 overrides the workspace default, but only for the notebooks you attach to it. Production keeps running on 1.3 while you test the same code on 2.0. Step 1: Create a test workspace Create a workspace such as WS-Runtime2-Test on the same capacity. If your production workspace is Git-connected, branch out to it, or simply export the 3–5 notebooks that matter most. Add a shortcut to the production lakehouse tables you read from, so you test against real data without copying it. Step 2: Create two Environment items Create env_rt13 and env_rt20 in the test workspace. Two environments, same libraries, only the runtime differs. That is what makes the comparison fair. New item → Environment, name it env_rt20. In the Runtime dropdown, select 2.0 (Spark 4.1, Delta 4.2). Under Spark compute → Acceleration, turn on Native Execution Engine. Copy your production library list into Public libraries (and custom JARs/wheels if you use them). Save, then Publish. Publishing is where library conflicts surface, so watch this step. Repeat for env_rt13 with 1.3 (Spark 3.5, Delta 3.2) and the same libraries. Step 3: Attach and run Open a copy of each notebook, and in the environment dropdown on the notebook ribbon, pick env_rt20. Run it. Then run the same notebook attached to env_rt13. A Spark job definition works the same way through its settings. Optional: the early access release channel There is a second mechanism worth knowing about for later. Release channels (preview) let you test the next set of updates within a runtime before they become default. You set it in the Environment's Spark properties: spark.fabric.pools.skipStarterPools=true spark.computeConf.runtime.releaseChannel=earlyAccess Early access does not use the Starter Pool, so session start is slower, and the setting is fixed for the lifetime of a session. Billing is the same. Use it after your 2.0 migration to catch library updates early. The test notebook I use in class Four cells. Run the whole notebook once on env_rt13 and once on env_rt20. It takes about 10 minutes, most of it session start. Cell 1: Prove which runtime you are on Don't trust the dropdown; print it. This also gives you the VHD image name that support will ask for if you raise a ticket. import sys from importlib.metadata import version, PackageNotFoundError jvm = spark.sparkContext._jvm def pkg(name): try: return version(name) except PackageNotFoundError: return "not installed" info = { "spark": spark.version, "python": sys.version.split()[0], "java": jvm.System.getProperty("java.version"), "scala": jvm.scala.util.Properties.versionNumberString(), "pandas": pkg("pandas"), "native engine": spark.conf.get("spark.native.enabled", "not set"), "ansi mode": spark.conf.get("spark.sql.ansi.enabled"), "vhd image": spark.conf.get("spark.synapse.vhd.name", ""), } for k, v in info.items(): print(f"{k:<15} {v}") Cell 2: The ANSI probe This is the cell that makes the room go quiet in a workshop. Same SQL, different behaviour. probes = { "bad cast": "SELECT CAST('abc' AS INT) AS v", "int overflow": "SELECT CAST(2147483647 AS INT) + 1 AS v", "bad date": "SELECT CAST('2026-02-30' AS DATE) AS v", "div by zero": "SELECT 1/0 AS v", } for name, sql in probes.items(): try: print(f"{name:<13} -> {spark.sql(sql).first()['v']}") except Exception as e: print(f"{name:<13} -> ERROR: {type(e).__name__}") On 1.3 you get None, -2147483648, None, None. Four silent problems. On 2.0 with default settings, all four raise errors. Now search your code for CAST(, arithmetic on IDs and date parsing from strings. Cell 3: Regression check on your real logic Paste one production transformation, write the result from each runtime, and compare row by row. tag = "rt20" if spark.version.startswith("4") else "rt13" result_df = spark.sql(""" -- paste one real production query here SELECT ... """) result_df.write.mode("overwrite").format("delta").saveAsTable(f"regress_{tag}") After both runs, compare from either environment: a = spark.table("regress_rt13") b = spark.table("regress_rt20") print("row counts :", a.count(), b.count()) print("only in 1.3 :", a.exceptAll(b).count()) print("only in 2.0 :", b.exceptAll(a).count()) Zero and zero is what you want. If you get differences on float columns, round them before comparing; tiny floating point drift is not a bug. Cell 4: Make sure other engines can still read the table A table written from Runtime 2.0 must stay readable by the SQL analytics endpoint and Direct Lake. Check the protocol: (spark.sql("DESCRIBE DETAIL regress_rt20") .select("minReaderVersion", "minWriterVersion", "tableFeatures") .show(truncate=False)) Compare it with regress_rt13. They should match. Then open the lakehouse's SQL analytics endpoint and run SELECT TOP 10 * FROM regress_rt20. If that works, your downstream reports are safe. Finally, note the duration of each run from the Monitor hub. With the Native Execution Engine on, 2.0 is often faster, but measure it on your workload rather than quoting benchmarks. Fixes, rollout plan and checklist These are the fixes I reach for when the test fails. Symptom Likely cause Fix Environment publish fails Library pinned for Python 3.11 Upgrade the pin to a version with a 3.13 wheel, or remove the pin NoClassDefFoundError / scala. errors JAR compiled for Scala 2.12 Rebuild the JAR for Scala 2.13 and Java 21 CAST_INVALID_INPUT, ARITHMETIC_OVERFLOW, DIVIDE_BY_ZERO ANSI mode Use try_cast, try_divide, try_add; clean the input SQL endpoint can't read a table Delta 4.x feature enabled on a shared table Don't enable 4.x-only features on tables other engines read Results differ, no error Changed default behaviour Diff with Cell 3, then read the Spark SQL migration guide On ANSI mode: you can set spark.sql.ansi.enabled=false in the Environment to get the old behaviour back. I use it only as a temporary bridge, never as the fix. Silent nulls are exactly the data quality issue ANSI mode exists to catch. My rollout plan test workspace, two Environments, run the four cells on your top 5 notebooks. fix libraries and ANSI failures; rerun until Cell 3 shows zero differences. attach env_rt20 to production notebooks one pipeline at a time. Keep env_rt13 so rollback is one dropdown change. once everything runs on 2.0, change the workspace default in Workspace settings → Data Engineering/Science → Spark settings → Environment → Runtime version. Before March 2027: nothing left on 1.3 when LTS ends. Checklist Test workspace created, production untouched env_rt13 and env_rt20 published with identical libraries Cell 1 confirms Spark 4.1 / Python 3.13 / Java 21 ANSI probe run and code searched for risky casts Cell 3 regression shows zero differences SQL analytics endpoint reads tables written by 2.0 Custom JARs rebuilt for Scala 2.13 Workspace default switched last If you do only one thing from this post, run Cell 2 on your busiest notebook today. Sources Runtime 2.0 in Fabric Runtime 1.3 in Fabric Lifecycle of Apache Spark runtimes in Fabric Fabric runtime release channels (preview) Delta Lake table format interoperability Spark Runtime releases and updates (GitHub)116Views0likes0CommentsConnect Fabric Data Agents to Copilot Studio with MCP (+ Assistants API Migration Fix)
Business users want answers where they already work, in Teams and Microsoft 365 Copilot, and those answers need to come from trusted, governed data. Fabric data agents in Copilot Studio do exactly that, and the integration is now generally available. There is also a deadline you may already have passed. If you connected to Fabric data agents through the OpenAI Assistants API, your integration has likely stopped working. This guide covers both the setup and the fix. What's new Fabric data agents now connect to Copilot Studio through a new tool-based experience: you select Add a tool, search for Fabric, and choose Fabric IQ Data MCP. Your Copilot Studio agent then calls the Fabric data agent like any other tool. Microsoft Community Three points matter for architects: Data stays governed. The Fabric data agent keeps running in Fabric, respects permissions on the underlying data sources, and returns answers grounded in governed enterprise data. Microsoft Community The orchestrator decides. The orchestrator chooses when to use enterprise data and when to use its other knowledge sources and tools. Microsoft Community You can combine agents. You can pair several Fabric data agents with other business systems to build richer solutions. A sales agent could query a sales data agent, a finance data agent and your CRM connector in a single conversation. Microsoft Community Behind this, Copilot Studio also introduced a redesigned authoring surface and runtime powered by the GitHub Copilot harness. Microsoft Community Why MCP matters Model Context Protocol (MCP) is becoming the standard way agents connect to tools and data. By exposing Fabric data agents through MCP, Microsoft makes them a reusable building block instead of a one-off integration. The same move is happening elsewhere: Fabric data agents in Microsoft Foundry are now easier to connect and monitor, and that integration is moving to MCP too. Microsoft Community You build a data agent once in Fabric, and Copilot Studio, Foundry and Microsoft 365 Copilot can all use it. Prerequisites Before you start, check that you have: A Microsoft Fabric workspace on Fabric capacity Data ready to use: a lakehouse, warehouse, semantic model or similar source with clean, well-named tables Permission to create items in the Fabric workspace Access to Copilot Studio with permission to build and publish agents Microsoft 365 Copilot or Teams, depending on where you want to publish Tip: the quality of your data agent depends on the quality of your data. Clear table and column names and descriptions make a big difference to answer accuracy. Step 1: Build and test your Fabric data agent In your Fabric workspace, create a new data agent item, then: Add data sources. Connect the lakehouse, warehouse or semantic model the agent should answer from. Start narrow, since one domain per agent works best. Write agent instructions. Explain the business context, key terms and rules. For example: "Revenue means net revenue after returns. Fiscal year starts in April. Always exclude internal test accounts." Add example questions and queries. Give sample questions with the correct query logic, so the agent learns your patterns. Test thoroughly. Ask the questions your business users actually ask and compare the answers with your certified reports. Publish the agent once the answers are reliable. Step 2: Add the data agent to Copilot Studio This is the new MCP-based flow. Create or open an agent in Copilot Studio, go to Tools, select Add a tool, search for Fabric, and add Fabric IQ Data MCP. Microsoft Community Then: Select the Fabric data agent you published in Step 1. Give the tool a clear description, such as "Answers questions about sales, revenue and pipeline from governed Fabric data." The orchestrator reads this to decide when to call the tool, so be specific. In your Copilot Studio agent's instructions, tell it when to use the tool, for example: "For any question about sales numbers or targets, use the Sales Data tool." Step 3: Test in Copilot Studio Test your agent in Preview before publishing. Check that: Microsoft Community Data questions trigger the Fabric tool, and general questions don't Answers match the results from Fabric directly A user without access to certain data gets no answer from it, since the data agent respects permissions on the underlying sources Follow-up questions ("and what about last quarter?") work as expected Step 4: Publish to Teams or Microsoft 365 Copilot When you're happy with the results, publish to Microsoft Teams or Microsoft 365 Copilot. Business users can now ask data questions in the tools they use every day, without opening a report. Microsoft Community Roll out to a small pilot group first, collect feedback, and improve the data agent's instructions before rolling out widely. The Assistants API retirement: what broke and how to fix it This part matters if you built programmatic integrations. What happened: The OpenAI Assistants API, which powered the orchestration layer of the Fabric data agent, was scheduled to be shut down by OpenAI on August 26, 2026. After that date, direct calls to the Assistants API stop working. Microsoft Community Who is affected: anyone who connects to a Fabric data agent programmatically through the Assistants API needs to migrate to the data agent MCP endpoint. Typical examples are custom apps, scripts, or third-party tools calling the data agent directly. Microsoft Community Who is mostly fine: SDK and Fabric portal users need little or no action, because Microsoft migrates those experiences internally, although conversation history may reset once. Microsoft Community What stays the same: existing agent data sources, instructions and tools remain unchanged. You don't need to rebuild your agents, only change how you connect to them. Microsoft Community Migration checklist Find your integrations. Search your code and configs for Assistants API calls that target Fabric data agents. Check custom apps, automation scripts, Azure Functions and any third-party tools. Check for failures since late August. Errors or silent failures after August 26 are a strong sign an integration is affected. Move to the MCP endpoint. Update each integration to connect through the data agent's MCP endpoint instead. Follow Microsoft's guide, "Prepare your Fabric Data Agent integrations for Assistants API retirement," for the current details. Consider Copilot Studio instead of custom code. If a custom app only exists to put a chat interface on a data agent, the Copilot Studio route above may replace it with less code to maintain. Tell users about history resets. If you use the portal or SDK, let users know earlier conversations may disappear once. Retest. Run your standard question set again after migrating to confirm the answers haven't changed. Best practices for production One domain per data agent. Separate agents for sales, finance and operations are easier to test, govern and improve than one agent that tries to do everything. Write tool descriptions carefully. In Copilot Studio, the tool description is how the orchestrator decides what to call. Vague descriptions lead to wrong tool choices. Treat instructions like code. Version them, review changes, and retest after every edit. Keep a test question set. Keep 20 to 30 real questions with expected answers, and run them after any change to data, instructions or model. Watch permissions. Answers respect underlying permissions, so a mistake in data security becomes a mistake in agent answers. Review access regularly. Start small and grow. Pilot with one team, measure how often answers are right, then expand. The bottom line Connecting Fabric data agents to Copilot Studio through MCP turns your governed Fabric data into a reusable tool that any agent can use, and it puts trusted answers directly into Teams and Microsoft 365 Copilot. The setup takes minutes once your data agent is solid. If you built integrations on the Assistants API, make migration your first job this week. Your agents, data sources and instructions don't need to change, only the connection. After that, you're on the platform Microsoft is building its agent strategy on.Fabric Capacity Cockpit: pause, resume and track the cost of your F-SKU from one window
Why I built it When my Fabric trial capacity expired, I moved my workspaces to a small pay-as-you-go F2. I use it for experiments and demos, not production, so the obvious way to keep costs down is to pause it whenever I'm not working with it. In practice, the information I need for that is spread across several places: Azure portal – capacity overview: status and pause/resume, but no cost Azure Cost Management: cost, but no status Fabric Capacity Metrics app: CU and storage usage, but neither cost nor status Workspace settings: which workspace is assigned to which capacity I wanted one small window that answers three questions at a glance (Is it running? What has it cost this month? Can I pause it now?) and does the pause/resume for me. The result is the Fabric Capacity Cockpit, now open source under the MIT license: 👉 https://github.com/Siebzehnundvier/FabricCockpit-GUI What it does Capacity picker: lists all Fabric capacities across all enabled subscriptions of the signed-in account. The list is cached, so the window is usable immediately while a fresh list loads in the background. Capacity info: name, subscription, resource group, region, SKU, state, provisioning state and month-to-date cost of the selected capacity. Cost split: the cost figure shows the OneLake storage share separately (e.g. 12.34 € (of which storage 0.87 €)), because that part keeps accruing while the capacity is paused. Pause / Resume: each action asks for confirmation, then shows live progress with an elapsed-time counter until the target state is reached. Auto-pause timer: choose 15, 30, 60 or 120 minutes. Two minutes before the deadline, a banner offers Pause now, Extend or Cancel. Tray icon: its colour follows the capacity state (green active, amber paused, blue transitioning). The right-click menu offers Open cockpit / Resume / Pause / Exit. Closing the window only hides it to the tray. Links: Azure capacity overview, cost analysis for the resource group, and your Capacity Metrics app. Log: every action and error is logged with a timestamp and can be copied to the clipboard. Under the hood The cockpit is plain Windows PowerShell 5.1 with WinForms, so it needs no installer, no compiled binaries and no admin rights. All Azure access goes through the Azure CLI: Capacity list, state, suspend and resume use the microsoft-fabric CLI extension (az fabric capacity ...). The cockpit installs the extension automatically on first start. Pause and resume are sent with --no-wait, and the state is then polled every 5 seconds. Without --no-wait, the CLI blocks until the operation finishes, which would freeze the UI. Cost comes from a Cost Management REST query at resource level, grouped by resource ID and meter. Stored-data meters count as storage; everything billed in CU counts as compute. Results are cached for 10 minutes, and throttling (HTTP 429) is handled with a retry. Sign-in and tokens are handled entirely by the Azure CLI. The tool stores no credentials. Its only local data is a settings file in %APPDATA%\FabricCockpit plus cost cache files in %TEMP%. There is no telemetry and no update check. I tested the state transitions (Paused → Resuming → Active and Active → Pausing → Paused) against Azure CLI 2.90.0 with microsoft-fabric 1.0.0b1. Getting started Install the Azure CLI and check that az --version works. Download the release zip from GitHub and unblock it (right-click → Properties → Unblock), then extract it anywhere. Double-click Start-Cockpit.cmd, click Sign in, and pick your capacity. Permissions: Viewing only: Reader on the capacity or resource group. Pause / Resume: additionally suspend/action and resume/action on the capacity, e.g. Contributor. Cost figure: Cost Management Reader. Without it, the field shows n/a and everything else still works. Before you pause anything The cockpit issues the same calls as the Azure portal, so the consequences are the same: Pausing makes every workspace on the capacity unavailable. Scheduled refreshes and pipeline runs fail while it is paused. Running work is aborted. The auto-pause timer does not check for running refreshes, notebooks or Spark jobs. Use it only on personal, dev or test capacities. Resuming starts billing immediately. OneLake storage is billed while paused. Reservations are not detected. If your capacity is covered by a reservation, pausing saves nothing. The cost figure is approximate. Azure cost data lags behind, typically can lag by a day or more. For invoices, use the portal. Trial capacities and Power BI Premium P-SKUs are not Azure resources and therefore don't appear in the list. How it was built I built the cockpit together with an AI assistant and verified the behaviour on my own capacity. The code is plain text and deliberately simple, so please review it before you run it in an environment with stricter policies. Feedback welcome This is version 0.52, a tool for my own daily use that I think others with small pay-as-you-go capacities might find handy. Issues, ideas and pull requests on GitHub are very welcome. I'd also be interested to hear how you manage pause/resume and cost tracking for your own capacities: automation runbooks, Logic Apps, or something else entirely?Row-Level Security in Direct Lake Models: The Undocumented Gotchas
When we moved our sales reporting from Import mode to Direct Lake, RLS looked like the easy part. We already had dynamic RLS working in the old Import model: a security table, a USERPRINCIPALNAME() filter, and a couple of relationships. We expected to copy it over and be done in an afternoon. It took two weeks. The RLS logic itself was fine. What caused the trouble was how Direct Lake interacts with workspace permissions, the SQL analytics endpoint, framing, and fallback. Most of it is technically documented, but spread across a dozen pages, and a few behaviors we only found by testing. This post covers what we hit, how we diagnosed it, and what we'd do differently next time. The setup Here's a simplified version of what we built: Lakehouse: LH_Sales with fact_sales, dim_region, dim_product, dim_date Security table: sec_user_region (UserEmail, RegionKey), one row per user per region Semantic model: Direct Lake on the SQL analytics endpoint Users: about 400 sales reps and managers, each seeing only their regions The RLS role was the standard dynamic pattern: // Role: RegionSecurity // Table: sec_user_region [UserEmail] = USERPRINCIPALNAME() The relationship sec_user_region[RegionKey] → dim_region[RegionKey] had "Apply security filter in both directions" enabled, and dim_region filtered fact_sales. It worked when I tested it as myself, and it failed in several different ways once real users got access. Gotcha #1: Workspace roles silently bypass your RLS This is the most common one, and it matters more in Direct Lake. Semantic model RLS only applies to users who have Read permission on the model. Anyone with Admin, Member, or Contributor in the workspace has write permission, so RLS does not apply to them. They see everything. That's the same as Import mode. The Direct Lake-specific problem is the next part. The Viewer role has a hole too. A workspace Viewer is subject to your semantic model RLS, but Viewers can also connect to the lakehouse's SQL analytics endpoint and query fact_sales directly from SSMS, Excel, or a notebook. Your DAX roles don't exist there, so they can read every row. We found this when a regional manager emailed us a pivot table with national numbers. He had connected Excel to the SQL endpoint because he found it faster. The relationship sec_user_region[RegionKey] → dim_region[RegionKey] had "Apply security filter in both directions" enabled, and dim_region filtered fact_sales. It worked when I tested it as myself, and it failed in several different ways once real users got access. What we changed: Removed every business user from workspace roles. Shared the report and semantic model through an App (or item-level sharing) with Read permission only, without "Build" unless it was genuinely needed. Did not grant lakehouse access to end users at all (Gotcha #2 explains how that still works). Rule of thumb: If a user can see the lakehouse, assume they can see all of it unless you've also secured it at the SQL/OneLake layer. Gotcha #2: SSO vs. fixed identity changes who needs lakehouse access By default, a Direct Lake model on the SQL endpoint uses single sign-on (SSO): the viewer's own identity is used to access the underlying data. That means every report viewer needs permission to read the lakehouse, which brings back the problem from Gotcha #1. The fix is to configure the model's data source connection to use a fixed identity through a shareable cloud connection (a service principal or a dedicated account). Then: The fixed identity reads the Delta tables. End users only need Read on the semantic model. Your DAX RLS still applies per user, because USERPRINCIPALNAME() still returns the viewer, not the fixed identity. The catch: once you switch to fixed identity, any RLS you defined on the SQL analytics endpoint (T-SQL security policies) is evaluated against the fixed identity, not the end user. Warehouse-level RLS effectively stops being per-user for this model. So pick one layer as the source of truth: Approach Security lives in Connection Users need lakehouse access? A Semantic model (DAX roles) Fixed identity No B SQL endpoint (T-SQL RLS) SSO Yes We chose A. Mixing both led to confusing results where the same user saw different numbers depending on the path the query took. Gotcha #3: SQL endpoint RLS quietly forces DirectQuery fallback We initially tried approach B, because the data engineering team liked having security defined once in T-SQL: CREATE FUNCTION sec.fn_region_filter(@RegionKey INT) RETURNS TABLE WITH SCHEMABINDING AS RETURN SELECT 1 AS allowed FROM dbo.sec_user_region s WHERE s.RegionKey = @RegionKey AND s.UserEmail = USER_NAME(); CREATE SECURITY POLICY sec.RegionPolicy ADD FILTER PREDICATE sec.fn_region_filter(RegionKey) ON dbo.fact_sales WITH (STATE = ON); Security worked. Performance got much worse. With Direct Lake on the SQL endpoint, any table with RLS (or OLS) defined at the SQL endpoint can't be read directly from Delta, because VertiPaq can't enforce the T-SQL policy. Queries against it fall back to DirectQuery. Our main visuals went from about 300 ms to 4–8 seconds. It fails quietly. Nothing errors. Reports just get slow. How to detect it: Run a trace in DAX Studio or SQL Profiler against the XMLA endpoint and look for DirectQuery Begin/End events. In pure Direct Lake you should see only VertiPaq scan events. In Performance Analyzer in Desktop, a "Direct query" duration on a visual is a clear sign of fallback. How to make it fail loudly instead: set the model's DirectLakeBehavior property (via Tabular Editor or TMDL) to DirectLakeOnly: model Model directLakeBehavior: directLakeOnly Now fallback-triggering queries error instead of silently degrading. We set this in dev and test so we find these issues before users do. In production you may prefer Automatic so reports keep working, but you should make that choice deliberately. Note: Direct Lake on OneLake (the newer flavour) reads Delta directly and doesn't use the SQL endpoint for queries, so SQL endpoint RLS isn't applied there at all. That's another reason not to rely on T-SQL RLS for Direct Lake models. Check the current docs on how OneLake security interacts with it, because that area is still changing. Gotcha #4: Your security table can't be a view or a calculated table In Import mode, our security table was a calculated table that unpivoted a hierarchy of managers and regions using DAX. In Direct Lake (on SQL endpoint), calculated tables and calculated columns over Direct Lake tables aren't supported. The obvious next step is a SQL view. But SQL views in a Direct Lake on SQL endpoint model always fall back to DirectQuery, because they aren't Delta tables. Your security table sits in the filter path of every query, so it pulls every query into DirectQuery with it. What worked: materialise the security table as a real Delta table in the lakehouse, rebuilt by a notebook in the same pipeline that loads the facts: from pyspark.sql import functions as F sec = ( spark.table("stg_user_access") .withColumn("UserEmail", F.lower(F.trim("UserEmail"))) .select("UserEmail", "RegionKey") .dropDuplicates() ) sec.write.mode("overwrite").format("delta").saveAsTable("sec_user_region") Keep it narrow (two columns) and deduplicated. The next gotcha explains why the lower() matters. Gotcha #5: Case sensitivity changes between Direct Lake and fallback This one took us a day to track down. A user could see data in one report page but got blanks on another. Same model, same role. The cause: USERPRINCIPALNAME() returned [email protected]. The security table had [email protected]. In Direct Lake (VertiPaq), string comparison is case-insensitive, so it matched. One page had a visual that fell back to DirectQuery (it hit a guardrail). In DirectQuery, the RLS filter is pushed down as SQL, and the Fabric SQL endpoint's default collation is case-sensitive, so it didn't match. Same user, same role, different result depending on the query path. Fix: normalise on both sides. // Role filter on sec_user_region [UserEmail] = LOWER ( USERPRINCIPALNAME () ) Also lowercase the data during ingestion (see the notebook above). Don't rely on the engine's collation. Gotcha #6: RLS changes don't take effect until the model reframes We removed a departing manager from sec_user_region, the notebook ran, and the Delta table was updated. He could still see his old regions two hours later. Direct Lake doesn't read "live" Delta. It reads the version of the table captured at the last framing operation (a refresh). If you've turned off "Keep your Direct Lake data up to date" (common when you want facts and dimensions to update together), the model keeps using the old security table until the next refresh. Security changes follow your refresh schedule, not your data load. What we do now: The pipeline that rebuilds sec_user_region ends with a semantic model refresh activity, so the model reframes immediately. For urgent revocations such as terminations, we trigger an on-demand refresh and don't wait for the schedule. We added a small "Security last refreshed" card to the report, driven by a LastUpdated column in the security table, so support can check it quickly. Gotcha #7: "Test as role" isn't the same as a real user Testing with Security → Test as role in the service, using "Other user" and typing an email, is useful but has limits: Under SSO, data access still happens with your identity. You may see rows the real user can't access at the lakehouse level, or the reverse. It doesn't show you workspace-role bypass (Gotcha #1), because you're testing the role, not the user's actual permissions. We now add a hidden debug page to every RLS model: Debug_UPN = USERPRINCIPALNAME () Debug_RegionCount = COUNTROWS ( VALUES ( dim_region[RegionKey] ) ) Debug_SecRows = COUNTROWS ( sec_user_region ) For UAT, real test accounts (not admins) open the app and screenshot this page. It shows exactly what UPN the engine sees, which also helped with a B2B guest user whose UPN didn't match the email in our HR feed. Always compare against what USERPRINCIPALNAME() actually returns, not what you expect it to return. Gotcha #8: Large security tables affect memory and cold-cache performance Direct Lake loads columns into memory on demand. The security table and its relationship columns are touched by every query for every user under that role. Our first version of sec_user_region expanded the org hierarchy down to store level: about 2.1 million rows. After a reframe, the first visual for each user was noticeably slow because the security columns had to be transcoded, and memory pressure caused more column evictions under load. Collapsing it to region level (about 3,800 rows) and moving store-level granularity into the dimension fixed the problem. Tips: Filter at the highest grain that meets the business requirement. Avoid high-cardinality string keys in the security relationship; use integer keys. Keep bidirectional filtering limited to the one security relationship, not across the whole model. Consider a warm-up query (a scheduled DAX query after refresh) for large models so the first user of the morning doesn't pay the cold-cache cost. Our checklist now Before any Direct Lake model with RLS goes to production, we check: End users have no workspace role and access through an App or item sharing only End users have no direct lakehouse / SQL endpoint access Connection uses a fixed identity (if security lives in the model) No T-SQL RLS/OLS on tables used by the model (or fallback is intentional) Security table is a physical Delta table, not a view or calculated table Emails are lowercased on both sides of the comparison DirectLakeBehavior = DirectLakeOnly in dev/test to surface fallback Security table rebuild triggers a model refresh Debug page verified by real non-admin test accounts, including guests Security table kept small, integer-keyed, and at the coarsest grain possible Final thoughts The RLS logic in Direct Lake is the same as in Import mode, so it's easy to assume nothing else changes. What does change is the environment around it: permissions are shared with the lakehouse, security can exist in two layers, data only updates when the model reframes, and queries can fall back to DirectQuery without warning. Most of our problems came from those areas, not from DAX. If you're planning a migration, test RLS with real user accounts early, run a trace to check for fallback, and decide at the start which layer owns security. Have you run into other RLS issues with Direct Lake? Please share them in the comments. I'm especially interested in how people are handling OneLake security alongside model RLS as that feature matures.An introduction to Fabric Apps
One of the shiny new items in Microsoft Fabric is Fabric Apps, and these open up a bunch of possibilities for what is possible in Fabric. Fabric Apps enabled functionality that previously would have needed to be custom workloads, which are a lot more complex to get up and running than Fabric Apps. What is a Fabric App? A Fabric app (built on the Rayfin SDK) is a Fabric item with two components, a backend service for data access and authentication, and a front end web app. When a Fabric App is created, Fabric automatically creates a SQL Database, an authentication server, and static content hosting for the front end web app. You provide data models for your application in TypeScript and Fabric Apps automatically creates a database schema and type-safe GraphQL API. Before you can create one, a tenant admin has to enable it. It's still in preview, so it needs to be explicitly enabled, and Fabric Apps are not available in all regions. Fabric Apps are currently available in the following regions (As of August 18th, 2026): US - Central US US - North Central US US - West US US - West US 2 Europe - West Europe France Central Italy North Norway East Switzerland North UAE North South Africa North Asia - East Asia Asia - Southeast Asia Australia East India - Central India Japan East Korea Central To enable the preview in your tenant: 1. Sign in to the Fabric admin portal (https://app.fabric.microsoft.com/admin-portal). 2. Go to Tenant settings. 3. Under Enable Fabric App Items (preview), toggle it to Enabled, scoped to your whole org or specific security groups. 4. Select Apply. Give it a few minutes to propagate. If you don't see App (preview) in your New item list, either this is why or your capacity is in an unsupported region. Part 1: Create and Deploy the Sample App Create the item in Fabric Creating the item itself is the same as any other Fabric item: 1. Open a workspace where you have contributor or higher access. 2. Select New item. 3. Search for App (preview), select it, give it a name, and select Create. This creates the item, and all the backend services I mentioned before. From here, you can select a blank app, or start with a sample To-Do App, or a sample Data App. For the sake of this post, I am going to select the To-Do App. Once selected, Fabric will start to deploy your app and present you with some instructions for getting the Rayfin CLI up and running. Create the project with npm From a terminal, run the command shown in the Fabric App: npm create Microsoft/rayfin@latest -- "<appitemname>" --template todoapp --workspace <workspacename> That one command creates a full project from the `todoapp` template and wires it to the workspace and item you just created, using the Rayfin CLI. Then change your working directory to the project directory that was just created cd <your project directory> Run it locally npm run dev This spins up the frontend and backend together against your Fabric backend, so you can test changes before anything goes live. By default it runs at `http://localhost:5173`. Deploy with npx npx rayfin up Under the hood, `rayfin up` does six things in order: 1. It creates (or reuses) the Fabric App item 2. It retrieves the publishable key 3. It syncs your `rayfin.yml` settings 4. It applies the database schema from your TypeScript models 5. It builds and deploys static content, 6. Finally, it writes the deployment details back to `rayfin.yml` and a `.env.fabric-<workspacename>` file. When it finishes, you get a live hosting URL, a Fabric portal link, and a deployment ID. Tip: Want to check what a deploy will do without actually running it? Use `npx rayfin up --dry-run`. To check current deployment state at any time, use `npx rayfin up status`. Part 2: The Database Objects Here's the part that surprised me the most coming from a data background: there's no SQL to write, and no separate database designer to open. Your data models are TypeScript classes, and Fabric Apps turns them directly into database tables. Defining an entity Entities or tables live in `rayfin/data/` and use the `@entity()` decorator from `@microsoft/rayfin-core`, as well as field decorators for each column: Each table needs to be defined in it's own typescript file. import { entity, uuid, text, boolean, date } from '@microsoft/rayfin-core'; @entity() export class Todo { @uuid() id!: string; text() title!: string; text({ optional: true }) description?: string; @boolean({ default: false }) isComplete!: boolean; date() createdAt!: Date; date() updatedAt!: Date; } Every entity gets a UUID `id` primary key. If you don't include it in your entity definition, Fabric will automatically add it. When records are inserted into your table, Fabric will generated a UUID server-side unless you supply your own. Composite keys and custom key names aren't supported. The full set of field decorators: `@uuid()`, `@text()`, `@int()`, `@decimal()`, `@boolean()`, `@date()`, `@email()`, and `@set()` for enumerated strings. Modifiers like `{ optional: true }`, `{ unique: true }`, `{ default: value }`, and `{ min, max }` add constraints to the columns. Note: The TypeScript `?` optional marker only affects the compile-time type. To actually make a database column nullable, you need `{ optional: true }` in the decorator itself. Relationships If your app has more than one table, use `@one()` and `@many()` to define relationships, and Fabric auto-generates the foreign key column for you following a `{property}_id` convention. One-to-many and many-to-one are supported; many-to-many is not, so model it with an explicit join table instead. Registering the schema Once every entity has been created, they then need to be added to `rayfin/data/schema.ts`: import { Todo } from './Todo.js'; export type TodoAppSchema = { Todo: Todo; }; export const schema = [Todo]; Applying schema changes Whenever you add or edit an entity: npx rayfin up db apply If a change would drop a column or rename a table, the CLI blocks it and warns you first. You can override with `--force`, but know that as soon as you add --force, you will be making a destructive change that cannot be undone. Warning: Using `--force` on a schema apply can cause data loss and cannot be undone. After being deployed, you are able to see the database in your workspace under the Fabric App. The SQL Database child item will allow you to access the Fabric SQL query editor where you can run SQL queries to read data. Don't change things in SQL here, as anything you change will be overwritten the next time you deploy the schema. Wrapping Up This post covers creating the sample app, deploying it, and how the database objects work. The full version, including how authentication (local email/password vs. Fabric SSO) and the front end (the RayfinClient data calls) work, is on the blog, linked below. Microsoft Learn references: Create your first Fabric app (https://learn.microsoft.com/en-us/fabric/apps/create-app) | Fabric Apps project structure (https://learn.microsoft.com/en-us/fabric/apps/project-structure) | Define data models for Fabric Apps (https://learn.microsoft.com/en-us/fabric/apps/data-models) | Deploy a Fabric app to Fabric (https://learn.microsoft.com/en-us/fabric/apps/deploy-app) Read more: Read the full post, including the Authentication and Front End sections, on the Fabric Field Notes blog: https://www.fabricfieldnotes.ca/blog/2026/08/18/an-introduction-to-fabric-apps296Views2likes0CommentsMicrosoft Fabric Reservations: Understanding What They Are and How to Create Them
Overview Over the past year, I’ve had the incredible opportunity to help Microsoft customers unlock the potential of Microsoft Fabric Capacities (FSKUs) by transitioning from Power BI Premium Capacities (PSKUs) or as they embark on their Fabric journey. During these engagements, we’ve tackled not just technical features, architecture design, deployment patterns, capacity planning, but also business considerations such as reservation purchasing strategies and scoping options to demystify how reservations work. Common questions I’ve frequently encountered are how to confidently navigate the reservation purchasing process, understand the scoping flexibility, and calculate consumption units effectively to maximize both value and efficiency. My hope is that this article will make your voyage smoother, offer clarity on these decisions, and empower you to make the most of your reservations when stepping into the world of Microsoft Fabric capacity planning!14KViews8likes2CommentsRebranding 40 Reports by Hand? Fabric's New Report JSON Functions Fix That
Ever had to open dozens of published reports one by one just to update a company name, a logo reference, or a text box label? Fabric now has official functions to pull a report's layout as JSON, make the change once, and push it back - without opening Power BI Desktop at all. Here's what that looks like with a real example.Governing the Flow: A Beginner’s Guide to Microsoft Fabric
“Governance is not about restricting access; it’s about providing the right access to the right people at the right time, while ensuring the data remains accurate and secure.” In the world of data, speed is often the enemy of stability. Microsoft Fabric is a powerhouse that allows data to move seamlessly from ingestion to AI, but without structure, that speed can quickly turn into “data sprawl.” Governance is the safety net that allows your team to innovate faster without fearing a security breach or a broken pipeline.90Views0likes0Comments