data engineering
731 TopicsAdd name and description to Item Schedule
A Fabric item can now have up to 20 schedules attached to it. These schedules can have different purposes, let's say one is for nightly runs and another one is for runs during working hours. If we want to update a schedule via API, we need the schedule ID. To get the ID, we can list the schedules. But, there's no easy way to pick the relevant schedule from the returned list or schedules - because there's no name or description associated with a schedule object. Please add the option to create a name or description for an item schedule. This could also show up in the run log of an item (run was triggered by schedule [Name]) which would provide useful context for the run. It could also be cool if an item could pick up the schedule name as a "TriggeredBy" property and potentially execute conditional logic depending on which schedule triggered the run.3.4KViews21likes5Comments- 9.2KViews139likes15Comments
Convert Classic Lakehouse to Schema-based
I created a lakehouse back in November last year. I believe it was while schema-enabled lakehouses were still in Preview. I understand that currently there isn't a way to convert the classic lakehouse to schema-based. Are there plans to create this feature? I really don't want to create a new lakehouse just so that it can be schema-enabled. Thank you!13Views0likes0CommentsData Pipelines - Run only selected activities
For debugging and testing pipeline activities during development, allow us to select one or multiple activities and run only the selected pipeline activities. For example, I'm working on editing a Notebook, and now I want to run the pipeline with that Notebook, but I don't want to run all the activities in the pipeline because that is time consuming. I'm aware of the Deactivate activity option that already exists, but that requires Editing and Saving the pipeline as well. Instead, I just want to select which activities to run interactively by highlighting those activities and click Run, without needing to Edit/Save the pipeline. An option to 'Run from selected activities' would also be great - select one or more activities, and run the selected activities and all subsequent activities.669Views14likes3CommentsEnable Interactive Editing and Fabric PySpark Execution for Git-Synced .Notebook Folders in VS Code
Allow *.Notebook/notebook-content.py from Fabric Git repositories to open in VS Code’s interactive notebook editor. Permit users to select a Fabric workspace, lakehouse and environment as runtime context without requiring the Git branch to be synchronized into that workspace.6Views1like0CommentsStreamlining Power BI Tenant Inventory Management Using a REST API Connector
Background & The Challenge In one of my projects as a Power BI Administrator, a common requirement was to maintain a comprehensive inventory of tenant assets, including workspaces, reports, datasets, dashboards, and capacity metrics. Historically, my team handled this using a legacy, multi-step extraction process: Data Extraction: Running manual PowerShell scripts to call various Power BI REST API endpoints. Storage: Dumping the raw JSON outputs into a SQL Server database. Analysis: Writing ad-hoc SQL queries against the database whenever we needed specific inventory information. Action: Using this queried data to perform quarterly clean-up tasks (e.g., deleting orphaned workspaces, updating capacity assignments, and removing unused reports). The Pain Point This legacy approach was highly inefficient. Every time we needed updated inventory data, an admin had to validate the PowerShell scripts, manually execute them to refresh the SQL database, and then write or run SQL queries to get actionable insights. The data was often stale, and the process was too reliant on manual intervention. The Proposed Solution To eliminate the manual overhead and the need for external database storage, I designed a streamlined solution: Consolidating the inventory tracking directly into a Power BI report using a Custom Connector. Instead of pushing API data out to SQL via scripts, the solution pulls the data directly into Power BI: Custom Power BI Connector: I developed a custom connector designed specifically to authenticate and call the required Power BI REST API endpoints (Workspaces, Datasets, Reports, etc.). Direct Reporting: I built an end-to-end Power BI report that consumes data directly from this custom connector, visualizing all necessary inventory and capacity metrics in one place. Automated Refresh: I configured an On-Premises Data Gateway connection to allow the dataset to refresh automatically on the Power BI Service. The Value Add Zero Manual Intervention: No more managing or executing PowerShell scripts. The data stays fresh via standard Power BI scheduled refreshes. Self-Service Access: Whenever stakeholders or admins need inventory details for quarterly cleanups, they simply open the shared Power BI report. Actionable Insights: Admins can instantly identify unused assets, capacity bottlenecks, and workspace bloat without writing a single SQL query. Key highlights: 🔐 Azure AD OAuth2 - no client secrets, no passwords 📄 Full inline documentation inside Power BI Desktop 🔄 Auto-pagination for large tenant environments ⚡ Works with both Desktop & On-premises Data Gateway Some Screenshots of POC I successfully built and tested this approach as a Proof of Concept (POC) end-to-end. I am sharing this idea with the Fabric community for others who might be struggling with tenant inventory management and looking for a more automated, native reporting approach. Hopefully, This will be help lot of members who is facing similar type issues if Microsoft team can officially launch this connector so everyone can just use and do whatever they need. Quick Update: I have created the custom conenctor for the Power BI as well as Fabic REST endpoints. As of not it covers... -> 56 Power BI REST API functions (Datasets, Reports, Dashboards, Pipelines, DAX execution & more...) -> 53 Microsoft Fabric functions (Lakehouse, Warehouse, Eventhouse, KQL & more...) Shoutout / Special Thanks: Boston Office D365 User Group www.youtube.com/@BostonOffice365UserGroup Note: AI was used to assist in formatting and refining this post for clarity.16Views0likes0CommentsSubscription Email Format
Can we please tailor that format of the subscription emails. They have a thumbnail as well as a (repeated) larger image. Remove too the 'Microsoft Power BI' reference. Also remove the 'Open report in Power BI' button. Also remove the 'You're receiving this email because . . . ' You have made what might be a useful mechanism into a verbose and clumsy-looking communication. Please could you clean it up as much as possible.9Views1like0CommentsBuilt a free tool to inventory & risk-rate legacy SSIS packages before a Fabric migration
Hi all, Sharing something I built that might save some of you the painful "open every .dtsx in SSDT one by one" phase of an SSIS → Fabric migration. The problem it addresses: most teams sitting on a legacy SSIS estate have dozens to hundreds of undocumented packages and no clear starting point. Before any actual migration work, someone has to answer: what does each package do, which components map cleanly to Fabric, and which need a full rewrite? What it does: Parses .dtsx package XML and extracts control flow tasks, data flow components, connection managers, variables, and precedence logic Maps each component against a Fabric equivalent — Data Flow Task → Dataflow Gen2, Execute SQL Task → Script/Stored Procedure Activity, Script Task → Notebook Activity, Foreach Loop → ForEach Activity, and so on Gives each package a Low/Medium/High risk rating based on concrete signals (Script Task usage, Fuzzy Lookup, deprecated CDC/Attunity components, nesting depth, dynamic connection strings) Flags every component individually as auto-mappable or needs-manual-redesign, not just the package as a whole Across a batch of packages, produces a portfolio rollup: risk tier counts, the blockers showing up across the most packages, deprecated components flagged separately, and a suggested migration order What it's built as: a Claude Skill (Anthropic's Claude AI), with the actual XML parsing done by a plain Python script with zero dependencies — not an LLM guessing at package structure. It's designed to run fully offline, no live Azure/Fabric API access assumed, though it flags anywhere that access would improve accuracy (resolving a dynamic connection string, for example). What it deliberately doesn't do yet: this covers the Inventory/Classify phase only. It won't draft actual Fabric pipeline JSON or claim a package is production-ready — every future draft artifact gets an explicit DRAFT/NEEDS REVIEW label. I'd rather it under-claim than over-claim. Repo's on GitHub with a sample package and example report included, so you can see the actual output before running it against your own packages: https://github.com/HBBH11/ssis-fabric-migration-assistant Genuinely interested in feedback from people who've done these migrations for real — particularly whether the risk heuristics hold up against messier packages than my test case, and whether the Fabric equivalence mappings match what you've found in practice. Happy to take suggestions on what a Phase 2 (draft pipeline generation) should prioritize too.10Views0likes0CommentsFabric REST API endpoint for dataflow transactions should GET *sortable* results for a time-range
Dataflow refresh APIs seem to be the buggiest; the list of transactions on a GET for Fabric-dataflow, are not sorted by the latest. OData query options such as $top, $filter, $orderby do not work with this specific endpoint. Endpoint accepts the parameter but ignores it. Endpoint implements its own pagination mechanism. If GET ../transactions returns 10 rows, I can't assume it will have the latest refresh in those 10. This has been confirmed by running at different times of the day.7Views0likes0CommentsTax on nested-pipeline orchestration patterns
Summary Fabric Data Pipelines' ExecutePipeline activity, when run with waitOnCompletion: true, has a documented but unresolved platform behavior: after the invoked child pipeline actually finishes, the parent pipeline's next activity can sit in a "queued" state for 1-5 minutes before it resumes. This delay is not configurable and is not attributable to capacity, connections, or activity content — it is purely platform dispatch/finalize overhead. This is already discussed in Microsoft Q&A threads, where Microsoft support has acknowledged the symptom as "strange" without offering a resolution or workaround. For any architecture that relies on nested ExecutePipeline calls with waitOnCompletion: true — which is the only way to build a fan-out/dynamic-dispatch orchestration pattern in Fabric today (see "Why this pattern exists" below) — this lag compounds directly with (a) the number of orchestration boundaries in a run and (b) the number of items processed within each boundary, and can end up being the majority of total wall-clock time even when the actual work is trivial. Why this pattern exists (context Microsoft's engineering team may not have visibility into) Fabric's ForEach activity does dynamic parallel iteration, but lane assignment is static: items are divided up front, so one lane can end up processing several slow items while other lanes sit idle once they've exhausted their assigned batch. There is no built-in dynamic work-stealing/pull-based queue primitive. To work around this, we built a custom pattern: an Until loop with N parallel lanes, where each lane independently and atomically dequeues the next available work item from a SQL table (UPDATE TOP (1) ... WITH (UPDLOCK) ... OUTPUT ...) the instant it finishes its current item. This correctly solves the load-imbalance problem ForEach can't. But it requires every single work item to be dispatched as its own ExecutePipeline call with waitOnCompletion: true (so a lane can actually block on and then immediately re-loop after each item) — and in our specific chain, that call passes through two nested levels (an outer pipeline-switch/dispatch pipeline, then the actual work pipeline), so every item pays the finalize lag twice. What we measured A real production batch run (internal reference MasterPipelineRunID = b24c1690-...) where every work item hit a fast-path skip (no actual data movement — a delta pre-check determined nothing had changed) still took ~22 minutes wall-clock, broken down as: Execution Group Entities Actual work Idle gap immediately after 1 (Bronze) 59 4m46s 5m45s 2 (Bronze SQL refresh) 1 20s 1m57s 3 (Silver pre-load refresh) 1 21s 1m49s 4 (Silver) 27 2m51s 3m46s 5 (Silver post-load refresh) 1 16s — (last group) ~8m54s of actual execution vs. ~13m17s of pure idle time between orchestration boundaries — roughly 60% of total runtime is platform overhead, not work, in the best case where nothing needed processing. The size of each gap scales with how many items were in the group that just finished (59- and 27-item groups produce multi-minute gaps; single-item groups still produce ~2-minute gaps), consistent with ForEach/Until performing internal aggregation/cleanup work on exit that scales with iteration count. What we're asking Acknowledge and document the actual mechanism, not just the symptom. Is this a polling interval on the orchestration service side, a queue-drain step, or something else? Right now customers have no way to reason about or budget for this cost because its cause isn't documented anywhere we've found beyond community Q&A threads. A configurable or reduced finalize interval for ExecutePipeline, at least for pipelines invoked at high frequency/short duration (our per-item work often completes in under a minute, so a 1-5 minute finalize tax is a 100-500%+ overhead multiplier on top of real work). A first-class dynamic work-queue primitive (a native "pull next available item" pattern) so customers don't have to build this ourselves via nested ExecutePipeline calls in the first place — ForEach with genuinely dynamic (not pre-assigned) lane scheduling would remove the need for this pattern, and the finalize-lag cost with it, entirely. We're happy to share the underlying pipeline JSON and the exact run telemetry above if it's useful for reproducing this on Microsoft's side.15Views1like0Comments