data science
96 TopicsEnable Interactive Editing and Fabric PySpark Execution for Git-Synced .Notebook Folders in VS Code
Allow *.Notebook/notebook-content.py from Fabric Git repositories to open in VS Code’s interactive notebook editor. Permit users to select a Fabric workspace, lakehouse and environment as runtime context without requiring the Git branch to be synchronized into that workspace.6Views1like0CommentsAllow us to rename fabric data agents published to m365
If we use deployment pipelines to promote fabric data agents between dev, test/UAT, and prod fabric workspaces, we need to keep the name of the fabric agent the same in each workspace. If we want to publish the data agent in both UAT and prod to m365, they're would be two agents with the same name. It would be helpful to be able to rename the published agent name in M365 to <agent>UAT when published so users can easily distinguish between the UAT and prod agents.208Views0likes2CommentsExpose Power Query as a Standalone Runtime and Command-Line Interface (CLI)
Summary Power Query has evolved into one of Microsoft's most powerful data transformation technologies and is now used across Excel, Power BI, Fabric, Dataflows, Power Platform, and other products. However, Power Query can only be executed through a host application, despite the existence of a mature M language and execution engine. I would like Microsoft to expose the Power Query engine as a first-class standalone runtime and provide an officially supported command-line interface (CLI) and API. The Problem Today, Power Query transformations are often embedded inside: Excel workbooks Power BI Desktop files Fabric Dataflows Power Platform Dataflows While this works well for interactive users, it creates challenges for enterprise-grade automation. Many organizations would like to: Schedule Power Query transformations without opening Excel. Run Power Query from PowerShell scripts. Integrate Power Query into CI/CD pipelines. Execute transformations on servers without Office dependencies. Reuse Power Query code across multiple solutions. Treat Power Query as a reusable transformation layer rather than a workbook artifact. Currently, Power Query feels like a language without an officially supported runtime, even though the engine already powers multiple Microsoft products.42Views0likes2CommentsExpose Data Agent schema, question interpretation, and generated query through MCP
I would like to request additional context and metadata capabilities for the Microsoft Fabric Data Agent MCP server. It would be very useful if the MCP server could expose more information about the Data Agent itself and how it processes a user's question. 1. Expose the Data Agent data schema It would be useful to retrieve through MCP the data schema available to the Data Agent, including information such as: Tables or semantic model entities Columns and fields Data types Relationships Measures Descriptions and metadata Other relevant schema information available to the Data Agent Ideally, this information could be exposed through a dedicated MCP resource or tool. This would allow an MCP client to understand what data the Data Agent has access to without having to maintain a separate copy of the schema. 2. Expose the interpreted or reformulated question It would also be very useful to retrieve information about how the Data Agent understood the user's question. For example: Original question: "How many customers did we lose last month?" Reformulated / interpreted question: "Calculate the number of customers who became inactive during the previous calendar month." If available, it would also be useful to expose structured information such as: Original user question Reformulated or normalized question Detected intent Relevant tables or entities Relevant columns or measures Filters and conditions identified from the question This would help developers understand how the Data Agent interpreted the request. 3. Expose the generated and executed query It would also be very valuable to expose the query generated by the Data Agent to answer the user's question. Depending on the underlying data source, this could include: SQL DAX KQL Other query languages supported by Fabric Data Agents For example, the MCP response or a dedicated MCP resource could expose: Query language: SQL Generated query: SELECT COUNT(*) FROM Customers WHERE ... Or: Query language: DAX Generated query: CALCULATE(...) It would also be useful to know: Which query language was selected The generated query The final query that was actually executed, if it differs The data source or semantic model targeted The tables, measures, or entities used by the query Execution status Query execution duration, when available Query errors or warnings, when applicable This information would be extremely useful for debugging, auditing, observability, and explaining how the Data Agent produced its answer. Why this would be useful Ideally, an MCP client could retrieve a structured execution context such as: Original question → Reformulated / interpreted question → Relevant schema and entities → Selected query language (SQL / DAX / KQL / etc.) → Generated query → Executed query → Result → Final Data Agent answer The goal is not to expose the model's private internal reasoning, but rather to expose the structured interpretation, generated query, schema context, and execution metadata that can safely be shared with the MCP client. These capabilities would make the Fabric Data Agent MCP server much more useful for building transparent, debuggable, auditable, and context-aware applications. Could Microsoft please consider exposing the Data Agent schema, structured question interpretation, and generated/executed queries through the MCP server?12Views1like0CommentsReal-time progress streaming for Fabric Data Agent MCP server
I would like to request true real-time progress streaming for the Microsoft Fabric Data Agent MCP server. Currently, when calling a published Fabric Data Agent through MCP using call_tool with a progress_callback, I receive progress events such as: analyze.database.fewshots.loading analyze.database.nl2code analyze.database.execute Message created However, these events are usually not delivered during the actual execution. For example, the request can take 15–20 seconds with no updates, and then several progress notifications arrive almost at the same time, immediately before the final response. It would be very useful if the Data Agent MCP server could emit and flush progress notifications incrementally as each step happens. Ideally, applications could display statuses such as: Understanding the question Generating the query Executing the query Processing the result Preparing the final answer This would allow developers to build a much better real-time user experience instead of showing only a loading indicator during long-running requests. The MCP transport already supports streaming and the client can receive progress notifications. The missing part appears to be real-time delivery of the Data Agent execution progress from the server. Could Microsoft please consider supporting true incremental progress streaming for the Fabric Data Agent MCP endpoint ?9Views0likes0CommentsAdd AI Smart Data Validation in Copy Job
I suggest adding an AI Smart Data Validation feature to Data Factory Copy Job. It should automatically detect missing values, duplicate records, data type errors, and mapping issues before copying. This will reduce errors, save time, improve data quality, and make Copy Job easier for all Microsoft Fabric users.22Views0likes0CommentsAI-Powered Data Insights in Power BI
I would like to suggest adding AI-powered data insights directly within Power BI to help users uncover trends, anomalies, and key takeaways faster. Key benefits: • Auto-generate insights and summaries using AI. • Detect anomalies and outliers in visuals. • Natural language explanations for charts and reports. • Save time and improve data-driven decision making. This feature will make Power BI even more powerful and user-friendly for all users.22Views0likes0CommentsAI-Powered ICU Mortality Prediction using MIMIC-IV
This project proposes an AI-powered ICU mortality prediction system using the MIMIC-IV dataset. The system analyzes patient demographics, vital signs, laboratory results, and ICU admission data to identify high-risk patients. Machine learning models such as XGBoost and Random Forest will be used to predict mortality risk. Explainable AI techniques will help clinicians understand the factors influencing predictions. An interactive Power BI dashboard will provide insights into patient risk levels, mortality trends, ICU stay duration, and clinical indicators. The goal is to support early intervention, improve resource allocation, and enhance patient outcomes through data-driven decision making. Technologies: Python, Pandas, Scikit-learn, XGBoost, Power BI, MIMIC-IV Dataset. Expected Outcome: Accurate mortality prediction, interpretable risk analysis, and a healthcare analytics dashboard for ICU patient monitoring.26Views0likes0CommentsDax functions that mimic List.Generate
Functions that could essentially iterate and expression to a suggested limit simliar to list.generate in power query. Iterate(SeedTable, NextState) that would mimic hard coding something like this State0 = SeedTable State1 = NextStep(State0) State2 = NextStep(State1) State3 = NextStep(State2) ... Could use this in dynamic scheduling, BOM explosions, heirarchy traversing, Fibinocci sequence etc. All can be done but require explicit calls to variables. Not true recursion but iterators. Please.47Views1like0Comments