Forum Discussion
Machine learning pipelines in Microsoft Fabric
- Anonymous2 years ago
hsn367 wrote:
I have a background in building machine learning pipelines in AzureML using AzureML SDK, this really helps us in orchestrating the end to end data science workflows. In workflows in our organization, we have ML pipelines written using AzureML SDK and then we have CI/CD pipelines that are supposed to publish these ML pipelines to AzureML studio.
Now moving into Fabric with this background, I have a couple of questions that I did not get answers to when going through the documentations.
1) How to orchestrate the data science workflow in Fabric. For instance we have multiple scripts for our end-to-end solution, we can easily build pipelines over it using AzureML SDK in AzureML studio but in fabric what is the alternative, how are we suppose to build ML pipelines?
2) Data drift monitoring in an important component of end-to-end data science solution, we can monitor drift of the model's data in AzureML but what is the alternative available in Fabric?
hsn367
An additional reply from the internal team for the above questions
We have pipelines in Fabric in the form of Data Factory, and you can run Notebooks with ML activities/code as part of those. Overall, we are working on strengthening our MLOps story. We have Model endpoints in PrPr and working on providing a better SDK. If you look for running MLOps in Production today, we recommend using AzureML with Fabric. AzureML has access to data in OneLake and working on improving that integration. Over time Fabric will become more complete on MLOps too, for data centric and analytics workloads. We focus on scenarios where you serve data to PowerBI today. And we are evolving into other scenarios gradually, like real time model endpoints for example.
We don't yet have drift monitoring in Fabric. On the roadmap. - 2 years ago
AnonymousThank you so much for all the support.
1. Orchestrating ML Workflows in FabricSince Fabric does not use the AzureML SDK for pipelines, you must reconstruct your workflow using the following integrated pieces:Fabric Data Pipelines: Use these to coordinate the sequence of your tasks. They replace the overall orchestration function of AzureML pipelines.Fabric Notebooks: Put your machine learning scripts inside PySpark or standard Python Notebooks.Activity Orchestration: Within a Data Pipeline, add a "Notebook Activity" to trigger your scripts in a specific order (e.g., Data Prep \(\rightarrow \) Feature Engineering \(\rightarrow \) Training \(\rightarrow \) Scoring).OneLake / Lakehouse: Use the built-in Microsoft Fabric Lakehouse to store your raw, clean, and processed datasets centrally.MLflow: Use Fabric's built-in MLflow integration to track your experiment runs, parameters, and model versions.2. Managing Data Drift MonitoringData drift means that the data entering your system today looks different from the data you used to train your model in the past. Fabric does not have an automatic button for this yet. You can handle it in two ways:The Code-Based Approach: Write a script inside a Fabric Notebook using open-source libraries like Evidently AI, Great Expectations, or Alibi. Run this notebook on a schedule using a Fabric Data Pipeline to compare your new production data against your baseline training data.The Hybrid Approach: Continue using Azure Machine Learning alongside Fabric. You can keep your data in Fabric's OneLake but route the model monitoring and governance workloads back to Azure Machine Learning to leverage its advanced MLOps tools.