Forum Discussion
Machine learning pipelines in Microsoft Fabric
- Anonymous2 years ago
hsn367 wrote:
I have a background in building machine learning pipelines in AzureML using AzureML SDK, this really helps us in orchestrating the end to end data science workflows. In workflows in our organization, we have ML pipelines written using AzureML SDK and then we have CI/CD pipelines that are supposed to publish these ML pipelines to AzureML studio.
Now moving into Fabric with this background, I have a couple of questions that I did not get answers to when going through the documentations.
1) How to orchestrate the data science workflow in Fabric. For instance we have multiple scripts for our end-to-end solution, we can easily build pipelines over it using AzureML SDK in AzureML studio but in fabric what is the alternative, how are we suppose to build ML pipelines?
2) Data drift monitoring in an important component of end-to-end data science solution, we can monitor drift of the model's data in AzureML but what is the alternative available in Fabric?
hsn367
An additional reply from the internal team for the above questions
We have pipelines in Fabric in the form of Data Factory, and you can run Notebooks with ML activities/code as part of those. Overall, we are working on strengthening our MLOps story. We have Model endpoints in PrPr and working on providing a better SDK. If you look for running MLOps in Production today, we recommend using AzureML with Fabric. AzureML has access to data in OneLake and working on improving that integration. Over time Fabric will become more complete on MLOps too, for data centric and analytics workloads. We focus on scenarios where you serve data to PowerBI today. And we are evolving into other scenarios gradually, like real time model endpoints for example.
We don't yet have drift monitoring in Fabric. On the roadmap. - 2 years ago
AnonymousThank you so much for all the support.
Hi hsn367
Thanks for using Fabric Community.
Transitioning from AzureML to Microsoft Fabric involves adapting to the tools and services that Fabric offers for machine learning and data science workflows.
Orchestrating Data Science Workflows in Fabric:
1) In Fabric, you can use Fabric notebooks for data science scenarios, which allow you to ingest data into a Fabric lakehouse using Apache Spark, load existing data from delta tables, and clean and transform data using Apache Spark and Python-based tools.
2) You can create experiments and runs to train different machine learning models within these notebooks.
3) For orchestrating workflows, you can construct data analytics workflows with Fabric Data Factory data pipelines, which provide a low-code solution for data integration and ETL projects.
4) The Data Factory in Fabric allows you to build automated workflows that combine different artifacts in your workspace, such as files, notebooks, and dataflows, to create an end-to-end data analytics workflow.
Please refer to these links:
Data science tutorial - get started - Microsoft Fabric | Microsoft Learn
Construct a data analytics workflow with a Fabric Data Factory data pipeline | Microsoft Fabric Blog | Microsoft Fabric
Data Drift Monitoring in Fabric:
1) Fabric doesn't have a built-in data drift monitoring tool like AzureML. However, you can leverage various options for drift detection
2) Monitoring in Fabric is centralized through the Monitoring hub, which enables users to monitor Fabric activities, including data pipelines, dataflows, lakehouses, notebooks, and semantic models.
3) While specific features for data drift monitoring like those in AzureML may not be directly mentioned, the Monitoring hub provides a comprehensive view of all activities and could be used to track changes and performance over time.
Use the Monitoring hub - Microsoft Fabric | Microsoft Learn
Hope this helps. Please let me know if you have any further questions.