Forum Discussion
Machine learning pipelines in Microsoft Fabric
- Anonymous2 years ago
hsn367 wrote:
I have a background in building machine learning pipelines in AzureML using AzureML SDK, this really helps us in orchestrating the end to end data science workflows. In workflows in our organization, we have ML pipelines written using AzureML SDK and then we have CI/CD pipelines that are supposed to publish these ML pipelines to AzureML studio.
Now moving into Fabric with this background, I have a couple of questions that I did not get answers to when going through the documentations.
1) How to orchestrate the data science workflow in Fabric. For instance we have multiple scripts for our end-to-end solution, we can easily build pipelines over it using AzureML SDK in AzureML studio but in fabric what is the alternative, how are we suppose to build ML pipelines?
2) Data drift monitoring in an important component of end-to-end data science solution, we can monitor drift of the model's data in AzureML but what is the alternative available in Fabric?
hsn367
An additional reply from the internal team for the above questions
We have pipelines in Fabric in the form of Data Factory, and you can run Notebooks with ML activities/code as part of those. Overall, we are working on strengthening our MLOps story. We have Model endpoints in PrPr and working on providing a better SDK. If you look for running MLOps in Production today, we recommend using AzureML with Fabric. AzureML has access to data in OneLake and working on improving that integration. Over time Fabric will become more complete on MLOps too, for data centric and analytics workloads. We focus on scenarios where you serve data to PowerBI today. And we are evolving into other scenarios gradually, like real time model endpoints for example.
We don't yet have drift monitoring in Fabric. On the roadmap. - 2 years ago
AnonymousThank you so much for all the support.
Hi Anonymous
Thank you so much for the detailed response. So here is what I got from your response.
1) AzureML pipelines alternative available in Fabric is Fabric Data Factory pipelines where we can orchestrate multiple python scripts of our data science solutions.
2) Data drift monitoring is not available in Fabric yet. You mentioned monitoring hub but that does not fulfil the needs of drift monitoring.
I have a couple of more questions regarding migrating to Fabric coming from AzureML background.
1) In AzureML, we were heavily relying on AzureML SDK's data asset management for data versioning of our data science solution or to version the data produced by the different components of the pipeline. And it was very easy to just use the latest version of the data or to use any previous version it was just a matter of specifying that version name. So migrating to Fabric how you think we can get the similar behavior there.
2) In our current workflows, we have three AML workspaces i.e. a separate one for development, test and prod environments. Now we develop the ML pipelines in dev workspace and then deploy them to test and prod workspaces via CI/CD pipelines. So is it possible to achieve the same behavior in Fabric?
- Anonymous2 years agoNot applicable
Hi hsn367
Data Versioning in Fabric:
Microsoft Fabric does not have the dataset concept as in Azure Data Factory
While Fabric’s approach to data versioning may differ from AzureML SDK’s data asset management, you can achieve similar behavior by leveraging OneLake and the data integration pipelines within Fabric.
You can use Notebooks and also Azure Devops to acheive version control in Fabric.
https://radacad.com/version-control-in-power-bi-and-fabric
https://www.linkedin.com/pulse/unraveling-past-empowering-future-versioning-timetravel-data/
CI/CD Pipeline Deployment Across Workspaces in Fabric:
Fabric’s lifecycle management tools, including Git integration and deployment pipelines, support a standardized system for collaboration and continuous delivery of updated content into production. Deployment pipelines in Fabric allow you to clone content from one stage to another, typically from development to test, and from test to production, maintaining the connections between copied items. You can have similar 3 workspaces in Fabric and achieve the same.
Introduction to the CI/CD process as part of the ALM cycle in Microsoft Fabric - Microsoft Fabric | Microsoft Learn
The Microsoft Fabric deployment pipelines process - Microsoft Fabric | Microsoft Learn
Hope this helps. Please let me know if you have any further questions.