Forum Discussion

codeautomation's avatar
codeautomation
New Member
9 days ago

Challenges in Deploying Machine Learning Models into Production Workflows

Hi everyone,

 

I am exploring how teams are moving machine learning models from experimentation into reliable production workflows.

 

In many projects, the challenge is not only training the model but also managing the complete lifecycle:

 

- Preparing and maintaining quality datasets

- Tracking experiments and model versions

- Deploying models for real-world usage

- Monitoring performance after deployment

- Handling model updates over time

 

I would like to understand how the community is approaching this with Microsoft Fabric.

 

A few questions:

 

What are the recommended patterns for managing ML model lifecycle in Fabric?

 

How are teams handling model versioning and experiment tracking?

 

Are you using Fabric notebooks, MLflow, or external platforms for managing production ML workflows?

 

What challenges have you faced when moving data science projects from development to production?

 

Would love to hear practical experiences and approaches from the community.

2 Replies

  • v-achippa's avatar
    v-achippa
    Icon for Community Support rankCommunity Support

    Hi codeautomation​,

    Thank you for reaching out to Microsoft Fabric Community.

    Thank you ShivekMaharaj​ for the prompt response.

    As we haven’t heard back from you, we wanted to kindly follow up to check if the solution provided by the user for the issue worked? or let us know if you need any further assistance.

    Thanks and regards,
    Anjan Kumar Chippa

  • ShivekMaharaj's avatar
    ShivekMaharaj
    Icon for Memorable Member rankMemorable Member

    Hi codeautomation​,

    For Fabric, I would separate the ML lifecycle into experiment tracking, model promotion, inference and monitoring rather than treating them as one deployment step.

    A pattern I would use is:

    Training data in Lakehouse
    -> Fabric notebook / training job
    -> MLflow experiment
    -> registered Fabric ML model/version
    -> validation/promotion gate
    -> batch or real-time inference
    -> monitoring and retraining

    Fabric has fairly strong native support for the first part. MLflow 3 in Fabric can track runs, parameters and metrics, and the newer LoggedModel entity links a model back to its source run, datasets, configuration and environment. A selected model can then be registered as a version of a Fabric ML model.

    For batch production workloads, I would normally use Fabric's PREDICT capability from a notebook/pipeline and write the scored results back to Delta tables.

    For real-time inference, Fabric now has ML model endpoints, but those are currently Preview and support a limited set of model flavors. I would therefore validate those limitations before using them for a production API with strict serving requirements.

    One MLOps detail that is easy to miss is CI/CD.

    Fabric now supports Git integration and deployment pipelines for ML experiments and models, but that integration is currently Preview, and Microsoft documents that experiment runs and model versions themselves are not promoted through the deployment pipeline. Only supported artifact metadata is synchronized.

    So I would not assume that moving a Fabric workspace from Dev -> Test -> Prod automatically promotes the trained model version. I would make model promotion an explicit controlled step.

    For production monitoring, I would also separate two things:

    • operational monitoring: training/scoring failures, endpoint traffic, latency, job health
    • model monitoring: data/feature drift, prediction distribution, ground-truth accuracy and business KPI degradation


    Fabric has built-in ML experiment/model monitoring capabilities, but those are also currently Preview, so I would persist the important model-quality signals myself rather than relying only on the UI.

    The production flow I normally aim for is:

    candidate model
    -> automated evaluation
    -> compare against approved production model
    -> promotion threshold / approval
    -> deploy or score
    -> capture prediction + outcome telemetry
    -> retrain when schedule or degradation criteria are met
    -> retain the previous approved version for rollback

    I would use Fabric notebooks + MLflow for most of this when the data and analytics platform is already in Fabric. I would bring in an external serving/MLOps platform when requirements such as unsupported model runtimes, stricter deployment controls, networking, or production serving guarantees exceed what the current Fabric-native capabilities provide.

    AI-assisted drafting: AI was used to help structure and phrase this response. I reviewed and validated the technical content before posting.