Forum Discussion
How are AI agents changing modern data engineering workflows?
I am exploring how AI agents can assist data engineering teams in managing complex data workflows and reducing repetitive operational tasks.
Some areas I am interested in:
- AI agents for monitoring data pipelines and detecting failures
- Automated data quality checks and anomaly detection
- Intelligent assistance for ETL/ELT workflow optimization
- Generating documentation for datasets and transformations
- Using AI with Fabric pipelines, notebooks, and data workflows
With the growth of AI-powered automation, I would like to understand how data engineers are approaching these patterns in real-world environments.
What approaches, architectures, or best practices are teams using to combine Microsoft Fabric capabilities with AI agents?
2 Replies
A practical approach is to treat AI agents as an intelligence layer around the data platform, rather than letting them directly control everything.
In a Fabric setup, the pattern I’ve seen making the most sense is:
- Fabric Pipelines handle orchestration and dependencies.
- Notebooks / Spark handle transformations and data processing.
- Data quality rules and monitoring detect failures, schema changes, volume anomalies, or unexpected data patterns.
- An AI agent sits on top of these signals, analyzes logs and pipeline metadata, and helps explain the issue or suggest the next action.
- For documentation, the agent can use metadata, schemas, notebooks, and transformation logic to generate and maintain dataset documentation.
For example, if a pipeline fails because of a schema change, the agent shouldn't blindly fix and rerun it. It can identify the failure, compare the new schema with the previous version, explain the impact, and recommend whether the pipeline or source needs attention.
The key best practice is human-in-the-loop automation: let agents investigate, summarize, recommend, and handle low-risk repetitive tasks, while keeping production-impacting changes under proper controls.
I think this approach gives teams the benefits of AI without turning the data platform into a black box.
- v-csrikanth
Community Support
Short version: "AI agent" currently means three different things in Fabric, and picking the right one is most of the answer.
- Copilot helps you author. Generates pipelines, queries and code; you review and commit. GA in Data Factory.
- Fabric data agent, lets people ask questions of governed data. GA, and read-only by design.
- Event-driven automation: job events, Activator, schedules. Not AI at all, but it does most of the real operational work.
The best practice that matters most: build the deterministic plumbing first, then add AI on top for diagnosis and authoring. Teams that start with an AI agent and no event backbone get something impressive in a demo and fragile in production.
Monitoring and failure detection. Start with Fabric job events, fabric raises ItemJobFailed when a job fails, including when it hangs or is cancelled, and Activator can alert or trigger a response. Cheap and predictable. On top of that, the Operations agent (preview) can investigate a failed pipeline run and surface likely root causes. Use it read-only first.
Data quality and anomaly detection. Three routes, depending on where your data sits: Anomaly detection in RTI (preview) for eventhouse tables, Purview Data Quality for lakehouse (separate product, separate licensing), or plain code checks in notebooks which most teams end up doing anyway. AI functions (ai.classify, ai.extract, …) are great for unstructured text, but they are not a statistical anomaly detector. Don't swap a distribution test for a per-row LLM call.
ETL/ELT optimization. Copilot for Data Factory (GA) is the most immediately useful item generating pipelines, explaining an inherited one, troubleshooting.
One trap: "Copilot doesn't produce a message for the skills that it doesn't support… it doesn't give an error message either." Silent no-ops. Always check the applied steps landed.Based on documentation copilot can suggest measure descriptions on semantic models, and there's AI Auto-Summary (preview) in the OneLake catalog. But there's no first-party way to auto-document lakehouse tables, columns or transformation logic. Teams doing this build it themselves: read schema in a notebook, draft descriptions with an AI function, write back via REST API.
The one thing I'd check before planning anything: anomaly detection and the operations agents are eventhouse-scoped today, while most data engineering teams live in the lakehouse. If that's you, these aren't drop-in you would need shortcuts into an Eventhouse first. Worth confirming before you design around them.
Two practical notes. Most of the agent features above are preview, and Microsoft is explicit that outputs "are probabilistic and can be incorrect" keep a human in the loop. And Copilot bills against your capacity per token, with output charged 4× input, so chatty agents get expensive; track it in the Fabric Capacity Metrics app. Minimum is F2, and it isn't available on trial capacities.
What are you trying to automate first failure alerting, quality gates, or documentation? And is your data mostly lakehouse or eventhouse? That answer changes the recommendation a lot.
References
- https://learn.microsoft.com/fabric/real-time-hub/explore-fabric-job-events
- https://learn.microsoft.com/fabric/data-factory/operations-agent-for-pipelines
- https://learn.microsoft.com/fabric/real-time-intelligence/operations-agent-limitations
- https://learn.microsoft.com/fabric/real-time-intelligence/anomaly-detection
- https://learn.microsoft.com/fabric/data-science/concept-data-agent
- https://learn.microsoft.com/fabric/data-science/ai-functions/overview
- https://learn.microsoft.com/fabric/fundamentals/copilot-fabric-data-factory
- https://learn.microsoft.com/fabric/fundamentals/copilot-fabric-consumption
- https://learn.microsoft.com/purview/unified-catalog-data-quality-fabric-lakehouse
Thanks,
C Srikanth
Community Support Team