Forum Discussion
How are AI agents changing modern data engineering workflows?
Short version: "AI agent" currently means three different things in Fabric, and picking the right one is most of the answer.
- Copilot helps you author. Generates pipelines, queries and code; you review and commit. GA in Data Factory.
- Fabric data agent, lets people ask questions of governed data. GA, and read-only by design.
- Event-driven automation: job events, Activator, schedules. Not AI at all, but it does most of the real operational work.
The best practice that matters most: build the deterministic plumbing first, then add AI on top for diagnosis and authoring. Teams that start with an AI agent and no event backbone get something impressive in a demo and fragile in production.
Monitoring and failure detection. Start with Fabric job events, fabric raises ItemJobFailed when a job fails, including when it hangs or is cancelled, and Activator can alert or trigger a response. Cheap and predictable. On top of that, the Operations agent (preview) can investigate a failed pipeline run and surface likely root causes. Use it read-only first.
Data quality and anomaly detection. Three routes, depending on where your data sits: Anomaly detection in RTI (preview) for eventhouse tables, Purview Data Quality for lakehouse (separate product, separate licensing), or plain code checks in notebooks which most teams end up doing anyway. AI functions (ai.classify, ai.extract, …) are great for unstructured text, but they are not a statistical anomaly detector. Don't swap a distribution test for a per-row LLM call.
ETL/ELT optimization. Copilot for Data Factory (GA) is the most immediately useful item generating pipelines, explaining an inherited one, troubleshooting.
One trap: "Copilot doesn't produce a message for the skills that it doesn't support… it doesn't give an error message either." Silent no-ops. Always check the applied steps landed.
Based on documentation copilot can suggest measure descriptions on semantic models, and there's AI Auto-Summary (preview) in the OneLake catalog. But there's no first-party way to auto-document lakehouse tables, columns or transformation logic. Teams doing this build it themselves: read schema in a notebook, draft descriptions with an AI function, write back via REST API.
The one thing I'd check before planning anything: anomaly detection and the operations agents are eventhouse-scoped today, while most data engineering teams live in the lakehouse. If that's you, these aren't drop-in you would need shortcuts into an Eventhouse first. Worth confirming before you design around them.
Two practical notes. Most of the agent features above are preview, and Microsoft is explicit that outputs "are probabilistic and can be incorrect" keep a human in the loop. And Copilot bills against your capacity per token, with output charged 4× input, so chatty agents get expensive; track it in the Fabric Capacity Metrics app. Minimum is F2, and it isn't available on trial capacities.
What are you trying to automate first failure alerting, quality gates, or documentation? And is your data mostly lakehouse or eventhouse? That answer changes the recommendation a lot.
References
- https://learn.microsoft.com/fabric/real-time-hub/explore-fabric-job-events
- https://learn.microsoft.com/fabric/data-factory/operations-agent-for-pipelines
- https://learn.microsoft.com/fabric/real-time-intelligence/operations-agent-limitations
- https://learn.microsoft.com/fabric/real-time-intelligence/anomaly-detection
- https://learn.microsoft.com/fabric/data-science/concept-data-agent
- https://learn.microsoft.com/fabric/data-science/ai-functions/overview
- https://learn.microsoft.com/fabric/fundamentals/copilot-fabric-data-factory
- https://learn.microsoft.com/fabric/fundamentals/copilot-fabric-consumption
- https://learn.microsoft.com/purview/unified-catalog-data-quality-fabric-lakehouse
Thanks,
C Srikanth
Community Support Team