data agent
25 TopicsFabric Data Agent on Ontology Fails for Simple and Multi‑Entity Queries (GROUP BY/Aggregation Error)
Hi, I’ve created a Fabric Data Agent using an Ontology as the data source, and I’m encountering consistent failures even when asking simple questions related to a single entity. Queries that span or join multiple entities also fail. Below are the details and the error output. Issue Summary When I ask a basic question (even involving only one entity), the agent returns an error. For multi-entity questions, it fails with the same pattern. The error indicates that the generated Ontology query includes invalid aggregation or grouping logic. Specifically, a field is referenced without being part of the GROUP BY clause or wrapped in an aggregation function. The underlying generated query seems to have a syntax or grouping issue and cannot execute. Query Output: Failed to execute step (RAID: 36eb1c04-9e56-4998-89ac-a9a56919482d). Error: Failed to execute Ontology query with error: "The query is invalid. Reason: BadRequest. Resource: Graph query (graphModelId=9231e6f0-87b9-44b4-9a36-5ddbe39b8d78). InternalCode: 42000. Message: syntax error or access rule violation. Cause: data exception; The identifier node_production_plant.plant_id cannot be used, as it is neither part of the GROUP BY nor an aggregation." I have already enabled "Support GROUP BY in GQL" in the Data Agent instructions. What I need help with: Has anyone seen similar Ontology-based Data Agent failures related to GROUP BY or aggregation? Is this a known limitation or bug when using Fabric Data Agent on Ontology models? Any best practices or modeling patterns to avoid such query-generation errors? Are there known workarounds to ensure the agent produces valid Ontology queries? I can share more examples or screenshots if needed. Thanks in advance for any guidance!1.9KViews1like4CommentsData Agent fails to query Ontology
Hello, I’m experiencing an error when using the Data Agent in Microsoft Fabric to query my Ontology. Even when I ask very simple questions, the Data Agent fails every time with the following error: Failed to execute step (RAID: "GUID"). Error: Failed to generate NL2Ontology query with error "{"code":"InternalError","subCode":0,"message":"An internal error occurred.","timeStamp":"2026-04-20T13:00:42.7677176Z","httpStatusCode":500,"hresult":-2147467259,"details":[{"code":"RootActivityId","message":"GUID"},{"code":"Param1","message":"Failed to translate NL query to ontology query."}]}" This was working correctly until last week, but for the past four days it has consistently failed with the error above, without any changes on my side. The Data Agent works perfectly when querying: a Lakehouse (NL2SQL), and a Semantic Model (NL2DAX). The issue occurs only when the Ontology is used as the data source. Also I have "Support GROUP BY in GQL" inside the data agent's instructions. Could someone please help me understand what might have changed recently or how to troubleshoot this issue? Thank you in advance.2.7KViews1like7CommentsMicrosoft Fabric Data Agent – HttpClient.Timeout after 300 seconds: where is it configured?
Hi everyone, this post is a follow-up to a previous discussion I opened about Fabric Data Agent queries being cancelled during aggregations: https://community.fabric.microsoft.com/t5/IQ/Fabric-Data-Agent-Query-cancelled-on-aggregations/td-p/5357096 After further testing, I believe I have isolated the issue more clearly, so I am opening a new post with a more specific title and description. The goal is also to make the issue easier to find for anyone searching for HttpClient.Timeout, 300 seconds, or Microsoft Fabric Data Agent timeout problems. The main issue is the following: Some Microsoft Fabric Data Agent requests fail after approximately 300 seconds because of HttpClient.Timeout. I am using a Fabric Data Agent with a Microsoft Fabric Ontology as its data source. The problem mainly appears on expensive queries, especially aggregations and queries involving larger portions of the ontology. However, further experiments suggest that this is not simply a query-performance or capacity issue. I have different Data Agents using the same underlying Ontology. One agent can execute queries for much longer than 5 minutes, in some cases close to 20 minutes. Another agent consistently fails after approximately 5 minutes with an HttpClient.Timeout. Fabric Capacity does not appear to be close to saturation when the timeout occurs. Individual entities can be queried correctly. This makes the fixed approximately 300-second behavior particularly confusing. The question I am now trying to answer is very specific: Which component in the Microsoft Fabric Data Agent architecture is enforcing this HttpClient.Timeout? For example, is the timeout defined in: the Data Agent orchestration layer, the MCP server associated with the published Data Agent, an internal Fabric service calling the ontology or graph engine, the client consuming the agent, or another intermediate HTTP layer? And, most importantly: Is this 300-second HttpClient.Timeout configurable anywhere? I have found other timeout-related settings in the surrounding architecture, but I have not found documentation identifying a setting that clearly corresponds to this specific timeout. There is also another important behavior that I would like to understand. If the published Data Agent is consumed programmatically, for example from a notebook or through its MCP endpoint, does the same 300-second HttpClient.Timeout still apply? Or does that use a different execution path with different timeout constraints? At this point I am specifically trying to identify the architectural source of the timeout rather than optimize the generated GQL query. The most useful clarification would therefore be: Where exactly is the 300-second HttpClient.Timeout enforced, and is there any supported way to configure it or use an execution path that does not have the same limit? This is related to my previous thread about query cancellation during aggregations, but the additional tests seem to narrow the issue down specifically to the HTTP timeout layer.34Views0likes1CommentFoundry Agent with Fabric data agent tool times out at 100s even in background mode
A Foundry agent with a Fabric data agent attached cancels any tool call that takes longer than 100 seconds, even with background mode enabled on a model that supports it. The documentation presents background mode as the supported path for MCP tool calls that exceed the synchronous timeout, but it does not lift the limit here. Setup Foundry project on Microsoft.CognitiveServices, West Europe Agent kind prompt, model gpt-5.6-luna (2026-07-09) metadata."microsoft.background-mode.enabled": "true" Requests to {project_endpoint}/openai/v1/responses with "background": true and an agent_reference Tested with both the fabric_iq_preview tool and the generic mcp tool, same server_url and same project_connection_id (authType: UserEntraToken) Behaviour Background mode is genuinely active: the response returns status: queued immediately and stays in_progress well past 100 seconds, so the model does support it. But the MCP tool call inside the run is still cancelled at exactly 100 seconds: { "code": "tool_user_error", "message": "TaskCanceledException encountered while invoking tool DataAgent_<name>: The request was canceled due to the configured HttpClient.Timeout of 100 seconds elapsing.. The remote MCP server did not complete the request within the configured timeout." } Same failure both ways — fabric_iq_preview failed at 115s, generic mcp at 108s. The response output shows a single mcp_call with status=failed. Why this looks like a Foundry-side gap The Fabric data agent MCP server advertises task support on initialize: "capabilities": { "tools": { "listChanged": false }, "tasks": { "list": {}, "cancel": {}, "requests": { "tools": { "call": {} } } } } and per tool in tools/list: "execution": { "taskSupport": "optional" } Calling that same server directly as a task, rather than through Foundry, completes the exact query Foundry cancels — in 81 to 173 seconds depending on the run. So the data agent can answer these questions; only the call made by Foundry is constrained. One detail that may explain it: the server advertises tasks as a top-level capability and accepts a task via a task field in the request params — an earlier form of the extension, rather than the current capabilities.extensions["io.modelcontextprotocol/tasks"] negotiation. A client following the current draft would not find the capability where it expects it and would fall back to a blocking call, which matches what we see. Questions Is background mode expected to lift the 100-second tool-call timeout for a Fabric data agent today, or is that combination not yet supported? Does Foundry's MCP client negotiate tasks with a server advertising them in this older form? If not, is alignment planned? Is there any setting, api-version, or tool property that makes Foundry request a task instead of blocking? Nothing in the agent definition or the portal changed the behaviour for us.51Views0likes1CommentHow can AI agents improve decision-making with Microsoft Fabric IQ?
I am exploring how AI agents can work with modern data platforms to help businesses move from traditional reporting toward proactive decision-making. Some possible use cases: AI agents analyzing business data and identifying important trends Automatically generating insights from Fabric data models Triggering workflows based on detected patterns or anomalies Helping teams interact with enterprise data using natural language Combining AI reasoning with governed data sources I would like to understand how the community is approaching AI-powered analytics with Microsoft Fabric IQ. Are teams using AI agents, Copilot experiences, or custom automation workflows on top of Fabric data solutions? What architecture patterns and best practices have you found useful?19Views1like1CommentUsing Data Agents for NL to KQL/GQL generation
Hi Team, I have a graph model and an eventhouse attached to a data agent and only want it to generate the correct KQL/GQL queries based on my natural language question, grounded in the schema object descriptions and example queries I've given for these sources. Is there a way I can get the data agent to emit just the query but not execute it? My use-case in NL to Code generation rather than execution and summarization. I tried fine tuning the agent instructions a bit, but I couldn't get the data agent to stop after the nl2code tool call, it ends up also executing the query every time. Kindly assist if someone has an idea. PS - I'm aware that Fabric has a real-time intelligence API for NL to KQL generation, but that is insufficient for us because we wanted both GQL and KQL support. Additionally, the RTI API only grounds itself in few-shot examples, unlike the data agent which has an advanced level of grounding using the schema object descriptions, etc. Therefore, I was more curious if this can be done using the data agent itself.Solved69Views2likes3CommentsFabric Data Agent: “Query cancelled” on aggregations
Good evening, I’m trying to build a Data Agent for my organization, even though the feature is still in preview. Its source is an ontology. I’m currently working with only six tables, but they are very large. I’m having an issue when asking the Data Agent questions that require grouping and aggregation, such as COUNT or SUM. In those cases, I often get the following error: "Query cancelled by upstream caller Status Code: Cancelled" The agent suggests that this may be caused by the amount of resources required by the query, which seems plausible. I also tried using smaller tables in the ontology. This improved things somewhat, and some queries now work, but aggregations are still extremely slow and I still frequently receive the same error. Is this expected behavior with large ontologies and aggregation queries, or could there be some configuration or setting that needs to be changed? Is there anything I can do to improve the performance of GROUP BY, COUNT, and SUM queries over an ontology-backed Data Agent? EDIT: I also found the following error in one of the question's execution log: 'The request was canceled due to the configured HttpClient.Timeout of 300 seconds elapsing.'158Views1like4CommentsFabric Agent consistently fails: backend-error "Bearer token not provided"
Hello, I have configured a Fabric agent using only a few tables from a single semantic model. I set up the agent instructions by following Microsoft's recommendations and added the Code Interpreter tool. Each time I ask a question ("Test the agent's responses" area, the agent is not published), even a very simple one, the same scenario occurs: The agent quickly starts an analysis and then proposes a DAX query that is correct. During this analysis phase, I can even see that the output is correct. However, it always ends with an error similar to the one below, which appears in the notifications: backend-error" Status code RequestTimeout: Request to https://agents.francecentral.hyena.infra.ai.azure.com/agents/v2.0/subscriptions/xxx/resourceGroups/AgentServiceProdEUVNet/providers/Microsoft.MachineLearningServices/workspaces/asfrclmln@asfrclmln@AML/openai/files?api-version=1.0 timed out" When I follow the URL from the notification, I get the following message: { "error": { "code": "UserError", "severity": null, "message": "Bearer token not provided.", "messageFormat": null, "messageParameters": null, "referenceCode": null, "detailsUri": null, "target": null, "details": [], "innerError": { "code": "AuthorizationError", "innerError": null }, "debugInfo": null, "additionalInfo": null }, "correlation": { "operation": "a47975e7212fb2eb4011d15654c85245", "request": "f98696fc72828139" }, "environment": "francecentral", "location": "francecentral", "time": "2026-08-28T07:52:40.3322903+00:00", "componentName": null, "statusCode": 401 } Thank you in advance for your help.Solved19Views0likes1CommentRetirement of Fabric data agent integration in Copilot in Power BI
https://community.fabric.microsoft.com/t5/Fabric-Updates-Blog/Retirement-of-Fabric-data-agent-integration-in-Copilot-in-Power/ba-p/5328344 I'm not sure I understand this announcement. I think it means that we will no longer be able to use Copilot from the screenshot below to access data agents/ IQ ontologies. Can someone please confirm? Any other thoughts around the implication of this announcement?Solved245Views1like5CommentsMultiple Semantic Models for one Ontology Layer
Hello everyone, When generating an ontology layer, I see that you can generate one from a semantic model. However, I see I can import multiple warehouses and lakehouses as part of an ontology layer if I wanted to start from scratch. Let's say I have four semantic models that each represent an important busines process (finance, sales, inventory, production/procurement). Each semantic model follows a well-defined star schema, and are mature enough to even be considered data products. If I want to eventually create an agent where I access data from both sales and inventory semantic model, would the best practice be to import/use both semantic models in the ontology layer? Or should I create one ontology layer per semantic model, and have my agent consume the information by attaching both ontology layers to said agent? For this example, let's say I want an agent that can see sales history, customer purchasing behavior history from the sales model, and can see inventory of products being ordered fro the inventory model. Back to the original question, what's the best practice ultimately for an agent to consume information from multiple semantic models?146Views2likes7Comments