Blog Post

Fabric Updates Blog
5 MIN READ

Data source routing in Microsoft Fabric data agents (Generally Available)

midesa's avatar
midesa
Icon for Microsoft Employee rankMicrosoft Employee
1 month ago

Ask a Microsoft Fabric data agent a question in plain English, such as “Which three regions had the highest revenue last quarter?”, and it responds in plain English. What you don’t see is that the agent may be connected to several data sources at once: a lakehouse, a warehouse, a Power BI semantic model, and a KQL database. Before it can answer, it must determine which source contains the information it needs.

 

That decision is called data source routing. It runs on every question, and when it works, you never notice it. That invisibility is exactly what makes it worth understanding. This post explains what routing does, why it can be more challenging than it sounds, and the steps that make it more reliable in your own agents.

 

The challenge of multiple data sources

 

Every source you connect answers more questions and adds one more place the agent can send a question by mistake. Consider an agent wired to two semantic models, a sales lakehouse, and a KQL database of application logs—then ask it how checkout performed last week. That single question could reasonably land in three places: the lakehouse, if “performance” means revenue and conversion; the logs, if it means latency and error rates; or a semantic model, if someone has already defined a curated measure for checkout.

 

A wrong turn is costly for two reasons. First, a bad route still returns an answer: a confident, well formatted result built on the wrong data. That’s harder to catch than an empty response, and it erodes trust faster.

 

Second, the agent does not read your entire schema before it decides. For speed, it works from a sample of each source’s metadata. On a small, tidy source that sample covers everything that matters, but on a large one, the table that answers the question may never appear in the sample. The agent cannot route to a table it does not know is there, so the question goes somewhere plausible but wrong, or comes back empty.

 

Routing has two jobs: choose correctly among the sources it can see and make sure the right part of a large schema is visible to choose from in the first place.

 

The solution: A routing step before any query

 

Every data agent has an orchestrator that plans each answer and selects the tools and sources to build it. When a question arrives, the orchestrator builds a plan, picks the source most likely to hold the answer, calls that source’s query-generation tool, and reviews what comes back. If it still needs more, it repeats with another source or another step.

 

The important part is that routing and query generation are two separate stages. This separation matters because generating the right query is only useful if the agent starts with the right source. Routing decides where a question should go. It compares the question against each source’s metadata (the name, description, selected schema, and example queries) and commits to the most likely source.

 

Only then does the matching query tool decide how. Each source type has its own tool: NL2SQL (natural language to SQL) for a lakehouse or warehouse, DAX for a Power BI semantic model, and KQL for an Eventhouse KQL database. That tool translates the question into the right language, then writes and runs the query.

 

To stay fast, the orchestrator usually routes from a subset of each source’s metadata rather than the full picture. When that subset is enough, the decision is immediate. When it is not, because the schema is large, the source names are similar, or the question is genuinely ambiguous, the orchestrator calls a dedicated routing tool that inspects the full schema and example queries before committing.

 

Technical drill-down: How the router decides

 

The router weighs a handful of signals together, and knowing what they are is what lets you shape its behavior.

 

Figure: The router compares an incoming question against each source’s signals (schema and metadata, description, and example queries) and routes to the source most likely to answer.

 

The first signal is schema and metadata. Each source exposes its tables, its columns, and, for semantic models, its measures, and the router uses them to reason about which source could even answer a given question. A question about “average ticket resolution time” will not route to a source that has no time or ticket fields. Because the selected schema is such a strong signal of what a source covers, a tight, well-named selection routes far better than a large, noisy one.

 

Figure: In the data agent authoring experience, you select which lakehouse tables the agent can use (left) and test a question in plain English (right).

 

The second signal is the description and example queries you attach to each source, the levers you control directly. A one-line description tells the router what a source is for, for example, transaction-level revenue for North America retail, including sales and returns. The router cannot infer that intent from column names alone.

 

Figure: Adding a data source description in the data agent authoring pane. The description explains what the source contains and when to use it, giving the router intent it can't infer from table and column names alone.

 

Example queries go a step further: they are representative questions a source is meant to answer, and the router matches each new question against them to find the closest source.

 

Figure: Adding example queries to a data source. Each pair links a natural-language question to the SQL that answers it, showing the router the kinds of questions the source is meant to handle.

 

Two shortcuts are worth knowing. When an agent has only one source, there is nothing to choose between, so the router skips routing and goes straight to query generation. Once the router commits to a source, it invokes that source’s query tool automatically.

 

How to improve routing

 

You don’t have to guess what the router did. After the agent answers, expand the run steps to see which source it routed to and what context shaped the choice. If the orchestrator had to call the routing tool, that appears as its own step and lists the metadata it reviewed before committing. That is a reliable sign the decision was ambiguous, and the first place to look when a question keeps routing the wrong way.

 

Figure: In the agent’s run steps, the routing tool appears as its own step, showing the source metadata the orchestrator reviewed before committing to a data source.

 

Everything the router reads is something you can edit, which means routing accuracy is largely under your control. Work through the following steps in order:

  1. Tighten your schema selection. The tables, views, and measures you select are the primary signal of what a source covers, so select only the entities the agent should consider, and give them descriptive names. Large or noisy selections make it harder for the router to tell what each source is for.
  2. Write a data source description. In a line or two, describe the source’s purpose rather than its structure: what questions is this the right home for? Keep it focused on the topics and entities the source covers, the way you would brief a new analyst. For every setting you can tune, review data agent configurations.
  3. Add example queries. Give each source a handful of representative questions it should answer, and prioritize the ones that routed to the wrong place before. Because the router matches new questions against these examples, they are the fastest way to correct a specific, repeatable mistake. To learn how to write them, review example queries.
  4. Add routing rules to your agent instructions. If a question still lands in the wrong place after the previous steps, declare explicit rules. For example:
    • Shipment delays, carrier performance, or logistics trends: use the logistics lakehouse.
    • Campaigns, ad spend, or channel performance: use the marketing warehouse.

Getting started

 

Good routing separates a multi-source agent that reliably answers from one that quietly guesses. The reader never sees the decision; they ask one question, and the right source responds. Because it is built from the schemas, descriptions, and rules you write, you can measure and tune it.

 

To improve routing in your own agent, start with the two highest-impact steps: tighten each source’s schema selection and add a focused description, then confirm the result in the run steps. For complete guidance, read improve data source routing.

 

If you are new to data agents, start with the Fabric data agent overview.

 

Updated 1 month ago
Version 1.0

1 Comment

  • bhpat_msft's avatar
    bhpat_msft
    Icon for Microsoft Employee rankMicrosoft Employee

    midesa​ Fantastic overview of the router mechanics. Following up on step 4 (agent instructions) and data source descriptions: how does Data Agent handle ontology integration attached as a data source? Especially if there are multiple of them attached (enterprise and domain specific ontologies as an example)