Forum Discussion
Staging query is getting executed for every sub query
- 8 years ago
Hi, I think I understand. You're point comes from the perspective of what M language does and how the compiler interprets duplicate vs reference.
My perspective comes from SQL DB optimization and reducing the total number of queries to the database. As over the wire latency maters much more than processor time.
What I found is... that Power Query does in fact send multiple SQL queries per sub query when the stagging query is set to 'Connection Only'.
Proof / Repro
- Set your stagging query to 'Connection only' as Ken recomended.
- Use a SQL connection or some other remote data source. I used SQL for this test (and easier to test as you'll see).
- Create multiple sub queries as 'reference' queries. It actually doesn't matter what these queries do.
- Open up network monitor of your choice. I used WireShark. Set your capture filter to your remote destination.
- Hit Refresh all in your Excel spreadsheet and inspect the wire.
I have detected exactly the same number of SQL Select statements hit the wire as there are sub queries.
Potential Workaround Found
- Set your stagging query to 'Load to Table'
- Create a new query that loads from the table you just created
let
Source = Excel.CurrentWorkbook(){[Name="yourTableNameHere"]}[Content]
in
Source - Set your sub queries to 'reference' the newly created query in step 2.
- Setup WireShark as before, and Hit Refresh All
Only one SQL query will hit the wire and all other sub queries will now wait for completion.
Further Discusion
Ideally, resultant set should be used, but M Language cannot tell where the data comes from. Is it being cached, is it dynamically created, or is it coming from a remote system? Without that, M Language cannot effectively provide step optimization techniques as claimed with 'duplicate' vs 'reference'. The combination of missunderstanding of the capabilities with M Language optimziation and bad architectual advice lead me to this state. Stanging Query concept only works if you 'cache' the resultant set somewhere. Either in the DataModel or in a table.
Hi jasbro,
Yes as the post that you showed says, the Duplicate will re-run the code while reference will just re-use the result set
FYR, I am just giving some sample M Code for Duplicate and Reference
M-Code of my Source:
let
Source = Table.FromRows(Json.Document(Binary.Decompress(Binary.FromText("dZBLCsIwEECvErIuNjOTpFmLay9QulC0KtIO1N4fU5oPjLhICO89wiR9r0+X9a4bfeYDj+rK/H7Nj08EKq5/7piOalx4UhPP61OtrG5bPjS9hhZaNBBiimjTZT9w2/cas7DOl1pAH0yqKQuyrtQCgkGXclsNlVxAIB9S7rLxrk4uIBrElPtsgGouIEJ8wZ532YSuvlRAMpRvD3VMLLmARF38x+EL", BinaryEncoding.Base64), Compression.Deflate)), let _t = ((type text) meta [Serialized.Text = true]) in type table [Input = _t, #" " = _t, Output = _t, #"(blank)" = _t, #"(blank).1" = _t, #"(blank).2" = _t]),
#"Changed Type" = Table.TransformColumnTypes(Source,{{"Input", type text}, {" ", type text}, {"Output", type text}, {"(blank)", type text}, {"(blank).1", type text}, {"(blank).2", type text}}),
in
#"Changed Type"M-Code for Duplicated Query:
let
Source = Table.FromRows(Json.Document(Binary.Decompress(Binary.FromText("dZBLCsIwEECvErIuNjOTpFmLay9QulC0KtIO1N4fU5oPjLhICO89wiR9r0+X9a4bfeYDj+rK/H7Nj08EKq5/7piOalx4UhPP61OtrG5bPjS9hhZaNBBiimjTZT9w2/cas7DOl1pAH0yqKQuyrtQCgkGXclsNlVxAIB9S7rLxrk4uIBrElPtsgGouIEJ8wZ532YSuvlRAMpRvD3VMLLmARF38x+EL", BinaryEncoding.Base64), Compression.Deflate)), let _t = ((type text) meta [Serialized.Text = true]) in type table [Input = _t, #" " = _t, Output = _t, #"(blank)" = _t, #"(blank).1" = _t, #"(blank).2" = _t]),
#"Changed Type" = Table.TransformColumnTypes(Source,{{"Input", type text}, {" ", type text}, {"Output", type text}, {"(blank)", type text}, {"(blank).1", type text}, {"(blank).2", type text}}),
#"Added Index" = Table.AddIndexColumn(#"Changed Type", "Index", 0, 1)
in
#"Added Index"M-Code for Referenced Query:
let
Source = Table1,
#"Added Index" = Table.AddIndexColumn(Source, "Index", 0, 1)
in
#"Added Index"As you can see, the duplicate Query runs through the source a second time apart from the actual Query. But in case of reference, it is just going to do the actions on top of the Source. So it is more likely that your duplicate query will run a separate query against your DB while you reference query will use the output of your source directly.
But when you do a manual Report Refresh, all your queries will be getting loading
The Other important thing that you can see when using a reference query is, when you make changes to your source query, Only your source query will be loaded again on giving apply changes, all the referenced queries will automatically reflect the changes. But here if you are using a duplicate query, you will have to make the changes again in all queries manually
Hi, I think I understand. You're point comes from the perspective of what M language does and how the compiler interprets duplicate vs reference.
My perspective comes from SQL DB optimization and reducing the total number of queries to the database. As over the wire latency maters much more than processor time.
What I found is... that Power Query does in fact send multiple SQL queries per sub query when the stagging query is set to 'Connection Only'.
Proof / Repro
- Set your stagging query to 'Connection only' as Ken recomended.
- Use a SQL connection or some other remote data source. I used SQL for this test (and easier to test as you'll see).
- Create multiple sub queries as 'reference' queries. It actually doesn't matter what these queries do.
- Open up network monitor of your choice. I used WireShark. Set your capture filter to your remote destination.
- Hit Refresh all in your Excel spreadsheet and inspect the wire.
I have detected exactly the same number of SQL Select statements hit the wire as there are sub queries.
Potential Workaround Found
- Set your stagging query to 'Load to Table'
- Create a new query that loads from the table you just created
let
Source = Excel.CurrentWorkbook(){[Name="yourTableNameHere"]}[Content]
in
Source - Set your sub queries to 'reference' the newly created query in step 2.
- Setup WireShark as before, and Hit Refresh All
Only one SQL query will hit the wire and all other sub queries will now wait for completion.
Further Discusion
Ideally, resultant set should be used, but M Language cannot tell where the data comes from. Is it being cached, is it dynamically created, or is it coming from a remote system? Without that, M Language cannot effectively provide step optimization techniques as claimed with 'duplicate' vs 'reference'. The combination of missunderstanding of the capabilities with M Language optimziation and bad architectual advice lead me to this state. Stanging Query concept only works if you 'cache' the resultant set somewhere. Either in the DataModel or in a table.
- Thejeswar8 years agoSuper User
Nice Observation and Explanation :smileyhappy: