Forum Discussion
Understanding why Table.Buffer makes a difference in dependency chain
I understand from Google that using Table.Buffer is often a good idea when multiple steps within a query will be referencing the same Table. But it seems from my experimentations that it's also a good idea when there are other downstream queries in the query dependency chain.
As per the question I posted here, I have a bunch of queries in a chain that are somewhat recursive in nature: I am finding matching between two data sources where the date doesn't always match, so I progressively match the tables on larger values of a 'tolerance' factor X using this condition:
Table1.Date = Table2.Date +/- X
And I'm effectiviely running this in a loop for values of X between 1 to 7, and removing successful matches on each pass leaving just unsuccessful ones. This lets me progressively increase the 'mismatch' tolerence on those dates, and remove matches at each pass leaving just the unmatched rows to do increasingly desperate matches on.
My query performs the following steps:
- Loads data from Table1 and Table2, and does an inner join on unique ID and Date
- Effectively removes these matches from Table1 and Table2 via an anti-join against each referenced table.
- Does an inner join on the remaining records, but this time offsets the dates in the Port table by +/- 1 day. This finds some more matches.
- Performs steps 2 and 3 over and over, increasing the date offset by 1 each time until I have captured all date mismatches up to a week.
The query dependency tree that results looks like so:
I've found that using Table.Buffer on both Tables whenever I do a join dramatically increases (whoops) decreases execution time.
In this file, with no Table.Buffer, it takes about 17 seconds to run through the chain:
Match to closest day_20190722 No Buffer.xlsx
In this file, with a Table.Buffer on every Table that gets joined on every step where a join occurs, it takes just 6 seconds:
Match to closest day_20190722 buffer3 Revised.xlsx
Can anyone shed light on why the Buffer works? Perhaps another interesting question for ImkeF
This thread contains a lot of information around how Table.Buffer works: https://social.technet.microsoft.com/Forums/en-US/34e454b5-3a18-4eef-b920-40703c93f390/tablebuffer-for-cashing-intermediate-query-results-or-how-workaround-unnecessary-queries-issue?forum=powerquery
7 Replies
- ImkeFCommunity Champion
This thread contains a lot of information around how Table.Buffer works: https://social.technet.microsoft.com/Forums/en-US/34e454b5-3a18-4eef-b920-40703c93f390/tablebuffer-for-cashing-intermediate-query-results-or-how-workaround-unnecessary-queries-issue?forum=powerquery
- JeffWeirAdvocate V
Wow....putting Table.Buffer around any steps that referenced previous queries reduced my load time from potentially hours to one and a half minutes. (I say potentially hours, because I'd never dared to load all 30k rows of data for both Tables at once, nor attempted to make as many as 7 recursive steps. Even just loading one tenth of the data for 5 steps was taking a good 10 minutes or more).
I take it that I want to buffer each input just once, as soon as it appears in the current query that is referencing previous ones?
- pb296Regular Visitor
Please can you re-post the link to the article you have shared. I have tried to access the link and the page can't be found
- Element115Memorable Member
https://chatgpt.com/share/6a4bc233-8278-83ea-8325-5a89fd6ac251
Basically, if you need a PhD to use it, then it's a mess AFAIC. But it is what it is unfortunately.