Forum Discussion
Remove duplicates by prioritizing rows based on another column
- 5 years ago
Okay, please try this. I have included comments that explain each step.
BEFORE:
The goal is to remove rows 1, 3, 6, 8 and to have all columns present in the result.
RESULT:
SCRIPT:
There are two different options in the grouped step along with scenarios for when one would apply as opposed to the other, based on the specifics of the source data. Currently, option 2 is in use. To switch to option 1, add two forward slashes in front of Table.LastN and remove the two slashes at the beginning of let varTable and Table.FirstN.
let Source = Table.FromRows(Json.Document(Binary.Decompress(Binary.FromText("jcy5DcAgEETRXjZGaHbBVwi+irDovw0vJiFAmOQjxGOeh0IIZIgB6BHiroW3YCsQ1kt+pGT+ISNXKiu9UTcEGaL1n40x5n/F7sfZGJ0q6HtwHoIMp10qOxV7XndjdB2CDK/dKKUX", BinaryEncoding.Base64), Compression.Deflate)), let _t = ((type nullable text) meta [Serialized.Text = true]) in type table [Latest = _t, #"NVE-Nr." = _t, Item = _t, Date = _t, #"LHM-Nr." = _t, Index = _t]), grouped = Table.Group( Source, // Column(s) containing the values from which you'd like duplicates removed. {"NVE-Nr."}, { { "Table", // Name of the new column. each //--------------------------------------------------------- // Option 1: Sort descending, then select the first result. //--------------------------------------------------------- // This will work if there is a maximum of only two rows per NVE-Nr. //let varTable = Table.Sort ( _, {{"LHM-Nr.", Order.Descending}}) in //Table.FirstN ( varTable, 1 ), //--------------------------------------------------------- // Option 2: Select the last LHM-Hr for each NVE-Nr. //--------------------------------------------------------- // This will work if the row to keep always appears last in the group. Table.LastN ( _, 1 ), type table } } ), expand = Table.ExpandTableColumn ( grouped, "Table", // Expand the tables in this column List.Difference ( // New column names Table.ColumnNames ( // are the column names Table.Combine ( grouped[Table] ) // in the nested tables ), Table.ColumnNames ( grouped ) // that do not appear in the grouped table. ) ) in expand
Hi, jennratten .
Thank you very much for your help. I don't seem to understand the source step. So far I always left it since I though it was irrelevant for my case. Could you please explain what it does and what this random characters do (do I also need to adjust it if I am applying to a table with a bigger size and if yes how do I do it?)
Regards
Sure thing - this value is generated by copying data and pasting it into Excel. For example, I converted your image into an Excel table. Then I copied the Excel data and in Power Query, Home> Enter Data > paste > OK (the options will be slightly different if you are using Excel). After doing so, a new query will be created and the only step that will be in the query is Source, with a similar script. See the image below. I To integrate the sample script I provided with your query, copy the grouped and expand steps from my script, open the Advanced Editor for your query, paste them after the step which generated the table in your first post (likely the last step in your query), then replace "Source" in the grouped step with the name of your last step.
- Anonymous5 years agoNot applicable
Hi,
I am now a bit confused since I import lots of pds into Power Query first and clean them first. The images/samples I provided are the ones after the cleanup, meaning I have multiple previous steps that are in Power Query. I could copy and paste an Excel table into the Power Query but the thing is, I will be expanding my table when I get new pdfs.
Is there any automatic way to remove duplicates without having to copy cleaned table in Excel and open it again in a new query but rather continue where I just left off processing my data?
ā
Thanks and regards,
- jennratten5 years agoSuper User
I apologize if my last response was confusing. I will separate it clearly.
Your question:
I don't seem to understand the source step. So far I always left it since I though it was irrelevant for my case. Could you please explain what it does and what this random characters do?
The first part of the response (answering your question).
This value is generated by copying data and pasting it into Excel. For example, I converted your image into an Excel table. Then I copied the Excel data and in Power Query, Home> Enter Data > paste > OK (the options will be slightly different if you are using Excel). After doing so, a new query will be created and the only step that will be in the query is Source, with a similar script.
The second part of the response (explains how to add this to your existing query).
To integrate the sample script I provided with your query, copy the grouped and expand steps from my script, open the Advanced Editor for your query, paste them after the step which generated the table in your first post (likely the last step in your query), then replace "Source" in the grouped step with the name of your last step.
- Anonymous5 years agoNot applicable
Still doesn't answer my question:
what happens if new data comes in and table expand? I will always have to copy and paste my whole table to adjust the code?
Second question:
At the moment I am not able to copy the whole table (it contains over 12 500 rows) and paste it into Data Editor within Power Query. How would you go about it?
Regards,