Forum Discussion
Measuring similarity between strings
- 3 years ago
Anonymous wrote:
I still have no information on what measure is used to measure the string similarity.
As I mentioned above, the measure used is Jaccard similarity.
Calculating the similarity between all pairs of 2 million strings means 4 trillion comparisons. This will take a lot of computing time and Power Query likely isn't the best tool for this particular job. You'll probably need some optimization techniques to make this more tractable. See this approach for example,
https://hpccsystems.com/blog/edit-distances-and-optimized-prefix-trees-big-data
Generating a trillion pairs is challenging no matter what tool you use. What are you trying to do that requires all possible pairs?
Microsoft documentation says that "Power Query uses the Jaccard similarity algorithm to measure the similarity between pairs of instances." Wiki: https://en.wikipedia.org/wiki/Jaccard_index
- Anonymous3 years agoNot applicable
I have a system with more than two 2 000 000 vendors that needs cleenup. A lot of these vendors are duplicates and I wish to find candidate duplicates by comparing their names and adresses to give them a "duplicate score". The top scoring vendors could be removed automatically while lesser scoring vedors should be checked manually.
- AlexisOlson3 years agoSuper User
As long as Power Query is the tool you're using, I think clustering is likely your best option. See also:
https://radacad.com/fuzzy-clustering-in-power-bi-using-power-query-finding-similar-values
- Anonymous3 years agoNot applicable
I have tried clustering but I still have no information on what measure is used to measure the string similarity. Also I initially don't want to reduce the data. I want to add a similarity score for each combination of all of the origial strings.
Right now I'm working on using Azure Data Bricks. Data bricks has a function for measuring the Levinsthein Distance.