Forum Discussion

chawong's avatar
chawong
New Member
9 years ago
Solved

Cluster algorithm

For the "automatically find cluster" capability in PowerBI desktop, what type of algorithm/logic is being used to arrive at the clusters? Trying to understand how the tool is arriving at the clusters based on parameters I input. Thanks

 

6 Replies

    • Greg_Deckler's avatar
      Greg_Deckler
      Community Champion

      v-sihou-msft - From the technical link in the link you provided:

      The Microsoft Clustering algorithm provides two methods for creating clusters and assigning data points to the clusters. The first, the K-means algorithm, is a hard clustering method. This means that a data point can belong to only one cluster, and that a single probability is calculated for the membership of each data point in that cluster. The second method, the Expectation Maximization(EM) method, is a soft clustering method. This means that a data point always belongs to multiple clusters, and that a probability is calculated for each combination of data point and cluster.+

      You can choose which algorithm to use by setting the CLUSTERING_METHOD parameter. The default method for clustering is scalable EM.

       

      https://docs.microsoft.com/en-us/sql/analysis-services/data-mining/microsoft-clustering-algorithm-technical-reference

       

      So, does Power BI use K-Means or EM? Sounds like it is likely EM if that is the default.

      • zkazimov's avatar
        zkazimov
        Helper I

        Hi Guys,

         

        I have customers clustered in power bi based on the margin. When I did clusterng, I had date filter = this month. Now, when I change date filter to this year, it is not reclustering. This month clustering grouped date into 5 groups and total of 461 customers. This year still shows 461 which I know is incorrect.

         

        See images below. Is there anyhting I can do to ensure it reclusters once filter changed to any other date?