tate-bowman's avatar
tate-bowman
Regular Visitor
5 months ago
Status:
New

Approximate Distinct Count in Import Mode via HyperLogLog (HLL)

Distinct counts are the bane of our existence in big data.  Distinct counts over DirectQuery/DirectLake are not fast or efficient enough, and are very expensive from a compute and cost standpoint.  We need a re-aggregable approximate distinct count in Import Mode using HyperLogLog (HLL).  The idea would be to enable the creation/ingestion of a HLL sketch at the same dimensional grain as other user-defined aggregations so that the sketches can be dynamically merged/aggregated to produce an approximate distinct count.  

 

Consider the following example where we need a distinct count of orders across channels and product groups.  Today, we would need to cache 4 unique intersections of data (1. Overall, 2. By Channel, 3. By Product Group, 4. By Channel & Product Group).  Not only is this compute intensive, but it forces us to create really terrible DAX to detect the user's reporting context and return the correct count.  In addition, it still doesn't allow for correct counts when users arbitrarily include/exclude members of a desired grouping (i.e. all product groups except "PRODUCT GROUP 02").  In an ideal world for this example, we would compute an HLL sketch at the Channel and Product Group level (image below), and either have a new DAX function like "HLLMERGE( <sketch> )" or enhance the APPROXIMATEDISTINCTCOUNT function to accept a column of sketchs, and this would dynamically merge the sketches in context to return the correct approximate distinct count.


This would be a huge unlock to be able to dynamically reaggregate approximate distinct counts, as they are very difficult to manage today, and Power BI is well-positioned to do this.  It would be a powerhouse feature for big data users.

No CommentsBe the first to comment

Recent ideas