Forum Discussion

DiscreetRaven's avatar
DiscreetRaven
Frequent Visitor
5 years ago
Solved

Find MAX within GROUP first, THEN filter resulting subset for desired value

This [TABLE1] is a simplified representation of my data. The TRUE data is 40 columns of 10M records. Contract Post Event Dollars Status 666666 1/1/2021 5 $5,000 good...
  • DiscreetRaven's avatar
    DiscreetRaven
    4 years ago

    Thanks to you, lbendlin, for taking the time. Your suggestion inspired me to look at the problem from another angle. I love the simplified approach but had to include a reference to [Event] to make it work. (I don't want the MAX [Status] prior to the selected date. I want the [Status] from the row with the MAX [Event] prior to the selected date.)

    GoodDollars =
    -- CHOOSE Post Date
    VAR P =
    ALLSELECTED ( 'Calendar'[Date] )
    -- DETERMINE the Single Most Recent Event Row for a Contract Prior or Equal to the Chosen Post Date
    VAR E =
    CALCULATE ( MAX ( Table1[Event] ), ALLEXCEPT ( Table1, Table1[Contract] ), Table1[Post] <= P )
    -- DETERMINE Status in the Row of the Most Recent Event (within All Rows for a Contract)
    VAR S =
    CALCULATE ( MAX ( Table1[Status] ), Table1[Event] = E )
    -- DETERMINE Dollars in the Row of the Most Recent Event
    VAR D =
    CALCULATE ( SUM ( Table1[Dollars] ), Table1[Event] = E )
    -- CHOOSE Dollars from the Row of the Most Recent Event ONLY IF the Status is Acceptable
    RETURN
    IF ( S IN { "good", "better" }, D )

    The result:

    This solution "works", but (as you noted) is not optimized. It's just fine for a sample table of two dozen records. Ten million plus - not so much. My first attempt was intended to create a much smaller 'temporary' table filtered to only those records remaining to be tested for their [Status]. 'MaxEvent' does just that. I want to believe I can directly reference the columns of ‘MaxEvent’ without having to recalculate the table a second time to determine the [Status], and a third time to determine [Dollars]. If that's possible, it's got to be more efficient. As the SQLBI article Anonymous pointed out says, the problem remains "the cardinality of the materialization required by the lowest level of context transition". OK. If you say so.