Forum Discussion

DennesTorres's avatar
DennesTorres
Icon for Power Participant rankPower Participant
3 years ago
Solved

Group by on eventstream's lakehouse target doesn't aggregate numeric fields

Hi, The problem is: The fields are not available to apply an aggregation function. I can make a count, but I can't make a sum, avg or other.   This was on the NYTaxi sample data, close the the...
  • Anonymous's avatar
    Anonymous
    2 years ago

    Hi DennesTorres ,

    As I understand you are facing issue in applying aggregate operations on the columns in the destination. Here we have two cases.

     

    1)In the first case the destination in the eventstream is set to be Lakehouse table. The sample data which you are using (i.e) Yellow Taxi is converted to string by default in this case as we are reading the data from a CSV file which does not have type. Hence all the columns will be converted to string. For lake house, it is string by default but you can change the type using event processor no code editor.

     

    2) In the second case the destination is set to be a KQL database. For kusto, it has the type as conversion by default. Hence the datatypes will not be changed and will be similar to that of the source datatypes as you have shown me in the screenshot.

     

    So if you want to use the destination for the eventstream as Lakehouse and want to apply aggregate operation on your columns, we have to change the datatype of the columns. We cannot apply aggregate operation on a string datatype column. This can be achieved by using event processor no code editor. I have created a repro at my end and attaching the screenshots for your reference. I am able to apply aggregate operation on other columns as well. 

    Hence it is not a bug.
    Please refer this link for changing the datatypes of the columns: link 



     

     

     

    Hope this helps. Do let us know if you have any further queries. Glad to help.

  • DennesTorres's avatar
    DennesTorres
    2 years ago

    Hi,

     

    This explain and works. But it doesn't change the fact it's a terrible idea to have transformations and types completely tied witht he target.

    Is bad for us, who will use the technology, because it's prone to make maintenance more difficult. It's bad for the developers because it multiply the need of development and creates the requirement to keep the development in sync in two different places - and it's not in sync now, the targets don't have the same features.

    This problem was not there before. When using streaming dataflows, you can notice by the image below, the transformations are done by the dataflow, independent of the target. I could add two targets receiving the same transformation result.


    When using Stream Analytics, the transformatiosn were done by the stream analytics query language, and we could drop the result in multiple targets

    The idea to keep the transformations and even the data types tied to the target doesn't seems a good idea at all and the previous technologies were not like this.

    I hope this feedback reaches the development team.

    Kind Regards,

    Dennes