v-chuncz-msft , thank you for the insight, anywhere in particular I can follow the progress, IcM is that Internal change Management and a issue-number? Where can I follow the status?
Note: I also have a support ticket open and a few supporters working on the issue, first feedback was to restart the capacity which I found a bit worrysome as your own documentation actually also states that Gen2 does not require/allow restart.
Anyways, the workaround I had to put into effect was to create a new capacity of same tier and generation and move all workspaces to this - it resolved the issue immediately, while the old capacity still continued at around 800+% load (without any workspaces associated) the new capacity ran splendidly fast and maxes at around 20-30% as normal.
It is not a sustainable issue as I manually have to monitor the load and move "many" workspaces whenever it happens - unhappy customers and increased cost for us.
I observed though, that around 9:45PM CET there is a drastic change in workload - i.e. old capacity dropped after having moved the workspaces to 0% and the new increase to ~20%, I can only deduct that there is a scheduled job around that time which does "something" to the cache and/or CPU load, and not always for the better.
PierreL I can see from your message that you experienced same issue and temp. resolution was the same.
Looking forward for a sustainable solution.