Forum Discussion
Power BI Gateway Issue
Hi all,
Sorry for not getting back sooner. Reason is that after a week of running agents and providing logs from servers back to MS. They came across the theory that our 8 VPU gateways do not have enough threads to support the number of requests been issued by the service and hence there is a timeout and the error. Just to note, these machines average about 4% utilization but we have followed the MS recommendation of 2 x 8VPU gateways. Anyway they are suggesting that each machine is limited by the number of threads and to remove/reduce the occurrence of the error we need to scale out .... Not easy to setup in a PRD environment. So the next best thing to do was to scale up. So for a period of a week we run 2 x 16VPU machines and noted the number of incidences occuring. Just for context we have 2 main errors cropping up in our PBI Service.
1) {"errorCode":"Gateway_Offline","errorDescription":"EnterpriseGateway_LongMessage_Gateway_Offline"}
2){"code":"DM_GWPipeline_Gateway_AdoNetProviderOpenConnectionTimeoutError","pbi.error":
The first error is still being looked at by MS... They can't understand that if all the infrastructure is within Azure how can a High Availability setup lose the gateways... more on that one when I get more out of MS.
Here is my analysis for the period. Showing that doubling the machine size improved but did not eliminate the issue for error 2 above.
Something doesn't seem right here..
FOR ::
{"code":"DM_GWPipeline_Gateway_AdoNetProviderOpenConnectionTimeoutError","pbi.error":
| VPU | Total Events | Unique Events |
27/01/2024 | 16 | 3 | 3 |
26/01/2024 | 16 | 9 | 6 |
25/01/2024 | 16 | 9 | 4 |
24/01/2024 | 16 | 21 | 9 |
23/01/2024 | 16 | 8 | 4 |
|
| 50 | 26 |
22/01/2024 | 8 | 19 | 7 |
21/01/2024 | 8 | 19 | 9 |
20/01/2024 | 8 | 15 | 10 |
19/01/2024 | 8 | 13 | 6 |
18/01/2024 | 8 | 24 | 9 |
|
| 90 | 41 |
Here are MS suggestions to us.
We would suggest below action plans:
- smooth the dataset refresh by adding 2-3 seconds interval, I do know that we discussed earlier you expect the refresh every 5 mins for each dataset , but if like dataset A refreshed at 00:00:00 and dataset B refreshed at 00:00:02 then dataset C refreshed at 00:00:04. And after 5 mins, they should start at 00:05:00 and 00:05:02 and 00:05:04, thus for each dataset, they should still refresh every 5 mins. And this should reduce the sudden spike on gateway side.
- Scale out the cluster with additional 1-2 nodes to handle the spike.
- Add the retry parameters in your request as we mentioned in earlier call
I hope this helps someone out there.
Will report back with updates on this.
Nick