Forum Discussion
Spark Cluster's scheduler is killing container for some reason
- 11 months ago
I finished the MT CSS support case (pro). The engineer is Chirag on Deepak's team in the Eastern US timezone.
They have a way to use kusto logs to retrieve yarn messages. Unfortunately they wouldn't share the kusto query syntax. And they say the telemetry logs are internal, in any case.
Below is the message that they say they retrieved. Obviously they are able to retrieve log data directly from yarn, unlike their customers. The following is verbatim from Chirag.
- When the memory limit is reached, the container is terminated.
2025-10-21 23:05:16,763 INFO org.apache.hadoop.yarn.server.resourcemanager.scheduler.capacity.ParentQueue: root, capacity=1.0, absoluteCapacity=1.0, maxCapacity=1.0, absoluteMaxCapacity=1.0, state=RUNNING, acls=SUBMIT_APP:*ADMINISTER_QUEUE:*, labels=*,
This indicates the capacity reached 100%.
Hopefully this is helpful. I'm still not satisfied that customers are blindfolded when we encounter yarn-related failures.
Hi dbeavon3,
Thank you for contacting the Microsoft Fabric Community Forum.
Based on my understanding, this behaviour can occur when YARN terminates the Spark executor container (exit code 137) due to excessive memory usage or node health problems. In Fabric, this may happen if the executor exceeds its allocated memory including Python or native memory or if a node becomes unstable, resulting in container removal and subsequent Delta write failures.
Please follow the steps below, which may help resolve the issue:
- Use a larger compute size, for example: Medium or Memory Optimised.
- Repartition the data or reduce shuffle pressure to lower per task memory usage.
- Enable Spark monitoring and diagnostics to capture executor metrics prior to failure.
Additionally, please refer to the following link:
Notebook contextual monitoring and debugging - Microsoft Fabric | Microsoft Learn
We hope the information provided helps to resolve the issue. Should you have any further queries, please feel free to contact the Microsoft Fabric Community.
Thank you.
- dbeavon311 months ago
Memorable Member
>>Based on my understanding, this behaviour can occur when YARN terminates the Spark executor container (exit code 137) due to excessive memory usage or node health problems
Can you share a source for this? There is no evidence of excessive memory usages or health problems. Else I would not have posted all the logs above. The executor is killed arbitrarily, according to the details I shared in this example.I understand that there are various theoretical reasons why an executor might need to be terminated. This is not helpful. I'm looking for evidence to show why this particular executor was deliberately terminated.
I'd rather have evidence than rely on guesswork.
This is the reason I shared so many logs and observations. If you believe the executor ran out of memory, then show me the OOM error or tell me where to find it. I'm not going to double the size of the nodes based on guesswork... that simply puts money into Microsoft's pocket and doesn't even guarantee that the problem will be permanently resolved.
FYI, The link you shared does not give a way to see the actual memory or CPU consumed. If you believe memory is the problem then please share an approach for monitoring the memory consumption of these drivers and executors.
- dbeavon311 months ago
Memorable Member
I finished the MT CSS support case (pro). The engineer is Chirag on Deepak's team in the Eastern US timezone.
They have a way to use kusto logs to retrieve yarn messages. Unfortunately they wouldn't share the kusto query syntax. And they say the telemetry logs are internal, in any case.
Below is the message that they say they retrieved. Obviously they are able to retrieve log data directly from yarn, unlike their customers. The following is verbatim from Chirag.
- When the memory limit is reached, the container is terminated.
2025-10-21 23:05:16,763 INFO org.apache.hadoop.yarn.server.resourcemanager.scheduler.capacity.ParentQueue: root, capacity=1.0, absoluteCapacity=1.0, maxCapacity=1.0, absoluteMaxCapacity=1.0, state=RUNNING, acls=SUBMIT_APP:*ADMINISTER_QUEUE:*, labels=*,
This indicates the capacity reached 100%.
Hopefully this is helpful. I'm still not satisfied that customers are blindfolded when we encounter yarn-related failures.
- dbeavon311 months ago
Memorable Member
I also wanted to mention that the main reason we keep running into memory problem is because of a feature taking effect when saving deltatable. It was a feature called "optimized" deltatable storage.
...It was super obnoxious and we found the setting to disable it, thereby saving massive amounts of ram in spark executors.
I'm told that my workspace was affected because it was created at the beginning of the year. Workspaces created after mid-2025 will no longer use this functionality by default and won't have as many memory issues.