Forum Discussion

dbeavon3's avatar
dbeavon3
Icon for Memorable Member rankMemorable Member
11 months ago
Solved

Spark Cluster's scheduler is killing container for some reason

Part-way thru the writing of a deltatable, there is a yarn container that is being intentionally killed.  Obviously this is causing problems for the work.   Can someone tell me why this is happenin...
  • dbeavon3's avatar
    dbeavon3
    11 months ago

    I finished the MT CSS support case (pro).  The engineer is Chirag on Deepak's team in the Eastern US timezone.

     

    They have a way to use kusto logs to retrieve yarn messages.  Unfortunately they wouldn't share the kusto query syntax.  And they say the telemetry logs are internal, in any case.

     

    Below is the message that they say they retrieved.  Obviously they are able to retrieve log data directly from yarn, unlike their customers.  The following is verbatim from Chirag.

     

     

    • When the memory limit is reached, the container is terminated.

    2025-10-21 23:05:16,763 INFO org.apache.hadoop.yarn.server.resourcemanager.scheduler.capacity.ParentQueue: root, capacity=1.0, absoluteCapacity=1.0, maxCapacity=1.0, absoluteMaxCapacity=1.0, state=RUNNING, acls=SUBMIT_APP:*ADMINISTER_QUEUE:*, labels=*,

     

    This indicates the capacity reached 100%.

     

     

    Hopefully this is helpful.  I'm still not satisfied that customers are blindfolded when we encounter yarn-related failures.