Forum Discussion

going_grey's avatar
going_grey
Frequent Visitor
5 months ago
Solved

Spark Jobs - Monitor Page

Hi

 

When viewing the history of a spark streaming job that was cancelled a couple of days ago, I've found a few queries that are still in a RUNNING status and the Duration time is still ticking over. Is this a quirk or are those queries actually still running?

To clarify, this is the navigation path I'm taking:
Monitor page --> Activity name (cancelled) --> Spark History Server (tab) --> Show incomplete applications (link) --> App ID (link) --> Structured Streaming (tab)

 

This is where I'm seeing RUNNING queries. If these are actually still running, then how can I stop it?

  • This is mostly a quirk of how the Spark History Server reports state, not that the queries are truly still running.

    When a streaming job is cancelled, the driver and executors are stopped, but the Structured Streaming UI (in the history server) is based on last known state written to event logs. If the query did not shut down cleanly (for example, abrupt cancel), its status can remain as “RUNNING” and the duration counter continues to tick because there is no final termination event recorded. So what you are seeing is a stale UI state, not an active computation.

    If they were actually running, you would see active Spark applications consuming cluster resources in the workspace or capacity metrics. To be sure, check the Fabric/cluster monitoring or active Spark sessions. There is nothing to stop at this point because the compute is already terminated. If you want to avoid this in future, ensure graceful shutdown of streaming queries (for example, stop() on the query or checkpoint-aware termination) so Spark can persist a proper “TERMINATED” state.

     

2 Replies

  • Hello going_grey 

     

    Spark History server is a reconstruction of Spark logs, and if the job was terminated without registering as complete, it may show as Running. The correct source of truth is Monitoring hub - if nothing is running on Monitoring hub, it is certainly not ticking in the background!  

     

    If the Spark application is still running on Monitoring hub, you can try cancelling it there. If you are unable to cancel from the UI, try the Fabric REST API endpoints. 

     

  • This is mostly a quirk of how the Spark History Server reports state, not that the queries are truly still running.

    When a streaming job is cancelled, the driver and executors are stopped, but the Structured Streaming UI (in the history server) is based on last known state written to event logs. If the query did not shut down cleanly (for example, abrupt cancel), its status can remain as “RUNNING” and the duration counter continues to tick because there is no final termination event recorded. So what you are seeing is a stale UI state, not an active computation.

    If they were actually running, you would see active Spark applications consuming cluster resources in the workspace or capacity metrics. To be sure, check the Fabric/cluster monitoring or active Spark sessions. There is nothing to stop at this point because the compute is already terminated. If you want to avoid this in future, ensure graceful shutdown of streaming queries (for example, stop() on the query or checkpoint-aware termination) so Spark can persist a proper “TERMINATED” state.