Forum Discussion

dbeavon3's avatar
dbeavon3
Icon for Memorable Member rankMemorable Member
10 months ago
Solved

SessionExpiredException occurs regularly in driver's stderr (org.apache.zookeeper.ClientCnxn)

My spark jobs have been failing regularly and the following seems to be one of the things that predicts an imminent failure:   2025-10-12 23:29:27,220 WARN ClientCnxn [Thread-62-SendThread(vm-6a61...
  • v-achippa's avatar
    v-achippa
    10 months ago

    Hi dbeavon3,

     

    Thank you for the response. In Fabric the spark coordination components(including ZooKeeper) are service-managed and containerized, not customer-managed.
    They are isolated at the capacity level, not at the individual workspace level.

    • On a dedicated capacity, these services run within resources allocated only to that capacity and are not shared with other tenants.
    • Multiple workspaces under the same capacity will share that capacity’s resources, but the service manages scheduling and isolation internally to prevent one job from affecting another.
    • On a shared capacity, the coordination layer is multi-tenant and managed by Microsoft’s service fabric layer, but each job runs in its own isolated session.

    You are right that these logs can appear even though the component is not customer-managed, they simply reflect transient coordination retries within the platform.

    If the timeout changes reduce failures, it confirms a transient condition, If not the Fabric Support can review the backend ZooKeeper health for your job timestamps.

     

    Thanks and regards,

    Anjan Kumar Chippa