Forum Discussion

ntimmerman's avatar
ntimmerman
Helper I
1 year ago

Notebook delta table merge operation error: Cannot seek after EOF

Hello Community,

 

For one of our data ingestion jobs, I am using a notebook to read MySQL data and I am executing a merge command to merge the source data with already existing data in the delta table. For some reason, I sometimes get the error message below.

I figured out that if I perform table maintenance on that table. the problem is solved. However, since I have quite a few tables this is difficult to do. Also, it requires manual intervention almost every day which is not ideal.

Any clue what's causing this error?

 

Notebook execution failed at Notebook service with http status code - '200', please check the Run logs on Notebook, additional details - 'Error name - Py4JJavaError, Error value - An error occurred while calling o6742.execute.
: org.apache.spark.SparkException: Job aborted due to stage failure: Task 1 in stage 49.0 failed 4 times, most recent failure: Lost task 1.3 in stage 49.0 (TID 411) (vm-0fc91660 executor 1): java.io.EOFException: Cannot seek after EOF
at org.apache.hadoop.fs.ChecksumFileSystem$FSDataBoundedInputStream.seek(ChecksumFileSystem.java:354)
at org.apache.parquet.hadoop.util.H1SeekableInputStream.seek(H1SeekableInputStream.java:46)
at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:555)
at org.apache.parquet.hadoop.ParquetFileReader.<init>(ParquetFileReader.java:799)
at org.apache.parquet.hadoop.ParquetFileReader.open(ParquetFileReader.java:666)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:85)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:71)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:66)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat.$anonfun$buildReaderWithPartitionValues$2(ParquetFileFormat.scala:263)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.org$apache$spark$sql$execution$datasources$FileScanRDD$$anon$$readCurrentFile(FileScanRDD.scala:252)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:316)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:162)
at org.apache.spark.sql.execution.FileSourceScanExec$$anon$1.hasNext(DataSourceScanExec.scala:614)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.columnartorow_nextBatch_0$(Unknown Source)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)
at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
at org.apache.spark.sql.execution.WholeStageCodegenEvaluatorFactory$WholeStageCodegenPartitionEvaluator$$anon$1.hasNext(WholeStageCodegenEvaluatorFactory.scala:43)
at org.apache.spark.sql.execution.SparkPlan.$anonfun$getByteArrayRdd$1(SparkPlan.scala:393)
at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:900)
at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:900)
at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:57)
at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:368)
at org.apache.spark.rdd.RDD.iterator(RDD.scala:332)
at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:93)
at org.apache.spark.TaskContext.runTaskWithListeners(TaskContext.scala:166)
at org.apache.spark.scheduler.Task.run(Task.scala:141)
at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$4(Executor.scala:620)
at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally(SparkErrorUtils.scala:64)
at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally$(SparkErrorUtils.scala:61)
at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:94)
at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:623)
at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
at java.base/java.lang.Thread.run(Thread.java:829)
 
Driver stacktrace:
at org.apache.spark.scheduler.DAGScheduler.failJobAndIndependentStages(DAGScheduler.scala:2935)
at org.apache.spark.scheduler.DAGScheduler.$anonfun$abortStage$2(DAGScheduler.scala:2871)
at org.apache.spark.scheduler.DAGScheduler.$anonfun$abortStage$2$adapted(DAGScheduler.scala:2870)
at scala.collection.mutable.ResizableArray.foreach(ResizableArray.scala:62)
at scala.collection.mutable.ResizableArray.foreach$(ResizableArray.scala:55)
at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:49)
at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:2870)
at org.apache.spark.scheduler.DAGScheduler.$anonfun$handleTaskSetFailed$1(DAGScheduler.scala:1264)
at org.apache.spark.scheduler.DAGScheduler.$anonfun$handleTaskSetFailed$1$adapted(DAGScheduler.scala:1264)
at scala.Option.foreach(Option.scala:407)
at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:1264)
at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:3139)
at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:3073)
at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:3062)
at org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:49)
at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:1000)
at org.apache.spark.SparkContext.runJob(SparkContext.scala:2563)
at org.apache.spark.SparkContext.runJob(SparkContext.scala:2584)
at org.apache.spark.SparkContext.runJob(SparkContext.scala:2603)
at org.apache.spark.SparkContext.runJob(SparkContext.scala:2628)
at org.apache.spark.rdd.RDD.$anonfun$collect$1(RDD.scala:1056)
at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)
at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:112)
at org.apache.spark.rdd.RDD.withScope(RDD.scala:411)
at org.apache.spark.rdd.RDD.collect(RDD.scala:1055)
at org.apache.spark.sql.execution.SparkPlan.executeCollectIterator(SparkPlan.scala:460)
at org.apache.spark.sql.execution.exchange.BroadcastExchangeExec.$anonfun$relationFuture$1(BroadcastExchangeExec.scala:140)
at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withThreadLocalCaptured$2(SQLExecution.scala:243)
at org.apache.spark.JobArtifactSet$.withActiveJobArtifactState(JobArtifactSet.scala:94)
at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withThreadLocalCaptured$1(SQLExecution.scala:238)
at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:264)
at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
at java.base/java.lang.Thread.run(Thread.java:829)
Caused by: java.io.EOFException: Cannot seek after EOF
at org.apache.hadoop.fs.ChecksumFileSystem$FSDataBoundedInputStream.seek(ChecksumFileSystem.java:354)
at org.apache.parquet.hadoop.util.H1SeekableInputStream.seek(H1SeekableInputStream.java:46)
at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:555)
at org.apache.parquet.hadoop.ParquetFileReader.<init>(ParquetFileReader.java:799)
at org.apache.parquet.hadoop.ParquetFileReader.open(ParquetFileReader.java:666)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:85)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:71)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:66)
at org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat.$anonfun$buildReaderWithPartitionValues$2(ParquetFileFormat.scala:263)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.org$apache$spark$sql$execution$datasources$FileScanRDD$$anon$$readCurrentFile(FileScanRDD.scala:252)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:316)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:162)
at org.apache.spark.sql.execution.FileSourceScanExec$$anon$1.hasNext(DataSourceScanExec.scala:614)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.columnartorow_nextBatch_0$(Unknown Source)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)
at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
at org.apache.spark.sql.execution.WholeStageCodegenEvaluatorFactory$WholeStageCodegenPartitionEvaluator$$anon$1.hasNext(WholeStageCodegenEvaluatorFactory.scala:43)
at org.apache.spark.sql.execution.SparkPlan.$anonfun$getByteArrayRdd$1(SparkPlan.scala:393)
at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:900)
at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:900)
at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:57)
at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:368)
at org.apache.spark.rdd.RDD.iterator(RDD.scala:332)
at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:93)
at org.apache.spark.TaskContext.runTaskWithListeners(TaskContext.scala:166)
at org.apache.spark.scheduler.Task.run(Task.scala:141)
at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$4(Executor.scala:620)
at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally(SparkErrorUtils.scala:64)
at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally$(SparkErrorUtils.scala:61)
at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:94)
at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:623)

9 Replies

  • Anonymous's avatar
    Anonymous
    Not applicable

    I don't know what's causing it, but I had exactly the same issue, and solved it in exactly the same way.  Fortunately, it hasn't failed since.

  • frithjof_v's avatar
    frithjof_v
    Community Champion

    Does it work if you automate the table maintenance?

     

    E.g. by scheduling a Notebook to run OPTIMIZE on the tables at a pre-defined interval?

    • ntimmerman's avatar
      ntimmerman
      Helper I

      Hi frithjof_v , Thanks for the suggestion. Yes, that's a good idea. I've implemented a table maintenance notebook that does this. However, it's difficult to determine when an optimize needs to happen. I had now done it when the last optimize was more than 7 days ago. however, sometimes I run into this issue less than a day after the previous time. So, do I need to schedule the table maintenance twice a day? That sounds like a lot of overhead and just not right. It's pretty frustrating to open the monitor every morning and find that jobs have failed again because of this issue...

      • frithjof_v's avatar
        frithjof_v
        Community Champion

        To be honest I don't have too much practical experience with it.

         

        How often is the merge operation running?

         

        (I'm wondering how often data is added/changed in your delta table)

    • ntimmerman's avatar
      ntimmerman
      Helper I

      Hi Anonymous , Sure! Here's the full error log which also contains the code that I'm executing. 

      Actually it's just a very simple merge operation with only maybe 10 records or so. The problem is that it's just very unreliable. Sometimes I get this error message, other times it's running fine. Here are the error details. Let me know if you have any other questions:

      Cell In[13], line 202, in load_duty_time_data(df_programs)
          199 if(df_duty_time.count() > 0):
          200     #merge data frame. don't worry about deleting as deleting is disabled in FS via a permission
          201     print(f"""merging {df_duty_time.count()} staff_duty_time records into delta table""")
      --> 202     DeltaTable.forPath(spark,'Tables/staff_duty_time').alias('target').merge(df_duty_time.alias('source'),"source.id = target.id and source.program_id = target.program_id").whenMatchedUpdateAll().whenNotMatchedInsertAll().execute()

      File /usr/hdp/current/spark3-client/jars/delta-spark_2.12-3.2.0.5.jar/delta/tables.py:1065, in DeltaMergeBuilder.execute(self)
         1058 @since(0.4)  # type: ignore[arg-type]
         1059 def execute(self) -> None:
         1060     """
         1061     Execute the merge operation based on the built matched and not matched actions.
         1062
         1063     See :py:class:`~delta.tables.DeltaMergeBuilder` for complete usage details.
         1064     """
      -> 1065     self._jbuilder.execute()

      File ~/cluster-env/trident_env/lib/python3.11/site-packages/py4j/java_gateway.py:1322, in JavaMember.__call__(self, *args)
         1316 command = proto.CALL_COMMAND_NAME +\
         1317     self.command_header +\
         1318     args_command +\
         1319     proto.END_COMMAND_PART
         1321 answer = self.gateway_client.send_command(command)
      -> 1322 return_value = get_return_value(
         1323     answer, self.gateway_client, self.target_id, self.name)
         1325 for temp_arg in temp_args:
         1326     if hasattr(temp_arg, "_detach"):

      File /opt/spark/python/lib/pyspark.zip/pyspark/errors/exceptions/captured.py:179, in capture_sql_exception.<locals>.deco(*a, **kw)
          177 def deco(*a: Any, **kw: Any) -> Any:
          178     try:
      --> 179         return f(*a, **kw)
          180     except Py4JJavaError as e:
          181         converted = convert_exception(e.java_exception)

      File ~/cluster-env/trident_env/lib/python3.11/site-packages/py4j/protocol.py:326, in get_return_value(answer, gateway_client, target_id, name)
          324 value = OUTPUT_CONVERTER[type](answer[2:], gateway_client)
          325 if answer[1] == REFERENCE_TYPE:
      --> 326     raise Py4JJavaError(
          327         "An error occurred while calling {0}{1}{2}.\n".
          328         format(target_id, ".", name), value)
          329 else:
          330     raise Py4JError(
          331         "An error occurred while calling {0}{1}{2}. Trace:\n{3}\n".
          332         format(target_id, ".", name, value))

      Py4JJavaError: An error occurred while calling o7542.execute.
      : org.apache.spark.SparkException: Job aborted due to stage failure: Task 1 in stage 371.0 failed 4 times, most recent failure: Lost task 1.3 in stage 371.0 (TID 2251) (vm-dad93454 executor 1): org.apache.spark.SparkException: Encountered error while reading file abfss://8c604c9b-3d05-46f3-a7a3-9fa42debcf1e@onelake.dfs.fabric.microsoft.com/f073b7a3-2275-4cc7-b1f1-8431a680313f/Tables/staff_duty_time/part-00000-e7ae0d9c-be99-4604-bc51-32401f102d06-c000.snappy.parquet. Details:
          at org.apache.spark.sql.errors.QueryExecutionErrors$.cannotReadFilesError(QueryExecutionErrors.scala:863)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:339)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:162)
          at org.apache.spark.sql.execution.FileSourceScanExec$$anon$1.hasNext(DataSourceScanExec.scala:614)
          at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.columnartorow_nextBatch_0$(Unknown Source)
          at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)
          at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
          at org.apache.spark.sql.execution.WholeStageCodegenEvaluatorFactory$WholeStageCodegenPartitionEvaluator$$anon$1.hasNext(WholeStageCodegenEvaluatorFactory.scala:43)
          at org.apache.spark.sql.execution.SparkPlan.$anonfun$getByteArrayRdd$1(SparkPlan.scala:393)
          at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:900)
          at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:900)
          at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:57)
          at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:368)
          at org.apache.spark.rdd.RDD.iterator(RDD.scala:332)
          at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:93)
          at org.apache.spark.TaskContext.runTaskWithListeners(TaskContext.scala:166)
          at org.apache.spark.scheduler.Task.run(Task.scala:141)
          at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$4(Executor.scala:620)
          at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally(SparkErrorUtils.scala:64)
          at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally$(SparkErrorUtils.scala:61)
          at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:94)
          at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:623)
          at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
          at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
          at java.base/java.lang.Thread.run(Thread.java:829)
      Caused by: java.io.EOFException: Cannot seek after EOF
          at org.apache.hadoop.fs.ChecksumFileSystem$FSDataBoundedInputStream.seek(ChecksumFileSystem.java:354)
          at org.apache.parquet.hadoop.util.H1SeekableInputStream.seek(H1SeekableInputStream.java:46)
          at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:555)
          at org.apache.parquet.hadoop.ParquetFileReader.<init>(ParquetFileReader.java:799)
          at org.apache.parquet.hadoop.ParquetFileReader.open(ParquetFileReader.java:666)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:85)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:71)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:66)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat.$anonfun$buildReaderWithPartitionValues$2(ParquetFileFormat.scala:263)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.org$apache$spark$sql$execution$datasources$FileScanRDD$$anon$$readCurrentFile(FileScanRDD.scala:252)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:316)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:162)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:329)
          ... 23 more

      Driver stacktrace:
          at org.apache.spark.scheduler.DAGScheduler.failJobAndIndependentStages(DAGScheduler.scala:2935)
          at org.apache.spark.scheduler.DAGScheduler.$anonfun$abortStage$2(DAGScheduler.scala:2871)
          at org.apache.spark.scheduler.DAGScheduler.$anonfun$abortStage$2$adapted(DAGScheduler.scala:2870)
          at scala.collection.mutable.ResizableArray.foreach(ResizableArray.scala:62)
          at scala.collection.mutable.ResizableArray.foreach$(ResizableArray.scala:55)
          at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:49)
          at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:2870)
          at org.apache.spark.scheduler.DAGScheduler.$anonfun$handleTaskSetFailed$1(DAGScheduler.scala:1264)
          at org.apache.spark.scheduler.DAGScheduler.$anonfun$handleTaskSetFailed$1$adapted(DAGScheduler.scala:1264)
          at scala.Option.foreach(Option.scala:407)
          at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:1264)
          at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:3139)
          at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:3073)
          at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:3062)
          at org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:49)
          at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:1000)
          at org.apache.spark.SparkContext.runJob(SparkContext.scala:2563)
          at org.apache.spark.SparkContext.runJob(SparkContext.scala:2584)
          at org.apache.spark.SparkContext.runJob(SparkContext.scala:2603)
          at org.apache.spark.SparkContext.runJob(SparkContext.scala:2628)
          at org.apache.spark.rdd.RDD.$anonfun$collect$1(RDD.scala:1056)
          at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)
          at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:112)
          at org.apache.spark.rdd.RDD.withScope(RDD.scala:411)
          at org.apache.spark.rdd.RDD.collect(RDD.scala:1055)
          at org.apache.spark.sql.execution.SparkPlan.executeCollectIterator(SparkPlan.scala:460)
          at org.apache.spark.sql.execution.exchange.BroadcastExchangeExec.$anonfun$relationFuture$1(BroadcastExchangeExec.scala:140)
          at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withThreadLocalCaptured$2(SQLExecution.scala:243)
          at org.apache.spark.JobArtifactSet$.withActiveJobArtifactState(JobArtifactSet.scala:94)
          at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withThreadLocalCaptured$1(SQLExecution.scala:238)
          at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:264)
          at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)
          at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)
          at java.base/java.lang.Thread.run(Thread.java:829)
      Caused by: org.apache.spark.SparkException: Encountered error while reading file abfss://8c604c9b-3d05-46f3-a7a3-9fa42debcf1e@onelake.dfs.fabric.microsoft.com/f073b7a3-2275-4cc7-b1f1-8431a680313f/Tables/staff_duty_time/part-00000-e7ae0d9c-be99-4604-bc51-32401f102d06-c000.snappy.parquet. Details:
          at org.apache.spark.sql.errors.QueryExecutionErrors$.cannotReadFilesError(QueryExecutionErrors.scala:863)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:339)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:162)
          at org.apache.spark.sql.execution.FileSourceScanExec$$anon$1.hasNext(DataSourceScanExec.scala:614)
          at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.columnartorow_nextBatch_0$(Unknown Source)
          at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)
          at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
          at org.apache.spark.sql.execution.WholeStageCodegenEvaluatorFactory$WholeStageCodegenPartitionEvaluator$$anon$1.hasNext(WholeStageCodegenEvaluatorFactory.scala:43)
          at org.apache.spark.sql.execution.SparkPlan.$anonfun$getByteArrayRdd$1(SparkPlan.scala:393)
          at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:900)
          at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:900)
          at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:57)
          at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:368)
          at org.apache.spark.rdd.RDD.iterator(RDD.scala:332)
          at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:93)
          at org.apache.spark.TaskContext.runTaskWithListeners(TaskContext.scala:166)
          at org.apache.spark.scheduler.Task.run(Task.scala:141)
          at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$4(Executor.scala:620)
          at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally(SparkErrorUtils.scala:64)
          at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally$(SparkErrorUtils.scala:61)
          at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:94)
          at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:623)
          ... 3 more
      Caused by: java.io.EOFException: Cannot seek after EOF
          at org.apache.hadoop.fs.ChecksumFileSystem$FSDataBoundedInputStream.seek(ChecksumFileSystem.java:354)
          at org.apache.parquet.hadoop.util.H1SeekableInputStream.seek(H1SeekableInputStream.java:46)
          at org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:555)
          at org.apache.parquet.hadoop.ParquetFileReader.<init>(ParquetFileReader.java:799)
          at org.apache.parquet.hadoop.ParquetFileReader.open(ParquetFileReader.java:666)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:85)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:71)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(ParquetFooterReader.java:66)
          at org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat.$anonfun$buildReaderWithPartitionValues$2(ParquetFileFormat.scala:263)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.org$apache$spark$sql$execution$datasources$FileScanRDD$$anon$$readCurrentFile(FileScanRDD.scala:252)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:316)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:162)
          at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:329)
          ... 23 more

      ERROR StatusConsoleListener Could not close org.apache.log4j.helpers.QuietWriter@71a0293c
       java.io.IOException: java.lang.InterruptedException
          at org.apache.hadoop.fs.azurebfs.services.AbfsOutputStream.shrinkWriteOperationQueue(AbfsOutputStream.java:709)
          at org.apache.hadoop.fs.azurebfs.services.AbfsOutputStream.uploadBlockAsync(AbfsOutputStream.java:361)
          at org.apache.hadoop.fs.azurebfs.services.AbfsOutputStream.uploadCurrentBlock(AbfsOutputStream.java:287)
          at org.apache.hadoop.fs.azurebfs.services.AbfsOutputStream.flushInternal(AbfsOutputStream.java:564)
          at org.apache.hadoop.fs.azurebfs.services.AbfsOutputStream.hflush(AbfsOutputStream.java:485)
          at org.apache.hadoop.fs.FSDataOutputStream.hflush(FSDataOutputStream.java:136)
          at org.apache.spark.microsoft.tools.logging.util.SecureStoreOutputStream.flush(SecureStoreOutputStream.java:56)
        

  • Anonymous's avatar
    Anonymous
    Not applicable

    Hi ntimmerman ,

    Did the above suggestions help with your scenario? if that is the case, you can consider Kudo or Accept the helpful suggestions to help others who faced similar requirements.

    Regards,

    Xiaoxin Sheng