Forum Discussion

lsabetta's avatar
lsabetta
Frequent Visitor
10 months ago
Solved

Spark Session configuration

Hi community,   Currently I'm working with an F32 based in Europe.  I have a notebook in which I am forced to use pyspark because I have to get tables from my lakehouse - if not I would totally us...
  • v-venuppu's avatar
    10 months ago

    Hi lsabetta ,

    The issue happens because your Lakehouse table is stored in Delta format, not plain Parquet.
    When you overwrite the table, new Parquet files are created, but old ones remain in the folder for versioning.
    If you read the folder directly as Parquet, it loads both the old and new files - that’s why you see duplicate or outdated records.

    To fix this, make sure you read the table as a Delta table instead of as raw Parquet.
    This ensures that only the latest valid version of the data is returned, without mixing older files.

    Thank you.