Forum Discussion

spencer_sa's avatar
spencer_sa
Impactful Individual
1 year ago
Solved

Nullable = False in schema not correctly preserved / written using pyspark

We've found some 'interesting' and potentially bugged behaviour when writing Delta tables in pyspark with nullable = False columns in the schema - specifically if you write it to a table and the re-r...
  • jennratten's avatar
    jennratten
    1 year ago

    Hello spencer_sa - I've done some more research on this.  Based on my findings, while PySpark does include the ability to specify the schema, I believe that when writing a delta table to a lakehouse in this way, the nullability specifications in the schema are not preserved - and all columns set to nullable so that it is optimized for schema-on-read operations.  Since the parameters are available that lead the user to think the schema can be explicitly specified, I think it would be good to get feedback from Microsoft so they can confirm.