Forum Discussion
Nullable = False in schema not correctly preserved / written using pyspark
- 1 year ago
Hello spencer_sa - I've done some more research on this. Based on my findings, while PySpark does include the ability to specify the schema, I believe that when writing a delta table to a lakehouse in this way, the nullability specifications in the schema are not preserved - and all columns set to nullable so that it is optimized for schema-on-read operations. Since the parameters are available that lead the user to think the schema can be explicitly specified, I think it would be good to get feedback from Microsoft so they can confirm.
Hi spencer_sa,
I tried your code on my side and I can reproduce this scenario. It seems like the 'write' delta table operations modify the field nullable property of the schema.
Current I checked the documents but not found them mentions these, perhaps you can try to report to dev team about this issues.
Regards,
Xiaoxin Sheng