Forum Discussion
Nullable = False in schema not correctly preserved / written using pyspark
- 1 year ago
Hello spencer_sa - I've done some more research on this. Based on my findings, while PySpark does include the ability to specify the schema, I believe that when writing a delta table to a lakehouse in this way, the nullability specifications in the schema are not preserved - and all columns set to nullable so that it is optimized for schema-on-read operations. Since the parameters are available that lead the user to think the schema can be explicitly specified, I think it would be good to get feedback from Microsoft so they can confirm.
Addressing both of your points I've amended the code and am still seeing the schema change issue.
Specifically I've readded the sample data I removed before I posted this the first time and I've added the overwriteSchema option to the write. (I also deleted the existing table to start from a clean slate)
from pyspark.sql.types import StructType, StructField, IntegerType, StringType
schema = StructType([
StructField("id", IntegerType(), nullable=False),
StructField("name", StringType(), nullable=True)
])
# Now with added data
data = [{'id': 1, 'name': 'Alice'},{'id': 2, 'name': 'Bob'}]
df = spark.createDataFrame(data,schema)
print(df.schema)
# And now overwriting the schema - theoretically
df.write.format('delta').mode('overwrite').option('overwriteSchema', 'true').save('Tables/pyspark_version')
df2 = spark.read.format('delta').load('Tables/pyspark_version')
print(df2.schema)results in;
StructType([StructField('id', IntegerType(), False), StructField('name', StringType(), True)])
StructType([StructField('id', IntegerType(), True), StructField('name', StringType(), True)])Hello spencer_sa - I've done some more research on this. Based on my findings, while PySpark does include the ability to specify the schema, I believe that when writing a delta table to a lakehouse in this way, the nullability specifications in the schema are not preserved - and all columns set to nullable so that it is optimized for schema-on-read operations. Since the parameters are available that lead the user to think the schema can be explicitly specified, I think it would be good to get feedback from Microsoft so they can confirm.
- spencer_sa1 year agoImpactful Individual
Already got a ticket in for this - can't get to it 'til I'm back from annual leave.
I probably found the same articles you've had - I just generated some sample code to reproduce it- asanz931 year agoNew Member
Hello, did you solve the issue? I don't see any clear approach