Forum Discussion
clusterBy does not work in dataframe API?
- 1 year ago
Hi smpa01 ,
the .clusterBy() method on DataFrameWriter is not supported because Fabric uses a customized Spark runtime that limits certain APIs to ensure simplicity and compatibility within its managed environment. Unlike Databricks, which offers extended Delta Lake features directly through the PySpark DataFrameWriter, Fabric restricts clustering capabilities to SQL DDL and the DeltaTable API
Thanks,
Prashanth Are
MS Fabric community support
Hi smpa01 ,
the .clusterBy() method on DataFrameWriter is not supported because Fabric uses a customized Spark runtime that limits certain APIs to ensure simplicity and compatibility within its managed environment. Unlike Databricks, which offers extended Delta Lake features directly through the PySpark DataFrameWriter, Fabric restricts clustering capabilities to SQL DDL and the DeltaTable API
Thanks,
Prashanth Are
MS Fabric community support
v-prasare without clusterBy in dataframe writer, I am guessing the clusterd files can't be written, if one intends to write only the raw files with the intention to create an external table.
df.write\
.format("delta")\
.mode("append")\
.clustrBy (cluster by fields)\
.save(file_path)