Forum Discussion
Writing to Lakehouse / OneLake using ABFS
- 4 months ago
Oh I see...
So, the asymmetry comes from what adlfs is built on. The abfs driver wraps BlobServiceClient, not DataLakeServiceClient. Reads happen to translate into requests OneLake's DFS endpoint can serve, so they go through.
Writes go through Block Blob operations (Put Block / Put Block List), which are part of the Blob API surface that OneLake does not expose. OneLake only implements the ADLS Gen2 DFS surface (create, append, flush).
Same account_host, different API underneath, that is why one direction works and the other does not.
For the on-the-fly case you do not actually need a mounted PyFileSystem. ParquetWriter accepts any file-like object with write and close, so you can wrap a DataLakeFileClient in a small append-only stream that calls append_data on every write and flush_data on close, and hand that straight to the writer.
Nothing hits disk and nothing buffers the full file in memory. Parquet writes sequentially and only finalizes the footer on close, so append-only is safe, no seek needed.
If you really need an fsspec-shaped object because something downstream insists on it, the same wrapper goes inside an AbstractFileSystem subclass returning it from _open(..., mode="wb"). For ParquetWriter itself that is overkill.
Sorry if this is a bit too long 😅 nevertheless I hope it kinda gives a good direction.
If that fixes it, a thumbs up and marking as the solution would be appreciated.
Best regards,
Shai Karmani
Oh I see...
So, the asymmetry comes from what adlfs is built on. The abfs driver wraps BlobServiceClient, not DataLakeServiceClient. Reads happen to translate into requests OneLake's DFS endpoint can serve, so they go through.
Writes go through Block Blob operations (Put Block / Put Block List), which are part of the Blob API surface that OneLake does not expose. OneLake only implements the ADLS Gen2 DFS surface (create, append, flush).
Same account_host, different API underneath, that is why one direction works and the other does not.
For the on-the-fly case you do not actually need a mounted PyFileSystem. ParquetWriter accepts any file-like object with write and close, so you can wrap a DataLakeFileClient in a small append-only stream that calls append_data on every write and flush_data on close, and hand that straight to the writer.
Nothing hits disk and nothing buffers the full file in memory. Parquet writes sequentially and only finalizes the footer on close, so append-only is safe, no seek needed.
If you really need an fsspec-shaped object because something downstream insists on it, the same wrapper goes inside an AbstractFileSystem subclass returning it from _open(..., mode="wb"). For ParquetWriter itself that is overkill.
Sorry if this is a bit too long 😅 nevertheless I hope it kinda gives a good direction.
If that fixes it, a thumbs up and marking as the solution would be appreciated.
Best regards,
Shai Karmani
Hi Shai_Karmani , thanks for the super input. I'll try wrapping your functionality into a custom Filesystem object with write and close. Will let you know how it goes 🙂
Best regards,
Ilias