Forum Discussion
Amir_m_h
1 year agoRegular Visitor
XML parsing fabric notebooks
Hi, TLDR: I have trouble in parsing an xml file in fabric lakehouse. Spark xml parser somehow misses _value of one of the tags. The xml file related to abstract looks like below and I have tro...
- 1 year ago
How I personally would approach loading 8k files with inconsistent schemas;
- If I can't identify the schema from the filename/metadata, process each file separately.
- Convert each file to flat pyspark/pandas data frames - the schemas I've see so far are hierarchical which can be a bit painful to merge.
- Once everything has been flattened I'd then consider writing to either a single table or a single pyspark data frame making sure I allow the process to add new columns (easier when appending to a table, but not hard)
- I should now have a table with 1 row per entry, and null values where columns don't exist in one row but do in another.
- From this point onwards it's now a data processing/analysis problem 😊
Anonymous
1 year agoNot applicable
Hi Amir_m_h
I wanted to check if you had the opportunity to review the information provided spencer_sa . Please feel free to contact us if you have any further questions. If your issue has been resolved ,please mark the helpful reply and accept it as the solution. This will be helpful for other community members who have similar problems to solve it faster.
Thank you.