Forum Discussion

Amir_m_h's avatar
Amir_m_h
Regular Visitor
1 year ago
Solved

XML parsing fabric notebooks

Hi, TLDR: I have trouble in parsing an xml file in fabric lakehouse. Spark xml parser somehow misses _value of one of the tags. The xml file related to abstract looks like below and I have tro...
  • spencer_sa's avatar
    spencer_sa
    1 year ago

    How I personally would approach loading 8k files with inconsistent schemas;

    1. If I can't identify the schema from the filename/metadata, process each file separately.
    2. Convert each file to flat pyspark/pandas data frames - the schemas I've see so far are hierarchical which can be a bit painful to merge.
    3. Once everything has been flattened I'd then consider writing to either a single table or a single pyspark data frame making sure I allow the process to add new columns (easier when appending to a table, but not hard)
    4. I should now have a table with 1 row per entry, and null values where columns don't exist in one row but do in another.
    5. From this point onwards it's now a data processing/analysis problem ðŸ˜Š