Forum Discussion
PySpark Notebook to process complex JSON
- Anonymous2 years ago
Hi ToddChitt ,
I tried to do some repro around your case, it is working perfectly fine.
Can you please find the code below,Sample Json:
[ { "id": "00000001-0000-0000-0000-000000000000", "positionData": { "manager": { "id": "00000002-0000-0000-0000-000000000000", "employeeNumber": "1234" } } }, { "id": "00000001-0000-0000-0000-000000000000", "positionData": { "manager": null } } ]
Sample Code:from pyspark.sql.types import StructType, StructField, StringType, ArrayType from pyspark.sql.functions import col, coalesce, lit # Define the schema for the nested objects schema = StructType([ StructField("id", StringType(), True), StructField("positionData", StructType([ StructField("manager", StructType([ StructField("id", StringType(), True), StructField("employeeNumber", StringType(), True) ]), True) ]), True) ]) # Read JSON data with multiline option and schema df = spark.read.option("multiline", "true").json("Files/testing.json", schema=schema) df = df.withColumn("ManagerId", coalesce(col("positionData.manager.id"), lit(None))) display(df)
Please try this and let me know if you have further queries.
Anonymous Unfortunately, the COALESCE function does the same thing as the WHEN / EXCEPT function: It evaluates all paths offered, even though it is only going to take ONE of those paths. In my case, one of the paths will result in an error as shown above, and even though the logic of the function is such that it will not return a certain element, it still needs to evaluate it.
a bit of 'airware' example:
COALESCE ( NULL, "some string not null", 1/0)
This will error out on the 1 devided by zero path even though the logic is to return "some string not null".
Any other suggestions?
Hi ToddChitt ,
Can you please share the output json, so I can try it at my end and may suggest you?