Forum Discussion
silly_bird
1 year agoRegular Visitor
Help extracting value from dict in a column
Hi all I'm working on API integration in PySpark notebook and there is a column with email & phone that is an array with random order contactMethods = [{'name': 'Email', 'value': 'e...
- 1 year ago
It looks like I managed to figure one solution.
Don't know good or bad, it is my only one
from pyspark.sql.functions import col, udf from pyspark.sql.types import StringType extract_email = udf(lambda cell: str(next(filter(lambda t: t["name"] == "Email", cell), {}).get("value", "")), StringType()) extract_mobile = udf(lambda cell: str(next(filter(lambda t: t["name"] == "Mobile", cell), {}).get("value", "")), StringType()) df = df.withColumn('email', extract_email(col("contactMethods"))).withColumn('mobile', extract_mobile(col("contactMethods"))) display(df)Profies, please advise!
silly_bird
1 year agoRegular Visitor
Update
If we read data from json, like that
df = spark.read.json(spark.sparkContext.parallelize([response.json()])).head(1)
.. then cell is a an array of Row objects, not an array of dict
I managed to workaround using asDict method on the row
extract_email = udf(lambda cell: None if cell is None else next(filter(lambda t: t["name"] == "Email", cell), Row(value=None)).asDict().get("value", None), StringType())