Forum Discussion
How do I unzip a .gz file?
I used a data pipeline to make a web call (http) and get a file.
The file has been downloaded to the Files area in the lakehouse. How can I uncompress this file?
I am using a PySpark notebook to unzip this file. The file itself is good, since I was able to download the file and uncompress it on my windows machine.
I tried this code, but that fails.
df = spark.read.format("json").option("multiLine", "false").load("Files/bronze/github-events-2025-01-15-12.json.gz")
# Display a preview
display(df.limit(5))
# Save as Delta table
df.write.mode("overwrite").format("delta").saveAsTable("github_events_bronze")
print(f"Successfully loaded {df.count()} records")Thanks.
My issue was slightly different.
My unzipping wasn't working because these were double zipped files.
This is the code I used:import gzip import os # Local path to bronze folder bronze_local_path = "/lakehouse/default/Files/bronze" # List .json.gz files gz_files = sorted([f for f in os.listdir(bronze_local_path) if f.endswith('.json.gz')]) print(f"Found {len(gz_files)} files to process\n") extracted_json_files = [] for gz_file in gz_files: gz_file_path = os.path.join(bronze_local_path, gz_file) json_file_path = gz_file_path.replace('.json.gz', '.json') print(f"Processing: {gz_file}") try: # Read and decompress with open(gz_file_path, "rb") as f: gz_data = f.read() decompressed = gzip.decompress(gz_data) # Check for double compression if decompressed[:2] == b'\x1f\x8b': print(f" Detected double-compression. Decompressing again...") decompressed = gzip.decompress(decompressed) # Write decompressed JSON locally with open(json_file_path, "wb") as f: f.write(decompressed) size_mb = len(decompressed) / (1024 ** 2) print(f" ✓ Decompressed: {size_mb:.2f} MB") extracted_json_files.append(json_file_path) except Exception as e: print(f" ✗ Error: {str(e)}") print(f"\n✓ Successfully extracted {len(extracted_json_files)} JSON files")
4 Replies
- sandeephijam
Helper II
Hi abhidotnet ,
Hi Abhi please see if you share the error message to understand the failure,
As an immediate check you can do the following to check
File existance : display(mssparkutils.fs.ls("Files")) and then display(mssparkutils.fs.ls("Files/bronze"))
Also check if we have attached the lakehouse
- tayloramy
Super User
Hi abhidotnet,
You can use the gzip package in Python to unzip a .gz file.
import gzip
import shutil# Define paths (e.g., in your Lakehouse Files section)
input_path = "/lakehouse/default/Files/my_data.csv.gz"
output_path = "/lakehouse/default/Files/my_data.csv"# Decompress the .gz file
with gzip.open(input_path, "rb") as f_in:
with open(output_path, "wb") as f_out:
shutil.copyfileobj(f_in, f_out)
print(f"Successfully unzipped to {output_path}") - chaitanyacd3005New Member
I think you should decompress the .gz file first using python's built-in gzip library and then read the decompressed json file using spark.
- abhidotnet
Advocate II
Thanks.
My issue was slightly different.
My unzipping wasn't working because these were double zipped files.
This is the code I used:import gzip import os # Local path to bronze folder bronze_local_path = "/lakehouse/default/Files/bronze" # List .json.gz files gz_files = sorted([f for f in os.listdir(bronze_local_path) if f.endswith('.json.gz')]) print(f"Found {len(gz_files)} files to process\n") extracted_json_files = [] for gz_file in gz_files: gz_file_path = os.path.join(bronze_local_path, gz_file) json_file_path = gz_file_path.replace('.json.gz', '.json') print(f"Processing: {gz_file}") try: # Read and decompress with open(gz_file_path, "rb") as f: gz_data = f.read() decompressed = gzip.decompress(gz_data) # Check for double compression if decompressed[:2] == b'\x1f\x8b': print(f" Detected double-compression. Decompressing again...") decompressed = gzip.decompress(decompressed) # Write decompressed JSON locally with open(json_file_path, "wb") as f: f.write(decompressed) size_mb = len(decompressed) / (1024 ** 2) print(f" ✓ Decompressed: {size_mb:.2f} MB") extracted_json_files.append(json_file_path) except Exception as e: print(f" ✗ Error: {str(e)}") print(f"\n✓ Successfully extracted {len(extracted_json_files)} JSON files")