Safely Extract Zip Files with Python
How to extract and validate zip archives securely using Python zipfile and shutil.
Overview
zf.extractall() on a user-supplied archive is one of those things that works fine in dev and then gets you a CVE. A malicious zip can carry path traversal entries (../../etc/passwd) or zip bombs that expand until they fill the disk. The zipfile module gives you everything needed to check an archive first, as long as you actually do it before writing a single byte.
When to Use
- Extracting zip files that users upload to your app
- Processing archives from sources you don’t control
- Validating contents first (file count, total uncompressed size)
- Pulling a few files out of an archive without unpacking all of it
Solution
Basic extraction
import zipfile
with zipfile.ZipFile("archive.zip", "r") as zf:
zf.extractall("output_dir")
Safe extraction with path traversal protection
import zipfile
import os
def safe_extract(zip_path, extract_to):
with zipfile.ZipFile(zip_path, "r") as zf:
for member in zf.namelist():
# Resolve the target path
target = os.path.realpath(os.path.join(extract_to, member))
# Ensure the target is inside the extraction directory
if not target.startswith(os.path.realpath(extract_to) + os.sep):
raise ValueError(f"Path traversal detected: {member}")
# Only extract after validation passes
zf.extractall(extract_to)
safe_extract("archive.zip", "output_dir")
Validate before extracting
import zipfile
def validate_zip(zip_path, max_files=1000, max_total_size_mb=500):
with zipfile.ZipFile(zip_path, "r") as zf:
files = zf.namelist()
if len(files) > max_files:
raise ValueError(f"Too many files: {len(files)} (max {max_files})")
total_size = sum(info.file_size for info in zf.infolist())
if total_size > max_total_size_mb * 1024 * 1024:
raise ValueError(f"Archive too large: {total_size / 1024 / 1024:.1f}MB")
# Check for suspicious entries
for member in files:
if member.startswith("/") or ".." in member:
raise ValueError(f"Unsafe path in archive: {member}")
return True
if validate_zip("archive.zip"):
with zipfile.ZipFile("archive.zip", "r") as zf:
zf.extractall("output_dir")
Extract specific files only
import zipfile
with zipfile.ZipFile("archive.zip", "r") as zf:
# List all files
for name in zf.namelist():
print(name)
# Extract only .csv files
csv_files = [f for f in zf.namelist() if f.endswith(".csv")]
for f in csv_files:
zf.extract(f, "csv_output/")
Extract to memory without writing to disk
import zipfile
with zipfile.ZipFile("archive.zip", "r") as zf:
with zf.open("data.json") as f:
content = f.read()
# Process content directly without writing to disk
print(content[:200])
Explanation
The key detail: zipfile can read archive metadata (names, sizes, compression) without extracting anything, so you can inspect an archive before it touches the disk.
Path traversal is the classic trick, and the one that bites people most often. An entry named ../../etc/passwd makes extractall() write outside the target directory, and it happily obliges. The safe version resolves each member path and rejects anything that escapes.
Zip bombs are the other attack: a file weighing 42KB on disk that decompresses to petabytes in memory. Sum file_size across all entries before extracting and you’ll catch them.
The pipeline, end to end:
Variants
| Approach | Safety | Use When |
|---|---|---|
| extractall() | None | Trusted archives only |
| Safe extract with path check | High | User uploads |
| Validate + extract | Highest | Untrusted sources |
| Extract to memory | High | Processing without disk I/O |
Guidelines
- Never call
extractall()on untrusted archives without validation. - Check total uncompressed size before extracting to avoid zip bombs.
- Resolving with
os.path.realpath()has a bonus: it also catches traversal through symlinks. - When you don’t need the file on disk,
zf.open()reads it straight into memory. - Cap the file count too: legitimate archives almost never carry 10,000 entries.
- Use
Path.resolve()instead ofos.path.realpath()for modern code.Path.resolve()handles symlinks and normalizes paths in one call, and works consistently across platforms:
from pathlib import Path
def is_safe_path(extract_dir: Path, member_name: str) -> bool:
"""Check if a zip member path is safe (no traversal)."""
target = (extract_dir / member_name).resolve()
try:
target.relative_to(extract_dir.resolve())
return True
except ValueError:
return False
- Quarantine suspicious archives instead of deleting them. Move suspicious zips to a quarantine directory for later analysis. That way the evidence is still around when incident response needs it:
import logging
import shutil
from pathlib import Path
logger = logging.getLogger(__name__)
QUARANTINE_DIR = Path("/app/quarantine")
def quarantine_zip(zip_path: str, reason: str) -> str:
"""Move a suspicious zip to quarantine. Returns quarantine path."""
QUARANTINE_DIR.mkdir(parents=True, exist_ok=True)
dest = QUARANTINE_DIR / Path(zip_path).name
shutil.move(zip_path, str(dest))
logger.warning(f"Quarantined {zip_path}: {reason}")
return str(dest)
- Log extraction metadata for audit trails. Record who extracted what, when, and the hashes of extracted files if your compliance posture requires it (SOC 2, PCI-DSS):
import json
from datetime import datetime, timezone
def log_extraction_audit(zip_path: str, result: dict, user_id: str) -> None:
"""Write extraction audit log as JSON."""
audit_entry = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"user_id": user_id,
"zip_path": zip_path,
"extracted_count": result["extracted_count"],
"total_uncompressed": result["total_uncompressed"],
"file_hashes": result["hashes"],
}
with open("/var/log/zip_extraction_audit.jsonl", "a") as f:
f.write(json.dumps(audit_entry) + "\n")
Common Mistakes
- Calling
extractall()directly on user uploads. This one accounts for most real-world zip extraction vulnerabilities. - Not checking
file_size(uncompressed). A 1MB archive can easily hide entries that balloon into the GB range once decompressed. - Trusting
member.startswith("..")checks alone. Symlinks and absolute paths slip past simple string checks. - Forgetting to handle password-protected archives.
zf.extractall(pwd=b"secret")raisesRuntimeErroron wrong passwords. - Not closing the ZipFile context. Use
withto ensure the file handle is released. - Not handling non-UTF8 filenames in zips. Zip files created on Windows may use CP437 or GBK encoding for filenames. Python’s
zipfileassumes UTF-8 names, so a CP437-encoded archive can throwUnicodeDecodeErroror hand you mojibake:
import zipfile
# Bad: default encoding may fail on non-UTF8 zips
# zf = zipfile.ZipFile("chinese_archive.zip", "r")
# names = zf.namelist() # May raise or return garbled names
# Good: handle encoding errors gracefully
def safe_namelist(zf: zipfile.ZipFile) -> list[str]:
"""Get zip filenames with encoding fallback."""
names = []
for info in zf.infolist():
try:
# Try UTF-8 first (flag bit 0x800 indicates UTF-8)
if info.flag_bits & 0x800:
names.append(info.filename)
else:
# Decode as CP437 and re-encode for display
raw = info.filename.encode("cp437")
names.append(raw.decode("utf-8", errors="replace"))
except Exception:
names.append(info.filename.encode("ascii", errors="replace").decode("ascii"))
return names
- Extracting zips from untrusted sources without a timeout. A malicious zip can cause extraction to hang indefinitely. Note that
signal.alarmonly exists on Unix; on Windows you’d get the same effect by running the extraction in a subprocess with a timeout:
import signal
import zipfile
class TimeoutError(Exception):
pass
def _timeout_handler(signum, frame):
raise TimeoutError("Zip extraction timed out")
def extract_with_timeout(zip_path: str, dest: str, timeout_sec: int = 60) -> int:
"""Extract zip with a timeout to prevent hangs. Unix only."""
signal.signal(signal.SIGALRM, _timeout_handler)
signal.alarm(timeout_sec)
try:
with zipfile.ZipFile(zip_path, "r") as zf:
zf.extractall(dest)
return len(zf.namelist())
finally:
signal.alarm(0) # Cancel the alarm
# extract_with_timeout("upload.zip", "/app/output", timeout_sec=30)
- Not checking for duplicate filenames across zip entries. Some malicious zips include the same filename multiple times. The last extraction wins, which can overwrite a safe file with a malicious one:
import zipfile
from collections import Counter
def check_duplicates(zip_path: str) -> list[str]:
"""Find duplicate filenames in a zip archive."""
with zipfile.ZipFile(zip_path, "r") as zf:
names = [m for m in zf.namelist() if not m.endswith("/")]
counts = Counter(names)
return [name for name, count in counts.items() if count > 1]
# dupes = check_duplicates("archive.zip")
# if dupes:
# raise ValueError(f"Duplicate entries found: {dupes}") Frequently Asked Questions
How do I extract a password-protected zip?
Pass the password as bytes: zf.extractall("output", pwd=b"mypassword"). AES-encrypted zips need pyzipper; the stdlib zipfile won't open them.
How do I detect a zip bomb?
Check the compression ratio. When the uncompressed size runs past ~100x the compressed size, something is definitely off. There's a simpler second line of defense too: cap the total uncompressed size at something reasonable, say 500MB.
Can I extract .tar.gz files with zipfile?
No. Reach for the tarfile module instead. The API looks nearly identical too: tarfile.open("file.tar.gz", "r:gz"). For gzip-compressed single files rather than archives, see Compress and Decompress Files or the gzip compression recipe.
How do I create a zip file in Python?
import zipfile
with zipfile.ZipFile("output.zip", "w", zipfile.ZIP_DEFLATED) as zf:
zf.write("file1.txt")
zf.write("file2.txt")
How do I handle AES-encrypted zip files?
Python's stdlib zipfile only supports legacy ZipCrypto encryption. For AES-256 encrypted zips, use pyzipper:
from pyzipper import AESZipFile
with AESZipFile("encrypted.zip", "r", compression=pyzipper.ZIP_LZMA, encryption=pyzipper.WZ_AES) as zf:
zf.setpassword(b"mypassword")
zf.extractall("output_dir")
How do I extract only files modified after a certain date?
Use ZipInfo.date_time to filter entries by modification time:
import zipfile
from datetime import datetime
def extract_after_date(zip_path: str, dest: str, after: datetime) -> list[str]:
"""Extract only files modified after the given date."""
extracted = []
with zipfile.ZipFile(zip_path, "r") as zf:
for info in zf.infolist():
if info.is_dir():
continue
file_date = datetime(*info.date_time)
if file_date > after:
zf.extract(info, dest)
extracted.append(info.filename)
return extracted
# recent = extract_after_date("archive.zip", "/app/output", datetime(2025, 1, 1))
Is this approach safe enough for production uploads?
The validation pattern here (resolve each path, reject anything outside the target, cap file count and total size) covers the standard zip extraction attacks described in OWASP and CWE-22. What it doesn't cover is malware inside the archive: if you accept uploads from the public, scan extracted files with your normal AV pipeline, and quarantine anything the validator rejects instead of deleting it. When handling untrusted uploads, combine this with the checks in file upload validation.
How do I debug zip extraction issues?
Start with whatever error you got. "Bad zip file" usually means the file isn't a zip at all, so confirm with python -c "import zipfile; zipfile.ZipFile('file.zip').testzip()". If the path check rejects a file that looks harmless, print both resolved paths (os.path.realpath(extract_to) and the member's resolved target) to see why it escaped.
Garbled filenames almost always mean a non-UTF8 zip: info.flag_bits & 0x800 tells you whether the entry claims UTF-8. A RuntimeError about encryption means the archive wants a password (pwd=b"..."), or pyzipper when it's AES. Hanging extraction calls for the timeout pattern above, and "disk full" is what the size check in validate_zip exists to prevent, so run it before extracting, not after. Corrupted archives sometimes still open with allowZip64=True, and jar xf file.zip tolerates damage that Python won't. And permission errors are the simplest case of all: os.access(extract_to, os.W_OK) tells you straight away if the target is writable.
Related Resources
Compress and Decompress Files
How to handle ZIP, GZIP, and TAR archives programmatically.
RecipeFile Upload Validation
How to handle file uploads securely with size, type, and content validation.
RecipeCompress and Decompress Files with Gzip and Brotli
How to reduce file sizes for APIs, static assets, and log files using Gzip, Brotli, and zlib with streaming compression, content negotiation, and what works.
RecipeCopy and Move Files Safely in Python, JS, Java, and Bash
Learn to copy and move files across platforms with Python, JavaScript, Java, and Bash. Includes atomic moves, checksums, symlinks, and batch patterns.
RecipeGenerate Temporary Files
How to create temporary files and directories safely with automatic cleanup across Python, Node.js, Java, and Bash.