StackPractices
beginner By Mathias Paulenko

Safely Extract Zip Files with Python

How to extract and validate zip archives securely using Python zipfile and shutil.

Overview

zf.extractall() on a user-supplied archive is one of those things that works fine in dev and then gets you a CVE. A malicious zip can carry path traversal entries (../../etc/passwd) or zip bombs that expand until they fill the disk. The zipfile module gives you everything needed to check an archive first, as long as you actually do it before writing a single byte.

When to Use

  • Extracting zip files that users upload to your app
  • Processing archives from sources you don’t control
  • Validating contents first (file count, total uncompressed size)
  • Pulling a few files out of an archive without unpacking all of it

Solution

Basic extraction

import zipfile

with zipfile.ZipFile("archive.zip", "r") as zf:
    zf.extractall("output_dir")

Safe extraction with path traversal protection

import zipfile
import os

def safe_extract(zip_path, extract_to):
    with zipfile.ZipFile(zip_path, "r") as zf:
        for member in zf.namelist():
            # Resolve the target path
            target = os.path.realpath(os.path.join(extract_to, member))

            # Ensure the target is inside the extraction directory
            if not target.startswith(os.path.realpath(extract_to) + os.sep):
                raise ValueError(f"Path traversal detected: {member}")

        # Only extract after validation passes
        zf.extractall(extract_to)

safe_extract("archive.zip", "output_dir")

Validate before extracting

import zipfile

def validate_zip(zip_path, max_files=1000, max_total_size_mb=500):
    with zipfile.ZipFile(zip_path, "r") as zf:
        files = zf.namelist()
        if len(files) > max_files:
            raise ValueError(f"Too many files: {len(files)} (max {max_files})")

        total_size = sum(info.file_size for info in zf.infolist())
        if total_size > max_total_size_mb * 1024 * 1024:
            raise ValueError(f"Archive too large: {total_size / 1024 / 1024:.1f}MB")

        # Check for suspicious entries
        for member in files:
            if member.startswith("/") or ".." in member:
                raise ValueError(f"Unsafe path in archive: {member}")

    return True

if validate_zip("archive.zip"):
    with zipfile.ZipFile("archive.zip", "r") as zf:
        zf.extractall("output_dir")

Extract specific files only

import zipfile

with zipfile.ZipFile("archive.zip", "r") as zf:
    # List all files
    for name in zf.namelist():
        print(name)

    # Extract only .csv files
    csv_files = [f for f in zf.namelist() if f.endswith(".csv")]
    for f in csv_files:
        zf.extract(f, "csv_output/")

Extract to memory without writing to disk

import zipfile

with zipfile.ZipFile("archive.zip", "r") as zf:
    with zf.open("data.json") as f:
        content = f.read()
        # Process content directly without writing to disk
        print(content[:200])

Explanation

The key detail: zipfile can read archive metadata (names, sizes, compression) without extracting anything, so you can inspect an archive before it touches the disk.

Path traversal is the classic trick, and the one that bites people most often. An entry named ../../etc/passwd makes extractall() write outside the target directory, and it happily obliges. The safe version resolves each member path and rejects anything that escapes.

Zip bombs are the other attack: a file weighing 42KB on disk that decompresses to petabytes in memory. Sum file_size across all entries before extracting and you’ll catch them.

The pipeline, end to end:

flowchart diagram: Zip received

Variants

ApproachSafetyUse When
extractall()NoneTrusted archives only
Safe extract with path checkHighUser uploads
Validate + extractHighestUntrusted sources
Extract to memoryHighProcessing without disk I/O

Guidelines

  • Never call extractall() on untrusted archives without validation.
  • Check total uncompressed size before extracting to avoid zip bombs.
  • Resolving with os.path.realpath() has a bonus: it also catches traversal through symlinks.
  • When you don’t need the file on disk, zf.open() reads it straight into memory.
  • Cap the file count too: legitimate archives almost never carry 10,000 entries.
  • Use Path.resolve() instead of os.path.realpath() for modern code. Path.resolve() handles symlinks and normalizes paths in one call, and works consistently across platforms:
from pathlib import Path

def is_safe_path(extract_dir: Path, member_name: str) -> bool:
    """Check if a zip member path is safe (no traversal)."""
    target = (extract_dir / member_name).resolve()
    try:
        target.relative_to(extract_dir.resolve())
        return True
    except ValueError:
        return False
  • Quarantine suspicious archives instead of deleting them. Move suspicious zips to a quarantine directory for later analysis. That way the evidence is still around when incident response needs it:
import logging
import shutil
from pathlib import Path

logger = logging.getLogger(__name__)
QUARANTINE_DIR = Path("/app/quarantine")

def quarantine_zip(zip_path: str, reason: str) -> str:
    """Move a suspicious zip to quarantine. Returns quarantine path."""
    QUARANTINE_DIR.mkdir(parents=True, exist_ok=True)
    dest = QUARANTINE_DIR / Path(zip_path).name
    shutil.move(zip_path, str(dest))
    logger.warning(f"Quarantined {zip_path}: {reason}")
    return str(dest)
  • Log extraction metadata for audit trails. Record who extracted what, when, and the hashes of extracted files if your compliance posture requires it (SOC 2, PCI-DSS):
import json
from datetime import datetime, timezone

def log_extraction_audit(zip_path: str, result: dict, user_id: str) -> None:
    """Write extraction audit log as JSON."""
    audit_entry = {
        "timestamp": datetime.now(timezone.utc).isoformat(),
        "user_id": user_id,
        "zip_path": zip_path,
        "extracted_count": result["extracted_count"],
        "total_uncompressed": result["total_uncompressed"],
        "file_hashes": result["hashes"],
    }
    with open("/var/log/zip_extraction_audit.jsonl", "a") as f:
        f.write(json.dumps(audit_entry) + "\n")

Common Mistakes

  • Calling extractall() directly on user uploads. This one accounts for most real-world zip extraction vulnerabilities.
  • Not checking file_size (uncompressed). A 1MB archive can easily hide entries that balloon into the GB range once decompressed.
  • Trusting member.startswith("..") checks alone. Symlinks and absolute paths slip past simple string checks.
  • Forgetting to handle password-protected archives. zf.extractall(pwd=b"secret") raises RuntimeError on wrong passwords.
  • Not closing the ZipFile context. Use with to ensure the file handle is released.
  • Not handling non-UTF8 filenames in zips. Zip files created on Windows may use CP437 or GBK encoding for filenames. Python’s zipfile assumes UTF-8 names, so a CP437-encoded archive can throw UnicodeDecodeError or hand you mojibake:
import zipfile

# Bad: default encoding may fail on non-UTF8 zips
# zf = zipfile.ZipFile("chinese_archive.zip", "r")
# names = zf.namelist()  # May raise or return garbled names

# Good: handle encoding errors gracefully
def safe_namelist(zf: zipfile.ZipFile) -> list[str]:
    """Get zip filenames with encoding fallback."""
    names = []
    for info in zf.infolist():
        try:
            # Try UTF-8 first (flag bit 0x800 indicates UTF-8)
            if info.flag_bits & 0x800:
                names.append(info.filename)
            else:
                # Decode as CP437 and re-encode for display
                raw = info.filename.encode("cp437")
                names.append(raw.decode("utf-8", errors="replace"))
        except Exception:
            names.append(info.filename.encode("ascii", errors="replace").decode("ascii"))
    return names
  • Extracting zips from untrusted sources without a timeout. A malicious zip can cause extraction to hang indefinitely. Note that signal.alarm only exists on Unix; on Windows you’d get the same effect by running the extraction in a subprocess with a timeout:
import signal
import zipfile

class TimeoutError(Exception):
    pass

def _timeout_handler(signum, frame):
    raise TimeoutError("Zip extraction timed out")

def extract_with_timeout(zip_path: str, dest: str, timeout_sec: int = 60) -> int:
    """Extract zip with a timeout to prevent hangs. Unix only."""
    signal.signal(signal.SIGALRM, _timeout_handler)
    signal.alarm(timeout_sec)
    try:
        with zipfile.ZipFile(zip_path, "r") as zf:
            zf.extractall(dest)
            return len(zf.namelist())
    finally:
        signal.alarm(0)  # Cancel the alarm

# extract_with_timeout("upload.zip", "/app/output", timeout_sec=30)
  • Not checking for duplicate filenames across zip entries. Some malicious zips include the same filename multiple times. The last extraction wins, which can overwrite a safe file with a malicious one:
import zipfile
from collections import Counter

def check_duplicates(zip_path: str) -> list[str]:
    """Find duplicate filenames in a zip archive."""
    with zipfile.ZipFile(zip_path, "r") as zf:
        names = [m for m in zf.namelist() if not m.endswith("/")]
    counts = Counter(names)
    return [name for name, count in counts.items() if count > 1]

# dupes = check_duplicates("archive.zip")
# if dupes:
#     raise ValueError(f"Duplicate entries found: {dupes}")

Frequently Asked Questions

How do I extract a password-protected zip?

Pass the password as bytes: zf.extractall("output", pwd=b"mypassword"). AES-encrypted zips need pyzipper; the stdlib zipfile won't open them.

How do I detect a zip bomb?

Check the compression ratio. When the uncompressed size runs past ~100x the compressed size, something is definitely off. There's a simpler second line of defense too: cap the total uncompressed size at something reasonable, say 500MB.

Can I extract .tar.gz files with zipfile?

No. Reach for the tarfile module instead. The API looks nearly identical too: tarfile.open("file.tar.gz", "r:gz"). For gzip-compressed single files rather than archives, see Compress and Decompress Files or the gzip compression recipe.

How do I create a zip file in Python?
import zipfile

with zipfile.ZipFile("output.zip", "w", zipfile.ZIP_DEFLATED) as zf:
    zf.write("file1.txt")
    zf.write("file2.txt")
How do I handle AES-encrypted zip files?

Python's stdlib zipfile only supports legacy ZipCrypto encryption. For AES-256 encrypted zips, use pyzipper:

from pyzipper import AESZipFile

with AESZipFile("encrypted.zip", "r", compression=pyzipper.ZIP_LZMA, encryption=pyzipper.WZ_AES) as zf:
    zf.setpassword(b"mypassword")
    zf.extractall("output_dir")
How do I extract only files modified after a certain date?

Use ZipInfo.date_time to filter entries by modification time:

import zipfile
from datetime import datetime

def extract_after_date(zip_path: str, dest: str, after: datetime) -> list[str]:
    """Extract only files modified after the given date."""
    extracted = []
    with zipfile.ZipFile(zip_path, "r") as zf:
        for info in zf.infolist():
            if info.is_dir():
                continue
            file_date = datetime(*info.date_time)
            if file_date > after:
                zf.extract(info, dest)
                extracted.append(info.filename)
    return extracted

# recent = extract_after_date("archive.zip", "/app/output", datetime(2025, 1, 1))
Is this approach safe enough for production uploads?

The validation pattern here (resolve each path, reject anything outside the target, cap file count and total size) covers the standard zip extraction attacks described in OWASP and CWE-22. What it doesn't cover is malware inside the archive: if you accept uploads from the public, scan extracted files with your normal AV pipeline, and quarantine anything the validator rejects instead of deleting it. When handling untrusted uploads, combine this with the checks in file upload validation.

How do I debug zip extraction issues?

Start with whatever error you got. "Bad zip file" usually means the file isn't a zip at all, so confirm with python -c "import zipfile; zipfile.ZipFile('file.zip').testzip()". If the path check rejects a file that looks harmless, print both resolved paths (os.path.realpath(extract_to) and the member's resolved target) to see why it escaped.

Garbled filenames almost always mean a non-UTF8 zip: info.flag_bits & 0x800 tells you whether the entry claims UTF-8. A RuntimeError about encryption means the archive wants a password (pwd=b"..."), or pyzipper when it's AES. Hanging extraction calls for the timeout pattern above, and "disk full" is what the size check in validate_zip exists to prevent, so run it before extracting, not after. Corrupted archives sometimes still open with allowZip64=True, and jar xf file.zip tolerates damage that Python won't. And permission errors are the simplest case of all: os.access(extract_to, os.W_OK) tells you straight away if the target is writable.