d2l-ai/d2l-zh · error · AssertionError

Only zip/tar files can be extracted

Error message

Only zip/tar files can be extracted

What it means

download_extract in contrib/to-rm-mx-contrib-text/d2lzh/text/embedding.py downloads a pretrained-embedding file with gluon.utils.download and then dispatches on its extension: '.zip' opens a ZipFile, '.tar'/'.gz' open a tarfile. Any other extension hits assert False, 'Only zip/tar files can be extracted' — the helper has no code path for non-archive files. It fires when the requested pretrained file, after download, is not one of the archive formats the function knows how to extract.

Solutions

  1. Check the actual downloaded file's extension with os.path.splitext; if it is a raw embedding file (.vec/.txt/.bin/.npy), do not call download_extract — use download() and consume the file directly, since there is nothing to extract.
  2. If the archive uses a format the branch misses (e.g. '.tar.bz2', '.tgz', '.ZIP'), normalize the check: use fname.lower() and cover '.zip', '.tar', '.gz', '.tgz', '.bz2', '.xz' with the matching tarfile.open mode ('r:*' handles all tar variants).
  3. If you control PRETRAINED_FILE, point the entry at a .zip or .tar/.gz archive of the embedding so download_extract's contract is satisfied.
  4. Verify the download actually succeeded and is not an HTML error page saved under a wrong name (gluon.utils.download without a matching sha1_hash can save a redirect page whose extension still trips the assert).
  5. Prefer the newer d2l package's token-embedding APIs over the legacy contrib/to-rm-mx-contrib-text module if you are not bound to old Gluon code.

Example fix

# before
fname = download(embedding_name, pretrained_file_name)
data_dir, ext = os.path.splitext(fname)
if ext == '.zip':
    fp = zipfile.ZipFile(fname, 'r')
elif ext in ('.tar', '.gz'):
    fp = tarfile.open(fname, 'r')
else:
    assert False, 'Only zip/tar files can be extracted'

# after (tolerant extension handling, tarfile 'r:*' auto-detects compression)
fname = download(embedding_name, pretrained_file_name)
data_dir, ext = os.path.splitext(fname)
ext = ext.lower()
if ext == '.zip':
    fp = zipfile.ZipFile(fname, 'r')
elif ext in ('.tar', '.gz', '.tgz', '.bz2', '.xz'):
    fp = tarfile.open(fname, 'r:*')
else:
    raise ValueError(f'Nothing to extract for non-archive file {fname}; use download() directly')
fp.extractall(os.path.dirname(fname))
Defensive patterns

Strategy: validation

Validate before calling

import os

EXTRACTABLE = {'.zip', '.tar', '.gz', '.tgz', '.bz2', '.xz'}

fname = download(embedding_name, pretrained_file_name)
_, ext = os.path.splitext(fname)
if ext.lower() not in EXTRACTABLE:
    # raw embedding file: nothing to extract, consume it directly
    embedding_path = fname
else:
    embedding_path = download_extract(embedding_name, pretrained_file_name)

Type guard

def is_extractable_archive(path) -> bool:
    _, ext = os.path.splitext(path)
    return ext.lower() in ('.zip', '.tar', '.gz', '.tgz', '.bz2', '.xz')

# usage
fname = download(embedding_name, pretrained_file_name)
if not is_extractable_archive(fname):
    raise ValueError(f'{fname} is a raw file; call download() and skip extraction')

Try / catch

try:
    out_dir = download_extract(embedding_name, pretrained_file_name)
except AssertionError as e:
    fname = download(embedding_name, pretrained_file_name)
    raise RuntimeError(
        f'download_extract cannot handle {os.path.splitext(fname)[1]!r}; '
        f'if it is a raw embedding file, use download() and read it directly') from e

Prevention

When it happens

Trigger: Calling download_extract(embedding_name, pretrained_file_name) where PRETRAINED_FILE[embedding_name] maps to a plain (non-archive) file such as a '.bin', '.vec', '.txt', or '.npy' embedding; or passing a custom pretrained_file_name whose extension is not .zip/.tar/.gz. Also triggered by case-sensitivity mistakes ('.ZIP', '.TAR') since os.path.splitext comparison is exact, and by compound extensions like '.tar.bz2' (ext is '.bz2', not in the allowed set).

Common situations: Trying to load fastText/GloVe/word2vec files shipped as raw binaries or text instead of zips; adding a new embedding to PRETRAINED_FILE without checking how the file is actually packaged; renaming or re-hosting a pretrained file so its extension changes; older d2lzh versions whose extension list ('.tar', '.gz') misses formats like .bz2/.xz that newer archives use; typos or uppercase extensions in the file name map.

Related errors


AI-assisted analysis of d2l-ai/d2l-zh@e6b18ccea7 (2026-08-14). Data as JSON: /api/errors/9ecde8ad691b458e. Report an issue: GitHub.

Appendix: source

Thrown at contrib/to-rm-mx-contrib-text/d2lzh/text/embedding.py:41

    if not os.path.exists(path):
        os.makedirs(path)

def download(embedding_name, pretrained_file_name, cache_dir=os.path.join('..', 'data')):
    url, sha1 = PRETRAINED_FILE[embedding_name][pretrained_file_name]
    mkdir_if_not_exist(cache_dir)
    return gluon.utils.download(url, cache_dir, sha1_hash=sha1)

def download_extract(embedding_name, pretrained_file_name, folder=None):
    """Download and extract a zip/tar file."""
    fname = download(embedding_name, pretrained_file_name)
    base_dir = os.path.dirname(fname) 
    data_dir, ext = os.path.splitext(fname)
    if ext == '.zip':
        fp = zipfile.ZipFile(fname, 'r')
    elif ext in ('.tar', '.gz'):
        fp = tarfile.open(fname, 'r')
    else:
        assert False, 'Only zip/tar files can be extracted'
    fp.extractall(base_dir)
    if folder:
        return os.path.join(base_dir, folder)
    else:
        return data_dir
    
def get_pretrained_file_names(embedding_name=None):
    if embedding_name is not None:
        return PRETRAINED_FILE[embedding_name].keys()
    else:
        return PRETRAINED_FILE
    
def create(embedding_name, pretrained_file_name, vocabulary=None):
    return TokenEmbedding(embedding_name, pretrained_file_name.lower(), vocabulary)
    
class TokenEmbedding:
    """Token Embedding."""
    def __init__(self, embedding_name, pretrained_file_name, vocabulary=None):

View on GitHub (pinned to e6b18ccea7)