pypa/pip · error · ValueError
Unknown format for %r
Error message
Unknown format for %r
What it means
Raised as ValueError by distlib.util.unarchive when no format was passed (format=None) and the archive filename does not end with any recognized extension: '.zip', '.whl', '.tar.gz', '.tgz', '.tar.bz2', '.tbz', or '.tar'. The function auto-detects format from the extension; an unknown suffix leaves it unable to choose a reader. It is the fallback else-branch (pragma: no cover) of the extension dispatch.
Solutions
- Pass the format explicitly: unarchive(path, dest, format='zip') or 'tar'/'tgz'/'tbz'.
- Rename the file to a recognized extension before calling, or copy to a temp path with a proper suffix.
- Validate the extension against ARCHIVE_EXTENSIONS before invoking unarchive.
Example fix
# before
unarchive('data.xz', '/tmp/out')
# after
import lzma, shutil, tarfile
# decompress xz first, then:
unarchive('data.tar', '/tmp/out') Defensive patterns
Strategy: validation
Validate before calling
from distlib.util import ARCHIVE_EXTENSIONS
def detect_format(path):
for ext, fmt in [('.zip','zip'),('.whl','zip'),('.tar.gz','tgz'),('.tgz','tgz'),('.tar.bz2','tbz'),('.tbz','tbz'),('.tar','tar')]:
if path.endswith(ext):
return fmt
return None Type guard
def has_known_archive_ext(path: str) -> bool:
return isinstance(path, str) and any(path.endswith(e) for e in ARCHIVE_EXTENSIONS) Try / catch
try:
unarchive(path, dest)
except ValueError as e:
if 'Unknown format' in str(e):
unarchive(path, dest, format='tar') # only if you know the real format Prevention
- Pass format= explicitly when extension is non-standard.
- Rename/copy to a recognized suffix before unarchive.
When it happens
Trigger: unarchive('archive.xz', '/tmp/out') where '.xz' is unsupported; unarchive('download', '/tmp/out') with no extension; renamed/corrupt archive files whose suffix was stripped. Provide format explicitly to override auto-detection.
Common situations: Downloading archives with non-standard extensions; CI caching files under content-hash names with no suffix; user-supplied URLs whose basename lacks a recognized suffix.
Related errors
- path outside destination: %r
- Cannot determine archive format of
- file '%r' does not exist
- Ill-formed name/version string
- invalid glob %r: mismatching set marker
AI-assisted analysis of pypa/pip@f399c37189 (2026-08-08).
Data as JSON: /api/errors/c7dd3c078c17e2a0.
Report an issue: GitHub.
Appendix: source
Thrown at src/pip/_vendor/distlib/util.py:1271
check_path(member.linkname, base=link_base)
dest_dir = os.path.abspath(dest_dir)
plen = len(dest_dir)
archive = None
if format is None:
if archive_filename.endswith(('.zip', '.whl')):
format = 'zip'
elif archive_filename.endswith(('.tar.gz', '.tgz')):
format = 'tgz'
mode = 'r:gz'
elif archive_filename.endswith(('.tar.bz2', '.tbz')):
format = 'tbz'
mode = 'r:bz2'
elif archive_filename.endswith('.tar'):
format = 'tar'
mode = 'r'
else: # pragma: no cover
raise ValueError('Unknown format for %r' % archive_filename)
try:
if format == 'zip':
archive = ZipFile(archive_filename, 'r')
if check:
names = archive.namelist()
for name in names:
check_path(name)
else:
archive = tarfile.open(archive_filename, mode)
if check:
for member in archive.getmembers():
check_path(member.name)
check_link(member)
if format != 'zip' and sys.version_info[0] < 3:
# See Python issue 17153. If the dest path contains Unicode,
# tarfile extraction fails on Python 2.x if a member path name
# contains non-ASCII characters - it leads to an implicit
# bytes -> unicode conversion using ASCII to decode.View on GitHub (pinned to f399c37189)