aio-libs/aiohttp · error · RuntimeError
Invalid default charset
Error message
Invalid default charset
What it means
In form-data processing, if the first part's field name is '_charset_', the reader treats its value as the default charset for subsequent parts. To prevent unbounded charset injection it reads at most 32 bytes; if the value exceeds 31 bytes, RuntimeError 'Invalid default charset' is raised.
Solutions
- Reject/validate multipart/form-data requests where the _charset_ field value exceeds a sane length (e.g. 31 bytes) before parsing.
- Ensure your producer only ever writes a real encoding label (utf-8, iso-8859-1, etc.) into _charset_.
- Catch RuntimeError around reader.next() and respond 400 Bad Request.
- Cap request body size and field sizes upstream so abuse is bounded.
Example fix
// before field name="_charset_" value = b"utf-8-very-long-invalid-label-xxxxxxxxxxxx" // after field name="_charset_" value = b"utf-8"
Defensive patterns
Strategy: validation
Validate before calling
MAX_CHARSET_LEN = 31
async def safe_next(reader):
part = await reader.next()
if (
isinstance(part, BodyPartReader)
and part.name == '_charset_'
):
probe = await part.read_chunk(MAX_CHARSET_LEN + 1)
if len(probe) > MAX_CHARSET_LEN:
raise web.HTTPBadRequest(text='_charset_ value too long')
part.unread_data(probe)
return part Type guard
def is_valid_charset_label(raw: bytes) -> bool:
return 0 < len(raw) <= 31 and raw.strip().isascii() Try / catch
try:
part = await reader.next()
except RuntimeError as e:
if 'Invalid default charset' in str(e):
return web.Response(status=400, text='Malformed _charset_ field')
raise Prevention
- Validate field names and sizes before parsing multipart/form-data.
- Only ever write a real encoding label into the _charset_ field on the producer side.
- Cap body and field sizes at the reverse proxy.
When it happens
Trigger: A multipart/form-data request whose first part is named '_charset_' and whose value (the declared charset) is 32 or more bytes long.
Common situations: Malicious or malformed requests exploiting the RFC 7578 _charset_ convention to inject a huge string; buggy producer writing the field name as the charset value.
Related errors
- data cannot be decoded with
- boundary %r is too long (70 chars max)
- Can not serialize value type: %r headers: %r value: %r
- content_type must be an instance of str. Got
- filename must be an instance of str. Got
AI-assisted analysis of aio-libs/aiohttp@d041d4d0fd (2026-08-11).
Data as JSON: /api/errors/1aa98f9eb00ca4f1.
Report an issue: GitHub.
Appendix: source
Thrown at aiohttp/multipart.py:775
await self._read_boundary()
if self._at_eof: # we just read the last boundary, nothing to do there
# https://github.com/python/mypy/issues/17537
return None # type: ignore[unreachable]
part = await self.fetch_next_part()
# https://datatracker.ietf.org/doc/html/rfc7578#section-4.6
if (
self._last_part is None
and self._mimetype.subtype == "form-data"
and isinstance(part, BodyPartReader)
):
_, params = parse_content_disposition(part.headers.get(CONTENT_DISPOSITION))
if params.get("name") == "_charset_":
# Longest encoding in https://encoding.spec.whatwg.org/encodings.json
# is 19 characters, so 32 should be more than enough for any valid encoding.
charset = await part.read_chunk(32)
if len(charset) > 31:
raise RuntimeError("Invalid default charset")
self._default_charset = charset.strip().decode()
part = await self.fetch_next_part()
self._last_part = part
return self._last_part
async def release(self) -> None:
"""Reads all the body parts to the void till the final boundary."""
while not self._at_eof:
item = await self.next()
if item is None:
break
await item.release()
async def fetch_next_part(
self,
) -> Union["MultipartReader", BodyPartReader]:
"""Returns the next body part reader."""
headers = await self._read_headers()View on GitHub (pinned to d041d4d0fd)