{"record":{"id":"51d6233dfaa8735b","repo":"aio-libs/aiohttp","slug":"1007-51d623","errorCode":"1007","errorMessage":"Invalid UTF-8 text message","messagePattern":"Invalid UTF-8 text message","errorType":"exception","errorClass":"WebSocketError","httpStatus":null,"severity":"error","filePath":"aiohttp/_websocket/reader_py.py","lineNumber":281,"sourceCode":"                        \"Compressed message has too many deflate members\",\n                    ) from exc\n                if self._max_msg_size and len(payload_merged) > self._max_msg_size:\n                    raise WebSocketError(\n                        WSCloseCode.MESSAGE_TOO_BIG,\n                        f\"Decompressed message exceeds size limit {self._max_msg_size}\",\n                    )\n            elif type(assembled_payload) is bytes:\n                payload_merged = assembled_payload\n            else:\n                payload_merged = bytes(assembled_payload)\n\n            size = len(payload_merged)\n            if opcode == OP_CODE_TEXT:\n                if self._decode_text:\n                    try:\n                        text = payload_merged.decode(\"utf-8\")\n                    except UnicodeDecodeError as exc:\n                        raise WebSocketError(\n                            WSCloseCode.INVALID_TEXT, \"Invalid UTF-8 text message\"\n                        ) from exc\n\n                    # XXX: The Text and Binary messages here can be a performance\n                    # bottleneck, so we use tuple.__new__ to improve performance.\n                    # This is not type safe, but many tests should fail in\n                    # test_client_ws_functional.py if this is wrong.\n                    msg = TUPLE_NEW(WSMessageText, (text, size, \"\", WS_MSG_TYPE_TEXT))\n                else:\n                    # Return raw bytes for TEXT messages when decode_text=False\n                    msg = TUPLE_NEW(\n                        WSMessageTextBytes, (payload_merged, size, \"\", WS_MSG_TYPE_TEXT)\n                    )\n            else:\n                msg = TUPLE_NEW(\n                    WSMessageBinary, (payload_merged, size, \"\", WS_MSG_TYPE_BINARY)\n                )\n","sourceCodeStart":263,"sourceCodeEnd":299,"githubUrl":"https://github.com/aio-libs/aiohttp/blob/d041d4d0fd48c3f0832084d33be16cf1c4835f85/aiohttp/_websocket/reader_py.py#L263-L299","documentation":"Raised as WebSocketError code 1007 (INVALID_TEXT) when a TEXT frame's assembled payload fails UTF-8 decoding (payload_merged.decode('utf-8') raises UnicodeDecodeError). RFC 6455 §8.1 mandates that the connection be closed with code 1007 when a text message contains invalid UTF-8. This path only runs when self._decode_text is True (the default); with decode_text=False the raw bytes are returned and no validation occurs here.","triggerScenarios":"A peer sends a TEXT frame (opcode 0x1) whose payload is not valid UTF-8 — e.g. a truncated multibyte sequence, lone surrogate bytes, or binary accidentally labelled as text. _handle_frame assembles the payload and calls .decode('utf-8'), which throws.","commonSituations":"A client serializes binary data into a TEXT frame by mistake; a frame is truncated by a buggy intermediary or network layer cutting a multibyte character; a custom encoder producing non-UTF-8 bytes; partial/fragmented text whose split point lands inside a multibyte sequence at the application layer (the reader reassembles correctly, but a hand-built sender can split badly).","solutions":["Ensure the peer only sends valid UTF-8 in TEXT frames; use BINARY frames (opcode 0x2) for non-text payloads.","If you are the sender, encode with a strict UTF-8 encoder (data.encode('utf-8')) and never split a multibyte character across fragments.","On the receiver side, catch WebSocketError and close with code 1007.","If you intentionally exchange raw bytes, use BINARY frames or set decode_text=False on the reader (Python-only, bypasses validation)."],"exampleFix":"// before: sender labels binary as text\nws.send_str(b'\\xff\\xfe ...')  // not valid utf-8\n// after: use bytes via send_bytes / BINARY\nws.send_bytes(raw_payload)\n// or ensure utf-8 on the text path\nws.send_str(text.decode('utf-8','replace') if unsure else text)","handlingStrategy":"try-catch","validationCode":"// Sender: ensure strict UTF-8 before sending text.\ntry:\n    text.encode('utf-8')  # validates on the send side\nexcept UnicodeEncodeError:\n    await ws.send_bytes(payload)  # use BINARY instead\n","typeGuard":"def is_valid_utf8(b: bytes) -> bool:\n    try:\n        b.decode('utf-8'); return True\n    except UnicodeDecodeError:\n        return False\n","tryCatchPattern":"try:\n    msg = await ws.receive()\nexcept WebSocketError as exc:\n    if exc.code == WSCloseCode.INVALID_TEXT:\n        await ws.close(code=WSCloseCode.INVALID_TEXT)\n","preventionTips":["Use BINARY frames for non-text payloads.","Encode all TEXT payloads with .encode('utf-8') and never split multibyte characters across fragments.","For raw-byte text channels, set decode_text=False deliberately and document why."],"tags":["websocket","utf-8","text-frame","rfc6455","validation","invalid-text"],"backgroundTag":null,"analyzedSha":"d041d4d0fd48c3f0832084d33be16cf1c4835f85","analyzedAt":"2026-08-11T20:44:15.550Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}