{"record":{"id":"13db22c40d5ffe11","repo":"vectordotdev/vector","slug":"unable-to-decode-message-as-utf8","errorCode":null,"errorMessage":"Unable to decode message as UTF8","messagePattern":"Unable to decode message as UTF8","errorType":"exception","errorClass":"LinesCodecError","httpStatus":null,"severity":"error","filePath":"lib/codecs/src/decoding/framing/octet_counting.rs","lineNumber":178,"sourceCode":"                if len > self.other.max_length() {\n                    // The length is greater than we want.\n                    //\n                    // We need to discard the entire message.\n                    self.octet_decoding = Some(State::Discarding(len));\n                    src.advance(space_pos + 1);\n\n                    Ok(None)\n                } else if let Some(msg) = src.get(from..to) {\n                    let bytes = match std::str::from_utf8(msg) {\n                        Ok(_) => Bytes::copy_from_slice(msg),\n                        Err(_) => {\n                            // The data was not valid UTF8 :-(.\n                            //\n                            // Advance the buffer past the erroneous bytes to\n                            // prevent us getting stuck in an infinite loop.\n                            src.advance(to);\n                            self.octet_decoding = None;\n                            return Err(LinesCodecError::Io(io::Error::new(\n                                io::ErrorKind::InvalidData,\n                                \"Unable to decode message as UTF8\",\n                            )));\n                        }\n                    };\n\n                    // We have managed to read the entire message as valid UTF8!\n                    src.advance(to);\n                    self.octet_decoding = None;\n                    Ok(Some(bytes))\n                } else {\n                    // We have an acceptable number of bytes in this message,\n                    // but not all the data was in the frame.\n                    //\n                    // Return `None` to indicate we want more data before we do\n                    // anything else.\n                    Ok(None)\n                }","sourceCodeStart":160,"sourceCodeEnd":196,"githubUrl":"https://github.com/vectordotdev/vector/blob/bdb87aeaa4c4ff27c0ba643c1c77b21bf2ef4013/lib/codecs/src/decoding/framing/octet_counting.rs#L160-L196","documentation":"After the octet-counting framer reads the declared length prefix, it decodes exactly that many bytes as UTF-8. This error is raised when the payload bytes are not valid UTF-8, so the message cannot be represented as a Rust string. The parser advances past the bad bytes to avoid stalling.","triggerScenarios":"A client sends the correct octet-counted length but the body contains binary/non-UTF-8 bytes (e.g. Latin-1 encoded logs, compressed data, or a miscounted length that cuts a multi-byte character in half).","commonSituations":"Legacy log shippers emitting non-UTF-8 encodings; truncation bugs in senders that miscount byte length and split multi-byte characters; forwarding binary payloads over a text-oriented protocol.","solutions":["Fix the sender to emit valid UTF-8 payloads and ensure the declared length matches the UTF-8 byte length exactly.","Re-encode data to UTF-8 before sending (e.g. iconv from Latin-1/Windows-1252).","Check for off-by-one length counting that splits multibyte characters across messages."],"exampleFix":"// before (sender, Latin-1 bytes with byte length)\n\"5 caf\\xe9\"\n// after (UTF-8 encoded)\n\"6 caf\\u{00e9}\" // \"café\" is 6 UTF-8 bytes","handlingStrategy":"validation","validationCode":"// ensure payload is valid UTF-8 and length matches byte length before sending\nfn frame_octet(msg: &str) -> String {\n    assert!(std::str::from_utf8(msg.as_bytes()).is_ok());\n    format!(\"{} {}\", msg.as_bytes().len(), msg)\n}","typeGuard":"fn is_utf8(bytes: &[u8]) -> bool {\n    std::str::from_utf8(bytes).is_ok()\n}","tryCatchPattern":"match framed_stream.next().await {\n    Err(e) if e.kind() == ErrorKind::InvalidData => {\n        // non-UTF-8 payload: reconfigure sender encoding to UTF-8\n    }\n    other => other?,\n}","preventionTips":["Encode all payloads as UTF-8 at the producer.","Count length in bytes, not characters, when writing length prefixes.","Never split multibyte characters across frames."],"tags":["codecs","framing","utf8","octet-counting"],"backgroundTag":"invalid-message-framing","analyzedSha":"bdb87aeaa4c4ff27c0ba643c1c77b21bf2ef4013","analyzedAt":"2026-09-16T02:53:35.741Z","contentChangedAt":"2026-09-16T02:53:35.741Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}