vectordotdev/vector · error · LinesCodecError
Unable to decode message as UTF8
Error message
Unable to decode message as UTF8
What it means
After the octet-counting framer reads the declared length prefix, it decodes exactly that many bytes as UTF-8. This error is raised when the payload bytes are not valid UTF-8, so the message cannot be represented as a Rust string. The parser advances past the bad bytes to avoid stalling.
Solutions
- Fix the sender to emit valid UTF-8 payloads and ensure the declared length matches the UTF-8 byte length exactly.
- Re-encode data to UTF-8 before sending (e.g. iconv from Latin-1/Windows-1252).
- Check for off-by-one length counting that splits multibyte characters across messages.
Example fix
// before (sender, Latin-1 bytes with byte length)
"5 caf\xe9"
// after (UTF-8 encoded)
"6 caf\u{00e9}" // "café" is 6 UTF-8 bytes Defensive patterns
Strategy: validation
Validate before calling
// ensure payload is valid UTF-8 and length matches byte length before sending
fn frame_octet(msg: &str) -> String {
assert!(std::str::from_utf8(msg.as_bytes()).is_ok());
format!("{} {}", msg.as_bytes().len(), msg)
} Type guard
fn is_utf8(bytes: &[u8]) -> bool {
std::str::from_utf8(bytes).is_ok()
} Try / catch
match framed_stream.next().await {
Err(e) if e.kind() == ErrorKind::InvalidData => {
// non-UTF-8 payload: reconfigure sender encoding to UTF-8
}
other => other?,
} Prevention
- Encode all payloads as UTF-8 at the producer.
- Count length in bytes, not characters, when writing length prefixes.
- Never split multibyte characters across frames.
When it happens
Trigger: A client sends the correct octet-counted length but the body contains binary/non-UTF-8 bytes (e.g. Latin-1 encoded logs, compressed data, or a miscounted length that cuts a multi-byte character in half).
Common situations: Legacy log shippers emitting non-UTF-8 encodings; truncation bugs in senders that miscount byte length and split multi-byte characters; forwarding binary payloads over a text-oriented protocol.
Understand the failure class
- Parsing and encoding errors: unexpected token, malformed input — why parsers reject input and how to find the real culprit.
Related errors
- Unable to decode message len as number
- Invalid chunk header with less than 10 bytes: 0x
- InvalidData
- Received chunk with message id
- Received chunk with message id
AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16).
Data as JSON: /api/errors/13db22c40d5ffe11.
Report an issue: GitHub.
Appendix: source
Thrown at lib/codecs/src/decoding/framing/octet_counting.rs:178
if len > self.other.max_length() {
// The length is greater than we want.
//
// We need to discard the entire message.
self.octet_decoding = Some(State::Discarding(len));
src.advance(space_pos + 1);
Ok(None)
} else if let Some(msg) = src.get(from..to) {
let bytes = match std::str::from_utf8(msg) {
Ok(_) => Bytes::copy_from_slice(msg),
Err(_) => {
// The data was not valid UTF8 :-(.
//
// Advance the buffer past the erroneous bytes to
// prevent us getting stuck in an infinite loop.
src.advance(to);
self.octet_decoding = None;
return Err(LinesCodecError::Io(io::Error::new(
io::ErrorKind::InvalidData,
"Unable to decode message as UTF8",
)));
}
};
// We have managed to read the entire message as valid UTF8!
src.advance(to);
self.octet_decoding = None;
Ok(Some(bytes))
} else {
// We have an acceptable number of bytes in this message,
// but not all the data was in the frame.
//
// Return `None` to indicate we want more data before we do
// anything else.
Ok(None)
}View on GitHub (pinned to bdb87aeaa4)