vectordotdev/vector · error · LinesCodecError

Unable to decode message as UTF8

Error message

Unable to decode message as UTF8

What it means

After the octet-counting framer reads the declared length prefix, it decodes exactly that many bytes as UTF-8. This error is raised when the payload bytes are not valid UTF-8, so the message cannot be represented as a Rust string. The parser advances past the bad bytes to avoid stalling.

Solutions

  1. Fix the sender to emit valid UTF-8 payloads and ensure the declared length matches the UTF-8 byte length exactly.
  2. Re-encode data to UTF-8 before sending (e.g. iconv from Latin-1/Windows-1252).
  3. Check for off-by-one length counting that splits multibyte characters across messages.

Example fix

// before (sender, Latin-1 bytes with byte length)
"5 caf\xe9"
// after (UTF-8 encoded)
"6 caf\u{00e9}" // "café" is 6 UTF-8 bytes
Defensive patterns

Strategy: validation

Validate before calling

// ensure payload is valid UTF-8 and length matches byte length before sending
fn frame_octet(msg: &str) -> String {
    assert!(std::str::from_utf8(msg.as_bytes()).is_ok());
    format!("{} {}", msg.as_bytes().len(), msg)
}

Type guard

fn is_utf8(bytes: &[u8]) -> bool {
    std::str::from_utf8(bytes).is_ok()
}

Try / catch

match framed_stream.next().await {
    Err(e) if e.kind() == ErrorKind::InvalidData => {
        // non-UTF-8 payload: reconfigure sender encoding to UTF-8
    }
    other => other?,
}

Prevention

When it happens

Trigger: A client sends the correct octet-counted length but the body contains binary/non-UTF-8 bytes (e.g. Latin-1 encoded logs, compressed data, or a miscounted length that cuts a multi-byte character in half).

Common situations: Legacy log shippers emitting non-UTF-8 encodings; truncation bugs in senders that miscount byte length and split multi-byte characters; forwarding binary payloads over a text-oriented protocol.

Understand the failure class

Related errors


AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16). Data as JSON: /api/errors/13db22c40d5ffe11. Report an issue: GitHub.

Appendix: source

Thrown at lib/codecs/src/decoding/framing/octet_counting.rs:178

                if len > self.other.max_length() {
                    // The length is greater than we want.
                    //
                    // We need to discard the entire message.
                    self.octet_decoding = Some(State::Discarding(len));
                    src.advance(space_pos + 1);

                    Ok(None)
                } else if let Some(msg) = src.get(from..to) {
                    let bytes = match std::str::from_utf8(msg) {
                        Ok(_) => Bytes::copy_from_slice(msg),
                        Err(_) => {
                            // The data was not valid UTF8 :-(.
                            //
                            // Advance the buffer past the erroneous bytes to
                            // prevent us getting stuck in an infinite loop.
                            src.advance(to);
                            self.octet_decoding = None;
                            return Err(LinesCodecError::Io(io::Error::new(
                                io::ErrorKind::InvalidData,
                                "Unable to decode message as UTF8",
                            )));
                        }
                    };

                    // We have managed to read the entire message as valid UTF8!
                    src.advance(to);
                    self.octet_decoding = None;
                    Ok(Some(bytes))
                } else {
                    // We have an acceptable number of bytes in this message,
                    // but not all the data was in the frame.
                    //
                    // Return `None` to indicate we want more data before we do
                    // anything else.
                    Ok(None)
                }

View on GitHub (pinned to bdb87aeaa4)