vectordotdev/vector · error

poisoned lock

Error message

poisoned lock

What it means

The chunked GELF decoder guards its shared reassembly state (`self.state: Mutex<...>`) with `.lock().expect("poisoned lock")`. A std::sync::Mutex is poisoned when another thread panicked while holding it, so after the first panic every subsequent caller panics with this message. The panic happens on the early-return path that checks an existing message's total-chunk count.

Solutions

  1. Find and fix the original panic that poisoned the mutex (look for the first panic in the logs)
  2. Upgrade to a non-poisoning lock (parking_lot::Mutex) or handle poisoning with `.lock().unwrap_or_else(|e| e.into_inner())` if the state is recoverable
  3. Restart the Vector process/workflow as a stopgap

Example fix

// before
let pending = self.state.lock().expect("poisoned lock");
// after
let pending = self.state.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
Defensive patterns

Strategy: try-catch

Validate before calling

// Rust: detect poisoned state before relying on the shared decoder
let healthy = !decoder.state.is_poisoned(); // std Mutex::is_poisoned (stabilized 1.86); else track manually

Type guard

fn lock_healthy<T>(m: &std::sync::Mutex<T>) -> bool {
    !m.is_poisoned()
}

Try / catch

// catch_unwind around decode calls, then rebuild the decoder
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| decoder.decode_message(bytes)));
match result {
    Ok(inner) => inner,
    Err(_) => { decoder = ChunkedGelfDecoder::new(config); /* retry once */ }
}

Prevention

When it happens

Trigger: `decode_chunk` (called from `decode_message`) locking the state mutex after a panic occurred elsewhere while the lock was held — e.g. a panic inside the timeout task or during chunk reassembly while another decoder clone held the lock.

Common situations: Any earlier panic in the chunked GELF decoder (or code holding that lock) in a long-running Vector process; subsequent chunks then all fail with this message, masking the original panic.

Related errors


AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16). Data as JSON: /api/errors/a1dce1ac4e9187cd. Report an issue: GitHub.

Appendix: source

Thrown at lib/codecs/src/decoding/framing/chunked_gelf.rs:452

        );

        let chunk_len = chunk.len();

        // A lone chunk is already complete, but it cannot reuse the ID of a pending message.
        if total_chunks == 1 {
            if chunk_len > self.max_length {
                return Err(ChunkedGelfDecoderError::MaxLengthExceed {
                    message_id,
                    sequence_number,
                    length: chunk_len,
                    max_length: self.max_length,
                });
            }

            // Copy before taking the shared-state lock, then keep the lock from the ID check
            // through return so another decoder clone cannot insert this ID between them.
            let chunk = Bytes::copy_from_slice(&chunk);
            let pending = self.state.lock().expect("poisoned lock");
            if let Some(message_state) = pending.messages.get(&message_id) {
                return Err(ChunkedGelfDecoderError::TotalChunksMismatch {
                    message_id,
                    sequence_number,
                    original_total_chunks: message_state.total_chunks,
                    received_total_chunks: total_chunks,
                });
            }
            return Ok(Some(chunk));
        }

        let mut pending = self.state.lock().expect("poisoned lock");

        let is_new_message = !pending.messages.contains_key(&message_id);
        // Settle every admission rule before creating state and a timeout task. The count limit
        // applies only on insert, since rejecting chunks of pending messages would stall them.
        if is_new_message {
            if chunk_len > self.max_length {

View on GitHub (pinned to bdb87aeaa4)