vectordotdev/vector · error
poisoned lock
Error message
poisoned lock
What it means
The chunked GELF decoder guards its shared reassembly state (`self.state: Mutex<...>`) with `.lock().expect("poisoned lock")`. A std::sync::Mutex is poisoned when another thread panicked while holding it, so after the first panic every subsequent caller panics with this message. The panic happens on the early-return path that checks an existing message's total-chunk count.
Solutions
- Find and fix the original panic that poisoned the mutex (look for the first panic in the logs)
- Upgrade to a non-poisoning lock (parking_lot::Mutex) or handle poisoning with `.lock().unwrap_or_else(|e| e.into_inner())` if the state is recoverable
- Restart the Vector process/workflow as a stopgap
Example fix
// before
let pending = self.state.lock().expect("poisoned lock");
// after
let pending = self.state.lock().unwrap_or_else(|poisoned| poisoned.into_inner()); Defensive patterns
Strategy: try-catch
Validate before calling
// Rust: detect poisoned state before relying on the shared decoder let healthy = !decoder.state.is_poisoned(); // std Mutex::is_poisoned (stabilized 1.86); else track manually
Type guard
fn lock_healthy<T>(m: &std::sync::Mutex<T>) -> bool {
!m.is_poisoned()
} Try / catch
// catch_unwind around decode calls, then rebuild the decoder
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| decoder.decode_message(bytes)));
match result {
Ok(inner) => inner,
Err(_) => { decoder = ChunkedGelfDecoder::new(config); /* retry once */ }
} Prevention
- Fix any panic occurring while the reassembly lock is held — it poisons the mutex for all later calls
- Prefer parking_lot::Mutex (non-poisoning) for shared decoder state
- Monitor logs for the first panic; 'poisoned lock' is always secondary
- Restart/rebuild decoder instances after an observed panic
When it happens
Trigger: `decode_chunk` (called from `decode_message`) locking the state mutex after a panic occurred elsewhere while the lock was held — e.g. a panic inside the timeout task or during chunk reassembly while another decoder clone held the lock.
Common situations: Any earlier panic in the chunked GELF decoder (or code holding that lock) in a long-running Vector process; subsequent chunks then all fail with this message, masking the original panic.
Related errors
AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16).
Data as JSON: /api/errors/a1dce1ac4e9187cd.
Report an issue: GitHub.
Appendix: source
Thrown at lib/codecs/src/decoding/framing/chunked_gelf.rs:452
);
let chunk_len = chunk.len();
// A lone chunk is already complete, but it cannot reuse the ID of a pending message.
if total_chunks == 1 {
if chunk_len > self.max_length {
return Err(ChunkedGelfDecoderError::MaxLengthExceed {
message_id,
sequence_number,
length: chunk_len,
max_length: self.max_length,
});
}
// Copy before taking the shared-state lock, then keep the lock from the ID check
// through return so another decoder clone cannot insert this ID between them.
let chunk = Bytes::copy_from_slice(&chunk);
let pending = self.state.lock().expect("poisoned lock");
if let Some(message_state) = pending.messages.get(&message_id) {
return Err(ChunkedGelfDecoderError::TotalChunksMismatch {
message_id,
sequence_number,
original_total_chunks: message_state.total_chunks,
received_total_chunks: total_chunks,
});
}
return Ok(Some(chunk));
}
let mut pending = self.state.lock().expect("poisoned lock");
let is_new_message = !pending.messages.contains_key(&message_id);
// Settle every admission rule before creating state and a timeout task. The count limit
// applies only on insert, since rejecting chunks of pending messages would stall them.
if is_new_message {
if chunk_len > self.max_length {View on GitHub (pinned to bdb87aeaa4)