vectordotdev/vector · critical
Failed to acquire write lock on recency map
Error message
Failed to acquire write lock on recency map
What it means
This is a Rust lock-poisoning panic. `set_expiry` replaces the recency map behind a `RwLock` using `self.recency.write()`. A `write()` call returns `Err` only when the lock is poisoned, i.e. another thread panicked while holding the write lock, and the code deliberately panics via `.expect("Failed to acquire write lock on recency map")` instead of recovering.
Solutions
- Find and fix the original panic that poisoned the lock: look earlier in the logs for the first panic in the metrics/recorder subsystem.
- Upgrade or patch the recorder code to use a poisoning-tolerant lock (e.g. `parking_lot::RwLock` or recover from the `PoisonError` with `into_inner()`).
- Restart the process; poisoning never clears at runtime, so the recorder cannot recover without a restart.
- Audit config values passed into the recency-map construction (global_timeout, expire_metrics_per_metric_set) that may trigger the initial panic.
Example fix
// before
*(self.recency.write()).expect("Failed to acquire write lock on recency map") = recency;
// after
let mut guard = self
.recency
.write()
.unwrap_or_else(|poisoned| poisoned.into_inner());
*guard = recency; Defensive patterns
Strategy: try-catch
Validate before calling
// Rust: cannot pre-validate poisoning; detect it before using the value let poisoned = matches!(self.recency.try_write(), Err(std::sync::TryLockError::Poisoned(_)));
Type guard
fn is_poisoned<T>(lock: &std::sync::RwLock<T>) -> bool {
matches!(lock.try_read(), Err(std::sync::TryLockError::Poisoned(_)))
} Try / catch
// Rust panics are not catchable with try/catch; isolate the call site
let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
recorder.set_expiry(global_timeout, expire_metrics);
}));
if result.is_err() {
// recorder lock poisoned: restart metrics subsystem or process
} Prevention
- Never panic while holding a lock; return Result from all code that mutates the recency map under the lock.
- Prefer parking_lot::RwLock, which does not poison.
- Monitor logs for any first panic in the metrics subsystem and treat it as a P1.
- Wrap metric-expiration updates in catch_unwind at task boundaries so one panic does not poison global state.
When it happens
Trigger: Calling `set_expiry` on the global metrics recorder when any prior thread/task panicked while holding the recency map's write lock (e.g. a panic inside the recency construction block or any earlier code that held the lock).
Common situations: A metrics-expiration update task panics once (bad timeout config, OOM-adjacent bug) and every subsequent `set_expiry` call then panics with this message; commonly surfaced in long-running Vector processes after an unrelated panic in the metrics subsystem.
Related errors
- Failed to acquire read lock on recency map
- a record with a next ID must have an event count
- a valid HTTP/1 URI is valid as an HTTP URI
- a validated HTTP endpoint is a valid `http 1` URI
- all json-file keys should be matched
AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16).
Data as JSON: /api/errors/095c44b65a013378.
Report an issue: GitHub.
Appendix: source
Thrown at lib/vector-core/src/metrics/recorder.rs:54
self.registry.clear();
}
pub(super) fn set_expiry(
&self,
global_timeout: Option<Duration>,
expire_metrics_per_metric_set: Vec<(MetricKeyMatcher, Duration)>,
) {
let recency = if global_timeout.is_none() && expire_metrics_per_metric_set.is_empty() {
None
} else {
Some(Recency::new(
Clock::new(),
MetricKindMask::ALL,
global_timeout,
expire_metrics_per_metric_set,
))
};
*(self.recency.write()).expect("Failed to acquire write lock on recency map") = recency;
}
pub(super) fn visit_metrics(&self) -> Vec<Metric> {
let timestamp = Utc::now();
let mut metrics = Vec::new();
let recency = self
.recency
.read()
.expect("Failed to acquire read lock on recency map");
let recency = recency.as_ref();
for (key, counter) in self.registry.get_counter_handles() {
if recency
.is_none_or(|recency| recency.should_store_counter(&key, &counter, &self.registry))
{
// NOTE this will truncate if the value is greater than 2**52.
#[allow(clippy::cast_precision_loss)]View on GitHub (pinned to bdb87aeaa4)