vectordotdev/vector · critical
Failed to acquire read lock on recency map
Error message
Failed to acquire read lock on recency map
What it means
Panic raised by `.expect()` on `self.recency.read()` inside `visit_metrics`. A std `RwLock` read acquisition only fails when the lock is poisoned — a writer panicked while holding or awaiting the lock — so this error means an earlier panic corrupted the shared recency-map state.
Solutions
- Locate the first panic in the logs that poisoned the lock; that root cause must be fixed, not this symptom.
- Replace std `RwLock` with `parking_lot::RwLock` (unpoisoned by design) or handle `PoisonError` via `into_inner()`.
- Restart the process — poisoned locks never recover in-process.
- Ensure metric-expiration config (`expire_metrics`, timeouts) does not trigger the upstream panic path.
Example fix
// before
let recency = self
.recency
.read()
.expect("Failed to acquire read lock on recency map");
// after
let recency = self
.recency
.read()
.unwrap_or_else(|poisoned| poisoned.into_inner()); Defensive patterns
Strategy: try-catch
Validate before calling
// Detect poisoning before the read let usable = matches!(self.recency.try_read(), Ok(_) | Err(std::sync::TryLockError::WouldBlock));
Type guard
fn recency_readable<T>(lock: &std::sync::RwLock<T>) -> bool {
!matches!(lock.try_read(), Err(std::sync::TryLockError::Poisoned(_)))
} Try / catch
// Contain the visit in a panic barrier and recover gracefully
let metrics = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
recorder.visit_metrics()
}))
.unwrap_or_default(); // return empty vec instead of crashing the scrape loop Prevention
- Fix any writer-side panic (e.g. in set_expiry) first; readers panic only after a writer poisoned the lock.
- Use unwrap_or_else(|p| p.into_inner()) to tolerate poisoning in read-only scrape paths.
- Avoid .expect() on lock acquisition in library code that runs on hot periodic paths.
- Add a health check that surfaces a poisoned metrics lock before data pipelines crash.
When it happens
Trigger: Calling `visit_metrics` (the periodic metric-flush/visit path) after any thread panicked while holding the recency map's write lock, e.g. the panic from `set_expiry`'s own expect or a panic during recency-map mutation.
Common situations: Vector's internal metrics scrape loop calling `visit_metrics` on every flush interval after a one-time panic poisoned the map; every subsequent scrape then panics, typically cascading into task shutdown.
Related errors
- Failed to acquire write lock on recency map
- a record with a next ID must have an event count
- a valid HTTP/1 URI is valid as an HTTP URI
- a validated HTTP endpoint is a valid `http 1` URI
- all json-file keys should be matched
AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16).
Data as JSON: /api/errors/cdfcd290d433d013.
Report an issue: GitHub.
Appendix: source
Thrown at lib/vector-core/src/metrics/recorder.rs:64
} else {
Some(Recency::new(
Clock::new(),
MetricKindMask::ALL,
global_timeout,
expire_metrics_per_metric_set,
))
};
*(self.recency.write()).expect("Failed to acquire write lock on recency map") = recency;
}
pub(super) fn visit_metrics(&self) -> Vec<Metric> {
let timestamp = Utc::now();
let mut metrics = Vec::new();
let recency = self
.recency
.read()
.expect("Failed to acquire read lock on recency map");
let recency = recency.as_ref();
for (key, counter) in self.registry.get_counter_handles() {
if recency
.is_none_or(|recency| recency.should_store_counter(&key, &counter, &self.registry))
{
// NOTE this will truncate if the value is greater than 2**52.
#[allow(clippy::cast_precision_loss)]
let value = counter.get_inner().load(Ordering::Relaxed) as f64;
let value = MetricValue::Counter { value };
metrics.push(Metric::from_metric_kv(&key, value, timestamp));
}
}
for (key, gauge) in self.registry.get_gauge_handles() {
if recency
.is_none_or(|recency| recency.should_store_gauge(&key, &gauge, &self.registry))
{
let value = gauge.get_inner().load(Ordering::Relaxed);View on GitHub (pinned to bdb87aeaa4)