vectordotdev/vector · critical

Failed to acquire read lock on recency map

Error message

Failed to acquire read lock on recency map

What it means

Panic raised by `.expect()` on `self.recency.read()` inside `visit_metrics`. A std `RwLock` read acquisition only fails when the lock is poisoned — a writer panicked while holding or awaiting the lock — so this error means an earlier panic corrupted the shared recency-map state.

Solutions

  1. Locate the first panic in the logs that poisoned the lock; that root cause must be fixed, not this symptom.
  2. Replace std `RwLock` with `parking_lot::RwLock` (unpoisoned by design) or handle `PoisonError` via `into_inner()`.
  3. Restart the process — poisoned locks never recover in-process.
  4. Ensure metric-expiration config (`expire_metrics`, timeouts) does not trigger the upstream panic path.

Example fix

// before
let recency = self
    .recency
    .read()
    .expect("Failed to acquire read lock on recency map");
// after
let recency = self
    .recency
    .read()
    .unwrap_or_else(|poisoned| poisoned.into_inner());
Defensive patterns

Strategy: try-catch

Validate before calling

// Detect poisoning before the read
let usable = matches!(self.recency.try_read(), Ok(_) | Err(std::sync::TryLockError::WouldBlock));

Type guard

fn recency_readable<T>(lock: &std::sync::RwLock<T>) -> bool {
    !matches!(lock.try_read(), Err(std::sync::TryLockError::Poisoned(_)))
}

Try / catch

// Contain the visit in a panic barrier and recover gracefully
let metrics = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
    recorder.visit_metrics()
}))
.unwrap_or_default(); // return empty vec instead of crashing the scrape loop

Prevention

When it happens

Trigger: Calling `visit_metrics` (the periodic metric-flush/visit path) after any thread panicked while holding the recency map's write lock, e.g. the panic from `set_expiry`'s own expect or a panic during recency-map mutation.

Common situations: Vector's internal metrics scrape loop calling `visit_metrics` on every flush interval after a one-time panic poisoned the map; every subsequent scrape then panics, typically cascading into task shutdown.

Related errors


AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16). Data as JSON: /api/errors/cdfcd290d433d013. Report an issue: GitHub.

Appendix: source

Thrown at lib/vector-core/src/metrics/recorder.rs:64

        } else {
            Some(Recency::new(
                Clock::new(),
                MetricKindMask::ALL,
                global_timeout,
                expire_metrics_per_metric_set,
            ))
        };
        *(self.recency.write()).expect("Failed to acquire write lock on recency map") = recency;
    }

    pub(super) fn visit_metrics(&self) -> Vec<Metric> {
        let timestamp = Utc::now();

        let mut metrics = Vec::new();
        let recency = self
            .recency
            .read()
            .expect("Failed to acquire read lock on recency map");
        let recency = recency.as_ref();

        for (key, counter) in self.registry.get_counter_handles() {
            if recency
                .is_none_or(|recency| recency.should_store_counter(&key, &counter, &self.registry))
            {
                // NOTE this will truncate if the value is greater than 2**52.
                #[allow(clippy::cast_precision_loss)]
                let value = counter.get_inner().load(Ordering::Relaxed) as f64;
                let value = MetricValue::Counter { value };
                metrics.push(Metric::from_metric_kv(&key, value, timestamp));
            }
        }
        for (key, gauge) in self.registry.get_gauge_handles() {
            if recency
                .is_none_or(|recency| recency.should_store_gauge(&key, &gauge, &self.registry))
            {
                let value = gauge.get_inner().load(Ordering::Relaxed);

View on GitHub (pinned to bdb87aeaa4)