block/buzz · error · anyhow::Error

another storage-snapshot worker already holds the lease

Error message

another storage-snapshot worker already holds the lease

What it means

Storage snapshots are serialized via a database lease (try_lock_storage_accounting). If another worker currently holds the lease, this invocation returns immediately with this error instead of blocking or queueing, so only one snapshot writer runs at a time.

Solutions

  1. Wait for the existing snapshot to finish and retry, or remove the duplicate cron/schedule overlap.
  2. Check for a stuck/stale lease from a crashed worker and clear it once confirmed no worker is alive.
  3. Add distributed scheduling (e.g. flock or leader election) so only one instance of the command is launched.
  4. Retry with backoff rather than tight-looping: the lease is transient if the other worker is healthy.

Example fix

// before: cron every minute, overlapping runs
* * * * * buzz-admin storage-snapshot --max-objects 1000000
// after: widen the interval or gate with flock
0 * * * * flock -n /tmp/buzz-snapshot.lock buzz-admin storage-snapshot --max-objects 1000000
Defensive patterns

Strategy: retry

Validate before calling

// Best-effort pre-check: ensure no overlapping schedule is due now
// (e.g. skip launch if a lockfile from a live run exists)
if [ -e /tmp/buzz-snapshot.lock ] && kill -0 "$(cat /tmp/buzz-snapshot.lock)" 2>/dev/null; then
  echo "snapshot already running"; exit 0
fi

Try / catch

match cmd_storage_snapshot(max_objects).await {
    Err(e) if e.to_string().contains("already holds the lease") => {
        // schedule retry with backoff; do not tight-loop
        tokio::time::sleep(Duration::from_secs(300)).await;
    }
    other => other?,
}

Prevention

When it happens

Trigger: Running `buzz-admin storage-snapshot` while a previous snapshot run (cron job, another operator, a stalled run that never released the lock) still holds the storage-accounting lease. try_lock_storage_accounting returns None and ok_or_else converts it to this error.

Common situations: Overlapping cron schedules, a second operator running the same command concurrently, or a crashed/killed run whose lease has not expired yet (stale lease).

Understand the failure class

Background: Conflicting config options: "cannot be used together" — configuration validation errors across open-source libraries — this error's family across 162 libraries.

Related errors


AI-assisted analysis of block/buzz@ef2aa1ae38 (2026-09-20). Data as JSON: /api/errors/29ccc6d120085894. Report an issue: GitHub.

Appendix: source

Thrown at crates/buzz-admin/src/main.rs:191

        } => cmd_list_product_feedback(limit).await,
        Command::Deletions { command } => deletions::run(command).await,
        Command::ReconcileChannels { channel, relay_key } => {
            reconcile_channels(channel, relay_key).await?;
            Ok(0)
        }
    }
}

async fn cmd_storage_snapshot(max_objects: u64) -> Result<i32> {
    let max_objects_db = i64::try_from(max_objects)
        .map_err(|_| anyhow::anyhow!("--max-objects must be at most {}", i64::MAX))?;
    if max_objects == 0 {
        return Err(anyhow::anyhow!("--max-objects must be greater than zero"));
    }

    let db = connect_db().await?;
    let mut leader = db.try_lock_storage_accounting().await?.ok_or_else(|| {
        anyhow::anyhow!("another storage-snapshot worker already holds the lease")
    })?;
    let storage = Arc::new(MediaStorage::new(&storage_config_from_env()?)?);
    let code_sha =
        std::env::var("BUZZ_STORAGE_SNAPSHOT_CODE_SHA").unwrap_or_else(|_| "unknown".to_string());
    if code_sha.is_empty() || code_sha.len() > 128 {
        return Err(anyhow::anyhow!(
            "BUZZ_STORAGE_SNAPSHOT_CODE_SHA must contain 1 to 128 bytes"
        ));
    }

    println!(
        "{}",
        serde_json::json!({
            "event": "storage_snapshot_started",
            "max_objects": max_objects,
            "code_sha": code_sha,
        })
    );

View on GitHub (pinned to ef2aa1ae38)