influxdata/influxdb · error · TableIndexCacheError

Table snapshot persistence task failed

Error message

Table snapshot persistence task failed

What it means

TableIndexCacheError::TableSnapshotPersistenceTaskFailed wraps a tokio::task::JoinError raised when a spawned task that persists table index snapshots (or updates table indexes) panicked or was cancelled. The cache fans out per-snapshot work with a JoinSet; a JoinError means one of those worker tasks died rather than returning its Result.

Solutions

  1. Check logs for a panic message from the spawned task; the JoinError's panic payload identifies the failing code path.
  2. If the panic stems from unwrap on externally supplied data (snapshot bytes, marker JSON), fix or validate the offending file and report/patch the unwrap.
  3. Re-run TableIndexCache::initialize after restart so the conversion re-processes any snapshots left unconverted by the panicked task.
  4. Avoid shutting the node down during first-run snapshot conversion; let the migration complete before stopping.

Example fix

// before: panic inside spawned task propagates as JoinError
let _permit = sem.acquire_owned().await.unwrap();
// after: handle acquisition failure gracefully
let permit = sem.acquire_owned().await.map_err(|_| TableIndexCacheError::Unexpected(anyhow::anyhow!("semaphore closed")))?;
Defensive patterns

Strategy: try-catch

Type guard

fn is_join_cancelled(e: &tokio::task::JoinError) -> bool {
    e.is_cancelled()
}

Try / catch

match res.map_err(TableIndexCacheError::TableSnapshotPersistenceTaskFailed)? {
    Err(TableIndexCacheError::TableSnapshotPersistenceTaskFailed(j)) if j.is_cancelled() => {
        info!("shutdown during conversion; will resume on next start");
    }
    other => other?,
}

Prevention

When it happens

Trigger: js.join_next() returns Err in split_persisted_snapshots_to_table_index_snapshots (mapped to TableSnapshotPersistenceTaskFailed at table_index_cache.rs:474) or in update_all_from_object_store worker sets — caused by a panic inside the spawned future (e.g. the unwrap() on semaphore acquire or a serde unwrap) or task cancellation at shutdown.

Common situations: A bug-triggered panic inside snapshot split/persist code (e.g. unwrap on malformed data), OOM killing a task, or server shutdown cancelling in-flight conversion tasks during startup.

Related errors


AI-assisted analysis of influxdata/influxdb@06200ef96b (2026-09-19). Data as JSON: /api/errors/161ec61920e545ba. Report an issue: GitHub.

Appendix: source

Thrown at influxdb3_write/src/table_index_cache.rs:60

    #[error("Object store operation failed")]
    ObjectStore(#[source] object_store::Error),

    #[error("JSON serialization/deserialization failed")]
    Json(#[source] serde_json::Error),

    #[error("TableIndex operation failed")]
    TableIndex(#[source] crate::table_index::TableIndexError),

    #[error("Object meta is missing filename")]
    MissingFilename,

    #[error("Failed to parse snapshot sequence number from filename")]
    InvalidSnapshotSequenceNumber,

    #[error("Unexpected error")]
    Unexpected(#[source] anyhow::Error),

    #[error("Table snapshot persistence task failed")]
    TableSnapshotPersistenceTaskFailed(#[source] tokio::task::JoinError),

    #[error("Failed to list table indices from object store")]
    ListIndices(#[source] object_store::Error),

    #[error("Failed to parse table index path from object store path")]
    TableIndexPath(#[source] crate::paths::PathError),

    #[error("Failed to list table index snapshots from object store")]
    ListSnapshots(#[source] object_store::Error),

    #[error("Failed to parse table index snapshot path from object store path")]
    TableIndexSnapshotPath(#[source] crate::paths::PathError),

    #[error("Failed to update table index: join task failed")]
    UpdateTaskFailed(#[source] tokio::task::JoinError),

    #[error("Failed to list object store metadata")]

View on GitHub (pinned to 06200ef96b)