influxdata/influxdb · error · TableIndexError

Table snapshot persistence task failed

Error message

Table snapshot persistence task failed

What it means

TableIndexError::TableSnapshotPersistenceTaskFailed is raised when the tokio task spawned to persist table snapshots is joined and returns a JoinError — the persistence task panicked or was cancelled. Persisting snapshots is done concurrently (bounded by a Semaphore) across tables; this error identifies which spawned work failed to complete.

Solutions

  1. Check JoinError::is_panic() vs is_cancelled(); for panics, extract the payload via into_panic() and match logs to the failing table.
  2. Avoid aborting/dropping persistence task handles; use graceful shutdown that drains in-flight persists.
  3. Retry the persistence for the affected tables — persistence is idempotent per snapshot sequence number.
  4. If reproducible, treat as a library bug and report with the snapshot data that triggers the panic.

Example fix

// before
for res in futures { res.await?; }
// after
for res in futures {
    res.await.map_err(|e| {
        if e.is_panic() { error!("snapshot persist panicked: {:?}", e.into_panic()); }
        TableIndexError::TableSnapshotPersistenceTaskFailed(e)
    })?;
}
Defensive patterns

Strategy: try-catch

Validate before calling

// bound concurrency so cancellation/panics are isolated per table
let sem = Arc::new(Semaphore::new(num_cpus));
let _permit = sem.acquire().await;

Type guard

fn is_cancelled_join(e: &tokio::task::JoinError) -> bool { e.is_cancelled() }

Try / catch

match handle.await {
    Ok(()) => {},
    Err(e) if e.is_cancelled() => {
        warn!("snapshot persist cancelled; will be retried next snapshot cycle");
    }
    Err(e) => return Err(TableIndexError::TableSnapshotPersistenceTaskFailed(e)),
}

Prevention

When it happens

Trigger: Awaiting JoinHandles from the FuturesOrdered of spawned snapshot-persist tasks; a task panics (serialization bug, poisoned lock) or the runtime/caller cancels it (drop of the handle, runtime shutdown).

Common situations: Panic inside serde serialization of a snapshot; runtime shutdown during shutdown-with-inflight-persists; timeouts implemented by aborting tasks; single-threaded runtime misconfiguration.

Related errors


AI-assisted analysis of influxdata/influxdb@06200ef96b (2026-09-19). Data as JSON: /api/errors/a501e38a6779fe49. Report an issue: GitHub.

Appendix: source

Thrown at influxdb3_write/src/table_index.rs:70

        actual: TableIndexId,
    },

    #[error("Failed to parse table index path from object store path")]
    TableIndexPath(#[source] crate::paths::PathError),

    #[error("Failed to parse table index snapshot path from object store path")]
    TableIndexSnapshotPath(#[source] crate::paths::PathError),

    #[error("Failed to update table index: join task failed")]
    UpdateTaskFailed(#[source] tokio::task::JoinError),

    #[error("Object meta is missing filename")]
    MissingFilename,

    #[error("Failed to parse snapshot sequence number from filename")]
    InvalidSnapshotSequenceNumber,

    #[error("Table snapshot persistence task failed")]
    TableSnapshotPersistenceTaskFailed(#[source] tokio::task::JoinError),

    #[error("Unexpected error")]
    Unexpected(#[source] anyhow::Error),

    #[error("Object store operation failed")]
    ObjectStore(#[source] object_store::Error),

    #[error("JSON serialization/deserialization failed")]
    Json(#[source] serde_json::Error),
}

pub type Result<T> = std::result::Result<T, TableIndexError>;

/// An partial, incremental index into a given database/table. Should not be used to make decisions
/// regarding gen1 file retention -- instead, all snapshots should be aggregated into a full
/// CoreTableIndex.
///

View on GitHub (pinned to 06200ef96b)