risingwavelabs/risingwave · critical

table Epoch <= committed_epoch

Error message

table {} Epoch {} <= committed_epoch {}

What it means

commit_epoch's sanity check found a state table whose epoch to commit is not strictly greater than the table's already-committed epoch in the current version. Committing would move a table's visible data backwards, violating the monotonic epoch invariant of Hummock's MVCC versioning.

Solutions

  1. Find the stale producer committing the old epoch (check which actor/node sent the SST or data) and ensure it refreshes from the latest hummock version before committing.
  2. Skip/idempotently handle re-commit of an already-committed epoch in the recovery path instead of failing.
  3. Audit epoch advancement logic (barrier collection) so a table's epoch never regresses; check for meta failover replaying finalized epochs.

Example fix

// before
return Err(anyhow!("table {} Epoch {} <= committed_epoch {}", table_id, committed_epoch, info.committed_epoch));
// after
if *committed_epoch <= info.committed_epoch {
    tracing::warn!(table_id, committed_epoch, prev = info.committed_epoch, "duplicate/stale epoch commit ignored");
    continue; // idempotent re-commit
}
Defensive patterns

Strategy: validation

Validate before calling

// before calling commit_epoch, check every table's epoch advances monotonically
for (table_id, epoch) in &tables_to_commit {
    if let Some(info) = current_version.state_table_info.info().get(table_id) {
        assert!(*epoch > info.committed_epoch,
            "table {} epoch {} must exceed committed {}", table_id, epoch, info.committed_epoch);
    }
}

Try / catch

match hummock_manager.commit_epoch(new_epoch, tables_to_commit).await {
    Err(e) if e.to_string().contains("<= committed_epoch") => {
        // stale/duplicate epoch: refresh from latest hummock version and re-drive the commit
        refresh_local_version().await;
    },
    other => other?,
}

Prevention

When it happens

Trigger: Calling commit_epoch with tables_to_commit containing (table_id, epoch) where the current version's state_table_info shows committed_epoch >= that epoch. Happens when a stale node replays an old epoch's data after a failover or when the same epoch is committed twice for a table.

Common situations: Meta failover causing duplicate epoch commits; a stream actor resending an old SST/data epoch after restart; recovery path replaying an already-committed Hummock version; versioning bugs after a version change.

Understand the failure class

Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/91b572acb1cdeb82. Report an issue: GitHub.

Appendix: source

Thrown at src/meta/src/hummock/manager/context.rs:246

                if *context_id == crate::manager::META_NODE_ID {
                    continue;
                }
            }
            if !context_info
                .check_context(*context_id, &self.metadata_manager)
                .await
            {
                return Err(Error::InvalidSst(*sst_id));
            }
        }
        drop(context_info);

        // sanity check on monotonically increasing table committed epoch
        for (table_id, committed_epoch) in tables_to_commit {
            if let Some(info) = current_version.state_table_info.info().get(table_id)
                && *committed_epoch <= info.committed_epoch
            {
                return Err(anyhow::anyhow!(
                    "table {} Epoch {} <= committed_epoch {}",
                    table_id,
                    committed_epoch,
                    info.committed_epoch,
                )
                .into());
            }
        }

        // HummockManager::now requires a write to the meta store. Thus, it should be avoided whenever feasible.
        if !sstables.is_empty() {
            // Sanity check to ensure SSTs to commit have not been full GCed yet.
            let now = self.now().await?;
            check_sst_retention(
                now,
                self.env.opts.min_sst_retention_time_sec,
                sstables
                    .iter()

View on GitHub (pinned to 6469eb736d)