risingwavelabs/risingwave · critical

iceberg pk-index sink commit failed for sink

Error message

iceberg pk-index sink commit failed for sink {} epoch {}

What it means

This error wraps the underlying iceberg commit failure (CommitError::Commit or CommitError::ReloadTable from commit_one_epoch) with context identifying the sink and epoch. After appending files to the Iceberg table (or reloading it afterwards) fails beyond commit_retry_num attempts, the coordinator surfaces this contextual error; the epoch stays in pending_sink_state so it can be retried.

Solutions

  1. Inspect the wrapped underlying error (context chain) for the real cause: commit conflict vs reload failure
  2. Retry the commit; pending state persisted in pending_sink_state makes it idempotent and recoverable on restart
  3. Increase commit_retry_num in iceberg config if conflicts with concurrent external writers are frequent
  4. Stop/coordinate external writers to the same table/branch, or write to a dedicated branch
  5. Verify catalog credentials and that the target branch exists and is writable

Example fix

// before
iceberg.commit_retry_num = 3 // conflicts with external writer
// after
iceberg.commit_retry_num = 10
// and/or: write to dedicated branch
iceberg.write_mode = ... // ensure commit_branch targets an RW-only branch
Defensive patterns

Strategy: retry

Try / catch

match coordinator.commit().await {
    Err(e) if e.to_string().contains("commit failed for sink") => {
        // inspect the source: Commit vs ReloadTable
        warn!(error = ?e.source(), "iceberg commit failed; will retry from pending state");
        retry_with_backoff(|| coordinator.commit()).await
    }
    other => other,
}

Prevention

When it happens

Trigger: Calling commit() (live or during recovery drain) when commit_one_epoch fails: the Iceberg table commit itself errors (commit conflict/exhausted retries) or reloading the table after commit fails.

Common situations: Concurrent external writers to the same Iceberg table causing repeated commit conflicts beyond retry count, catalog connectivity/auth failures during commit or table reload, snapshot expiration racing the commit, branch missing (wrong branch config) making reload fail.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/eee5c7b058185a37. Report an issue: GitHub.

Appendix: source

Thrown at src/meta/src/manager/iceberg_pk_index_sink/coordinator.rs:203

    pub async fn commit(&mut self) -> Result<()> {
        let Some(commit) = self.waiting_commit.take() else {
            return Ok(());
        };

        let refreshed_table = commit_one_epoch(
            self.catalog.clone(),
            self.table.identifier().clone(),
            self.target_branch.clone(),
            self.sink_id,
            &commit,
            self.retry_num,
        )
        .await
        .map_err(|err| {
            let err_report = match err {
                CommitError::Commit(e) | CommitError::ReloadTable(e) => e,
            };
            anyhow!(err_report).context(format!(
                "iceberg pk-index sink commit failed for sink {} epoch {}",
                self.sink_id, commit.epoch
            ))
        })?;
        self.table = refreshed_table;

        commit_and_prune_epoch(
            &self.db,
            self.sink_id,
            commit.epoch,
            self.prev_committed_epoch,
        )
        .await
        .with_context(|| {
            format!(
                "iceberg pk-index sink mark_committed failed for sink {} epoch {}",
                self.sink_id, commit.epoch
            )

View on GitHub (pinned to 6469eb736d)