apache/iceberg · error · RuntimeException
Interrupted during commit
Error message
Interrupted during commit
What it means
If the commit thread is interrupted while waiting or communicating with HMS, doCommit re-interrupts the thread and throws a RuntimeException 'Interrupted during commit'. The commit outcome may be unknown, so callers must check the table state before retrying.
Solutions
- Allow the commit to complete before cancelling the job; avoid killing tasks mid-commit
- After interruption, verify the table's metadata_location to see if the commit landed before retrying
- Increase timeouts so commits aren't interrupted by supervising frameworks
- Re-run the operation on a fresh (refreshed) table snapshot if the commit did not land
Example fix
// before future.cancel(true); // interrupts an in-flight commit // after future.get(10, TimeUnit.MINUTES); // wait for commit completion before teardown
Defensive patterns
Strategy: try-catch
Validate before calling
// avoid committing inside interruptible cancellation scopes; check Thread.interrupted() before starting a long commit
Try / catch
try {
table.newAppend().appendFile(f).commit();
} catch (RuntimeException e) {
if (e.getMessage() != null && e.getMessage().contains("Interrupted during commit")) {
verifyCommitLanded(catalog.loadTable(ident)); // check metadata_location before retry
} else throw e;
} Prevention
- Don't cancel/kill tasks mid-commit; use graceful completion
- Give commits enough time budget in Spark/Flink
- Restore interrupt status and verify table state before any retry
- Use idempotent commit verification after interruptions
When it happens
Trigger: doCommit interrupted via Thread.interrupt() — e.g. task cancellation in Spark/Flink, executor shutdown, or query timeout killing the writing task while it is inside the HMS commit path.
Common situations: Spark stage cancellation/kill; Flink job cancel; application shutdown hooks; speculative-execution cleanup interrupting a slow commit.
Understand the failure class
Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.
Related errors
- Cannot commit: Base metadata location
- Interrupted during commit
- Interrupted during refresh
- Interrupted during refresh
- Interrupted in call to check table existence of
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/1ec696aaecb3712a.
Report an issue: GitHub.
Appendix: source
Thrown at hive-metastore/src/main/java/org/apache/iceberg/hive/HiveTableOperations.java:420
commitStatus = checkCommitStatus(newMetadataLocation, tableMetadata);
}
switch (commitStatus) {
case SUCCESS:
break;
case FAILURE:
throw e;
case UNKNOWN:
throw new CommitStateUnknownException(e);
}
}
} catch (TException e) {
throw new RuntimeException(
String.format("Metastore operation failed for %s.%s", database, tableName), e);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new RuntimeException("Interrupted during commit", e);
} catch (LockException e) {
throw new CommitFailedException(e);
} finally {
HiveOperationsBase.cleanupMetadataAndUnlock(io(), commitStatus, newMetadataLocation, lock);
}
LOG.info(
"Committed to table {} with the new metadata location {}", fullName, newMetadataLocation);
}
@Override
public long maxHiveTablePropertySize() {
return maxHiveTablePropertySize;
}
@OverrideView on GitHub (pinned to 86d9c8fc54)