apache/iceberg · warning
The iceberg transaction has been committed, but we failed…
Error message
The iceberg transaction has been committed, but we failed to clean the temporary flink manifests: {} What it means
FlinkManifestUtil logs this warning when an Iceberg snapshot has already been committed successfully but deletion of the temporary Flink-written manifest files (or manifest lists) failed. The commit itself is durable, so this is a cleanup/housekeeping problem, not a data-loss problem; orphan files remain in the table location until expired by snapshot expiration.
Solutions
- Verify the table data is correct — the commit succeeded, no action is needed for correctness.
- Check FileIO/storage credentials and permissions allow deletion in the table metadata directory.
- Run expireSnapshots / removeOrphanFiles to garbage-collect the leftover manifests.
- Retry the job; if the warning persists, inspect the wrapped exception (logged as 'e') for the storage-side cause.
Defensive patterns
Strategy: retry
Validate before calling
// ensure FileIO can delete in the table location before committing io.deleteFile(testManifestPath); // or check permissions via a dry-run delete of a temp file
Try / catch
try { FlinkManifestUtil.deleteCommittedManifests(...); } catch (Exception e) { LOG.warn("manifest cleanup deferred; run expireSnapshots", e); } Prevention
- Grant delete permissions on the table metadata location to the Flink job's storage credentials
- Schedule regular expireSnapshots/removeOrphanFiles to reclaim orphan manifests
- Monitor this warn and alert only if it repeats, indicating persistent storage issues
When it happens
Trigger: deleteCommittedManifests is invoked after a checkpoint/jobgraph commit and the underlying FileIO.delete of a manifest or manifest-list path throws (e.g. transient S3/HDFS outage, permission denied on the table location, concurrent snapshot expiration already removed the file).
Common situations: Object-store eventual-consistency or throttling (S3 503) during delete; IAM/storage permissions lacking delete on the metadata dir; job crash between commit and cleanup; files already deleted by expireSnapshots running concurrently.
Understand the failure class
Background: "failed to write file", "Could not save figure", "Error saving remote file" — file write failed: causes and fixes across languages and libraries — this error's family across 38 libraries.
Related errors
- An error occurred while aborting the stream
- Bulk deletion failed
- Fail to deserialize aggregated statistics
- Failed on snapshot while reading manifest list
- Failed to close equality delta writer
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/9ac19f105f1e16a1.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/sink/FlinkManifestUtil.java:181
static void deleteCommittedManifests(
String tableName,
FileIO io,
List<ManifestFile> manifestsPath,
String newFlinkJobId,
long checkpointId) {
for (ManifestFile manifest : manifestsPath) {
try {
io.deleteFile(manifest.path());
} catch (Exception e) {
// The flink manifests cleaning failure shouldn't abort the completed checkpoint.
String details =
MoreObjects.toStringHelper(FlinkManifestUtil.class)
.add("tableName", tableName)
.add("flinkJobId", newFlinkJobId)
.add("checkpointId", checkpointId)
.add("manifestPath", manifest)
.toString();
LOG.warn(
"The iceberg transaction has been committed, but we failed to clean the temporary flink manifests: {}",
details,
e);
}
}
}
}
View on GitHub (pinned to 86d9c8fc54)