apache/iceberg · warning
Merging duplicate DVs for data file in table .
Error message
Merging {} duplicate DVs for data file {} in table {}. What it means
mergeDVs() groups delete-vector (DV) DeleteFiles by the data file they reference. When more than one DV targets the same data file, they are merged and this WARN reports the count, data file path, and table name. Duplicate DVs usually indicate the same data file received delete vectors more than once (repeated delete operations), and merging is done to keep the delete manifest valid.
Solutions
- This is a handled warning (DVs are merged automatically) — verify results are correct; no action strictly required.
- Compact data files (rewrite_data_files / RewriteDataFiles) so future deletes start from fresh files and DVs do not accumulate.
- Investigate job orchestration if the same delete operation is being applied multiple times (duplicate task retries).
- Consider enabling copy-on-write or frequent compaction for hot files repeatedly receiving DVs.
Example fix
// before: repeated DV deletes accumulate on hot files
spark.sql("DELETE FROM t WHERE id = 1");
// after: schedule compaction to consolidate files/DVs between delete batches
spark.sql("CALL catalog.system.rewrite_data_files(table => 't')"); Defensive patterns
Strategy: validation
Validate before calling
// detect hot data files accumulating multiple DVs before commit
if (dvsByReferencedFile.values().stream().anyMatch(l -> l.size() > 1)) {
scheduleCompaction(table); // rewrite files so future deletes start fresh
} Type guard
static boolean hasDuplicateDVs(Map<String, List<DeleteFile>> dvsByFile) {
return dvsByFile.values().stream().anyMatch(l -> l.size() > 1);
} Prevention
- Schedule rewrite_data_files compaction for files repeatedly receiving DVs.
- Avoid re-running failed delete jobs that may reapply the same DV deletes.
- Monitor this WARN as a signal of duplicate delete application or missing compaction.
- Prefer copy-on-write mode for tables with heavy delete workloads on hot files.
When it happens
Trigger: Multiple DV-producing delete operations target the same data file (e.g. equality/position deletes producing DVs, then another delete on the same file without rewriting it), so dvsByReferencedFile maps one data file to several DeleteFile entries when the new delete manifest is built.
Common situations: Repeated small delete/merge-on-read workloads hitting the same data files; jobs re-run after partial failure issuing duplicate DV deletes; engine configurations that do not compact/rewrite files between deletes (delete.effective-hour/no DV coalescing).
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- Cannot get value for invalid index:
- Cannot get value for invalid index
- Cannot rewrite manifests in a
- Deleted rows scan task is not supported yet
- Deleted rows scan task is not supported yet
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/720e567d5d2948fc.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/MergingSnapshotProducer.java:1212
newDeleteFilesBySpec.forEach(
(specId, deleteFiles) -> {
PartitionSpec spec = ops().current().spec(specId);
deleteFiles.forEach(file -> addedDeleteFilesSummary.addedFile(spec, file));
List<ManifestFile> newDeleteManifests = writeDeleteManifests(deleteFiles, spec);
cachedNewDeleteManifests.addAll(newDeleteManifests);
});
this.hasNewDeleteFiles = false;
}
return cachedNewDeleteManifests;
}
private List<DeleteFile> mergeDVs() {
for (Map.Entry<String, List<DeleteFile>> entry : dvsByReferencedFile.entrySet()) {
if (entry.getValue().size() > 1) {
LOG.warn(
"Merging {} duplicate DVs for data file {} in table {}.",
entry.getValue().size(),
entry.getKey(),
tableName);
}
}
FileIO fileIO = EncryptingFileIO.combine(ops().io(), ops().encryption());
String dvOutputLocation =
ops()
.locationProvider()
.newDataLocation(
FileFormat.PUFFIN.addExtension(
String.format(
"merged-dvs-%s-%s", snapshotId(), dvMergeAttempt.incrementAndGet())));
return DVUtil.mergeAndWriteDVsIfRequired(View on GitHub (pinned to 86d9c8fc54)