{"record":{"id":"720e567d5d2948fc","repo":"apache/iceberg","slug":"merging-duplicate-dvs-for-data-file-in-table","errorCode":null,"errorMessage":"Merging {} duplicate DVs for data file {} in table {}.","messagePattern":"Merging (.+?) duplicate DVs for data file (.+?) in table (.+?)\\.","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"core/src/main/java/org/apache/iceberg/MergingSnapshotProducer.java","lineNumber":1212,"sourceCode":"\n      newDeleteFilesBySpec.forEach(\n          (specId, deleteFiles) -> {\n            PartitionSpec spec = ops().current().spec(specId);\n            deleteFiles.forEach(file -> addedDeleteFilesSummary.addedFile(spec, file));\n            List<ManifestFile> newDeleteManifests = writeDeleteManifests(deleteFiles, spec);\n            cachedNewDeleteManifests.addAll(newDeleteManifests);\n          });\n\n      this.hasNewDeleteFiles = false;\n    }\n\n    return cachedNewDeleteManifests;\n  }\n\n  private List<DeleteFile> mergeDVs() {\n    for (Map.Entry<String, List<DeleteFile>> entry : dvsByReferencedFile.entrySet()) {\n      if (entry.getValue().size() > 1) {\n        LOG.warn(\n            \"Merging {} duplicate DVs for data file {} in table {}.\",\n            entry.getValue().size(),\n            entry.getKey(),\n            tableName);\n      }\n    }\n\n    FileIO fileIO = EncryptingFileIO.combine(ops().io(), ops().encryption());\n\n    String dvOutputLocation =\n        ops()\n            .locationProvider()\n            .newDataLocation(\n                FileFormat.PUFFIN.addExtension(\n                    String.format(\n                        \"merged-dvs-%s-%s\", snapshotId(), dvMergeAttempt.incrementAndGet())));\n\n    return DVUtil.mergeAndWriteDVsIfRequired(","sourceCodeStart":1194,"sourceCodeEnd":1230,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/core/src/main/java/org/apache/iceberg/MergingSnapshotProducer.java#L1194-L1230","documentation":"mergeDVs() groups delete-vector (DV) DeleteFiles by the data file they reference. When more than one DV targets the same data file, they are merged and this WARN reports the count, data file path, and table name. Duplicate DVs usually indicate the same data file received delete vectors more than once (repeated delete operations), and merging is done to keep the delete manifest valid.","triggerScenarios":"Multiple DV-producing delete operations target the same data file (e.g. equality/position deletes producing DVs, then another delete on the same file without rewriting it), so dvsByReferencedFile maps one data file to several DeleteFile entries when the new delete manifest is built.","commonSituations":"Repeated small delete/merge-on-read workloads hitting the same data files; jobs re-run after partial failure issuing duplicate DV deletes; engine configurations that do not compact/rewrite files between deletes (delete.effective-hour/no DV coalescing).","solutions":["This is a handled warning (DVs are merged automatically) — verify results are correct; no action strictly required.","Compact data files (rewrite_data_files / RewriteDataFiles) so future deletes start from fresh files and DVs do not accumulate.","Investigate job orchestration if the same delete operation is being applied multiple times (duplicate task retries).","Consider enabling copy-on-write or frequent compaction for hot files repeatedly receiving DVs."],"exampleFix":"// before: repeated DV deletes accumulate on hot files\nspark.sql(\"DELETE FROM t WHERE id = 1\");\n// after: schedule compaction to consolidate files/DVs between delete batches\nspark.sql(\"CALL catalog.system.rewrite_data_files(table => 't')\");","handlingStrategy":"validation","validationCode":"// detect hot data files accumulating multiple DVs before commit\nif (dvsByReferencedFile.values().stream().anyMatch(l -> l.size() > 1)) {\n  scheduleCompaction(table); // rewrite files so future deletes start fresh\n}","typeGuard":"static boolean hasDuplicateDVs(Map<String, List<DeleteFile>> dvsByFile) {\n  return dvsByFile.values().stream().anyMatch(l -> l.size() > 1);\n}","tryCatchPattern":null,"preventionTips":["Schedule rewrite_data_files compaction for files repeatedly receiving DVs.","Avoid re-running failed delete jobs that may reapply the same DV deletes.","Monitor this WARN as a signal of duplicate delete application or missing compaction.","Prefer copy-on-write mode for tables with heavy delete workloads on hot files."],"tags":["deletes","delete-vectors","manifests","duplicates"],"backgroundTag":"invalid-state-transition","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}