apache/iceberg · warning

Failed to delete file

Error message

Failed to delete file: {}

What it means

A logged warning emitted per-file by DeleteOrphanFilesSparkAction.deleteNonBulk when the table's FileIO does not implement SupportsBulkOperations and individual deleteFile calls fail. Tasks.foreach with suppressFailureWhenFinished and noRetry collects each failure in the onFailure handler so one bad file does not stop the sweep; failed files are simply left behind.

Solutions

  1. Re-run deleteOrphanFiles after fixing the cause; previously failed files are retried.
  2. Grant the job's principal delete permission on the table/data directory (HDFS ACLs, POSIX perms).
  3. Check each chained exception: 'file does not exist' means a concurrent job already deleted it (benign).
  4. Switch to a FileIO with bulk support (e.g., S3FileIO) to reduce partial-failure windows.
  5. Serialize orphan sweeps with other maintenance jobs on the same table.

Example fix

// before
// HadoopFileIO user lacks delete perms -> per-file warnings
hdfs dfs -chmod -R o-rwx /warehouse/db/table  # wrong direction
// after
hdfs dfs -chmod -R 775 /warehouse/db/table
hdfs dfs -chown -R hive:hive /warehouse/db/table
// then re-run deleteOrphanFiles
Defensive patterns

Strategy: retry

Validate before calling

// pre-verify delete permission on the data directory
FileSystem fs = new Path(tableLocation).getFileSystem(conf);
FsPermission perm = fs.getFileStatus(new Path(tableLocation)).getPermission();

Try / catch

try { deleteNonBulk(paths); } catch (Exception e) { /* per-file failures already logged; re-run */ }

Prevention

When it happens

Trigger: Running DeleteOrphanFiles on a FileIO without bulk operations (e.g., HadoopFileIO or a custom IO) when table.io().deleteFile(file) throws per file — HDFS permission errors, files already removed by a concurrent job, or transient NameNode/storage errors.

Common situations: HDFS clusters where the job user lacks delete permission on the table directory; orphan sweep racing with ExpireSnapshots; files deleted between listing and delete (ENOENT); misconfigured custom FileIO.

Understand the failure class

Background: "failed to write file", "Could not save figure", "Error saving remote file" — file write failed: causes and fixes across languages and libraries — this error's family across 38 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/7d7bf150bb4ffc18. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/actions/DeleteOrphanFilesSparkAction.java:336

  private void deleteBulk(SupportsBulkOperations io, List<String> paths) {
    try {
      io.deleteFiles(paths);
      LOG.info("Deleted {} files using bulk deletes", paths.size());
    } catch (BulkDeletionFailureException e) {
      int deletedFilesCount = paths.size() - e.numberFailedObjects();
      LOG.warn(
          "Deleted only {} of {} files using bulk deletes", deletedFilesCount, paths.size(), e);
    }
  }

  private void deleteNonBulk(List<String> paths) {
    Tasks.Builder<String> deleteTasks =
        Tasks.foreach(paths)
            .noRetry()
            .executeWith(deleteExecutorService)
            .suppressFailureWhenFinished()
            .onFailure((file, exc) -> LOG.warn("Failed to delete file: {}", file, exc));

    if (deleteFunc == null) {
      LOG.info(
          "Table IO {} does not support bulk operations. Using non-bulk deletes.",
          table.io().getClass().getName());
      deleteTasks.run(table.io()::deleteFile);
    } else {
      LOG.info("Custom delete function provided. Using non-bulk deletes");
      deleteTasks.run(deleteFunc::accept);
    }
  }

  @VisibleForTesting
  static Dataset<String> findOrphanFiles(
      Dataset<FileURI> actualFileIdentDS,
      Dataset<FileURI> validFileIdentDS,
      PrefixMismatchMode prefixMismatchMode) {

View on GitHub (pinned to 86d9c8fc54)