apache/iceberg · info

Failed to get block locations for path

Error message

Failed to get block locations for path {}

What it means

Util.blockLocations() resolves host lists for the blocks backing a CombineFileSplit so tasks can be scheduled locality-aware. When fetching block locations from the FileSystem throws IOException, it logs this warning, skips that file's hosts, and returns the hosts gathered so far. The computation proceeds with degraded locality rather than failing.

Solutions

  1. Re-run the query so splits are re-planned against current file list
  2. Ensure snapshot expiration/compaction jobs are not deleting files concurrently with reads
  3. Fix HDFS/cloud storage connectivity issues indicated by the IOException
  4. Ignore if infrequent — this only affects data locality, not correctness

Example fix

// before: files deleted mid-query by expireSnapshots running concurrently
// after: quiesce maintenance jobs before long-running reads, or pin to a snapshot
Table table = catalog.loadTable(...); table.refresh();
// re-plan the scan after maintenance completes
Defensive patterns

Strategy: try-catch

Try / catch

// framework-level; users cannot catch directly.
// Guard by ensuring input splits reference live files:
try {
  dataset.scan().planTasks();
} catch (RuntimeException e) { /* replan after maintenance */ }

Prevention

When it happens

Trigger: A task configuration with CombineFileSplit input where fs.getFileBlockLocations(path, start, length) throws IOException — e.g. the file was deleted or renamed between split planning and task execution, or a transient HDFS/S3 error occurred.

Common situations: Reading a table where a compaction/rewrite job deleted data files after the query plan was created; HDFS NameNode hiccup; file not found at task time due to expired snapshot cleanup.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/1e58304d1265c9b0. Report an issue: GitHub.

Appendix: source

Thrown at core/src/main/java/org/apache/iceberg/hadoop/Util.java:71

  public static FileSystem getFs(Path path, Configuration conf) {
    try {
      return path.getFileSystem(conf);
    } catch (IOException e) {
      throw new RuntimeIOException(e, "Failed to get file system for path: %s", path);
    }
  }

  public static String[] blockLocations(ScanTaskGroup<FileScanTask> taskGroup, Configuration conf) {
    Set<String> locationSets = Sets.newHashSet();
    for (FileScanTask f : taskGroup.tasks()) {
      Path path = new Path(f.file().location());
      try {
        FileSystem fs = path.getFileSystem(conf);
        for (BlockLocation b : fs.getFileBlockLocations(path, f.start(), f.length())) {
          locationSets.addAll(Arrays.asList(b.getHosts()));
        }
      } catch (IOException ioe) {
        LOG.warn("Failed to get block locations for path {}", path, ioe);
      }
    }

    return locationSets.toArray(new String[0]);
  }

  public static String[] blockLocations(FileIO io, ScanTaskGroup<?> taskGroup) {
    Set<String> locations = Sets.newHashSet();

    for (ScanTask task : taskGroup.tasks()) {
      if (task instanceof ContentScanTask) {
        Collections.addAll(locations, blockLocations(io, (ContentScanTask<?>) task));
      }
    }

    return locations.toArray(HadoopInputFile.NO_LOCATION_PREFERENCE);
  }

View on GitHub (pinned to 86d9c8fc54)