apache/iceberg · info
Failed to get block locations for path
Error message
Failed to get block locations for path {} What it means
Util.blockLocations() resolves host lists for the blocks backing a CombineFileSplit so tasks can be scheduled locality-aware. When fetching block locations from the FileSystem throws IOException, it logs this warning, skips that file's hosts, and returns the hosts gathered so far. The computation proceeds with degraded locality rather than failing.
Solutions
- Re-run the query so splits are re-planned against current file list
- Ensure snapshot expiration/compaction jobs are not deleting files concurrently with reads
- Fix HDFS/cloud storage connectivity issues indicated by the IOException
- Ignore if infrequent — this only affects data locality, not correctness
Example fix
// before: files deleted mid-query by expireSnapshots running concurrently // after: quiesce maintenance jobs before long-running reads, or pin to a snapshot Table table = catalog.loadTable(...); table.refresh(); // re-plan the scan after maintenance completes
Defensive patterns
Strategy: try-catch
Try / catch
// framework-level; users cannot catch directly.
// Guard by ensuring input splits reference live files:
try {
dataset.scan().planTasks();
} catch (RuntimeException e) { /* replan after maintenance */ } Prevention
- Don't run expireSnapshots/orphan cleanup concurrently with long reads
- Pin reads to a snapshot and hold retention (max-snapshot-age-ms) above read duration
- Monitor HDFS/cloud storage errors for root cause
- Treat as benign unless locality-sensitive performance regresses
When it happens
Trigger: A task configuration with CombineFileSplit input where fs.getFileBlockLocations(path, start, length) throws IOException — e.g. the file was deleted or renamed between split planning and task execution, or a transient HDFS/S3 error occurred.
Common situations: Reading a table where a compaction/rewrite job deleted data files after the query plan was created; HDFS NameNode hiccup; file not found at task time due to expired snapshot cleanup.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Create namespace failed
- Error reading version hint file
- Error trying to recover the latest version number for
- Failed to create file
- Failed to create Parquet input file for
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/1e58304d1265c9b0.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/hadoop/Util.java:71
public static FileSystem getFs(Path path, Configuration conf) {
try {
return path.getFileSystem(conf);
} catch (IOException e) {
throw new RuntimeIOException(e, "Failed to get file system for path: %s", path);
}
}
public static String[] blockLocations(ScanTaskGroup<FileScanTask> taskGroup, Configuration conf) {
Set<String> locationSets = Sets.newHashSet();
for (FileScanTask f : taskGroup.tasks()) {
Path path = new Path(f.file().location());
try {
FileSystem fs = path.getFileSystem(conf);
for (BlockLocation b : fs.getFileBlockLocations(path, f.start(), f.length())) {
locationSets.addAll(Arrays.asList(b.getHosts()));
}
} catch (IOException ioe) {
LOG.warn("Failed to get block locations for path {}", path, ioe);
}
}
return locationSets.toArray(new String[0]);
}
public static String[] blockLocations(FileIO io, ScanTaskGroup<?> taskGroup) {
Set<String> locations = Sets.newHashSet();
for (ScanTask task : taskGroup.tasks()) {
if (task instanceof ContentScanTask) {
Collections.addAll(locations, blockLocations(io, (ContentScanTask<?>) task));
}
}
return locations.toArray(HadoopInputFile.NO_LOCATION_PREFERENCE);
}
View on GitHub (pinned to 86d9c8fc54)