{"record":{"id":"1e58304d1265c9b0","repo":"apache/iceberg","slug":"failed-to-get-block-locations-for-path","errorCode":null,"errorMessage":"Failed to get block locations for path {}","messagePattern":"Failed to get block locations for path (.+?)","errorType":"console","errorClass":null,"httpStatus":null,"severity":"info","filePath":"core/src/main/java/org/apache/iceberg/hadoop/Util.java","lineNumber":71,"sourceCode":"  public static FileSystem getFs(Path path, Configuration conf) {\n    try {\n      return path.getFileSystem(conf);\n    } catch (IOException e) {\n      throw new RuntimeIOException(e, \"Failed to get file system for path: %s\", path);\n    }\n  }\n\n  public static String[] blockLocations(ScanTaskGroup<FileScanTask> taskGroup, Configuration conf) {\n    Set<String> locationSets = Sets.newHashSet();\n    for (FileScanTask f : taskGroup.tasks()) {\n      Path path = new Path(f.file().location());\n      try {\n        FileSystem fs = path.getFileSystem(conf);\n        for (BlockLocation b : fs.getFileBlockLocations(path, f.start(), f.length())) {\n          locationSets.addAll(Arrays.asList(b.getHosts()));\n        }\n      } catch (IOException ioe) {\n        LOG.warn(\"Failed to get block locations for path {}\", path, ioe);\n      }\n    }\n\n    return locationSets.toArray(new String[0]);\n  }\n\n  public static String[] blockLocations(FileIO io, ScanTaskGroup<?> taskGroup) {\n    Set<String> locations = Sets.newHashSet();\n\n    for (ScanTask task : taskGroup.tasks()) {\n      if (task instanceof ContentScanTask) {\n        Collections.addAll(locations, blockLocations(io, (ContentScanTask<?>) task));\n      }\n    }\n\n    return locations.toArray(HadoopInputFile.NO_LOCATION_PREFERENCE);\n  }\n","sourceCodeStart":53,"sourceCodeEnd":89,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/core/src/main/java/org/apache/iceberg/hadoop/Util.java#L53-L89","documentation":"Util.blockLocations() resolves host lists for the blocks backing a CombineFileSplit so tasks can be scheduled locality-aware. When fetching block locations from the FileSystem throws IOException, it logs this warning, skips that file's hosts, and returns the hosts gathered so far. The computation proceeds with degraded locality rather than failing.","triggerScenarios":"A task configuration with CombineFileSplit input where fs.getFileBlockLocations(path, start, length) throws IOException — e.g. the file was deleted or renamed between split planning and task execution, or a transient HDFS/S3 error occurred.","commonSituations":"Reading a table where a compaction/rewrite job deleted data files after the query plan was created; HDFS NameNode hiccup; file not found at task time due to expired snapshot cleanup.","solutions":["Re-run the query so splits are re-planned against current file list","Ensure snapshot expiration/compaction jobs are not deleting files concurrently with reads","Fix HDFS/cloud storage connectivity issues indicated by the IOException","Ignore if infrequent — this only affects data locality, not correctness"],"exampleFix":"// before: files deleted mid-query by expireSnapshots running concurrently\n// after: quiesce maintenance jobs before long-running reads, or pin to a snapshot\nTable table = catalog.loadTable(...); table.refresh();\n// re-plan the scan after maintenance completes","handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"// framework-level; users cannot catch directly.\n// Guard by ensuring input splits reference live files:\ntry {\n  dataset.scan().planTasks();\n} catch (RuntimeException e) { /* replan after maintenance */ }","preventionTips":["Don't run expireSnapshots/orphan cleanup concurrently with long reads","Pin reads to a snapshot and hold retention (max-snapshot-age-ms) above read duration","Monitor HDFS/cloud storage errors for root cause","Treat as benign unless locality-sensitive performance regresses"],"tags":["hadoop","io","locality","splits"],"backgroundTag":"file-read-failed","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}