{"record":{"id":"34d82b2b61437785","repo":"apache/iceberg","slug":"failed-to-close-table-scan-s","errorCode":null,"errorMessage":"Failed to close table scan: %s","messagePattern":"Failed to close table scan: (.+?)","errorType":"exception","errorClass":"UncheckedIOException","httpStatus":null,"severity":"error","filePath":"mr/src/main/java/org/apache/iceberg/mr/mapreduce/IcebergInputFormat.java","lineNumber":149,"sourceCode":"    }\n\n    // TODO add a filter parser to get rid of Serialization\n    Expression filter =\n        SerializationUtil.deserializeFromBase64(conf.get(InputFormatConfig.FILTER_EXPRESSION));\n    if (filter != null) {\n      scan = scan.filter(filter);\n    }\n\n    List<InputSplit> splits = Lists.newArrayList();\n    scan = scan.planWith(workerPool);\n    try (CloseableIterable<CombinedScanTask> tasksIterable = scan.planTasks()) {\n      Table serializableTable = SerializableTable.copyOf(table);\n      tasksIterable.forEach(\n          task -> {\n            splits.add(new IcebergSplit(serializableTable, conf, task));\n          });\n    } catch (IOException e) {\n      throw new UncheckedIOException(String.format(\"Failed to close table scan: %s\", scan), e);\n    }\n\n    // if enabled, do not serialize FileIO hadoop config to decrease split size\n    // However, do not skip serialization for metatable queries, because some metadata tasks cache\n    // the IO object and we\n    // wouldn't be able to inject the config into these tasks on the deserializer-side, unlike for\n    // standard queries\n    if (scan instanceof DataTableScan) {\n      checkAndSkipIoConfigSerialization(conf, table);\n    }\n\n    return splits;\n  }\n\n  /**\n   * If enabled, it ensures that the FileIO's hadoop configuration will not be serialized. This\n   * might be desirable for decreasing the overall size of serialized table objects.\n   *","sourceCodeStart":131,"sourceCodeEnd":167,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/mr/src/main/java/org/apache/iceberg/mr/mapreduce/IcebergInputFormat.java#L131-L167","documentation":"IcebergInputFormat.planInputSplits closes the TableScan via Tasks and if closing throws IOException it wraps it, with the scan in the message, into an UncheckedIOException. The split planning itself succeeded or partially succeeded; the failure is in releasing scan resources.","triggerScenarios":"TableScan.close() throwing IOException during getSplits() — typically an underlying FileIO/Hadoop filesystem error while releasing resources after listing manifests.","commonSituations":"HDFS/S3 connectivity blips during job submission; NameNode unreachable; credentials expiring mid-planning when the scan holds open resources.","solutions":["Inspect the chained cause (getCause()) for the real IOException from the filesystem/FileIO and fix connectivity/credentials.","Retry getSplits(); planning is typically retryable.","Verify Hadoop/Hive conf and access permissions for the table location on the submitting client.","Check for known FileIO issues with your storage plugin version."],"exampleFix":"// before: failing silently surfaces wrapped scan-close error\nList<InputSplit> splits = inputFormat.getSplits(jobContext);\n// after: add retry/cause inspection\ntry {\n  splits = inputFormat.getSplits(jobContext);\n} catch (UncheckedIOException e) {\n  LOG.error(\"scan close failed: {}\", e.getCause(), e);\n  throw e;\n}","handlingStrategy":"try-catch","validationCode":"// pre-check storage reachability before submitting the job\nfs = new Path(tableLocation).getFileSystem(conf);\nif (!fs.exists(new Path(tableLocation))) throw new IOException(\"table location unreachable\");","typeGuard":null,"tryCatchPattern":"try {\n  splits = inputFormat.getSplits(jobContext);\n} catch (UncheckedIOException e) {\n  LOG.error(\"scan close failed: {}\", e.getCause(), e);\n  throw new IOException(e.getCause());\n}","preventionTips":["Inspect getCause() — the wrap hides the real FileIO/FileSystem error","Ensure Hadoop credentials/kerberos are valid for the job submitter","Retry split planning on transient storage errors","Check NameNode/S3 endpoint reachability from the client node"],"tags":["mapreduce","io-exception","table-scan"],"backgroundTag":"file-read-failed","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}