apache/iceberg · error · RuntimeException

Unable to read Avro file:

Error message

Unable to read Avro file: 

What it means

TableMigrationUtil.getAvroMetrics reads an Avro file's row count via HadoopInputFile and Avro.rowCount. If reading throws UncheckedIOException (unreadable/corrupt file, IO error), it is rethrown as a RuntimeException 'Unable to read Avro file: <path>'. Only the row count metric is computed; the failure means the file could not be opened/read at all.

Solutions

  1. Open the chained cause and check the file directly (permissions, existence, footer integrity).
  2. Repair or re-export the corrupt/truncated Avro file before migrating.
  3. Exclude the damaged file from migration and re-add it after repair.
  4. Retry on transient storage errors (throttling, connection resets).

Example fix

// before: one bad Avro file aborts metrics for the partition
Metrics m = TableMigrationUtil.metrics(path, conf, spec, "avro", metricsSpec, mapping);

// after: pre-check readability
if (fs.exists(path) && fs.open(path) != null) {
  Metrics m = TableMigrationUtil.metrics(path, conf, spec, "avro", metricsSpec, mapping);
}
Defensive patterns

Strategy: validation

Validate before calling

// pre-check the Avro file is readable before computing metrics
try (DataFileReader<Object, Object> r =
         new DataFileReader<>(new File(path.toUri()), new GenericDatumReader<>())) {
  r.getBlockCount();
}

Try / catch

try {
  Metrics m = TableMigrationUtil.metrics(path, conf, spec, "avro", metricsSpec, mapping);
} catch (RuntimeException e) {
  if (e.getMessage().startsWith("Unable to read Avro file")) {
    quarantine(path); // repair or exclude the corrupt file
  } else { throw e; }
}

Prevention

When it happens

Trigger: Calling metrics()/listPartition on an Avro data file that is corrupt, truncated, missing, or unreadable due to permissions/filesystem errors — Avro.rowCount's UncheckedIOException is converted here.

Common situations: Migrating legacy Avro tables with damaged or partially-written files; files deleted between listing and metric computation; permissions changed mid-migration; HDFS/S3 transient errors while reading the footer.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/2d9120a19af5e313. Report an issue: GitHub.

Appendix: source

Thrown at data/src/main/java/org/apache/iceberg/data/TableMigrationUtil.java:220

        throw new UnsupportedOperationException("Unknown partition format: " + format);
      }
      return Arrays.asList(datafiles);
    } catch (IOException e) {
      throw new UncheckedIOException("Unable to list files in partition: " + partitionUri, e);
    } finally {
      if (service != null) {
        service.shutdown();
      }
    }
  }

  private static Metrics getAvroMetrics(Path path, Configuration conf) {
    try {
      InputFile file = HadoopInputFile.fromPath(path, conf);
      long rowCount = Avro.rowCount(file);
      return new Metrics(rowCount, null, null, null, null);
    } catch (UncheckedIOException e) {
      throw new RuntimeException("Unable to read Avro file: " + path, e);
    }
  }

  private static Metrics getParquetMetrics(
      Path path, Configuration conf, MetricsConfig metricsSpec, NameMapping mapping) {
    try {
      InputFile file = HadoopInputFile.fromPath(path, conf);
      return ParquetUtil.fileMetrics(file, metricsSpec, mapping);
    } catch (UncheckedIOException e) {
      throw new RuntimeException("Unable to read the metrics of the Parquet file: " + path, e);
    }
  }

  private static Metrics getOrcMetrics(
      Path path, Configuration conf, MetricsConfig metricsSpec, NameMapping mapping) {
    try {
      return OrcMetrics.fromInputFile(HadoopInputFile.fromPath(path, conf), metricsSpec, mapping);
    } catch (UncheckedIOException e) {

View on GitHub (pinned to 86d9c8fc54)