apache/iceberg · warning

Using full compute as previous statistics file is corrupted…

Error message

Using full compute as previous statistics file is corrupted for incremental compute.

What it means

PartitionStatsHandler.computeAndWriteStatsFile tries to compute partition statistics incrementally by merging the previously written statistics file. If the previous statistics file is invalid (InvalidStatsFileException, e.g. unreadable or corrupt), it falls back to a full recompute from all manifests and logs this warning. The operation still succeeds; only efficiency is lost.

Solutions

  1. Let the fallback run: the partition stats file is recomputed fully and the next incremental compute will use the new valid file.
  2. Check the previous statistics file (snapshot.statisticsFiles(io)) for readability/corruption and delete stale entries via removeStatistics / expireSnapshots.
  3. Rewrite all historical statistics files with a full compute once to restore incremental merging.
  4. Upgrade Iceberg if the corrupt file was produced by an older writer version.
Defensive patterns

Strategy: fallback

Validate before calling

StatisticsFile prev = latestStatisticsFile(table, prevSnapshotId);
boolean usable = prev != null && table.io().newInputFile(prev.path()).exists();

Try / catch

try { computeIncremental(...); } catch (InvalidStatsFileException e) { computeFull(...); }

Prevention

When it happens

Trigger: Calling computeAndWriteStatsFile (or via Spark procedures writing partition stats) when the statistics file referenced by the previous snapshot's statistics list is corrupt, truncated, written by a buggy/older writer, or missing required content.

Common situations: Statistics files lost or truncated by incomplete uploads to object storage, metadata files copied between buckets without RewriteTablePathUtil, tables written by an older Iceberg version whose stats format differs, and interrupted write jobs that left a partial Parquet stats file registered.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/1627b443926d284a. Report an issue: GitHub.

Appendix: source

Thrown at core/src/main/java/org/apache/iceberg/PartitionStatsHandler.java:122

    Collection<PartitionStatistics> stats;
    PartitionStatisticsFile statisticsFile = latestStatsFile(table, snapshot.snapshotId());
    if (statisticsFile == null) {
      LOG.info(
          "Using full compute as previous statistics file is not present for incremental compute.");
      stats =
          computeStats(table, snapshot.allManifests(table.io()), false /* incremental */).values();
    } else {
      if (statisticsFile.snapshotId() == snapshotId) {
        // no-op
        LOG.info("Returning existing statistics file for snapshot {}", snapshotId);
        return statisticsFile;
      }

      try {
        stats = computeAndMergeStatsIncremental(table, snapshot, statisticsFile.snapshotId());
      } catch (InvalidStatsFileException exception) {
        LOG.warn(
            "Using full compute as previous statistics file is corrupted for incremental compute.");
        stats =
            computeStats(table, snapshot.allManifests(table.io()), false /* incremental */)
                .values();
      }
    }

    if (stats.isEmpty()) {
      // empty branch case
      return null;
    }

    StructType partitionType = Partitioning.partitionType(table);
    List<PartitionStatistics> sortedStats = sortStatsByPartition(stats, partitionType);
    return writePartitionStatsFile(
        table,
        snapshot.snapshotId(),
        PartitionStatistics.schema(partitionType, TableUtil.formatVersion(table)),

View on GitHub (pinned to 86d9c8fc54)