apache/iceberg · warning

File at offset contains records, exceeding…

Error message

File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.

What it means

AsyncSparkMicroBatchPlanner.computeLimitedOffset caps micro-batch input at maxRecordsPerMicroBatch, but when a single file alone exceeds the limit it cannot be split here, so it is planned whole to guarantee forward progress. The warning tells the user the memory-safety limit was exceeded and that memory usage may spike. The batch still runs correctly.

Solutions

  1. Increase the maxRecordsPerMicroBatch limit to exceed the typical file's record count.
  2. Rewrite/compact the source with smaller files so no single file exceeds the limit.
  3. Accept the warning if occasional large-file batches are tolerable; it is safe but uses more memory.

Example fix

// before
spark.sql('ALTER TABLE db.t SET TBLPROPERTIES (\'maxRecordsPerMicroBatch\'=\'1000\')')
// after
spark.sql('ALTER TABLE db.t SET TBLPROPERTIES (\'maxRecordsPerMicroBatch\'=\'1000000\')')
Defensive patterns

Strategy: validation

Validate before calling

long maxRecords = Long.parseLong(sparkConf.get("maxRecordsPerMicroBatch", "10000"));
if (largestFileRowsInStream(table) > maxRecords) {
  // raise limit or compact source files
}

Prevention

When it happens

Trigger: A data file in the stream offset range contains more rows than the configured maxRecordsPerMicroBatch; detected in latestOffset while computing the limited batch offset.

Common situations: Streaming reads over tables written with very large compaction outputs or unpartitioned bulk backfills; too-low maxRecordsPerMicroBatch setting relative to file sizes.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/06e4dd2edb50f161. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.0/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java:328

        if (filesSeen == 0) {
          return null;
        }
        LOG.debug(
            "latestOffset hit file limit at {}, rows: {}, files: {}",
            elem.first(),
            rowsSeen,
            filesSeen);
        return elem.first();
      }

      // Soft limit on rows - include file FIRST, then check
      rowsSeen += fileRows;
      filesSeen += 1;

      // Check if we've hit the row limit after including this file
      if (rowsSeen >= unpackedLimits.getMaxRows()) {
        if (filesSeen == 1 && rowsSeen > unpackedLimits.getMaxRows()) {
          LOG.warn(
              "File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
                  + "This file will be processed entirely to guarantee forward progress. "
                  + "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
              elem.second().file().location(),
              elem.first(),
              fileRows,
              unpackedLimits.getMaxRows());
        }
        // Return the offset of the NEXT element (or synthesize tail+1)
        if (i + 1 < queueSnapshot.size()) {
          LOG.debug(
              "latestOffset hit row limit at {}, rows: {}, files: {}",
              queueSnapshot.get(i + 1).first(),
              rowsSeen,
              filesSeen);
          return queueSnapshot.get(i + 1).first();
        } else {
          // This is the last element - return tail+1

View on GitHub (pinned to 86d9c8fc54)