apache/iceberg · warning

File at offset contains records, exceeding…

Error message

File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.

What it means

AsyncSparkMicroBatchPlanner limits each micro-batch by maxRecordsPerMicroBatch. If a single file alone exceeds the row limit, including it wholly is the only way to guarantee forward progress; the planner logs this warning and processes the oversized file in one batch, potentially raising memory usage.

Solutions

  1. Increase maxRecordsPerMicroBatch so it comfortably exceeds the largest expected file's row count
  2. Configure upstream writers to write smaller files (e.g. smaller target file size or more frequent compaction)
  3. Compact/rewrite large files so future batches can respect the limit
  4. Monitor memory during such batches and size executors accordingly

Example fix

// before
.option("maxRecordsPerMicroBatch", "1000")   // files contain >1000 rows each
// after
.option("maxRecordsPerMicroBatch", "500000") // above largest file size
Defensive patterns

Strategy: validation

Validate before calling

long maxRows = Long.parseLong(limits.get("maxRecordsPerMicroBatch"));
long largestFileRows = files.stream().mapToLong(f -> f.recordCount()).max().orElse(0);
if (largestFileRows > maxRows) { /* raise maxRecordsPerMicroBatch above largestFileRows */ }

Prevention

When it happens

Trigger: computeLimitedOffset (called from latestOffset) encounters a first file (filesSeen == 1) whose row count exceeds unpackedLimits.getMaxRows() from the maxRecordsPerMicroBatch limit.

Common situations: Downstream writers producing very large files relative to a small maxRecordsPerMicroBatch setting; users lowering the limit for latency without considering file sizes.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/5330cf297632d697. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.1/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java:328

        if (filesSeen == 0) {
          return null;
        }
        LOG.debug(
            "latestOffset hit file limit at {}, rows: {}, files: {}",
            elem.first(),
            rowsSeen,
            filesSeen);
        return elem.first();
      }

      // Soft limit on rows - include file FIRST, then check
      rowsSeen += fileRows;
      filesSeen += 1;

      // Check if we've hit the row limit after including this file
      if (rowsSeen >= unpackedLimits.getMaxRows()) {
        if (filesSeen == 1 && rowsSeen > unpackedLimits.getMaxRows()) {
          LOG.warn(
              "File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
                  + "This file will be processed entirely to guarantee forward progress. "
                  + "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
              elem.second().file().location(),
              elem.first(),
              fileRows,
              unpackedLimits.getMaxRows());
        }
        // Return the offset of the NEXT element (or synthesize tail+1)
        if (i + 1 < queueSnapshot.size()) {
          LOG.debug(
              "latestOffset hit row limit at {}, rows: {}, files: {}",
              queueSnapshot.get(i + 1).first(),
              rowsSeen,
              filesSeen);
          return queueSnapshot.get(i + 1).first();
        } else {
          // This is the last element - return tail+1

View on GitHub (pinned to 86d9c8fc54)