apache/iceberg · warning

File at offset contains records, exceeding…

Error message

File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.

What it means

A LOG.warn in AsyncSparkMicroBatchPlanner.computeLimitedOffset: a single source file contains more records than maxRecordsPerMicroBatch. Because a file cannot be split mid-batch here, the entire file is processed anyway to guarantee forward progress, potentially exceeding the intended batch size and memory budget.

Solutions

  1. Increase maxRecordsPerMicroBatch so typical files fit within the limit.
  2. Make the producer write smaller files (smaller target-file-size / more frequent commits) so no single file exceeds the batch limit.
  3. Accept the behavior if forward progress matters more than strict batch sizing; just provision executor memory for the largest file.
  4. Combine with maxBytesPerMicroBatch or file-size-based limits if available to bound memory more predictably.

Example fix

// before
option("maxRecordsPerMicroBatch", 10000) // files contain 500k rows
// after
option("maxRecordsPerMicroBatch", 1000000) // >= largest expected file's row count
Defensive patterns

Strategy: fallback

Validate before calling

// Compare configured row limit against the largest expected file size
long approxRowsPerFile = targetFileSizeBytes / avgRowSizeBytes;
if (approxRowsPerFile > maxRecordsPerMicroBatch) {
  LOG.warn("maxRecordsPerMicroBatch below single-file row count; oversized files will bypass the limit");
}

Prevention

When it happens

Trigger: Streaming query using the async/micro-batch planner where one Parquet data file's record count pushes rowsSeen past unpackedLimits.getMaxRows() while it is the first (filesSeen == 1) file considered in the offset computation.

Common situations: Upstream batch jobs writing very large files into a table consumed by a streaming read with a low maxRecordsPerMicroBatch; misconfigured size-based limits assuming many small files; memory pressure or long micro-batches observed on the consumer.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/3ad29ef41ca8848a. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java:328

        if (filesSeen == 0) {
          return null;
        }
        LOG.debug(
            "latestOffset hit file limit at {}, rows: {}, files: {}",
            elem.first(),
            rowsSeen,
            filesSeen);
        return elem.first();
      }

      // Soft limit on rows - include file FIRST, then check
      rowsSeen += fileRows;
      filesSeen += 1;

      // Check if we've hit the row limit after including this file
      if (rowsSeen >= unpackedLimits.getMaxRows()) {
        if (filesSeen == 1 && rowsSeen > unpackedLimits.getMaxRows()) {
          LOG.warn(
              "File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
                  + "This file will be processed entirely to guarantee forward progress. "
                  + "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
              elem.second().file().location(),
              elem.first(),
              fileRows,
              unpackedLimits.getMaxRows());
        }
        // Return the offset of the NEXT element (or synthesize tail+1)
        if (i + 1 < queueSnapshot.size()) {
          LOG.debug(
              "latestOffset hit row limit at {}, rows: {}, files: {}",
              queueSnapshot.get(i + 1).first(),
              rowsSeen,
              filesSeen);
          return queueSnapshot.get(i + 1).first();
        } else {
          // This is the last element - return tail+1

View on GitHub (pinned to 86d9c8fc54)