apache/iceberg · warning

File contains records, exceeding maxRecordsPerMicroBatch…

Error message

File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.

What it means

A WARN logged by SyncSparkMicroBatchPlanner.latestOffset when a single data file's recordCount exceeds the maxRecordsPerMicroBatch limit (spark.message streaming-max-records-per-micro-batch). Since a micro-batch cannot split a file, the whole file is included anyway to guarantee forward progress; memory usage per batch may exceed the configured limit.

Solutions

  1. Increase streaming-max-rows-per-micro-batch to exceed your largest file's record count
  2. Write smaller files upstream (target file size, more frequent compaction)
  3. Run rewriteDataFiles to shrink existing oversized files
  4. Raise the limit only as much as executor memory allows

Example fix

// before
.option("streaming-max-rows-per-micro-batch", 100000)
// after
.option("streaming-max-rows-per-micro-batch", 5000000) // > largest file's record count
Defensive patterns

Strategy: validation

Validate before calling

// check table file sizes against the configured limit before streaming
long maxRows = 5_000_000L; // streaming-max-rows-per-micro-batch
TableScan scan = table.newScan();
for (File f : scan.planFiles()) {
  if (f.recordCount() > maxRows) {
    throw new IllegalArgumentException("File " + f.path() + " has " + f.recordCount() + " > limit " + maxRows);
  }
}

Prevention

When it happens

Trigger: Streaming micro-batch planning encounters an added file whose task.file().recordCount() pushes curRecordCount past maxRows while curFilesAdded == 1 (first oversized file in the batch).

Common situations: Compaction disabled or small upstream writers producing very large files; STREAMING_MAX_ROWS_PER_MICRO_BATCH set too low relative to table file sizes; backfill/initial batch reading large files.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/0ec5b752f7859a81. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java:191

            CloseableIterator<FileScanTask> taskIter = taskIterable.iterator()) {
          while (taskIter.hasNext()) {
            FileScanTask task = taskIter.next();
            if (curPos >= startPosOfSnapOffset) {
              if ((curFilesAdded + 1) > maxFiles) {
                // On including the file it might happen that we might exceed, the configured
                // soft limit on the number of records, since this is a soft limit its acceptable.
                shouldContinueReading = false;
                break;
              }

              curFilesAdded += 1;
              curRecordCount += task.file().recordCount();

              if (curRecordCount >= maxRows) {
                // we included the file, so increment the number of files
                // read in the current snapshot.
                if (curFilesAdded == 1 && curRecordCount > maxRows) {
                  LOG.warn(
                      "File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
                          + "This file will be processed entirely to guarantee forward progress. "
                          + "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
                      task.file().location(),
                      task.file().recordCount(),
                      maxRows);
                }
                ++curPos;
                shouldContinueReading = false;
                break;
              }
            }
            ++curPos;
          }
        } catch (IOException ioe) {
          LOG.warn("Failed to close task iterable", ioe);
        }
      }

View on GitHub (pinned to 86d9c8fc54)