apache/iceberg · warning
File contains records, exceeding maxRecordsPerMicroBatch…
Error message
File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage. What it means
A WARN logged by SyncSparkMicroBatchPlanner.latestOffset when a single data file's recordCount exceeds the maxRecordsPerMicroBatch limit (spark.message streaming-max-records-per-micro-batch). Since a micro-batch cannot split a file, the whole file is included anyway to guarantee forward progress; memory usage per batch may exceed the configured limit.
Solutions
- Increase streaming-max-rows-per-micro-batch to exceed your largest file's record count
- Write smaller files upstream (target file size, more frequent compaction)
- Run rewriteDataFiles to shrink existing oversized files
- Raise the limit only as much as executor memory allows
Example fix
// before
.option("streaming-max-rows-per-micro-batch", 100000)
// after
.option("streaming-max-rows-per-micro-batch", 5000000) // > largest file's record count Defensive patterns
Strategy: validation
Validate before calling
// check table file sizes against the configured limit before streaming
long maxRows = 5_000_000L; // streaming-max-rows-per-micro-batch
TableScan scan = table.newScan();
for (File f : scan.planFiles()) {
if (f.recordCount() > maxRows) {
throw new IllegalArgumentException("File " + f.path() + " has " + f.recordCount() + " > limit " + maxRows);
}
} Prevention
- Set streaming-max-rows-per-micro-batch above your largest file's record count
- Keep files near target-file-size with regular compaction (rewriteDataFiles)
- Monitor max file record counts in production tables
When it happens
Trigger: Streaming micro-batch planning encounters an added file whose task.file().recordCount() pushes curRecordCount past maxRows while curFilesAdded == 1 (first oversized file in the batch).
Common situations: Compaction disabled or small upstream writers producing very large files; STREAMING_MAX_ROWS_PER_MICRO_BATCH set too low relative to table file sizes; backfill/initial batch reading large files.
Related errors
- File at offset contains records, exceeding…
- File at offset contains records, exceeding…
- File at offset contains records, exceeding…
- File at offset contains records, exceeding…
- File contains records, exceeding maxRecordsPerMicroBatch…
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/0ec5b752f7859a81.
Report an issue: GitHub.
Appendix: source
Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java:191
CloseableIterator<FileScanTask> taskIter = taskIterable.iterator()) {
while (taskIter.hasNext()) {
FileScanTask task = taskIter.next();
if (curPos >= startPosOfSnapOffset) {
if ((curFilesAdded + 1) > maxFiles) {
// On including the file it might happen that we might exceed, the configured
// soft limit on the number of records, since this is a soft limit its acceptable.
shouldContinueReading = false;
break;
}
curFilesAdded += 1;
curRecordCount += task.file().recordCount();
if (curRecordCount >= maxRows) {
// we included the file, so increment the number of files
// read in the current snapshot.
if (curFilesAdded == 1 && curRecordCount > maxRows) {
LOG.warn(
"File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
+ "This file will be processed entirely to guarantee forward progress. "
+ "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
task.file().location(),
task.file().recordCount(),
maxRows);
}
++curPos;
shouldContinueReading = false;
break;
}
}
++curPos;
}
} catch (IOException ioe) {
LOG.warn("Failed to close task iterable", ioe);
}
}View on GitHub (pinned to 86d9c8fc54)