apache/iceberg · warning
File contains records, exceeding maxRecordsPerMicroBatch…
Error message
File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage. What it means
In SyncSparkMicroBatchPlanner.latestOffset (Iceberg Spark micro-batch streaming), the planner accumulates files until maxRecordsPerMicroBatch is reached. If a single file alone exceeds the limit, the warning fires: the whole oversized file must be included in one batch to guarantee forward progress (a file cannot be split), so the batch will exceed the configured row cap.
Solutions
- Increase spark.sql.iceberg.maxRecordsPerMicroBatch, or accept larger batches.
- Compact the table (RewriteDataFiles action) so individual files stay below the limit.
- Fix upstream write sizing (target-file-size-bytes / write fanout) to prevent huge files.
- If memory pressure results, raise executor memory or reduce task parallelism accordingly.
Example fix
// before
spark.conf.set("spark.sql.iceberg.maxRecordsPerMicroBatch", "10000")
// after: limit above typical file record counts
spark.conf.set("spark.sql.iceberg.maxRecordsPerMicroBatch", "500000") Defensive patterns
Strategy: validation
Validate before calling
// Check for oversized files relative to the micro-batch limit before streaming
long maxRows = Long.parseLong(spark.conf().get("spark.sql.iceberg.maxRecordsPerMicroBatch", "10000000"));
table.newScan().planFiles().forEach(t -> {
if (t.file().recordCount() >= maxRows) {
LOG.warn("File {} has {} records >= limit {}", t.file().location(), t.file().recordCount(), maxRows);
}
}); Prevention
- Run RewriteDataFiles regularly to keep files small
- Set write.target-file-size-bytes appropriately at write time
- Set maxRecordsPerMicroBatch above typical file record counts
- Size executor memory for the largest expected batch
When it happens
Trigger: Structured Streaming reads from an Iceberg table where a parquet/data file's recordCount is >= maxRecordsPerMicroBatch (spark.sql.iceberg.maxRecordsPerMicroBatch) when the planner adds it in latestOffset.
Common situations: Tables with very large, uncompacted data files (big compaction backlog) combined with a small maxRecordsPerMicroBatch to limit memory; upstream writers producing huge files.
Understand the failure class
Background: "File too large" / "file size exceeds limit" errors: why libraries cap file sizes and how to fix them — this error's family across 46 libraries.
Related errors
- File at offset contains records, exceeding…
- File at offset contains records, exceeding…
- File at offset contains records, exceeding…
- File at offset contains records, exceeding…
- File contains records, exceeding maxRecordsPerMicroBatch…
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/c5b1b58e97b8ccb8.
Report an issue: GitHub.
Appendix: source
Thrown at spark/v3.5/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java:191
CloseableIterator<FileScanTask> taskIter = taskIterable.iterator()) {
while (taskIter.hasNext()) {
FileScanTask task = taskIter.next();
if (curPos >= startPosOfSnapOffset) {
if ((curFilesAdded + 1) > maxFiles) {
// On including the file it might happen that we might exceed, the configured
// soft limit on the number of records, since this is a soft limit its acceptable.
shouldContinueReading = false;
break;
}
curFilesAdded += 1;
curRecordCount += task.file().recordCount();
if (curRecordCount >= maxRows) {
// we included the file, so increment the number of files
// read in the current snapshot.
if (curFilesAdded == 1 && curRecordCount > maxRows) {
LOG.warn(
"File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
+ "This file will be processed entirely to guarantee forward progress. "
+ "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
task.file().location(),
task.file().recordCount(),
maxRows);
}
++curPos;
shouldContinueReading = false;
break;
}
}
++curPos;
}
} catch (IOException ioe) {
LOG.warn("Failed to close task iterable", ioe);
}
}View on GitHub (pinned to 86d9c8fc54)