apache/iceberg · warning
File at offset contains records, exceeding…
Error message
File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage. What it means
AsyncSparkMicroBatchPlanner limits each micro-batch by maxRecordsPerMicroBatch. If a single file alone exceeds the row limit, including it wholly is the only way to guarantee forward progress; the planner logs this warning and processes the oversized file in one batch, potentially raising memory usage.
Solutions
- Increase maxRecordsPerMicroBatch so it comfortably exceeds the largest expected file's row count
- Configure upstream writers to write smaller files (e.g. smaller target file size or more frequent compaction)
- Compact/rewrite large files so future batches can respect the limit
- Monitor memory during such batches and size executors accordingly
Example fix
// before
.option("maxRecordsPerMicroBatch", "1000") // files contain >1000 rows each
// after
.option("maxRecordsPerMicroBatch", "500000") // above largest file size Defensive patterns
Strategy: validation
Validate before calling
long maxRows = Long.parseLong(limits.get("maxRecordsPerMicroBatch"));
long largestFileRows = files.stream().mapToLong(f -> f.recordCount()).max().orElse(0);
if (largestFileRows > maxRows) { /* raise maxRecordsPerMicroBatch above largestFileRows */ } Prevention
- Keep maxRecordsPerMicroBatch well above the largest data file's record count
- Configure writers with smaller target file sizes
- Compact oversized files proactively
- Alert on file record counts relative to streaming batch limits
When it happens
Trigger: computeLimitedOffset (called from latestOffset) encounters a first file (filesSeen == 1) whose row count exceeds unpackedLimits.getMaxRows() from the maxRecordsPerMicroBatch limit.
Common situations: Downstream writers producing very large files relative to a small maxRecordsPerMicroBatch setting; users lowering the limit for latency without considering file sizes.
Understand the failure class
Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.
Related errors
- File at offset contains records, exceeding…
- File at offset contains records, exceeding…
- File contains records, exceeding maxRecordsPerMicroBatch…
- File at offset contains records, exceeding…
- File contains records, exceeding maxRecordsPerMicroBatch…
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/5330cf297632d697.
Report an issue: GitHub.
Appendix: source
Thrown at spark/v4.1/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java:328
if (filesSeen == 0) {
return null;
}
LOG.debug(
"latestOffset hit file limit at {}, rows: {}, files: {}",
elem.first(),
rowsSeen,
filesSeen);
return elem.first();
}
// Soft limit on rows - include file FIRST, then check
rowsSeen += fileRows;
filesSeen += 1;
// Check if we've hit the row limit after including this file
if (rowsSeen >= unpackedLimits.getMaxRows()) {
if (filesSeen == 1 && rowsSeen > unpackedLimits.getMaxRows()) {
LOG.warn(
"File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
+ "This file will be processed entirely to guarantee forward progress. "
+ "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
elem.second().file().location(),
elem.first(),
fileRows,
unpackedLimits.getMaxRows());
}
// Return the offset of the NEXT element (or synthesize tail+1)
if (i + 1 < queueSnapshot.size()) {
LOG.debug(
"latestOffset hit row limit at {}, rows: {}, files: {}",
queueSnapshot.get(i + 1).first(),
rowsSeen,
filesSeen);
return queueSnapshot.get(i + 1).first();
} else {
// This is the last element - return tail+1View on GitHub (pinned to 86d9c8fc54)