apache/iceberg · warning

File contains records, exceeding maxRecordsPerMicroBatch…

Error message

File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.

What it means

In SyncSparkMicroBatchPlanner.latestOffset (Iceberg Spark micro-batch streaming), the planner accumulates files until maxRecordsPerMicroBatch is reached. If a single file alone exceeds the limit, the warning fires: the whole oversized file must be included in one batch to guarantee forward progress (a file cannot be split), so the batch will exceed the configured row cap.

Solutions

  1. Increase spark.sql.iceberg.maxRecordsPerMicroBatch, or accept larger batches.
  2. Compact the table (RewriteDataFiles action) so individual files stay below the limit.
  3. Fix upstream write sizing (target-file-size-bytes / write fanout) to prevent huge files.
  4. If memory pressure results, raise executor memory or reduce task parallelism accordingly.

Example fix

// before
spark.conf.set("spark.sql.iceberg.maxRecordsPerMicroBatch", "10000")

// after: limit above typical file record counts
spark.conf.set("spark.sql.iceberg.maxRecordsPerMicroBatch", "500000")
Defensive patterns

Strategy: validation

Validate before calling

// Check for oversized files relative to the micro-batch limit before streaming
long maxRows = Long.parseLong(spark.conf().get("spark.sql.iceberg.maxRecordsPerMicroBatch", "10000000"));
table.newScan().planFiles().forEach(t -> {
  if (t.file().recordCount() >= maxRows) {
    LOG.warn("File {} has {} records >= limit {}", t.file().location(), t.file().recordCount(), maxRows);
  }
});

Prevention

When it happens

Trigger: Structured Streaming reads from an Iceberg table where a parquet/data file's recordCount is >= maxRecordsPerMicroBatch (spark.sql.iceberg.maxRecordsPerMicroBatch) when the planner adds it in latestOffset.

Common situations: Tables with very large, uncompacted data files (big compaction backlog) combined with a small maxRecordsPerMicroBatch to limit memory; upstream writers producing huge files.

Understand the failure class

Background: "File too large" / "file size exceeds limit" errors: why libraries cap file sizes and how to fix them — this error's family across 46 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/c5b1b58e97b8ccb8. Report an issue: GitHub.

Appendix: source

Thrown at spark/v3.5/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java:191

            CloseableIterator<FileScanTask> taskIter = taskIterable.iterator()) {
          while (taskIter.hasNext()) {
            FileScanTask task = taskIter.next();
            if (curPos >= startPosOfSnapOffset) {
              if ((curFilesAdded + 1) > maxFiles) {
                // On including the file it might happen that we might exceed, the configured
                // soft limit on the number of records, since this is a soft limit its acceptable.
                shouldContinueReading = false;
                break;
              }

              curFilesAdded += 1;
              curRecordCount += task.file().recordCount();

              if (curRecordCount >= maxRows) {
                // we included the file, so increment the number of files
                // read in the current snapshot.
                if (curFilesAdded == 1 && curRecordCount > maxRows) {
                  LOG.warn(
                      "File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
                          + "This file will be processed entirely to guarantee forward progress. "
                          + "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
                      task.file().location(),
                      task.file().recordCount(),
                      maxRows);
                }
                ++curPos;
                shouldContinueReading = false;
                break;
              }
            }
            ++curPos;
          }
        } catch (IOException ioe) {
          LOG.warn("Failed to close task iterable", ioe);
        }
      }

View on GitHub (pinned to 86d9c8fc54)