apache/iceberg · warning

File at offset contains records, exceeding…

Error message

File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.

What it means

During micro-batch offset planning, a single data file already exceeds the maxRecordsPerMicroBatch row limit. The planner cannot split a file, so it processes it whole to guarantee forward progress and logs this warning.

Solutions

  1. Increase maxRecordsPerMicroBatch above the largest existing data file's record count
  2. Rewrite/compact input files into smaller files (rewrite_data_files with target-file-size-bytes)
  3. Lower the writer's target file size so new files fit within the batch limit

Example fix

// before
option("read.limit", "1000")
// after
option("read.limit", "500000") // or compact large files with rewrite_data_files
Defensive patterns

Strategy: validation

Validate before calling

// pre-check file sizes before setting the limit
long maxFileRecords = files.stream().mapToLong(DataFile::recordCount).max().orElse(0L);
if (maxFileRecords > maxRecordsPerMicroBatch) { /* raise limit or compact first */ }

Prevention

When it happens

Trigger: Streaming read with read.limit/max-records-per-microbatch configured smaller than the row count of at least one existing data file; latestOffset -> computeLimitedOffset hits a single file with rowsSeen > maxRows.

Common situations: Compaction-less writers producing large files, or users lowering maxRecordsPerMicroBatch for latency without checking file sizes.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/532fb091f12aed00. Report an issue: GitHub.

Appendix: source

Thrown at spark/v3.5/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java:328

        if (filesSeen == 0) {
          return null;
        }
        LOG.debug(
            "latestOffset hit file limit at {}, rows: {}, files: {}",
            elem.first(),
            rowsSeen,
            filesSeen);
        return elem.first();
      }

      // Soft limit on rows - include file FIRST, then check
      rowsSeen += fileRows;
      filesSeen += 1;

      // Check if we've hit the row limit after including this file
      if (rowsSeen >= unpackedLimits.getMaxRows()) {
        if (filesSeen == 1 && rowsSeen > unpackedLimits.getMaxRows()) {
          LOG.warn(
              "File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. "
                  + "This file will be processed entirely to guarantee forward progress. "
                  + "Consider increasing the limit or writing smaller files to avoid unexpected memory usage.",
              elem.second().file().location(),
              elem.first(),
              fileRows,
              unpackedLimits.getMaxRows());
        }
        // Return the offset of the NEXT element (or synthesize tail+1)
        if (i + 1 < queueSnapshot.size()) {
          LOG.debug(
              "latestOffset hit row limit at {}, rows: {}, files: {}",
              queueSnapshot.get(i + 1).first(),
              rowsSeen,
              filesSeen);
          return queueSnapshot.get(i + 1).first();
        } else {
          // This is the last element - return tail+1

View on GitHub (pinned to 86d9c8fc54)