{"record":{"id":"0ec5b752f7859a81","repo":"apache/iceberg","slug":"file-contains-records-exceeding-maxrecordsp-0ec5b7","errorCode":null,"errorMessage":"File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.","messagePattern":"File (.+?) contains (.+?) records, exceeding maxRecordsPerMicroBatch limit of (.+?)\\. This file will be processed entirely to guarantee forward progress\\. Consider increasing the limit or writing smaller files to avoid unexpected memory usage\\.","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java","lineNumber":191,"sourceCode":"            CloseableIterator<FileScanTask> taskIter = taskIterable.iterator()) {\n          while (taskIter.hasNext()) {\n            FileScanTask task = taskIter.next();\n            if (curPos >= startPosOfSnapOffset) {\n              if ((curFilesAdded + 1) > maxFiles) {\n                // On including the file it might happen that we might exceed, the configured\n                // soft limit on the number of records, since this is a soft limit its acceptable.\n                shouldContinueReading = false;\n                break;\n              }\n\n              curFilesAdded += 1;\n              curRecordCount += task.file().recordCount();\n\n              if (curRecordCount >= maxRows) {\n                // we included the file, so increment the number of files\n                // read in the current snapshot.\n                if (curFilesAdded == 1 && curRecordCount > maxRows) {\n                  LOG.warn(\n                      \"File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. \"\n                          + \"This file will be processed entirely to guarantee forward progress. \"\n                          + \"Consider increasing the limit or writing smaller files to avoid unexpected memory usage.\",\n                      task.file().location(),\n                      task.file().recordCount(),\n                      maxRows);\n                }\n                ++curPos;\n                shouldContinueReading = false;\n                break;\n              }\n            }\n            ++curPos;\n          }\n        } catch (IOException ioe) {\n          LOG.warn(\"Failed to close task iterable\", ioe);\n        }\n      }","sourceCodeStart":173,"sourceCodeEnd":209,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java#L173-L209","documentation":"A WARN logged by SyncSparkMicroBatchPlanner.latestOffset when a single data file's recordCount exceeds the maxRecordsPerMicroBatch limit (spark.message streaming-max-records-per-micro-batch). Since a micro-batch cannot split a file, the whole file is included anyway to guarantee forward progress; memory usage per batch may exceed the configured limit.","triggerScenarios":"Streaming micro-batch planning encounters an added file whose task.file().recordCount() pushes curRecordCount past maxRows while curFilesAdded == 1 (first oversized file in the batch).","commonSituations":"Compaction disabled or small upstream writers producing very large files; STREAMING_MAX_ROWS_PER_MICRO_BATCH set too low relative to table file sizes; backfill/initial batch reading large files.","solutions":["Increase streaming-max-rows-per-micro-batch to exceed your largest file's record count","Write smaller files upstream (target file size, more frequent compaction)","Run rewriteDataFiles to shrink existing oversized files","Raise the limit only as much as executor memory allows"],"exampleFix":"// before\n.option(\"streaming-max-rows-per-micro-batch\", 100000)\n// after\n.option(\"streaming-max-rows-per-micro-batch\", 5000000) // > largest file's record count","handlingStrategy":"validation","validationCode":"// check table file sizes against the configured limit before streaming\nlong maxRows = 5_000_000L; // streaming-max-rows-per-micro-batch\nTableScan scan = table.newScan();\nfor (File f : scan.planFiles()) {\n  if (f.recordCount() > maxRows) {\n    throw new IllegalArgumentException(\"File \" + f.path() + \" has \" + f.recordCount() + \" > limit \" + maxRows);\n  }\n}","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Set streaming-max-rows-per-micro-batch above your largest file's record count","Keep files near target-file-size with regular compaction (rewriteDataFiles)","Monitor max file record counts in production tables"],"tags":["spark","streaming","micro-batch","memory"],"backgroundTag":"config-value-exceeds-limit","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}