{"record":{"id":"c5b1b58e97b8ccb8","repo":"apache/iceberg","slug":"file-contains-records-exceeding-maxrecordsp","errorCode":null,"errorMessage":"File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.","messagePattern":"File (.+?) contains (.+?) records, exceeding maxRecordsPerMicroBatch limit of (.+?)\\. This file will be processed entirely to guarantee forward progress\\. Consider increasing the limit or writing smaller files to avoid unexpected memory usage\\.","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"spark/v3.5/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java","lineNumber":191,"sourceCode":"            CloseableIterator<FileScanTask> taskIter = taskIterable.iterator()) {\n          while (taskIter.hasNext()) {\n            FileScanTask task = taskIter.next();\n            if (curPos >= startPosOfSnapOffset) {\n              if ((curFilesAdded + 1) > maxFiles) {\n                // On including the file it might happen that we might exceed, the configured\n                // soft limit on the number of records, since this is a soft limit its acceptable.\n                shouldContinueReading = false;\n                break;\n              }\n\n              curFilesAdded += 1;\n              curRecordCount += task.file().recordCount();\n\n              if (curRecordCount >= maxRows) {\n                // we included the file, so increment the number of files\n                // read in the current snapshot.\n                if (curFilesAdded == 1 && curRecordCount > maxRows) {\n                  LOG.warn(\n                      \"File {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. \"\n                          + \"This file will be processed entirely to guarantee forward progress. \"\n                          + \"Consider increasing the limit or writing smaller files to avoid unexpected memory usage.\",\n                      task.file().location(),\n                      task.file().recordCount(),\n                      maxRows);\n                }\n                ++curPos;\n                shouldContinueReading = false;\n                break;\n              }\n            }\n            ++curPos;\n          }\n        } catch (IOException ioe) {\n          LOG.warn(\"Failed to close task iterable\", ioe);\n        }\n      }","sourceCodeStart":173,"sourceCodeEnd":209,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v3.5/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java#L173-L209","documentation":"In SyncSparkMicroBatchPlanner.latestOffset (Iceberg Spark micro-batch streaming), the planner accumulates files until maxRecordsPerMicroBatch is reached. If a single file alone exceeds the limit, the warning fires: the whole oversized file must be included in one batch to guarantee forward progress (a file cannot be split), so the batch will exceed the configured row cap.","triggerScenarios":"Structured Streaming reads from an Iceberg table where a parquet/data file's recordCount is >= maxRecordsPerMicroBatch (spark.sql.iceberg.maxRecordsPerMicroBatch) when the planner adds it in latestOffset.","commonSituations":"Tables with very large, uncompacted data files (big compaction backlog) combined with a small maxRecordsPerMicroBatch to limit memory; upstream writers producing huge files.","solutions":["Increase spark.sql.iceberg.maxRecordsPerMicroBatch, or accept larger batches.","Compact the table (RewriteDataFiles action) so individual files stay below the limit.","Fix upstream write sizing (target-file-size-bytes / write fanout) to prevent huge files.","If memory pressure results, raise executor memory or reduce task parallelism accordingly."],"exampleFix":"// before\nspark.conf.set(\"spark.sql.iceberg.maxRecordsPerMicroBatch\", \"10000\")\n\n// after: limit above typical file record counts\nspark.conf.set(\"spark.sql.iceberg.maxRecordsPerMicroBatch\", \"500000\")","handlingStrategy":"validation","validationCode":"// Check for oversized files relative to the micro-batch limit before streaming\nlong maxRows = Long.parseLong(spark.conf().get(\"spark.sql.iceberg.maxRecordsPerMicroBatch\", \"10000000\"));\ntable.newScan().planFiles().forEach(t -> {\n  if (t.file().recordCount() >= maxRows) {\n    LOG.warn(\"File {} has {} records >= limit {}\", t.file().location(), t.file().recordCount(), maxRows);\n  }\n});","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Run RewriteDataFiles regularly to keep files small","Set write.target-file-size-bytes appropriately at write time","Set maxRecordsPerMicroBatch above typical file record counts","Size executor memory for the largest expected batch"],"tags":["spark","streaming","memory","file-size"],"backgroundTag":"file-size-limit-exceeded","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}