{"record":{"id":"3ad29ef41ca8848a","repo":"apache/iceberg","slug":"file-at-offset-contains-records-exceedin-3ad29e","errorCode":null,"errorMessage":"File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. This file will be processed entirely to guarantee forward progress. Consider increasing the limit or writing smaller files to avoid unexpected memory usage.","messagePattern":"File (.+?) at offset (.+?) contains (.+?) records, exceeding maxRecordsPerMicroBatch limit of (.+?)\\. This file will be processed entirely to guarantee forward progress\\. Consider increasing the limit or writing smaller files to avoid unexpected memory usage\\.","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java","lineNumber":328,"sourceCode":"        if (filesSeen == 0) {\n          return null;\n        }\n        LOG.debug(\n            \"latestOffset hit file limit at {}, rows: {}, files: {}\",\n            elem.first(),\n            rowsSeen,\n            filesSeen);\n        return elem.first();\n      }\n\n      // Soft limit on rows - include file FIRST, then check\n      rowsSeen += fileRows;\n      filesSeen += 1;\n\n      // Check if we've hit the row limit after including this file\n      if (rowsSeen >= unpackedLimits.getMaxRows()) {\n        if (filesSeen == 1 && rowsSeen > unpackedLimits.getMaxRows()) {\n          LOG.warn(\n              \"File {} at offset {} contains {} records, exceeding maxRecordsPerMicroBatch limit of {}. \"\n                  + \"This file will be processed entirely to guarantee forward progress. \"\n                  + \"Consider increasing the limit or writing smaller files to avoid unexpected memory usage.\",\n              elem.second().file().location(),\n              elem.first(),\n              fileRows,\n              unpackedLimits.getMaxRows());\n        }\n        // Return the offset of the NEXT element (or synthesize tail+1)\n        if (i + 1 < queueSnapshot.size()) {\n          LOG.debug(\n              \"latestOffset hit row limit at {}, rows: {}, files: {}\",\n              queueSnapshot.get(i + 1).first(),\n              rowsSeen,\n              filesSeen);\n          return queueSnapshot.get(i + 1).first();\n        } else {\n          // This is the last element - return tail+1","sourceCodeStart":310,"sourceCodeEnd":346,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java#L310-L346","documentation":"A LOG.warn in AsyncSparkMicroBatchPlanner.computeLimitedOffset: a single source file contains more records than maxRecordsPerMicroBatch. Because a file cannot be split mid-batch here, the entire file is processed anyway to guarantee forward progress, potentially exceeding the intended batch size and memory budget.","triggerScenarios":"Streaming query using the async/micro-batch planner where one Parquet data file's record count pushes rowsSeen past unpackedLimits.getMaxRows() while it is the first (filesSeen == 1) file considered in the offset computation.","commonSituations":"Upstream batch jobs writing very large files into a table consumed by a streaming read with a low maxRecordsPerMicroBatch; misconfigured size-based limits assuming many small files; memory pressure or long micro-batches observed on the consumer.","solutions":["Increase maxRecordsPerMicroBatch so typical files fit within the limit.","Make the producer write smaller files (smaller target-file-size / more frequent commits) so no single file exceeds the batch limit.","Accept the behavior if forward progress matters more than strict batch sizing; just provision executor memory for the largest file.","Combine with maxBytesPerMicroBatch or file-size-based limits if available to bound memory more predictably."],"exampleFix":"// before\noption(\"maxRecordsPerMicroBatch\", 10000) // files contain 500k rows\n// after\noption(\"maxRecordsPerMicroBatch\", 1000000) // >= largest expected file's row count","handlingStrategy":"fallback","validationCode":"// Compare configured row limit against the largest expected file size\nlong approxRowsPerFile = targetFileSizeBytes / avgRowSizeBytes;\nif (approxRowsPerFile > maxRecordsPerMicroBatch) {\n  LOG.warn(\"maxRecordsPerMicroBatch below single-file row count; oversized files will bypass the limit\");\n}","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Set maxRecordsPerMicroBatch above the largest expected file's row count","Control producer file sizes (target-file-size) when consuming via streaming","Provision executor memory to absorb at least one oversized file per batch","Monitor batch durations and spill metrics for signs of oversized batches"],"tags":["spark","streaming","micro-batch","memory"],"backgroundTag":"batch-size-limit-exceeded","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}