{"record":{"id":"74485681570c732c","repo":"apache/seatunnel","slug":"the-inversesamplingrate-is-which-is-greater-th","errorCode":null,"errorMessage":"The inverseSamplingRate is {}, which is greater than chunkSize {}, so we set inverseSamplingRate to chunkSize","messagePattern":"The inverseSamplingRate is (.+?), which is greater than chunkSize (.+?), so we set inverseSamplingRate to chunkSize","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"seatunnel-connectors-v2/connector-jdbc/src/main/java/org/apache/seatunnel/connectors/seatunnel/jdbc/source/DynamicChunkSplitter.java","lineNumber":487,"sourceCode":"            Object min,\n            Object max,\n            int chunkSize,\n            TablePath tablePath,\n            int sampleShardingThreshold,\n            boolean sampleShardingAllow,\n            long approximateRowCnt)\n            throws Exception {\n        int shardCount = (int) (approximateRowCnt / chunkSize);\n        int inverseSamplingRate = config.getSplitInverseSamplingRate();\n        if (sampleShardingAllow && sampleShardingThreshold < shardCount) {\n            // It is necessary to ensure that the number of data rows sampled by the\n            // sampling rate is greater than the number of shards.\n            // Otherwise, if the sampling rate is too low, it may result in an insufficient\n            // number of data rows for the shards, leading to an inadequate number of\n            // shards.\n            // Therefore, inverseSamplingRate should be less than chunkSize\n            if (inverseSamplingRate > chunkSize) {\n                log.warn(\n                        \"The inverseSamplingRate is {}, which is greater than chunkSize {}, so we set inverseSamplingRate to chunkSize\",\n                        inverseSamplingRate,\n                        chunkSize);\n                inverseSamplingRate = chunkSize;\n            }\n            log.info(\n                    \"Use sampling sharding for table {}, the sampling rate is {}\",\n                    tablePath,\n                    inverseSamplingRate);\n            Object[] sample =\n                    jdbcDialect.sampleDataFromColumn(\n                            getOrEstablishConnection(),\n                            table,\n                            splitColumnName,\n                            inverseSamplingRate,\n                            config.getFetchSize());\n            log.info(\n                    \"Sample data from table {} end, the sample size is {}\",","sourceCodeStart":469,"sourceCodeEnd":505,"githubUrl":"https://github.com/apache/seatunnel/blob/cf67b549a7a6c35fa0beb12d83c62892427ea919/seatunnel-connectors-v2/connector-jdbc/src/main/java/org/apache/seatunnel/connectors/seatunnel/jdbc/source/DynamicChunkSplitter.java#L469-L505","documentation":"During JDBC chunk splitting with sampling-based uneven data distribution, the inverseSamplingRate (1-in-N row sampling factor) exceeds the computed chunkSize. Since sampling every N rows with N larger than the rows-per-chunk target would yield too few sample rows per shard, the splitter caps the sampling rate at chunkSize and emits this warning before proceeding.","triggerScenarios":"getChunkRangesWithUnevenlyData is invoked via charsetBasedColumnSplitChunks or evenlyColumnSplitChunks when sampling-based split is enabled (e.g. shard.sampling-strategy) and the configured/approximate inverseSamplingRate value is larger than the chunkKeyQuantity/chunkSize derived from the table row count and chunk size config.","commonSituations":"Users set a very large inverseSamplingRate (sparse sampling) on a small table, or table row count is small so computed chunkSize falls below the sampling rate; common after switching to uneven-split on tables with few rows.","solutions":["Do nothing if acceptable: the splitter auto-clamps inverseSamplingRate to chunkSize; the job proceeds correctly.","Lower the inverseSamplingRate config value (e.g. to 100 or 1000) so it is <= chunkSize and sampling remains meaningful.","Increase the table's target chunk size config or add more rows so chunkSize naturally exceeds the sampling rate.","Suppress concern by verifying in logs that only this warning appears and split counts are as expected."],"exampleFix":"// before\nurl = \"jdbc:mysql://host/db?table=id&shard.sampling-strategy=SAMPLE&inverse.sampling.rate=100000\"\n// after\nurl = \"jdbc:mysql://host/db?table=id&shard.sampling-strategy=SAMPLE&inverse.sampling.rate=1000\" // <= chunkSize","handlingStrategy":"validation","validationCode":"long chunkSize = /* computed from config/row count */;\nlong inverseSamplingRate = /* from config */;\nif (inverseSamplingRate > chunkSize) {\n    inverseSamplingRate = chunkSize; // match splitter's clamp\n}","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Keep inverseSamplingRate <= expected rows per chunk","Prefer small sampling rates (100-10000) for small tables","Review chunk size estimates against table row counts before enabling sampling"],"tags":["jdbc","chunk-splitting","sampling","configuration"],"backgroundTag":"invalid-config-value","analyzedSha":"cf67b549a7a6c35fa0beb12d83c62892427ea919","analyzedAt":"2026-09-10T21:44:55.265Z","contentChangedAt":"2026-09-10T21:44:55.265Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}