apache/seatunnel · warning
The inverseSamplingRate is
Error message
The inverseSamplingRate is {}, which is greater than chunkSize {}, so we set inverseSamplingRate to chunkSize What it means
During JDBC chunk splitting with sampling-based uneven data distribution, the inverseSamplingRate (1-in-N row sampling factor) exceeds the computed chunkSize. Since sampling every N rows with N larger than the rows-per-chunk target would yield too few sample rows per shard, the splitter caps the sampling rate at chunkSize and emits this warning before proceeding.
Solutions
- Do nothing if acceptable: the splitter auto-clamps inverseSamplingRate to chunkSize; the job proceeds correctly.
- Lower the inverseSamplingRate config value (e.g. to 100 or 1000) so it is <= chunkSize and sampling remains meaningful.
- Increase the table's target chunk size config or add more rows so chunkSize naturally exceeds the sampling rate.
- Suppress concern by verifying in logs that only this warning appears and split counts are as expected.
Example fix
// before url = "jdbc:mysql://host/db?table=id&shard.sampling-strategy=SAMPLE&inverse.sampling.rate=100000" // after url = "jdbc:mysql://host/db?table=id&shard.sampling-strategy=SAMPLE&inverse.sampling.rate=1000" // <= chunkSize
Defensive patterns
Strategy: validation
Validate before calling
long chunkSize = /* computed from config/row count */;
long inverseSamplingRate = /* from config */;
if (inverseSamplingRate > chunkSize) {
inverseSamplingRate = chunkSize; // match splitter's clamp
} Prevention
- Keep inverseSamplingRate <= expected rows per chunk
- Prefer small sampling rates (100-10000) for small tables
- Review chunk size estimates against table row counts before enabling sampling
When it happens
Trigger: getChunkRangesWithUnevenlyData is invoked via charsetBasedColumnSplitChunks or evenlyColumnSplitChunks when sampling-based split is enabled (e.g. shard.sampling-strategy) and the configured/approximate inverseSamplingRate value is larger than the chunkKeyQuantity/chunkSize derived from the table row count and chunk size config.
Common situations: Users set a very large inverseSamplingRate (sparse sampling) on a small table, or table row count is small so computed chunkSize falls below the sampling rate; common after switching to uneven-split on tables with few rows.
Understand the failure class
Background: "Invalid value" and "allowed values are" config errors: what your library rejected and how to fix it — this error's family across 41 libraries.
Related errors
- Either table path or query must be specified in source…
- Exactly once is enabled, but not found primary key or…
- Failed to build the split data read statement.
- Failed to split chunks for table " + tableId
- Generate Splits for table
AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10).
Data as JSON: /api/errors/74485681570c732c.
Report an issue: GitHub.
Appendix: source
Thrown at seatunnel-connectors-v2/connector-jdbc/src/main/java/org/apache/seatunnel/connectors/seatunnel/jdbc/source/DynamicChunkSplitter.java:487
Object min,
Object max,
int chunkSize,
TablePath tablePath,
int sampleShardingThreshold,
boolean sampleShardingAllow,
long approximateRowCnt)
throws Exception {
int shardCount = (int) (approximateRowCnt / chunkSize);
int inverseSamplingRate = config.getSplitInverseSamplingRate();
if (sampleShardingAllow && sampleShardingThreshold < shardCount) {
// It is necessary to ensure that the number of data rows sampled by the
// sampling rate is greater than the number of shards.
// Otherwise, if the sampling rate is too low, it may result in an insufficient
// number of data rows for the shards, leading to an inadequate number of
// shards.
// Therefore, inverseSamplingRate should be less than chunkSize
if (inverseSamplingRate > chunkSize) {
log.warn(
"The inverseSamplingRate is {}, which is greater than chunkSize {}, so we set inverseSamplingRate to chunkSize",
inverseSamplingRate,
chunkSize);
inverseSamplingRate = chunkSize;
}
log.info(
"Use sampling sharding for table {}, the sampling rate is {}",
tablePath,
inverseSamplingRate);
Object[] sample =
jdbcDialect.sampleDataFromColumn(
getOrEstablishConnection(),
table,
splitColumnName,
inverseSamplingRate,
config.getFetchSize());
log.info(
"Sample data from table {} end, the sample size is {}",View on GitHub (pinned to cf67b549a7)