apache/seatunnel · warning

The inverseSamplingRate is

Error message

The inverseSamplingRate is {}, which is greater than chunkSize {}, so we set inverseSamplingRate to chunkSize

What it means

During JDBC chunk splitting with sampling-based uneven data distribution, the inverseSamplingRate (1-in-N row sampling factor) exceeds the computed chunkSize. Since sampling every N rows with N larger than the rows-per-chunk target would yield too few sample rows per shard, the splitter caps the sampling rate at chunkSize and emits this warning before proceeding.

Solutions

  1. Do nothing if acceptable: the splitter auto-clamps inverseSamplingRate to chunkSize; the job proceeds correctly.
  2. Lower the inverseSamplingRate config value (e.g. to 100 or 1000) so it is <= chunkSize and sampling remains meaningful.
  3. Increase the table's target chunk size config or add more rows so chunkSize naturally exceeds the sampling rate.
  4. Suppress concern by verifying in logs that only this warning appears and split counts are as expected.

Example fix

// before
url = "jdbc:mysql://host/db?table=id&shard.sampling-strategy=SAMPLE&inverse.sampling.rate=100000"
// after
url = "jdbc:mysql://host/db?table=id&shard.sampling-strategy=SAMPLE&inverse.sampling.rate=1000" // <= chunkSize
Defensive patterns

Strategy: validation

Validate before calling

long chunkSize = /* computed from config/row count */;
long inverseSamplingRate = /* from config */;
if (inverseSamplingRate > chunkSize) {
    inverseSamplingRate = chunkSize; // match splitter's clamp
}

Prevention

When it happens

Trigger: getChunkRangesWithUnevenlyData is invoked via charsetBasedColumnSplitChunks or evenlyColumnSplitChunks when sampling-based split is enabled (e.g. shard.sampling-strategy) and the configured/approximate inverseSamplingRate value is larger than the chunkKeyQuantity/chunkSize derived from the table row count and chunk size config.

Common situations: Users set a very large inverseSamplingRate (sparse sampling) on a small table, or table row count is small so computed chunkSize falls below the sampling rate; common after switching to uneven-split on tables with few rows.

Understand the failure class

Background: "Invalid value" and "allowed values are" config errors: what your library rejected and how to fix it — this error's family across 41 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/74485681570c732c. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-connectors-v2/connector-jdbc/src/main/java/org/apache/seatunnel/connectors/seatunnel/jdbc/source/DynamicChunkSplitter.java:487

            Object min,
            Object max,
            int chunkSize,
            TablePath tablePath,
            int sampleShardingThreshold,
            boolean sampleShardingAllow,
            long approximateRowCnt)
            throws Exception {
        int shardCount = (int) (approximateRowCnt / chunkSize);
        int inverseSamplingRate = config.getSplitInverseSamplingRate();
        if (sampleShardingAllow && sampleShardingThreshold < shardCount) {
            // It is necessary to ensure that the number of data rows sampled by the
            // sampling rate is greater than the number of shards.
            // Otherwise, if the sampling rate is too low, it may result in an insufficient
            // number of data rows for the shards, leading to an inadequate number of
            // shards.
            // Therefore, inverseSamplingRate should be less than chunkSize
            if (inverseSamplingRate > chunkSize) {
                log.warn(
                        "The inverseSamplingRate is {}, which is greater than chunkSize {}, so we set inverseSamplingRate to chunkSize",
                        inverseSamplingRate,
                        chunkSize);
                inverseSamplingRate = chunkSize;
            }
            log.info(
                    "Use sampling sharding for table {}, the sampling rate is {}",
                    tablePath,
                    inverseSamplingRate);
            Object[] sample =
                    jdbcDialect.sampleDataFromColumn(
                            getOrEstablishConnection(),
                            table,
                            splitColumnName,
                            inverseSamplingRate,
                            config.getFetchSize());
            log.info(
                    "Sample data from table {} end, the sample size is {}",

View on GitHub (pinned to cf67b549a7)