apache/seatunnel · error · IllegalArgumentException

Binary chunk size too large (max 100MB), got

Error message

Binary chunk size too large (max 100MB), got: %s

What it means

BinaryReadStrategy.init caps binary_chunk_size at 100MB (100 * 1024 * 1024 bytes) to prevent oversized in-memory buffers. A larger configured value throws IllegalArgumentException with the configured number. This protects worker heap memory from a single huge allocation.

Solutions

  1. Reduce binary_chunk_size to at most 104857600 bytes.
  2. Remove the option to use the default chunk size.
  3. If throughput is the goal, increase source parallelism instead of chunk size.

Example fix

// before
binary_chunk_size = 268435456  // 256MB
// after
binary_chunk_size = 104857600  // max 100MB
Defensive patterns

Strategy: validation

Validate before calling

// Java
int chunk = config.getInt("binary_chunk_size", DEFAULT_CHUNK);
int MAX = 100 * 1024 * 1024;
if (chunk > MAX) {
    throw new IllegalArgumentException("binary_chunk_size max is 100MB, got: " + chunk);
}

Prevention

When it happens

Trigger: Setting binary_chunk_size greater than 104857600, e.g. binary_chunk_size = 268435456 (256MB), when initializing the binary read strategy.

Common situations: Users reading very large binary files assuming a bigger chunk is faster; unit confusion (setting a GB value thinking it is MB); copied configs from other tools without the cap.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/41fefac9b959e229. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/reader/BinaryReadStrategy.java:82

            basePathIsFile = hadoopFileSystemProxy.isFile(basePath);
        } catch (IOException e) {
            throw new FileConnectorException(
                    SeaTunnelAPIErrorCode.CONFIG_VALIDATION_FAILED,
                    "Failed to determine whether file source path is a file or directory: "
                            + basePath,
                    e);
        }

        // Load binary chunk size configuration
        if (pluginConfig.hasPath(FileBaseSourceOptions.BINARY_CHUNK_SIZE.key())) {
            binaryChunkSize = pluginConfig.getInt(FileBaseSourceOptions.BINARY_CHUNK_SIZE.key());
            // Validate chunk size - should be positive and reasonable
            if (binaryChunkSize <= 0) {
                throw new IllegalArgumentException(
                        "Binary chunk size must be positive, got: " + binaryChunkSize);
            }
            if (binaryChunkSize > 100 * 1024 * 1024) { // 100MB limit
                throw new IllegalArgumentException(
                        "Binary chunk size too large (max 100MB), got: " + binaryChunkSize);
            }
        }

        // Load complete file mode configuration
        if (pluginConfig.hasPath(FileBaseSourceOptions.BINARY_COMPLETE_FILE_MODE.key())) {
            completeFileMode =
                    pluginConfig.getBoolean(FileBaseSourceOptions.BINARY_COMPLETE_FILE_MODE.key());
        }
    }

    @Override
    public void read(String path, String tableId, Collector<SeaTunnelRow> output)
            throws IOException, FileConnectorException {
        MessageDigest digest = createSha256Digest();
        lastReadFingerprint = null;
        try (InputStream inputStream =
                new DigestTrackingInputStream(hadoopFileSystemProxy.getInputStream(path), digest)) {

View on GitHub (pinned to cf67b549a7)