apache/iceberg · error · java.lang.UnsupportedOperationException
Expected number of buckets to be tinyint, shortint or int
Error message
Expected number of buckets to be tinyint, shortint or int
What it means
Iceberg's bucket(x, n) Spark function validates, at binding time, that the numBuckets argument is one of Spark's ByteType, ShortType, or IntegerType. Any other literal or column type for the bucket count (e.g. long, string, decimal) cannot be used to compute a bucket, so bind() throws this UnsupportedOperationException. The error is thrown during query analysis, before execution.
Solutions
- Cast the bucket count to int: bucket(col, CAST(128 AS INT)).
- Pass a plain integer literal without a suffix, e.g. bucket(id, 128) instead of 128L.
- Check the numBuckets argument position; it must be the second argument to bucket().
- If the count comes from a column or expression, ensure its declared type is byte, short, or int.
Example fix
// before SELECT bucket(id, 128L) FROM t // after SELECT bucket(id, CAST(128 AS INT)) FROM t
Defensive patterns
Strategy: validation
Validate before calling
// Spark Scala, before calling bucket()
val n = 128
require(n.isInstanceOf[Integer] || n.isInstanceOf[java.lang.Byte] || n.isInstanceOf[java.lang.Short],
s"bucket count must be int/short/tinyint, got: ${if (n == null) "null" else n.getClass.getSimpleName}")
// SQL: ensure literal is int: bucket(id, CAST(? AS INT)) Type guard
def isSupportedNumBucketsType(dt: org.apache.spark.sql.types.DataType): Boolean = dt == org.apache.spark.sql.types.ByteType || dt == org.apache.spark.sql.types.ShortType || dt == org.apache.spark.sql.types.IntegerType
Try / catch
try {
df.select(functions.call("bucket", col("id"), lit(n)))
} catch {
case e: UnsupportedOperationException if e.getMessage.contains("number of buckets") =>
throw new IllegalArgumentException("bucket() requires an int-typed bucket count", e)
} Prevention
- Always pass plain integer literals (128, not 128L) as the bucket count
- Cast user/config-provided bucket counts to INT before use in SQL
- Check argument order: bucket(value, numBuckets)
- Validate partition spec transforms when building CREATE TABLE statements programmatically
When it happens
Trigger: Calling bucket(value, n) where n is a LongType, StringType, DecimalType, or any type outside ByteType/ShortType/IntegerType — e.g. bucket(id, 128L), a string literal, or a column of type long passed as the bucket count.
Common situations: Writing bucket transforms in CREATE TABLE partition specs where the bucket count literal is typed as long; Spark SQL widening literals to LongType; passing a computed expression (e.g. CAST or a column) as the bucket count with the wrong type.
Understand the failure class
Background: Type mismatch errors: IllegalArgumentException, TypeError and type guards across 150 open-source libraries — this error's family across 150 libraries.
Related errors
- Expected truncation width to be tinyint, shortint or int
- Expected column to be date, tinyint, smallint, int, bigint…
- Expected column to be date, tinyint, smallint, int, bigint…
- Expected number of buckets to be tinyint, shortint or int
- Expected truncation col to be tinyint, shortint, int…
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/4884373d585225e7.
Report an issue: GitHub.
Appendix: source
Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java:79
private static final int NUM_BUCKETS_ORDINAL = 0;
private static final int VALUE_ORDINAL = 1;
private static final Set<DataType> SUPPORTED_NUM_BUCKETS_TYPES =
ImmutableSet.of(DataTypes.ByteType, DataTypes.ShortType, DataTypes.IntegerType);
@Override
@SuppressWarnings("checkstyle:CyclomaticComplexity")
public BoundFunction bind(StructType inputType) {
if (inputType.size() != 2) {
throw new UnsupportedOperationException(
"Wrong number of inputs (expected numBuckets and value)");
}
StructField numBucketsField = inputType.fields()[NUM_BUCKETS_ORDINAL];
StructField valueField = inputType.fields()[VALUE_ORDINAL];
if (!SUPPORTED_NUM_BUCKETS_TYPES.contains(numBucketsField.dataType())) {
throw new UnsupportedOperationException(
"Expected number of buckets to be tinyint, shortint or int");
}
DataType type = valueField.dataType();
if (type instanceof DateType) {
return new BucketInt(type);
} else if (type instanceof ByteType
|| type instanceof ShortType
|| type instanceof IntegerType) {
return new BucketInt(DataTypes.IntegerType);
} else if (type instanceof LongType) {
return new BucketLong(type);
} else if (type instanceof TimestampType) {
return new BucketLong(type);
} else if (type instanceof TimestampNTZType) {
return new BucketLong(type);
} else if (type instanceof DecimalType) {
return new BucketDecimal(type);View on GitHub (pinned to 86d9c8fc54)