apache/iceberg · error · UnsupportedOperationException
Expected column to be date, tinyint, smallint, int, bigint…
Error message
Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary
What it means
After validating the bucket width, BucketFunction.bind() inspects the value column's data type and dispatches to a typed BucketFunction implementation. Only date, tinyint, smallint, int, bigint, decimal, timestamp, string, and binary are supported; any other type (float, double, boolean, array, struct, map, timestamp_ntz in older versions) is rejected with this UnsupportedOperationException.
Solutions
- Cast the value to a supported type before bucketing: bucket(16, cast(price AS DECIMAL(38, 8))) or bucket(16, cast(col AS STRING)).
- Use a different transform for the column type, or bucket on an integer surrogate key instead of a float.
- Check the column type with DESC TABLE or typeof(col) to confirm which type needs conversion.
Example fix
// before SELECT iceberg.bucket(16, price) FROM t; -- price DOUBLE // after SELECT iceberg.bucket(16, CAST(price AS DECIMAL(38, 8))) FROM t;
Defensive patterns
Strategy: type-guard
Validate before calling
val t = col.dataType val supported = Seq(DateType, TimestampType, StringType, BinaryType) ++ Seq(ByteType, ShortType, IntegerType, LongType) :+ DecimalType.SYSTEM_DEFAULT require(supported.exists(_.acceptsType(t)), s"bucket value type $t unsupported")
Type guard
def isBucketable(t: DataType): Boolean = t match {
case _: DecimalType => true
case DateType | TimestampType | StringType | BinaryType |
ByteType | ShortType | IntegerType | LongType => true
case _ => false
} Try / catch
try {
df.select(functions.callUDF("iceberg.bucket", lit(16), col("v")))
} catch {
case e: UnsupportedOperationException if e.getMessage.contains("Expected column") =>
df.select(functions.callUDF("iceberg.bucket", lit(16), col("v").cast(StringType)))
} Prevention
- Never bucket float/double columns directly — cast to DECIMAL or bucket an integer key instead
- Check column types with DESC TABLE before writing bucket transforms in SQL
When it happens
Trigger: Calling iceberg.bucket(width, col) where col's type is DOUBLE, FLOAT, BOOLEAN, ARRAY, STRUCT, or MAP — e.g. bucket(16, price) where price is DOUBLE.
Common situations: Bucketing on a floating-point column such as a price or score; mistakenly passing a struct/array column from a nested schema; SQL literal of unsupported type (e.g. a DOUBLE literal like 1.5 is DECIMAL but 1.0e2 is DOUBLE).
Understand the failure class
Background: Type mismatch errors: IllegalArgumentException, TypeError and type guards across 150 open-source libraries — this error's family across 150 libraries.
Related errors
- Expected truncation col to be tinyint, shortint, int…
- Expected value to be date or timestamp
- Expected value to be date or timestamp
- Expected value to be timestamp
- Cannot bind: does not accept arguments
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/8eb0809ce1fe056f.
Report an issue: GitHub.
Appendix: source
Thrown at spark/v4.0/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java:103
return new BucketInt(type);
} else if (type instanceof ByteType
|| type instanceof ShortType
|| type instanceof IntegerType) {
return new BucketInt(DataTypes.IntegerType);
} else if (type instanceof LongType) {
return new BucketLong(type);
} else if (type instanceof TimestampType) {
return new BucketLong(type);
} else if (type instanceof TimestampNTZType) {
return new BucketLong(type);
} else if (type instanceof DecimalType) {
return new BucketDecimal(type);
} else if (type instanceof StringType) {
return new BucketString();
} else if (type instanceof BinaryType) {
return new BucketBinary();
} else {
throw new UnsupportedOperationException(
"Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary");
}
}
@Override
public String description() {
return name()
+ "(numBuckets, col) - Call Iceberg's bucket transform\n"
+ " numBuckets :: number of buckets to divide the rows into, e.g. bucket(100, 34) -> 79 (must be a tinyint, smallint, or int)\n"
+ " col :: column to bucket (must be a date, integer, long, timestamp, decimal, string, or binary)";
}
@Override
public String name() {
return "bucket";
}
public abstract static class BucketBase extends BaseScalarFunction<Integer>View on GitHub (pinned to 86d9c8fc54)