apache/iceberg · error · UnsupportedOperationException

Expected column to be date, tinyint, smallint, int, bigint…

Error message

Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary

What it means

After validating the bucket width, BucketFunction.bind() inspects the value column's data type and dispatches to a typed BucketFunction implementation. Only date, tinyint, smallint, int, bigint, decimal, timestamp, string, and binary are supported; any other type (float, double, boolean, array, struct, map, timestamp_ntz in older versions) is rejected with this UnsupportedOperationException.

Solutions

  1. Cast the value to a supported type before bucketing: bucket(16, cast(price AS DECIMAL(38, 8))) or bucket(16, cast(col AS STRING)).
  2. Use a different transform for the column type, or bucket on an integer surrogate key instead of a float.
  3. Check the column type with DESC TABLE or typeof(col) to confirm which type needs conversion.

Example fix

// before
SELECT iceberg.bucket(16, price) FROM t; -- price DOUBLE
// after
SELECT iceberg.bucket(16, CAST(price AS DECIMAL(38, 8))) FROM t;
Defensive patterns

Strategy: type-guard

Validate before calling

val t = col.dataType
val supported = Seq(DateType, TimestampType, StringType, BinaryType) ++
  Seq(ByteType, ShortType, IntegerType, LongType) :+ DecimalType.SYSTEM_DEFAULT
require(supported.exists(_.acceptsType(t)), s"bucket value type $t unsupported")

Type guard

def isBucketable(t: DataType): Boolean = t match {
  case _: DecimalType => true
  case DateType | TimestampType | StringType | BinaryType |
       ByteType | ShortType | IntegerType | LongType => true
  case _ => false
}

Try / catch

try {
  df.select(functions.callUDF("iceberg.bucket", lit(16), col("v")))
} catch {
  case e: UnsupportedOperationException if e.getMessage.contains("Expected column") =>
    df.select(functions.callUDF("iceberg.bucket", lit(16), col("v").cast(StringType)))
}

Prevention

When it happens

Trigger: Calling iceberg.bucket(width, col) where col's type is DOUBLE, FLOAT, BOOLEAN, ARRAY, STRUCT, or MAP — e.g. bucket(16, price) where price is DOUBLE.

Common situations: Bucketing on a floating-point column such as a price or score; mistakenly passing a struct/array column from a nested schema; SQL literal of unsupported type (e.g. a DOUBLE literal like 1.5 is DECIMAL but 1.0e2 is DOUBLE).

Understand the failure class

Background: Type mismatch errors: IllegalArgumentException, TypeError and type guards across 150 open-source libraries — this error's family across 150 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/8eb0809ce1fe056f. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.0/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java:103

      return new BucketInt(type);
    } else if (type instanceof ByteType
        || type instanceof ShortType
        || type instanceof IntegerType) {
      return new BucketInt(DataTypes.IntegerType);
    } else if (type instanceof LongType) {
      return new BucketLong(type);
    } else if (type instanceof TimestampType) {
      return new BucketLong(type);
    } else if (type instanceof TimestampNTZType) {
      return new BucketLong(type);
    } else if (type instanceof DecimalType) {
      return new BucketDecimal(type);
    } else if (type instanceof StringType) {
      return new BucketString();
    } else if (type instanceof BinaryType) {
      return new BucketBinary();
    } else {
      throw new UnsupportedOperationException(
          "Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary");
    }
  }

  @Override
  public String description() {
    return name()
        + "(numBuckets, col) - Call Iceberg's bucket transform\n"
        + "  numBuckets :: number of buckets to divide the rows into, e.g. bucket(100, 34) -> 79 (must be a tinyint, smallint, or int)\n"
        + "  col :: column to bucket (must be a date, integer, long, timestamp, decimal, string, or binary)";
  }

  @Override
  public String name() {
    return "bucket";
  }

  public abstract static class BucketBase extends BaseScalarFunction<Integer>

View on GitHub (pinned to 86d9c8fc54)