apache/iceberg · error · java.lang.UnsupportedOperationException

Expected column to be date, tinyint, smallint, int, bigint…

Error message

Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary

What it means

Iceberg's bucket(x, n) Spark function only supports value columns of date, tinyint, smallint, int, bigint, decimal, timestamp (and timestamp_ntz), string, or binary. If the value column has any other Spark data type (e.g. float, double, boolean, struct, map, array), bind() cannot select a bucketing implementation and throws this UnsupportedOperationException during query analysis.

Solutions

  1. Cast the value column to a supported type, e.g. bucket(CAST(score AS DECIMAL(10,4)), 16).
  2. Use a different column of a supported type for bucketing.
  3. For floats/doubles, convert to decimal or bigint (e.g. scaled integer) before bucketing.
  4. Check the table's partition spec if this arises from a CREATE TABLE — change the transform or column type.

Example fix

// before
SELECT bucket(score, 16) FROM t  -- score is DOUBLE
// after
SELECT bucket(CAST(score AS DECIMAL(10,4)), 16) FROM t
Defensive patterns

Strategy: validation

Validate before calling

// Spark Scala, before calling bucket()
val supported = Set("datetype","tinyint","smallint","int","bigint","decimal","timestamp","timestamp_ntz","string","binary")
val dt = df.schema("value_col").dataType.catalogString.toLowerCase
require(supported.exists(s => dt.startsWith(s)), s"bucket() does not support value type: $dt")

Type guard

def isBucketableType(dt: org.apache.spark.sql.types.DataType): Boolean = dt match {
  case _: org.apache.spark.sql.types.DecimalType => true
  case t => Set(
    org.apache.spark.sql.types.DateType,
    org.apache.spark.sql.types.TimestampType,
    org.apache.spark.sql.types.StringType,
    org.apache.spark.sql.types.BinaryType).contains(t)
}

Try / catch

try {
  df.select(expr(s"bucket(value_col, $n)"))
} catch {
  case e: UnsupportedOperationException if e.getMessage.contains("Expected column to be") =>
    throw new IllegalArgumentException("bucket() value column has unsupported type; cast to int/string/decimal/etc.", e)
}

Prevention

When it happens

Trigger: Calling bucket(value, n) where value is a DoubleType, FloatType, BooleanType, or complex type (struct/array/map) column, e.g. bucket(score, 16) where score is double.

Common situations: Bucketing floating-point columns in a partition spec; bucketing complex or boolean columns; schema drift where a column's type changed from string to double after the transform was written.

Understand the failure class

Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/f112f9610c7cfd4f. Report an issue: GitHub.

Appendix: source

Thrown at spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java:103

      return new BucketInt(type);
    } else if (type instanceof ByteType
        || type instanceof ShortType
        || type instanceof IntegerType) {
      return new BucketInt(DataTypes.IntegerType);
    } else if (type instanceof LongType) {
      return new BucketLong(type);
    } else if (type instanceof TimestampType) {
      return new BucketLong(type);
    } else if (type instanceof TimestampNTZType) {
      return new BucketLong(type);
    } else if (type instanceof DecimalType) {
      return new BucketDecimal(type);
    } else if (type instanceof StringType) {
      return new BucketString();
    } else if (type instanceof BinaryType) {
      return new BucketBinary();
    } else {
      throw new UnsupportedOperationException(
          "Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary");
    }
  }

  @Override
  public String description() {
    return name()
        + "(numBuckets, col) - Call Iceberg's bucket transform\n"
        + "  numBuckets :: number of buckets to divide the rows into, e.g. bucket(100, 34) -> 79 (must be a tinyint, smallint, or int)\n"
        + "  col :: column to bucket (must be a date, integer, long, timestamp, decimal, string, or binary)";
  }

  @Override
  public String name() {
    return "bucket";
  }

  public abstract static class BucketBase extends BaseScalarFunction<Integer>

View on GitHub (pinned to 86d9c8fc54)