{"record":{"id":"f112f9610c7cfd4f","repo":"apache/iceberg","slug":"expected-column-to-be-date-tinyint-smallint-int-f112f9","errorCode":null,"errorMessage":"Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary","messagePattern":"Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary","errorType":"exception","errorClass":"java.lang.UnsupportedOperationException","httpStatus":null,"severity":"error","filePath":"spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java","lineNumber":103,"sourceCode":"      return new BucketInt(type);\n    } else if (type instanceof ByteType\n        || type instanceof ShortType\n        || type instanceof IntegerType) {\n      return new BucketInt(DataTypes.IntegerType);\n    } else if (type instanceof LongType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampNTZType) {\n      return new BucketLong(type);\n    } else if (type instanceof DecimalType) {\n      return new BucketDecimal(type);\n    } else if (type instanceof StringType) {\n      return new BucketString();\n    } else if (type instanceof BinaryType) {\n      return new BucketBinary();\n    } else {\n      throw new UnsupportedOperationException(\n          \"Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary\");\n    }\n  }\n\n  @Override\n  public String description() {\n    return name()\n        + \"(numBuckets, col) - Call Iceberg's bucket transform\\n\"\n        + \"  numBuckets :: number of buckets to divide the rows into, e.g. bucket(100, 34) -> 79 (must be a tinyint, smallint, or int)\\n\"\n        + \"  col :: column to bucket (must be a date, integer, long, timestamp, decimal, string, or binary)\";\n  }\n\n  @Override\n  public String name() {\n    return \"bucket\";\n  }\n\n  public abstract static class BucketBase extends BaseScalarFunction<Integer>","sourceCodeStart":85,"sourceCodeEnd":121,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java#L85-L121","documentation":"Iceberg's bucket(x, n) Spark function only supports value columns of date, tinyint, smallint, int, bigint, decimal, timestamp (and timestamp_ntz), string, or binary. If the value column has any other Spark data type (e.g. float, double, boolean, struct, map, array), bind() cannot select a bucketing implementation and throws this UnsupportedOperationException during query analysis.","triggerScenarios":"Calling bucket(value, n) where value is a DoubleType, FloatType, BooleanType, or complex type (struct/array/map) column, e.g. bucket(score, 16) where score is double.","commonSituations":"Bucketing floating-point columns in a partition spec; bucketing complex or boolean columns; schema drift where a column's type changed from string to double after the transform was written.","solutions":["Cast the value column to a supported type, e.g. bucket(CAST(score AS DECIMAL(10,4)), 16).","Use a different column of a supported type for bucketing.","For floats/doubles, convert to decimal or bigint (e.g. scaled integer) before bucketing.","Check the table's partition spec if this arises from a CREATE TABLE — change the transform or column type."],"exampleFix":"// before\nSELECT bucket(score, 16) FROM t  -- score is DOUBLE\n// after\nSELECT bucket(CAST(score AS DECIMAL(10,4)), 16) FROM t","handlingStrategy":"validation","validationCode":"// Spark Scala, before calling bucket()\nval supported = Set(\"datetype\",\"tinyint\",\"smallint\",\"int\",\"bigint\",\"decimal\",\"timestamp\",\"timestamp_ntz\",\"string\",\"binary\")\nval dt = df.schema(\"value_col\").dataType.catalogString.toLowerCase\nrequire(supported.exists(s => dt.startsWith(s)), s\"bucket() does not support value type: $dt\")","typeGuard":"def isBucketableType(dt: org.apache.spark.sql.types.DataType): Boolean = dt match {\n  case _: org.apache.spark.sql.types.DecimalType => true\n  case t => Set(\n    org.apache.spark.sql.types.DateType,\n    org.apache.spark.sql.types.TimestampType,\n    org.apache.spark.sql.types.StringType,\n    org.apache.spark.sql.types.BinaryType).contains(t)\n}","tryCatchPattern":"try {\n  df.select(expr(s\"bucket(value_col, $n)\"))\n} catch {\n  case e: UnsupportedOperationException if e.getMessage.contains(\"Expected column to be\") =>\n    throw new IllegalArgumentException(\"bucket() value column has unsupported type; cast to int/string/decimal/etc.\", e)\n}","preventionTips":["Inspect df.schema before applying bucket() to a column","Avoid bucketing float/double/boolean/complex columns — cast to decimal or bigint first","Use days()/hours() for temporal bucketing of dates and timestamps","Validate table partition specs against supported transform types at write time"],"tags":["spark","sql-function","unsupported-type","bind"],"backgroundTag":"unsupported-operation","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}