{"record":{"id":"8eb0809ce1fe056f","repo":"apache/iceberg","slug":"expected-column-to-be-date-tinyint-smallint-int-8eb080","errorCode":null,"errorMessage":"Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary","messagePattern":"Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary","errorType":"exception","errorClass":"UnsupportedOperationException","httpStatus":null,"severity":"error","filePath":"spark/v4.0/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java","lineNumber":103,"sourceCode":"      return new BucketInt(type);\n    } else if (type instanceof ByteType\n        || type instanceof ShortType\n        || type instanceof IntegerType) {\n      return new BucketInt(DataTypes.IntegerType);\n    } else if (type instanceof LongType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampNTZType) {\n      return new BucketLong(type);\n    } else if (type instanceof DecimalType) {\n      return new BucketDecimal(type);\n    } else if (type instanceof StringType) {\n      return new BucketString();\n    } else if (type instanceof BinaryType) {\n      return new BucketBinary();\n    } else {\n      throw new UnsupportedOperationException(\n          \"Expected column to be date, tinyint, smallint, int, bigint, decimal, timestamp, string, or binary\");\n    }\n  }\n\n  @Override\n  public String description() {\n    return name()\n        + \"(numBuckets, col) - Call Iceberg's bucket transform\\n\"\n        + \"  numBuckets :: number of buckets to divide the rows into, e.g. bucket(100, 34) -> 79 (must be a tinyint, smallint, or int)\\n\"\n        + \"  col :: column to bucket (must be a date, integer, long, timestamp, decimal, string, or binary)\";\n  }\n\n  @Override\n  public String name() {\n    return \"bucket\";\n  }\n\n  public abstract static class BucketBase extends BaseScalarFunction<Integer>","sourceCodeStart":85,"sourceCodeEnd":121,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v4.0/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java#L85-L121","documentation":"After validating the bucket width, BucketFunction.bind() inspects the value column's data type and dispatches to a typed BucketFunction implementation. Only date, tinyint, smallint, int, bigint, decimal, timestamp, string, and binary are supported; any other type (float, double, boolean, array, struct, map, timestamp_ntz in older versions) is rejected with this UnsupportedOperationException.","triggerScenarios":"Calling iceberg.bucket(width, col) where col's type is DOUBLE, FLOAT, BOOLEAN, ARRAY, STRUCT, or MAP — e.g. bucket(16, price) where price is DOUBLE.","commonSituations":"Bucketing on a floating-point column such as a price or score; mistakenly passing a struct/array column from a nested schema; SQL literal of unsupported type (e.g. a DOUBLE literal like 1.5 is DECIMAL but 1.0e2 is DOUBLE).","solutions":["Cast the value to a supported type before bucketing: bucket(16, cast(price AS DECIMAL(38, 8))) or bucket(16, cast(col AS STRING)).","Use a different transform for the column type, or bucket on an integer surrogate key instead of a float.","Check the column type with DESC TABLE or typeof(col) to confirm which type needs conversion."],"exampleFix":"// before\nSELECT iceberg.bucket(16, price) FROM t; -- price DOUBLE\n// after\nSELECT iceberg.bucket(16, CAST(price AS DECIMAL(38, 8))) FROM t;","handlingStrategy":"type-guard","validationCode":"val t = col.dataType\nval supported = Seq(DateType, TimestampType, StringType, BinaryType) ++\n  Seq(ByteType, ShortType, IntegerType, LongType) :+ DecimalType.SYSTEM_DEFAULT\nrequire(supported.exists(_.acceptsType(t)), s\"bucket value type $t unsupported\")","typeGuard":"def isBucketable(t: DataType): Boolean = t match {\n  case _: DecimalType => true\n  case DateType | TimestampType | StringType | BinaryType |\n       ByteType | ShortType | IntegerType | LongType => true\n  case _ => false\n}","tryCatchPattern":"try {\n  df.select(functions.callUDF(\"iceberg.bucket\", lit(16), col(\"v\")))\n} catch {\n  case e: UnsupportedOperationException if e.getMessage.contains(\"Expected column\") =>\n    df.select(functions.callUDF(\"iceberg.bucket\", lit(16), col(\"v\").cast(StringType)))\n}","preventionTips":["Never bucket float/double columns directly — cast to DECIMAL or bucket an integer key instead","Check column types with DESC TABLE before writing bucket transforms in SQL"],"tags":["spark","sql-function-binding","type-mismatch"],"backgroundTag":"type-mismatch","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}