{"record":{"id":"4884373d585225e7","repo":"apache/iceberg","slug":"expected-number-of-buckets-to-be-tinyint-shortint-488437","errorCode":null,"errorMessage":"Expected number of buckets to be tinyint, shortint or int","messagePattern":"Expected number of buckets to be tinyint, shortint or int","errorType":"exception","errorClass":"java.lang.UnsupportedOperationException","httpStatus":null,"severity":"error","filePath":"spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java","lineNumber":79,"sourceCode":"  private static final int NUM_BUCKETS_ORDINAL = 0;\n  private static final int VALUE_ORDINAL = 1;\n\n  private static final Set<DataType> SUPPORTED_NUM_BUCKETS_TYPES =\n      ImmutableSet.of(DataTypes.ByteType, DataTypes.ShortType, DataTypes.IntegerType);\n\n  @Override\n  @SuppressWarnings(\"checkstyle:CyclomaticComplexity\")\n  public BoundFunction bind(StructType inputType) {\n    if (inputType.size() != 2) {\n      throw new UnsupportedOperationException(\n          \"Wrong number of inputs (expected numBuckets and value)\");\n    }\n\n    StructField numBucketsField = inputType.fields()[NUM_BUCKETS_ORDINAL];\n    StructField valueField = inputType.fields()[VALUE_ORDINAL];\n\n    if (!SUPPORTED_NUM_BUCKETS_TYPES.contains(numBucketsField.dataType())) {\n      throw new UnsupportedOperationException(\n          \"Expected number of buckets to be tinyint, shortint or int\");\n    }\n\n    DataType type = valueField.dataType();\n    if (type instanceof DateType) {\n      return new BucketInt(type);\n    } else if (type instanceof ByteType\n        || type instanceof ShortType\n        || type instanceof IntegerType) {\n      return new BucketInt(DataTypes.IntegerType);\n    } else if (type instanceof LongType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampNTZType) {\n      return new BucketLong(type);\n    } else if (type instanceof DecimalType) {\n      return new BucketDecimal(type);","sourceCodeStart":61,"sourceCodeEnd":97,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java#L61-L97","documentation":"Iceberg's bucket(x, n) Spark function validates, at binding time, that the numBuckets argument is one of Spark's ByteType, ShortType, or IntegerType. Any other literal or column type for the bucket count (e.g. long, string, decimal) cannot be used to compute a bucket, so bind() throws this UnsupportedOperationException. The error is thrown during query analysis, before execution.","triggerScenarios":"Calling bucket(value, n) where n is a LongType, StringType, DecimalType, or any type outside ByteType/ShortType/IntegerType — e.g. bucket(id, 128L), a string literal, or a column of type long passed as the bucket count.","commonSituations":"Writing bucket transforms in CREATE TABLE partition specs where the bucket count literal is typed as long; Spark SQL widening literals to LongType; passing a computed expression (e.g. CAST or a column) as the bucket count with the wrong type.","solutions":["Cast the bucket count to int: bucket(col, CAST(128 AS INT)).","Pass a plain integer literal without a suffix, e.g. bucket(id, 128) instead of 128L.","Check the numBuckets argument position; it must be the second argument to bucket().","If the count comes from a column or expression, ensure its declared type is byte, short, or int."],"exampleFix":"// before\nSELECT bucket(id, 128L) FROM t\n// after\nSELECT bucket(id, CAST(128 AS INT)) FROM t","handlingStrategy":"validation","validationCode":"// Spark Scala, before calling bucket()\nval n = 128\nrequire(n.isInstanceOf[Integer] || n.isInstanceOf[java.lang.Byte] || n.isInstanceOf[java.lang.Short],\n  s\"bucket count must be int/short/tinyint, got: ${if (n == null) \"null\" else n.getClass.getSimpleName}\")\n// SQL: ensure literal is int: bucket(id, CAST(? AS INT))","typeGuard":"def isSupportedNumBucketsType(dt: org.apache.spark.sql.types.DataType): Boolean =\n  dt == org.apache.spark.sql.types.ByteType ||\n  dt == org.apache.spark.sql.types.ShortType ||\n  dt == org.apache.spark.sql.types.IntegerType","tryCatchPattern":"try {\n  df.select(functions.call(\"bucket\", col(\"id\"), lit(n)))\n} catch {\n  case e: UnsupportedOperationException if e.getMessage.contains(\"number of buckets\") =>\n    throw new IllegalArgumentException(\"bucket() requires an int-typed bucket count\", e)\n}","preventionTips":["Always pass plain integer literals (128, not 128L) as the bucket count","Cast user/config-provided bucket counts to INT before use in SQL","Check argument order: bucket(value, numBuckets)","Validate partition spec transforms when building CREATE TABLE statements programmatically"],"tags":["spark","sql-function","type-mismatch","bind"],"backgroundTag":"type-mismatch","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}