{"record":{"id":"e202a8e8260ed73f","repo":"apache/iceberg","slug":"expected-number-of-buckets-to-be-tinyint-shortint-e202a8","errorCode":null,"errorMessage":"Expected number of buckets to be tinyint, shortint or int","messagePattern":"Expected number of buckets to be tinyint, shortint or int","errorType":"exception","errorClass":"UnsupportedOperationException","httpStatus":null,"severity":"error","filePath":"spark/v4.0/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java","lineNumber":79,"sourceCode":"  private static final int NUM_BUCKETS_ORDINAL = 0;\n  private static final int VALUE_ORDINAL = 1;\n\n  private static final Set<DataType> SUPPORTED_NUM_BUCKETS_TYPES =\n      ImmutableSet.of(DataTypes.ByteType, DataTypes.ShortType, DataTypes.IntegerType);\n\n  @Override\n  @SuppressWarnings(\"checkstyle:CyclomaticComplexity\")\n  public BoundFunction bind(StructType inputType) {\n    if (inputType.size() != 2) {\n      throw new UnsupportedOperationException(\n          \"Wrong number of inputs (expected numBuckets and value)\");\n    }\n\n    StructField numBucketsField = inputType.fields()[NUM_BUCKETS_ORDINAL];\n    StructField valueField = inputType.fields()[VALUE_ORDINAL];\n\n    if (!SUPPORTED_NUM_BUCKETS_TYPES.contains(numBucketsField.dataType())) {\n      throw new UnsupportedOperationException(\n          \"Expected number of buckets to be tinyint, shortint or int\");\n    }\n\n    DataType type = valueField.dataType();\n    if (type instanceof DateType) {\n      return new BucketInt(type);\n    } else if (type instanceof ByteType\n        || type instanceof ShortType\n        || type instanceof IntegerType) {\n      return new BucketInt(DataTypes.IntegerType);\n    } else if (type instanceof LongType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampType) {\n      return new BucketLong(type);\n    } else if (type instanceof TimestampNTZType) {\n      return new BucketLong(type);\n    } else if (type instanceof DecimalType) {\n      return new BucketDecimal(type);","sourceCodeStart":61,"sourceCodeEnd":97,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/spark/v4.0/spark/src/main/java/org/apache/iceberg/spark/functions/BucketFunction.java#L61-L97","documentation":"Iceberg's bucket table function (days/hours style transform exposed as a Spark function) binds its arguments at planning time. The second argument's first input field is the bucket count, which must be one of ByteType, ShortType, or IntegerType so it can be widened to an int internally. Spark's bind() throws UnsupportedOperationException when the numBuckets field has any other type (e.g. long, string, decimal, null).","triggerScenarios":"Calling iceberg.bucket(width, col) where the width expression is not a tinyint/short/int literal or column — e.g. bucket(1024L, id), bucket(cast('10' as string), id), or bucket(someBigIntCol, id).","commonSituations":"Passing the bucket width as a BIGINT constant from another column or computed expression; typing the width in SQL as a plain number that Spark parses to BIGINT (literals default to INT only below certain sizes); schema evolution changing the width column type.","solutions":["Cast the bucket width expression to INT: bucket(cast(width AS INT), col) — Spark literal 1024 is already INT.","If the width comes from a BIGINT column, wrap it: bucket(int(widthCol), col).","Check the actual type of the first argument with typeof() in Spark SQL to confirm the mismatch.","If passing a NULL, cast it explicitly: cast(NULL AS INT)."],"exampleFix":"// before\nSELECT iceberg.bucket(1024L, id) FROM t;\n// after\nSELECT iceberg.bucket(1024, id) FROM t; -- INT literal\n-- or\nSELECT iceberg.bucket(CAST(width_col AS INT), id) FROM t;","handlingStrategy":"validation","validationCode":"val widthType = widthExpr.dataType\nrequire(widthType == ByteType || widthType == ShortType || widthType == IntegerType,\n  s\"bucket width must be tinyint/short/int, got $widthType\")","typeGuard":"def isSupportedWidth(t: DataType): Boolean =\n  t == ByteType || t == ShortType || t == IntegerType","tryCatchPattern":null,"preventionTips":["Use plain integer literals for bucket widths (1024, 256), never L-suffixed or BIGINT-typed values","Cast any width derived from a column or config string to INT before calling iceberg.bucket"],"tags":["spark","sql-function-binding","unsupported-operation"],"backgroundTag":"unsupported-operation","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}