{"record":{"id":"75773180ef0f22cd","repo":"apache/iceberg","slug":"byte-stream-split-encoding-is-not-supported-for-ty","errorCode":null,"errorMessage":"Byte stream split encoding is not supported for type \" + type.getPrimitiveTypeName()","messagePattern":"Byte stream split encoding is not supported for type \" \\+ type\\.getPrimitiveTypeName\\(\\)","errorType":"exception","errorClass":"UnsupportedOperationException","httpStatus":null,"severity":"error","filePath":"arrow/src/main/java/org/apache/iceberg/arrow/vectorized/parquet/VectorizedPageIterator.java","lineNumber":397,"sourceCode":"              batchSize,\n              holder,\n              dictionaryEncodedValuesReader,\n              dictionary);\n    }\n  }\n\n  private static int byteStreamSplitElementSize(PrimitiveType type) {\n    switch (type.getPrimitiveTypeName()) {\n      case INT32:\n      case FLOAT:\n        return VectorizedValuesReader.INT_SIZE;\n      case INT64:\n      case DOUBLE:\n        return VectorizedValuesReader.LONG_SIZE;\n      case FIXED_LEN_BYTE_ARRAY:\n        return type.getTypeLength();\n      default:\n        throw new UnsupportedOperationException(\n            \"Byte stream split encoding is not supported for type \" + type.getPrimitiveTypeName());\n    }\n  }\n\n  private int getActualBatchSize(int expectedBatchSize) {\n    return Math.min(expectedBatchSize, triplesCount - triplesRead);\n  }\n\n  class FixedSizeBinaryPageReader extends BasePageReader {\n    @Override\n    protected void nextVal(\n        FieldVector vector, int batchSize, int numVals, int typeWidth, NullabilityHolder holder) {\n      vectorizedDefinitionLevelReader\n          .fixedSizeBinaryReader()\n          .nextBatch(vector, numVals, typeWidth, batchSize, holder, valuesReader);\n    }\n\n    @Override","sourceCodeStart":379,"sourceCodeEnd":415,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/arrow/src/main/java/org/apache/iceberg/arrow/vectorized/parquet/VectorizedPageIterator.java#L379-L415","documentation":"Byte stream split (BSS) encoding in Parquet is only implemented for fixed-width types with known element sizes: INT32, INT64, FLOAT, DOUBLE, and FIXED_LEN_BYTE_ARRAY. byteStreamSplitElementSize returns the element byte size for these types and throws UnsupportedOperationException for any other primitive (e.g. BYTE_ARRAY, INT96) because the vectorized BSS decoder cannot determine element boundaries.","triggerScenarios":"Reading a Parquet column encoded with BYTE_STREAM_SPLIT encoding whose primitive type is not one of INT32/INT64/FLOAT/DOUBLE/FIXED_LEN_BYTE_ARRAY — detected in initDataReader when selecting the BSS read path.","commonSituations":"Files written by other engines (e.g. certain Arrow/Parquet C++ writers, DuckDB, IoT time-series encoders) that use BYTE_STREAM_SPLIT on variable-length or unsupported types, then read via Iceberg's vectorized Spark/Arrow reader.","solutions":["Rewrite the data without BYTE_STREAM_SPLIT encoding (use PLAIN, DELTA_BINARY_PACKED, or dictionary).","Disable vectorized reads (read.parquet.vectorization.enabled=false) so the generic path handles the file.","Extend byteStreamSplitElementSize support only if the type has a fixed width."],"exampleFix":"// before: vectorized read of BSS-encoded BYTE_ARRAY column fails\n// after\ntable.newScan().option(\"read.parquet.vectorization.enabled\", \"false\");","handlingStrategy":"validation","validationCode":"ParquetMetadata footer = ParquetFileReader.readFooter(conf, path);\nColumnChunkMetaData col = footer.getBlocks().get(0).getColumns().get(0);\nboolean bss = col.getEncodings().contains(Encoding.BYTE_STREAM_SPLIT);\nboolean supported = Set.of(INT32, INT64, FLOAT, DOUBLE, FIXED_LEN_BYTE_ARRAY)\n    .contains(col.getPrimitiveType().getPrimitiveTypeName());\nboolean vectorize = !(bss && !supported);","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Inspect Parquet footers for BYTE_STREAM_SPLIT encoding on unsupported types before enabling vectorized reads.","Disable vectorized reads when reading files produced by non-Java engines.","Rewrite/compact files with standard encodings.","Track Iceberg release notes for newly supported BSS types."],"tags":["parquet","encoding","unsupported-type","vectorized-reads"],"backgroundTag":"unsupported-operation","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}