apache/iceberg · error · UnsupportedOperationException

Byte stream split encoding is not supported for type " +…

Error message

Byte stream split encoding is not supported for type " + type.getPrimitiveTypeName()

What it means

Byte stream split (BSS) encoding in Parquet is only implemented for fixed-width types with known element sizes: INT32, INT64, FLOAT, DOUBLE, and FIXED_LEN_BYTE_ARRAY. byteStreamSplitElementSize returns the element byte size for these types and throws UnsupportedOperationException for any other primitive (e.g. BYTE_ARRAY, INT96) because the vectorized BSS decoder cannot determine element boundaries.

Solutions

  1. Rewrite the data without BYTE_STREAM_SPLIT encoding (use PLAIN, DELTA_BINARY_PACKED, or dictionary).
  2. Disable vectorized reads (read.parquet.vectorization.enabled=false) so the generic path handles the file.
  3. Extend byteStreamSplitElementSize support only if the type has a fixed width.

Example fix

// before: vectorized read of BSS-encoded BYTE_ARRAY column fails
// after
table.newScan().option("read.parquet.vectorization.enabled", "false");
Defensive patterns

Strategy: validation

Validate before calling

ParquetMetadata footer = ParquetFileReader.readFooter(conf, path);
ColumnChunkMetaData col = footer.getBlocks().get(0).getColumns().get(0);
boolean bss = col.getEncodings().contains(Encoding.BYTE_STREAM_SPLIT);
boolean supported = Set.of(INT32, INT64, FLOAT, DOUBLE, FIXED_LEN_BYTE_ARRAY)
    .contains(col.getPrimitiveType().getPrimitiveTypeName());
boolean vectorize = !(bss && !supported);

Prevention

When it happens

Trigger: Reading a Parquet column encoded with BYTE_STREAM_SPLIT encoding whose primitive type is not one of INT32/INT64/FLOAT/DOUBLE/FIXED_LEN_BYTE_ARRAY — detected in initDataReader when selecting the BSS read path.

Common situations: Files written by other engines (e.g. certain Arrow/Parquet C++ writers, DuckDB, IoT time-series encoders) that use BYTE_STREAM_SPLIT on variable-length or unsupported types, then read via Iceberg's vectorized Spark/Arrow reader.

Understand the failure class

Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/75773180ef0f22cd. Report an issue: GitHub.

Appendix: source

Thrown at arrow/src/main/java/org/apache/iceberg/arrow/vectorized/parquet/VectorizedPageIterator.java:397

              batchSize,
              holder,
              dictionaryEncodedValuesReader,
              dictionary);
    }
  }

  private static int byteStreamSplitElementSize(PrimitiveType type) {
    switch (type.getPrimitiveTypeName()) {
      case INT32:
      case FLOAT:
        return VectorizedValuesReader.INT_SIZE;
      case INT64:
      case DOUBLE:
        return VectorizedValuesReader.LONG_SIZE;
      case FIXED_LEN_BYTE_ARRAY:
        return type.getTypeLength();
      default:
        throw new UnsupportedOperationException(
            "Byte stream split encoding is not supported for type " + type.getPrimitiveTypeName());
    }
  }

  private int getActualBatchSize(int expectedBatchSize) {
    return Math.min(expectedBatchSize, triplesCount - triplesRead);
  }

  class FixedSizeBinaryPageReader extends BasePageReader {
    @Override
    protected void nextVal(
        FieldVector vector, int batchSize, int numVals, int typeWidth, NullabilityHolder holder) {
      vectorizedDefinitionLevelReader
          .fixedSizeBinaryReader()
          .nextBatch(vector, numVals, typeWidth, batchSize, holder, valuesReader);
    }

    @Override

View on GitHub (pinned to 86d9c8fc54)