apache/iceberg · error · UnsupportedOperationException

Batch reading is not supported in non-vectorized reader

Error message

Batch reading is not supported in non-vectorized reader

What it means

Iceberg's ORC reader can read either rows (generic ORC records) or columnar batches (VectorizedRowBatch). recordsPerBatch() only makes sense in batch/vectorized mode; the builder tracks an isBatchReader flag and throws UnsupportedOperationException when batch reading was not requested. It is thrown immediately, before any rows are read, so no scan is started with an incompatible mode.

Solutions

  1. Configure a batch reader on the ReadBuilder (e.g. .createBatchedReaderFunc(...) / use the vectorized reader factory) before calling recordsPerBatch.
  2. If batch semantics are not needed, drop the recordsPerBatch call and consume the row iterator instead.
  3. Check the reader configuration/API version you are using; ensure the ORC format model supports batch reading for your table's schema (e.g. no unsupported types forcing the non-vectorized fallback).

Example fix

// before
CloseableIterable<Record> rows = ORC.read(file)
    .project(schema)
    .recordsPerBatch(1024)  // UnsupportedOperationException
    .build();

// after
ORC.ReadBuilder<ColumnarBatch> batches = ORC.read(file)
    .project(schema)
    .createBatchedReaderFunc((type, batch) -> new MyOrcBatchReader(type, batch))
    .recordsPerBatch(1024);
Defensive patterns

Strategy: validation

Validate before calling

// ensure a batch reader is configured before requesting batches
if (!readBuilderUsesBatchedReader) {
  throw new IllegalStateException("recordsPerBatch requires a vectorized/batch reader");
}

Prevention

When it happens

Trigger: Calling ORC.reads(...).project(schema).recordsPerBatch(n) (e.g. via .recordsPerBatch() on the ReadBuilder) without first calling .createReaderFunc(...)/selecting a batch reader (ORC.read(...).createBatchedReaderFunc). Any code path that requests batched reads on a reader built with the non-vectorized row factory.

Common situations: Migrating code from row-based scans to vectorized batch reads by just adding recordsPerBatch; copying Spark-style batch APIs; forgetting that the ORC ReadBuilder needs createReaderFunc with a batch-producing reader for recordsPerBatch to be valid.

Understand the failure class

Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/6f5b0eee9c27bbd2. Report an issue: GitHub.

Appendix: source

Thrown at orc/src/main/java/org/apache/iceberg/orc/ORCFormatModel.java:273

      return this;
    }

    @Override
    public ReadBuilder<D, S> set(String key, String value) {
      internal.config(key, value);
      return this;
    }

    @Override
    public ReadBuilder<D, S> reuseContainers() {
      this.reuseContainers = true;
      return this;
    }

    @Override
    public ReadBuilder<D, S> recordsPerBatch(int numRowsPerBatch) {
      if (!isBatchReader) {
        throw new UnsupportedOperationException(
            "Batch reading is not supported in non-vectorized reader");
      }

      internal.recordsPerBatch(numRowsPerBatch);
      return this;
    }

    @Override
    public ReadBuilder<D, S> idToConstant(Map<Integer, ?> newIdToConstant) {
      internal.constantFieldIds(newIdToConstant.keySet());
      this.idToConstant = newIdToConstant;
      return this;
    }

    @Override
    public ReadBuilder<D, S> withNameMapping(NameMapping nameMapping) {
      internal.withNameMapping(nameMapping);
      return this;

View on GitHub (pinned to 86d9c8fc54)