apache/iceberg · error · UnsupportedOperationException
Batch reading is not supported in non-vectorized reader
Error message
Batch reading is not supported in non-vectorized reader
What it means
Iceberg's ORC reader can read either rows (generic ORC records) or columnar batches (VectorizedRowBatch). recordsPerBatch() only makes sense in batch/vectorized mode; the builder tracks an isBatchReader flag and throws UnsupportedOperationException when batch reading was not requested. It is thrown immediately, before any rows are read, so no scan is started with an incompatible mode.
Solutions
- Configure a batch reader on the ReadBuilder (e.g. .createBatchedReaderFunc(...) / use the vectorized reader factory) before calling recordsPerBatch.
- If batch semantics are not needed, drop the recordsPerBatch call and consume the row iterator instead.
- Check the reader configuration/API version you are using; ensure the ORC format model supports batch reading for your table's schema (e.g. no unsupported types forcing the non-vectorized fallback).
Example fix
// before
CloseableIterable<Record> rows = ORC.read(file)
.project(schema)
.recordsPerBatch(1024) // UnsupportedOperationException
.build();
// after
ORC.ReadBuilder<ColumnarBatch> batches = ORC.read(file)
.project(schema)
.createBatchedReaderFunc((type, batch) -> new MyOrcBatchReader(type, batch))
.recordsPerBatch(1024); Defensive patterns
Strategy: validation
Validate before calling
// ensure a batch reader is configured before requesting batches
if (!readBuilderUsesBatchedReader) {
throw new IllegalStateException("recordsPerBatch requires a vectorized/batch reader");
} Prevention
- Always pair recordsPerBatch with createBatchedReaderFunc/vectorized reader setup
- Do not copy batch-read APIs across engines without checking the ORC ReadBuilder contract
- Read the ReadBuilder docs: batch mode is opt-in for ORC
When it happens
Trigger: Calling ORC.reads(...).project(schema).recordsPerBatch(n) (e.g. via .recordsPerBatch() on the ReadBuilder) without first calling .createReaderFunc(...)/selecting a batch reader (ORC.read(...).createBatchedReaderFunc). Any code path that requests batched reads on a reader built with the non-vectorized row factory.
Common situations: Migrating code from row-based scans to vectorized batch reads by just adding recordsPerBatch; copying Spark-style batch APIs; forgetting that the ORC ReadBuilder needs createReaderFunc with a batch-producing reader for recordsPerBatch to be valid.
Understand the failure class
Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.
Related errors
- Backup table name cannot be specified
- Bucket transform is not supported
- Can't get Stripe's length from the file writer with path
- Can't modify an empty struct
- Can't retrieve values from an empty struct
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/6f5b0eee9c27bbd2.
Report an issue: GitHub.
Appendix: source
Thrown at orc/src/main/java/org/apache/iceberg/orc/ORCFormatModel.java:273
return this;
}
@Override
public ReadBuilder<D, S> set(String key, String value) {
internal.config(key, value);
return this;
}
@Override
public ReadBuilder<D, S> reuseContainers() {
this.reuseContainers = true;
return this;
}
@Override
public ReadBuilder<D, S> recordsPerBatch(int numRowsPerBatch) {
if (!isBatchReader) {
throw new UnsupportedOperationException(
"Batch reading is not supported in non-vectorized reader");
}
internal.recordsPerBatch(numRowsPerBatch);
return this;
}
@Override
public ReadBuilder<D, S> idToConstant(Map<Integer, ?> newIdToConstant) {
internal.constantFieldIds(newIdToConstant.keySet());
this.idToConstant = newIdToConstant;
return this;
}
@Override
public ReadBuilder<D, S> withNameMapping(NameMapping nameMapping) {
internal.withNameMapping(nameMapping);
return this;View on GitHub (pinned to 86d9c8fc54)