apache/iceberg · error · RuntimeIOException
Failed to create reader for dictionary page
Error message
Failed to create reader for dictionary page
What it means
Thrown as a RuntimeIOException when initializing the dictionary reader for a column's dictionary page (page.getEncoding().initDictionary) throws IOException. The dictionary page exists but its content cannot be decoded into a usable Dictionary object for the column's encoding.
Solutions
- Verify/re-download the Parquet file and check checksums — corruption is the most common cause.
- Rewrite the file with a standard tool (Spark/Iceberg) to regenerate valid dictionary pages.
- Catch RuntimeIOException in the filter path and fall back to full record evaluation.
- Upgrade Iceberg/parquet-mr in case of an encoding initialization bug.
Example fix
// before
dict = page.getEncoding().initDictionary(col, page);
// after
try {
dict = page.getEncoding().initDictionary(col, page);
} catch (RuntimeIOException e) {
return null; // fall back to non-dictionary filtering
} Defensive patterns
Strategy: try-catch
Try / catch
try {
result = filter.shouldRead(rowGroup, ...);
} catch (RuntimeIOException e) {
LOG.warn("Dictionary page unreadable, evaluating rows directly", e);
return true; // conservative: evaluate rows
} Prevention
- Validate file integrity (checksums, complete uploads) before reading Parquet data.
- Avoid partial/truncated file sources; retry downloads on failure.
- Keep Iceberg/parquet-mr patched for encoding initialization bugs.
When it happens
Trigger: Corrupt or truncated dictionary page bytes in the Parquet file; an encoding whose initDictionary fails for the given column descriptor (e.g. mismatched page encoding vs column type); underlying file/IO errors surfaced as IOException from the page source.
Common situations: Reading damaged files (network truncation, partial upload); files written by buggy or exotic encoders; encodings not matching expected initialization.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Cannot decode dictionary of type
- could not decode the dictionary for
- could not read page in col " + desc
- could not read page in col
- could not read page " + valueCount + " in col " + desc
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/6586fc2993d15dad.
Report an issue: GitHub.
Appendix: source
Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ParquetDictionaryRowGroupFilter.java:438
Set<?> cached = dictCache.get(id);
if (cached != null) {
return (Set<T>) cached;
}
ColumnDescriptor col = cols.get(id);
DictionaryPage page = dictionaries.readDictionaryPage(col);
// may not be dictionary-encoded
if (page == null) {
throw new IllegalStateException("Failed to read required dictionary page for id: " + id);
}
Function<Object, Object> conversion = conversions.get(id);
Dictionary dict;
try {
dict = page.getEncoding().initDictionary(col, page);
} catch (IOException e) {
throw new RuntimeIOException("Failed to create reader for dictionary page");
}
Set<T> dictSet = Sets.newTreeSet(comparator);
for (int i = 0; i <= dict.getMaxId(); i++) {
switch (col.getPrimitiveType().getPrimitiveTypeName()) {
case FIXED_LEN_BYTE_ARRAY:
dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));
break;
case BINARY:
dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));
break;
case INT32:
dictSet.add((T) conversion.apply(dict.decodeToInt(i)));
break;
case INT64:
dictSet.add((T) conversion.apply(dict.decodeToLong(i)));
break;View on GitHub (pinned to 86d9c8fc54)