apache/iceberg · error · RuntimeIOException

Failed to create reader for dictionary page

Error message

Failed to create reader for dictionary page

What it means

Thrown as a RuntimeIOException when initializing the dictionary reader for a column's dictionary page (page.getEncoding().initDictionary) throws IOException. The dictionary page exists but its content cannot be decoded into a usable Dictionary object for the column's encoding.

Solutions

  1. Verify/re-download the Parquet file and check checksums — corruption is the most common cause.
  2. Rewrite the file with a standard tool (Spark/Iceberg) to regenerate valid dictionary pages.
  3. Catch RuntimeIOException in the filter path and fall back to full record evaluation.
  4. Upgrade Iceberg/parquet-mr in case of an encoding initialization bug.

Example fix

// before
dict = page.getEncoding().initDictionary(col, page);
// after
try {
  dict = page.getEncoding().initDictionary(col, page);
} catch (RuntimeIOException e) {
  return null; // fall back to non-dictionary filtering
}
Defensive patterns

Strategy: try-catch

Try / catch

try {
  result = filter.shouldRead(rowGroup, ...);
} catch (RuntimeIOException e) {
  LOG.warn("Dictionary page unreadable, evaluating rows directly", e);
  return true; // conservative: evaluate rows
}

Prevention

When it happens

Trigger: Corrupt or truncated dictionary page bytes in the Parquet file; an encoding whose initDictionary fails for the given column descriptor (e.g. mismatched page encoding vs column type); underlying file/IO errors surfaced as IOException from the page source.

Common situations: Reading damaged files (network truncation, partial upload); files written by buggy or exotic encoders; encodings not matching expected initialization.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/6586fc2993d15dad. Report an issue: GitHub.

Appendix: source

Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ParquetDictionaryRowGroupFilter.java:438

      Set<?> cached = dictCache.get(id);
      if (cached != null) {
        return (Set<T>) cached;
      }

      ColumnDescriptor col = cols.get(id);
      DictionaryPage page = dictionaries.readDictionaryPage(col);
      // may not be dictionary-encoded
      if (page == null) {
        throw new IllegalStateException("Failed to read required dictionary page for id: " + id);
      }

      Function<Object, Object> conversion = conversions.get(id);

      Dictionary dict;
      try {
        dict = page.getEncoding().initDictionary(col, page);
      } catch (IOException e) {
        throw new RuntimeIOException("Failed to create reader for dictionary page");
      }

      Set<T> dictSet = Sets.newTreeSet(comparator);

      for (int i = 0; i <= dict.getMaxId(); i++) {
        switch (col.getPrimitiveType().getPrimitiveTypeName()) {
          case FIXED_LEN_BYTE_ARRAY:
            dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));
            break;
          case BINARY:
            dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));
            break;
          case INT32:
            dictSet.add((T) conversion.apply(dict.decodeToInt(i)));
            break;
          case INT64:
            dictSet.add((T) conversion.apply(dict.decodeToLong(i)));
            break;

View on GitHub (pinned to 86d9c8fc54)