apache/iceberg · error · IllegalStateException

Failed to read required dictionary page for id

Error message

Failed to read required dictionary page for id: <id>

What it means

Thrown when the row-group dictionary filter expects a dictionary for a column id (i.e. the filter's metrics said dictionary data exists) but readDictionaryPage returned null, meaning the column's data pages are not dictionary-encoded. The filter cannot evaluate predicates against a dictionary set without the page, so it fails fast with IllegalStateException.

Solutions

  1. Ensure the code handles null dictionary pages by falling back to non-dictionary (record-level) filtering instead of requiring a dictionary.
  2. Tune writer settings (parquet.dictionary.page.size / fallback thresholds) so dictionary encoding is consistent, or disable dictionary encoding for high-cardinality columns.
  3. Update Iceberg — newer versions treat a missing dictionary page as 'cannot use dictionary filter' and fall back gracefully.

Example fix

// before
if (page == null) {
  throw new IllegalStateException("Failed to read required dictionary page for id: " + id);
}
// after
if (page == null) {
  return null; // fall back to non-dictionary filtering
}
Defensive patterns

Strategy: fallback

Validate before calling

// before using the dictionary filter, confirm dictionary encoding is present in row-group metadata
ColumnChunkMeta meta = ...;
boolean dictUsable = meta.getEncodings().contains(Encoding.PLAIN_DICTIONARY)
    || meta.getEncodings().contains(Encoding.RLE_DICTIONARY);

Try / catch

try {
  Dictionary d = filter.dictionary(id);
} catch (IllegalStateException e) {
  LOG.warn("No dictionary page for column {}, falling back to record filtering", id);
  return ParquetDictionaryRowGroupFilter.ROWS_MATCH;
}

Prevention

When it happens

Trigger: Calling dictionary()/dict() for a column whose row group pages are plain-encoded (writer fell back after too many distinct values) while the filter still demands a dictionary for that column.

Common situations: High-cardinality columns that writers encoded with a dictionary page in some row groups but not others; corrupt footer metadata claiming dictionary encoding; using the dictionary filter on columns that frequently exceed the dictionary size limit.

Understand the failure class

Background: "This is a bug, please report it": internal invariant violations, unreachable panics, and SNH errors explained — this error's family across 47 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/e39561e054cf8c07. Report an issue: GitHub.

Appendix: source

Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ParquetDictionaryRowGroupFilter.java:429

      }

      return ROWS_CANNOT_MATCH;
    }

    @SuppressWarnings("unchecked")
    private <T> Set<T> dict(int id, Comparator<T> comparator) {
      Preconditions.checkNotNull(dictionaries, "Dictionary is required");

      Set<?> cached = dictCache.get(id);
      if (cached != null) {
        return (Set<T>) cached;
      }

      ColumnDescriptor col = cols.get(id);
      DictionaryPage page = dictionaries.readDictionaryPage(col);
      // may not be dictionary-encoded
      if (page == null) {
        throw new IllegalStateException("Failed to read required dictionary page for id: " + id);
      }

      Function<Object, Object> conversion = conversions.get(id);

      Dictionary dict;
      try {
        dict = page.getEncoding().initDictionary(col, page);
      } catch (IOException e) {
        throw new RuntimeIOException("Failed to create reader for dictionary page");
      }

      Set<T> dictSet = Sets.newTreeSet(comparator);

      for (int i = 0; i <= dict.getMaxId(); i++) {
        switch (col.getPrimitiveType().getPrimitiveTypeName()) {
          case FIXED_LEN_BYTE_ARRAY:
            dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));
            break;

View on GitHub (pinned to 86d9c8fc54)