{"record":{"id":"e39561e054cf8c07","repo":"apache/iceberg","slug":"failed-to-read-required-dictionary-page-for-id-i","errorCode":null,"errorMessage":"Failed to read required dictionary page for id: <id>","messagePattern":"Failed to read required dictionary page for id: <id>","errorType":"exception","errorClass":"IllegalStateException","httpStatus":null,"severity":"error","filePath":"parquet/src/main/java/org/apache/iceberg/parquet/ParquetDictionaryRowGroupFilter.java","lineNumber":429,"sourceCode":"      }\n\n      return ROWS_CANNOT_MATCH;\n    }\n\n    @SuppressWarnings(\"unchecked\")\n    private <T> Set<T> dict(int id, Comparator<T> comparator) {\n      Preconditions.checkNotNull(dictionaries, \"Dictionary is required\");\n\n      Set<?> cached = dictCache.get(id);\n      if (cached != null) {\n        return (Set<T>) cached;\n      }\n\n      ColumnDescriptor col = cols.get(id);\n      DictionaryPage page = dictionaries.readDictionaryPage(col);\n      // may not be dictionary-encoded\n      if (page == null) {\n        throw new IllegalStateException(\"Failed to read required dictionary page for id: \" + id);\n      }\n\n      Function<Object, Object> conversion = conversions.get(id);\n\n      Dictionary dict;\n      try {\n        dict = page.getEncoding().initDictionary(col, page);\n      } catch (IOException e) {\n        throw new RuntimeIOException(\"Failed to create reader for dictionary page\");\n      }\n\n      Set<T> dictSet = Sets.newTreeSet(comparator);\n\n      for (int i = 0; i <= dict.getMaxId(); i++) {\n        switch (col.getPrimitiveType().getPrimitiveTypeName()) {\n          case FIXED_LEN_BYTE_ARRAY:\n            dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));\n            break;","sourceCodeStart":411,"sourceCodeEnd":447,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/parquet/src/main/java/org/apache/iceberg/parquet/ParquetDictionaryRowGroupFilter.java#L411-L447","documentation":"Thrown when the row-group dictionary filter expects a dictionary for a column id (i.e. the filter's metrics said dictionary data exists) but readDictionaryPage returned null, meaning the column's data pages are not dictionary-encoded. The filter cannot evaluate predicates against a dictionary set without the page, so it fails fast with IllegalStateException.","triggerScenarios":"Calling dictionary()/dict() for a column whose row group pages are plain-encoded (writer fell back after too many distinct values) while the filter still demands a dictionary for that column.","commonSituations":"High-cardinality columns that writers encoded with a dictionary page in some row groups but not others; corrupt footer metadata claiming dictionary encoding; using the dictionary filter on columns that frequently exceed the dictionary size limit.","solutions":["Ensure the code handles null dictionary pages by falling back to non-dictionary (record-level) filtering instead of requiring a dictionary.","Tune writer settings (parquet.dictionary.page.size / fallback thresholds) so dictionary encoding is consistent, or disable dictionary encoding for high-cardinality columns.","Update Iceberg — newer versions treat a missing dictionary page as 'cannot use dictionary filter' and fall back gracefully."],"exampleFix":"// before\nif (page == null) {\n  throw new IllegalStateException(\"Failed to read required dictionary page for id: \" + id);\n}\n// after\nif (page == null) {\n  return null; // fall back to non-dictionary filtering\n}","handlingStrategy":"fallback","validationCode":"// before using the dictionary filter, confirm dictionary encoding is present in row-group metadata\nColumnChunkMeta meta = ...;\nboolean dictUsable = meta.getEncodings().contains(Encoding.PLAIN_DICTIONARY)\n    || meta.getEncodings().contains(Encoding.RLE_DICTIONARY);","typeGuard":null,"tryCatchPattern":"try {\n  Dictionary d = filter.dictionary(id);\n} catch (IllegalStateException e) {\n  LOG.warn(\"No dictionary page for column {}, falling back to record filtering\", id);\n  return ParquetDictionaryRowGroupFilter.ROWS_MATCH;\n}","preventionTips":["Don't assume dictionary encoding for high-cardinality columns; prefer filters that degrade gracefully.","Enable dictionary consistently in writers or disable it for such columns.","Use an Iceberg version where a missing dictionary page disables the filter instead of throwing."],"tags":["parquet","dictionary","filter","row-group"],"backgroundTag":"internal-invariant-violation","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}