{"record":{"id":"6586fc2993d15dad","repo":"apache/iceberg","slug":"failed-to-create-reader-for-dictionary-page","errorCode":null,"errorMessage":"Failed to create reader for dictionary page","messagePattern":"Failed to create reader for dictionary page","errorType":"exception","errorClass":"RuntimeIOException","httpStatus":null,"severity":"error","filePath":"parquet/src/main/java/org/apache/iceberg/parquet/ParquetDictionaryRowGroupFilter.java","lineNumber":438,"sourceCode":"      Set<?> cached = dictCache.get(id);\n      if (cached != null) {\n        return (Set<T>) cached;\n      }\n\n      ColumnDescriptor col = cols.get(id);\n      DictionaryPage page = dictionaries.readDictionaryPage(col);\n      // may not be dictionary-encoded\n      if (page == null) {\n        throw new IllegalStateException(\"Failed to read required dictionary page for id: \" + id);\n      }\n\n      Function<Object, Object> conversion = conversions.get(id);\n\n      Dictionary dict;\n      try {\n        dict = page.getEncoding().initDictionary(col, page);\n      } catch (IOException e) {\n        throw new RuntimeIOException(\"Failed to create reader for dictionary page\");\n      }\n\n      Set<T> dictSet = Sets.newTreeSet(comparator);\n\n      for (int i = 0; i <= dict.getMaxId(); i++) {\n        switch (col.getPrimitiveType().getPrimitiveTypeName()) {\n          case FIXED_LEN_BYTE_ARRAY:\n            dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));\n            break;\n          case BINARY:\n            dictSet.add((T) conversion.apply(dict.decodeToBinary(i)));\n            break;\n          case INT32:\n            dictSet.add((T) conversion.apply(dict.decodeToInt(i)));\n            break;\n          case INT64:\n            dictSet.add((T) conversion.apply(dict.decodeToLong(i)));\n            break;","sourceCodeStart":420,"sourceCodeEnd":456,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/parquet/src/main/java/org/apache/iceberg/parquet/ParquetDictionaryRowGroupFilter.java#L420-L456","documentation":"Thrown as a RuntimeIOException when initializing the dictionary reader for a column's dictionary page (page.getEncoding().initDictionary) throws IOException. The dictionary page exists but its content cannot be decoded into a usable Dictionary object for the column's encoding.","triggerScenarios":"Corrupt or truncated dictionary page bytes in the Parquet file; an encoding whose initDictionary fails for the given column descriptor (e.g. mismatched page encoding vs column type); underlying file/IO errors surfaced as IOException from the page source.","commonSituations":"Reading damaged files (network truncation, partial upload); files written by buggy or exotic encoders; encodings not matching expected initialization.","solutions":["Verify/re-download the Parquet file and check checksums — corruption is the most common cause.","Rewrite the file with a standard tool (Spark/Iceberg) to regenerate valid dictionary pages.","Catch RuntimeIOException in the filter path and fall back to full record evaluation.","Upgrade Iceberg/parquet-mr in case of an encoding initialization bug."],"exampleFix":"// before\ndict = page.getEncoding().initDictionary(col, page);\n// after\ntry {\n  dict = page.getEncoding().initDictionary(col, page);\n} catch (RuntimeIOException e) {\n  return null; // fall back to non-dictionary filtering\n}","handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"try {\n  result = filter.shouldRead(rowGroup, ...);\n} catch (RuntimeIOException e) {\n  LOG.warn(\"Dictionary page unreadable, evaluating rows directly\", e);\n  return true; // conservative: evaluate rows\n}","preventionTips":["Validate file integrity (checksums, complete uploads) before reading Parquet data.","Avoid partial/truncated file sources; retry downloads on failure.","Keep Iceberg/parquet-mr patched for encoding initialization bugs."],"tags":["parquet","dictionary","io","corruption"],"backgroundTag":"file-read-failed","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}