{"record":{"id":"2ad8c7a9849d3dad","repo":"apache/iceberg","slug":"failed-to-read-required-bloom-filter-for-id-s","errorCode":null,"errorMessage":"Failed to read required bloom filter for id: %s","messagePattern":"Failed to read required bloom filter for id: (.+?)","errorType":"exception","errorClass":"IllegalStateException","httpStatus":null,"severity":"error","filePath":"parquet/src/main/java/org/apache/iceberg/parquet/ParquetBloomRowGroupFilter.java","lineNumber":270,"sourceCode":"    public <T> Boolean startsWith(BoundReference<T> ref, Literal<T> lit) {\n      // bloom filter is based on hash and cannot eliminate based on startsWith\n      return ROWS_MIGHT_MATCH;\n    }\n\n    @Override\n    public <T> Boolean notStartsWith(BoundReference<T> ref, Literal<T> lit) {\n      // bloom filter is based on hash and cannot eliminate based on startsWith\n      return ROWS_MIGHT_MATCH;\n    }\n\n    private BloomFilter loadBloomFilter(int id) {\n      if (bloomCache.containsKey(id)) {\n        return bloomCache.get(id);\n      } else {\n        ColumnChunkMetaData columnChunkMetaData = columnMetaMap.get(id);\n        BloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);\n        if (bloomFilter == null) {\n          throw new IllegalStateException(\"Failed to read required bloom filter for id: \" + id);\n        } else {\n          bloomCache.put(id, bloomFilter);\n        }\n\n        return bloomFilter;\n      }\n    }\n\n    private <T> boolean shouldRead(\n        PrimitiveType primitiveType, T value, BloomFilter bloom, Type type) {\n      long hashValue;\n      switch (primitiveType.getPrimitiveTypeName()) {\n        case INT32:\n          switch (type.typeId()) {\n            case DECIMAL:\n              BigDecimal decimalValue = (BigDecimal) value;\n              hashValue = bloom.hash(decimalValue.unscaledValue().intValue());\n              return bloom.findHash(hashValue);","sourceCodeStart":252,"sourceCodeEnd":288,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/parquet/src/main/java/org/apache/iceberg/parquet/ParquetBloomRowGroupFilter.java#L252-L288","documentation":"ParquetBloomRowGroupFilter.loadBloomFilter calls the Parquet bloom filter reader for a column chunk, which returned null even though the filter is required for row-group pruning. The library throws IllegalStateException because proceeding would mean evaluating the predicate without an expected bloom filter, risking incorrect skip/keep decisions.","triggerScenarios":"Row-group filtering with a predicate on a column whose bloom filter is expected (bloom filter enabled, column chunk metadata present) but readBloomFilter returns null — e.g., a corrupted footer/chunk, truncated file, or a writer that did not actually serialize the filter.","commonSituations":"Files written with bloom filters partially disabled or by buggy writers; damaged/corrupt Parquet files; mismatch between table property write.parquet.bloom-filter-enabled and actual file contents.","solutions":["Verify the Parquet file integrity (read the footer with parquet-tools and inspect the bloom filter offsets)","Rewrite/compact the table so all row groups are produced by a writer that actually emits bloom filters","Disable bloom-filter-based filtering (remove/adjust write.parquet.bloom-filter-enabled.column.* or the filter pushdown that requires it) or upgrade the Iceberg version, as a null filter should be treated as absent"],"exampleFix":"// before\nBloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);\n// after (library-side guard, upgrade to fixed version):\nBloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);\nif (bloomFilter == null) { return null; /* treat as absent, don't prune */ }","handlingStrategy":"try-catch","validationCode":"// before scanning, confirm file integrity\nParquetMetadata footer = ParquetFileReader.readFooter(conf, path);\nboolean bloomOffsetsValid = footer.getBlocks().stream()\n    .flatMap(b -> b.getColumns().stream())\n    .allMatch(cc -> cc.getBloomFilterOffset() >= 0 || cc.getBloomFilterOffset() == -1);","typeGuard":null,"tryCatchPattern":"try { parquetTask.execute(); } catch (IllegalStateException e) { if (e.getMessage().startsWith(\"Failed to read required bloom filter\")) { // rewrite file or retry with bloom filtering disabled\n  conf.set(\"read.parquet.bloom-filter-enabled\", \"false\"); retry(); } else { throw e; } }","preventionTips":["Keep Iceberg/parquet-mr versions aligned on both write and read paths","Compact tables after partial or interrupted writes to remove suspect row groups","Don't enable bloom filters only on some writers while expecting them on read","Validate files after crash recovery before running predicate-pushdown scans"],"tags":["parquet","bloom-filter","corruption","row-group-filter"],"backgroundTag":"file-read-failed","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}