apache/iceberg · error · IllegalStateException

Failed to read required bloom filter for id

Error message

Failed to read required bloom filter for id: %s

What it means

ParquetBloomRowGroupFilter.loadBloomFilter calls the Parquet bloom filter reader for a column chunk, which returned null even though the filter is required for row-group pruning. The library throws IllegalStateException because proceeding would mean evaluating the predicate without an expected bloom filter, risking incorrect skip/keep decisions.

Solutions

  1. Verify the Parquet file integrity (read the footer with parquet-tools and inspect the bloom filter offsets)
  2. Rewrite/compact the table so all row groups are produced by a writer that actually emits bloom filters
  3. Disable bloom-filter-based filtering (remove/adjust write.parquet.bloom-filter-enabled.column.* or the filter pushdown that requires it) or upgrade the Iceberg version, as a null filter should be treated as absent

Example fix

// before
BloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);
// after (library-side guard, upgrade to fixed version):
BloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);
if (bloomFilter == null) { return null; /* treat as absent, don't prune */ }
Defensive patterns

Strategy: try-catch

Validate before calling

// before scanning, confirm file integrity
ParquetMetadata footer = ParquetFileReader.readFooter(conf, path);
boolean bloomOffsetsValid = footer.getBlocks().stream()
    .flatMap(b -> b.getColumns().stream())
    .allMatch(cc -> cc.getBloomFilterOffset() >= 0 || cc.getBloomFilterOffset() == -1);

Try / catch

try { parquetTask.execute(); } catch (IllegalStateException e) { if (e.getMessage().startsWith("Failed to read required bloom filter")) { // rewrite file or retry with bloom filtering disabled
  conf.set("read.parquet.bloom-filter-enabled", "false"); retry(); } else { throw e; } }

Prevention

When it happens

Trigger: Row-group filtering with a predicate on a column whose bloom filter is expected (bloom filter enabled, column chunk metadata present) but readBloomFilter returns null — e.g., a corrupted footer/chunk, truncated file, or a writer that did not actually serialize the filter.

Common situations: Files written with bloom filters partially disabled or by buggy writers; damaged/corrupt Parquet files; mismatch between table property write.parquet.bloom-filter-enabled and actual file contents.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/2ad8c7a9849d3dad. Report an issue: GitHub.

Appendix: source

Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ParquetBloomRowGroupFilter.java:270

    public <T> Boolean startsWith(BoundReference<T> ref, Literal<T> lit) {
      // bloom filter is based on hash and cannot eliminate based on startsWith
      return ROWS_MIGHT_MATCH;
    }

    @Override
    public <T> Boolean notStartsWith(BoundReference<T> ref, Literal<T> lit) {
      // bloom filter is based on hash and cannot eliminate based on startsWith
      return ROWS_MIGHT_MATCH;
    }

    private BloomFilter loadBloomFilter(int id) {
      if (bloomCache.containsKey(id)) {
        return bloomCache.get(id);
      } else {
        ColumnChunkMetaData columnChunkMetaData = columnMetaMap.get(id);
        BloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);
        if (bloomFilter == null) {
          throw new IllegalStateException("Failed to read required bloom filter for id: " + id);
        } else {
          bloomCache.put(id, bloomFilter);
        }

        return bloomFilter;
      }
    }

    private <T> boolean shouldRead(
        PrimitiveType primitiveType, T value, BloomFilter bloom, Type type) {
      long hashValue;
      switch (primitiveType.getPrimitiveTypeName()) {
        case INT32:
          switch (type.typeId()) {
            case DECIMAL:
              BigDecimal decimalValue = (BigDecimal) value;
              hashValue = bloom.hash(decimalValue.unscaledValue().intValue());
              return bloom.findHash(hashValue);

View on GitHub (pinned to 86d9c8fc54)