apache/iceberg · error · IllegalStateException
Failed to read required bloom filter for id
Error message
Failed to read required bloom filter for id: %s
What it means
ParquetBloomRowGroupFilter.loadBloomFilter calls the Parquet bloom filter reader for a column chunk, which returned null even though the filter is required for row-group pruning. The library throws IllegalStateException because proceeding would mean evaluating the predicate without an expected bloom filter, risking incorrect skip/keep decisions.
Solutions
- Verify the Parquet file integrity (read the footer with parquet-tools and inspect the bloom filter offsets)
- Rewrite/compact the table so all row groups are produced by a writer that actually emits bloom filters
- Disable bloom-filter-based filtering (remove/adjust write.parquet.bloom-filter-enabled.column.* or the filter pushdown that requires it) or upgrade the Iceberg version, as a null filter should be treated as absent
Example fix
// before
BloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);
// after (library-side guard, upgrade to fixed version):
BloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);
if (bloomFilter == null) { return null; /* treat as absent, don't prune */ } Defensive patterns
Strategy: try-catch
Validate before calling
// before scanning, confirm file integrity
ParquetMetadata footer = ParquetFileReader.readFooter(conf, path);
boolean bloomOffsetsValid = footer.getBlocks().stream()
.flatMap(b -> b.getColumns().stream())
.allMatch(cc -> cc.getBloomFilterOffset() >= 0 || cc.getBloomFilterOffset() == -1); Try / catch
try { parquetTask.execute(); } catch (IllegalStateException e) { if (e.getMessage().startsWith("Failed to read required bloom filter")) { // rewrite file or retry with bloom filtering disabled
conf.set("read.parquet.bloom-filter-enabled", "false"); retry(); } else { throw e; } } Prevention
- Keep Iceberg/parquet-mr versions aligned on both write and read paths
- Compact tables after partial or interrupted writes to remove suspect row groups
- Don't enable bloom filters only on some writers while expecting them on read
- Validate files after crash recovery before running predicate-pushdown scans
When it happens
Trigger: Row-group filtering with a predicate on a column whose bloom filter is expected (bloom filter enabled, column chunk metadata present) but readBloomFilter returns null — e.g., a corrupted footer/chunk, truncated file, or a writer that did not actually serialize the filter.
Common situations: Files written with bloom filters partially disabled or by buggy writers; damaged/corrupt Parquet files; mismatch between table property write.parquet.bloom-filter-enabled and actual file contents.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Failed to create reader for dictionary page
- Skipping bloom filter config for missing field
- AlwaysFalse is a placeholder only
- AlwaysTrue is a placeholder only
- Avro writer does not support variant types
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/2ad8c7a9849d3dad.
Report an issue: GitHub.
Appendix: source
Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ParquetBloomRowGroupFilter.java:270
public <T> Boolean startsWith(BoundReference<T> ref, Literal<T> lit) {
// bloom filter is based on hash and cannot eliminate based on startsWith
return ROWS_MIGHT_MATCH;
}
@Override
public <T> Boolean notStartsWith(BoundReference<T> ref, Literal<T> lit) {
// bloom filter is based on hash and cannot eliminate based on startsWith
return ROWS_MIGHT_MATCH;
}
private BloomFilter loadBloomFilter(int id) {
if (bloomCache.containsKey(id)) {
return bloomCache.get(id);
} else {
ColumnChunkMetaData columnChunkMetaData = columnMetaMap.get(id);
BloomFilter bloomFilter = bloomReader.readBloomFilter(columnChunkMetaData);
if (bloomFilter == null) {
throw new IllegalStateException("Failed to read required bloom filter for id: " + id);
} else {
bloomCache.put(id, bloomFilter);
}
return bloomFilter;
}
}
private <T> boolean shouldRead(
PrimitiveType primitiveType, T value, BloomFilter bloom, Type type) {
long hashValue;
switch (primitiveType.getPrimitiveTypeName()) {
case INT32:
switch (type.typeId()) {
case DECIMAL:
BigDecimal decimalValue = (BigDecimal) value;
hashValue = bloom.hash(decimalValue.unscaledValue().intValue());
return bloom.findHash(hashValue);View on GitHub (pinned to 86d9c8fc54)