apache/iceberg · critical · RuntimeException
Failed to discover new splits
Error message
Failed to discover new splits
What it means
ContinuousIcebergEnumerator.processDiscoveredSplits wraps the underlying planning error in a RuntimeException once the number of consecutive incremental-scan failures exceeds scanContext.maxAllowedPlanningFailures. Until the threshold is reached the error is only logged, after which the enumerator fails the source instead of silently retrying forever.
Solutions
- Inspect the cause (the wrapped 'error') for the real root cause — catalog/auth/IO failure.
- Increase 'max_allowed_planning-failures' (or the scanContext's maxAllowedPlanningFailures) to tolerate transient planning failures.
- Verify catalog credentials and that the table location is reachable from the job manager.
- Check whether snapshot expiration is concurrently deleting snapshots the planner needs; coordinate expiry schedules.
Example fix
// before
TableLoader loader = ...; // default low failure tolerance
// after: tolerate transient failures
.tableLoader(loader)
.monitorInterval(Duration.ofSeconds(30))
.streaming(true)
// and set property
properties.put("max_allowed_planning-failures", "10"); Defensive patterns
Strategy: retry
Validate before calling
// validate catalog access before starting the job
tableLoader.open(); Table table = tableLoader.loadTable();
table.refresh(); // throws early if metadata is unreachable
// configure tolerated failures
long failures = Long.parseLong(properties.getProperty("max_allowed_planning-failures", "-1")); Prevention
- Set max_allowed_planning-failures generously for streaming jobs.
- Avoid scheduling snapshot expiration while streaming scans are active on the same table.
- Monitor catalog/object-store health and credentials expiry.
When it happens
Trigger: Continuous (streaming) enumeration repeatedly fails split discovery (table metadata reads, manifest scans) and consecutiveFailures exceeds maxAllowedPlanningFailures (default unlimited when negative).
Common situations: Table metadata unreadable due to permissions/object-store outages; concurrent table expiry deleting snapshots mid-scan; misconfigured catalog credentials; setting maxAllowedPlanningFailures to a small value like 0-3 makes transient outages fatal.
Related errors
- Failed to plan files for main index
- [For table with [ ] at ]: Failed to plan data file rewrite…
- [For table with [ ] at ]: Failed to plan data file rewrite…
- Received invalid default split request event from subtask
- Unsupported status
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/10f0e1ba1f522574.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.1/flink/src/main/java/org/apache/iceberg/flink/source/enumerator/ContinuousIcebergEnumerator.java:184
result.toPosition());
} else {
LOG.info(
"No new splits discovered between ({}, {}]",
result.fromPosition(),
result.toPosition());
}
// update the enumerator position even if there is no split discovered
// or the toPosition is empty (e.g. for empty table).
enumeratorPosition.set(result.toPosition());
LOG.info("Update enumerator position to {}", result.toPosition());
}
} else {
consecutiveFailures++;
if (scanContext.maxAllowedPlanningFailures() < 0
|| consecutiveFailures <= scanContext.maxAllowedPlanningFailures()) {
LOG.error("Failed to discover new splits", error);
} else {
throw new RuntimeException("Failed to discover new splits", error);
}
}
}
}
View on GitHub (pinned to 86d9c8fc54)