apache/iceberg · error
Planning failed for plan ID
Error message
Planning failed for plan ID: {} What it means
RESTTableScan polls the REST catalog for an asynchronous server-side plan result; each poll that sees an incomplete plan throws NotCompleteException and Tasks retries with exponential backoff. When retries are exhausted (plan not completed within max wait) this warning is logged, plan resources are cleaned up, and the failure is rethrown.
Solutions
- Increase table property 'rest.plan.max-wait-ms' (and check MIN/MAX polling intervals) to allow more time.
- Investigate REST catalog server load/health and its planning backend latency.
- Retry the query after cleanup; the plan resources were released.
- Disable server-side planning if the catalog does not reliably support it (use default client-side planning).
Example fix
// before
ALTER TABLE t SET TBLPROPERTIES ('rest.plan.max-wait-ms'='5000');
// after
ALTER TABLE t SET TBLPROPERTIES ('rest.plan.max-wait-ms'='60000'); Defensive patterns
Strategy: retry
Validate before calling
long waitMs = Long.parseLong(table.properties().getOrDefault("rest.plan.max-wait-ms", "30000"));
if (tableSizeEstimateLarge && waitMs < 60000) { /* raise wait or avoid server-side planning */ } Try / catch
try { closeableIterable = scan.planFiles(); } catch (NotCompleteException | RESTException e) { LOG.warn("plan did not finish in time; retrying", e); scan.planFiles(); } Prevention
- Raise rest.plan.max-wait-ms for large tables
- Monitor catalog server planning latency
- Fall back to client-side planning if server planning is unreliable
When it happens
Trigger: Executing a scan against a catalog using server-side planning (planning-enabled) where the plan does not complete within rest.plan.max-wait-ms and MAX_RETRIES polls.
Common situations: Overloaded REST catalog or slow underlying object store, network latency between client and server, too-small rest.plan.max-wait-ms for large tables, server having evicted/dropped the plan resources.
Understand the failure class
Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.
Related errors
- Remote scan planning for planId
- Server error
- Cannot call commit on temporary table operations
- Cannot call refresh on temporary table operations
- Cannot commit due to unexpected exception
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/e11af7030521f866.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/rest/RESTTableScan.java:269
PropertyUtil.propertyAsLong(
catalogProperties,
RESTCatalogProperties.REST_SCAN_PLANNING_POLL_TIMEOUT_MS,
RESTCatalogProperties.REST_SCAN_PLANNING_POLL_TIMEOUT_MS_DEFAULT);
Preconditions.checkArgument(
maxWaitTimeMs > 0,
"Invalid value for %s: %s (must be positive)",
RESTCatalogProperties.REST_SCAN_PLANNING_POLL_TIMEOUT_MS,
maxWaitTimeMs);
AtomicReference<FetchPlanningResultResponse> result = new AtomicReference<>();
try {
Tasks.foreach(planId)
.exponentialBackoff(MIN_SLEEP_MS, MAX_SLEEP_MS, maxWaitTimeMs, SCALE_FACTOR)
.retry(MAX_RETRIES)
.onlyRetryOn(NotCompleteException.class)
.onFailure(
(id, err) -> {
LOG.warn("Planning failed for plan ID: {}", id, err);
cleanupPlanResources();
})
.throwFailureWhenFinished()
.run(
id -> {
FetchPlanningResultResponse response =
client.get(
resourcePaths.plan(tableIdentifier, id),
headers,
FetchPlanningResultResponse.class,
headers,
ErrorHandlers.planErrorHandler(),
parserContext);
switch (response.planStatus()) {
case COMPLETED:
result.set(response);
break;View on GitHub (pinned to 86d9c8fc54)