apache/iceberg · critical · RuntimeException

Failed to discover new splits

Error message

Failed to discover new splits

What it means

ContinuousIcebergEnumerator.processDiscoveredSplits wraps the underlying planning error in a RuntimeException once the number of consecutive incremental-scan failures exceeds scanContext.maxAllowedPlanningFailures. Until the threshold is reached the error is only logged, after which the enumerator fails the source instead of silently retrying forever.

Solutions

  1. Inspect the cause (the wrapped 'error') for the real root cause — catalog/auth/IO failure.
  2. Increase 'max_allowed_planning-failures' (or the scanContext's maxAllowedPlanningFailures) to tolerate transient planning failures.
  3. Verify catalog credentials and that the table location is reachable from the job manager.
  4. Check whether snapshot expiration is concurrently deleting snapshots the planner needs; coordinate expiry schedules.

Example fix

// before
TableLoader loader = ...; // default low failure tolerance
// after: tolerate transient failures
.tableLoader(loader)
.monitorInterval(Duration.ofSeconds(30))
.streaming(true)
// and set property
properties.put("max_allowed_planning-failures", "10");
Defensive patterns

Strategy: retry

Validate before calling

// validate catalog access before starting the job
tableLoader.open(); Table table = tableLoader.loadTable();
table.refresh(); // throws early if metadata is unreachable
// configure tolerated failures
long failures = Long.parseLong(properties.getProperty("max_allowed_planning-failures", "-1"));

Prevention

When it happens

Trigger: Continuous (streaming) enumeration repeatedly fails split discovery (table metadata reads, manifest scans) and consecutiveFailures exceeds maxAllowedPlanningFailures (default unlimited when negative).

Common situations: Table metadata unreadable due to permissions/object-store outages; concurrent table expiry deleting snapshots mid-scan; misconfigured catalog credentials; setting maxAllowedPlanningFailures to a small value like 0-3 makes transient outages fatal.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/10f0e1ba1f522574. Report an issue: GitHub.

Appendix: source

Thrown at flink/v2.1/flink/src/main/java/org/apache/iceberg/flink/source/enumerator/ContinuousIcebergEnumerator.java:184

              result.toPosition());
        } else {
          LOG.info(
              "No new splits discovered between ({}, {}]",
              result.fromPosition(),
              result.toPosition());
        }
        // update the enumerator position even if there is no split discovered
        // or the toPosition is empty (e.g. for empty table).
        enumeratorPosition.set(result.toPosition());
        LOG.info("Update enumerator position to {}", result.toPosition());
      }
    } else {
      consecutiveFailures++;
      if (scanContext.maxAllowedPlanningFailures() < 0
          || consecutiveFailures <= scanContext.maxAllowedPlanningFailures()) {
        LOG.error("Failed to discover new splits", error);
      } else {
        throw new RuntimeException("Failed to discover new splits", error);
      }
    }
  }
}

View on GitHub (pinned to 86d9c8fc54)