apache/iceberg · warning
Failed to close ZkLockFactory for lockId
Error message
Failed to close ZkLockFactory for lockId: {} What it means
ZkLockFactory.closeQuietly is called from open() when initialization partially fails: it attempts close() and, if closing the ZooKeeper-based lock factory throws, logs this warning with the lockId instead of propagating. The original failure that triggered the cleanup remains the primary error; this warning only reports that resource cleanup also failed.
Solutions
- Find and fix the original exception that caused open() to fail (usually a ZooKeeper connectivity or configuration problem).
- Verify ZooKeeper connect string, session timeout, and that the ZK ensemble is reachable from the Flink cluster.
- Ensure the Curator client is not closed twice (check lifecycle management of ZkLockFactory in the maintenance job).
Example fix
// before
lockFactoryBuilder = ZkLockFactory.builder()
.setZkServers("wrong-host:2181")
.setLockId("maintenance");
// after
lockFactoryBuilder = ZkLockFactory.builder()
.setZkServers("zk1:2181,zk2:2181,zk3:2181")
.setLockId("maintenance"); Defensive patterns
Strategy: try-catch
Validate before calling
// Pre-flight: confirm ZK reachable before opening the factory
try (Socket s = new Socket()) {
s.connect(new InetSocketAddress(zkHost, zkPort), 5000);
} Try / catch
try {
lockFactory.open();
} catch (Exception e) {
LOG.error("ZkLockFactory init failed (see also 'Failed to close ZkLockFactory' if cleanup also failed)", e);
throw e;
} Prevention
- Validate the ZooKeeper connect string and reachability before starting the maintenance job.
- Avoid double-closing the Curator client; manage lifecycle in one place.
- Monitor the 'Failed to close ZkLockFactory' warning as a symptom of a deeper init failure.
When it happens
Trigger: ZkLockFactory.open() (or initialization) fails partway and calls closeQuietly(); close() throws (e.g. Curator/ZooKeeper client already closed or connection broken), producing this warning.
Common situations: ZooKeeper unreachable at startup so factory init fails and cleanup also throws; double-close of the Curator framework client; timeout during shutdown.
Related errors
- Failed to close ZkLockFactory for lockId
- Connection to Zookeeper timed out
- Connection to Zookeeper timed out
- Failed to acquire Zookeeper lock
- Failed to acquire Zookeeper lock
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/ba185fb434799357.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/maintenance/api/ZkLockFactory.java:156
} catch (Exception e) {
closeQuietly();
throw new RuntimeException("Failed to initialize SharedCount", e);
}
}
private String getTaskSharePath() {
return LOCK_BASE_PATH + lockId + "/task";
}
private String getRecoverySharedPath() {
return LOCK_BASE_PATH + lockId + "/recovery";
}
private void closeQuietly() {
try {
close();
} catch (Exception e) {
LOG.warn("Failed to close ZkLockFactory for lockId: {}", lockId, e);
}
}
@Override
public Lock createLock() {
return new ZkLock(getTaskSharePath(), taskSharedCount);
}
@Override
public Lock createRecoveryLock() {
return new ZkLock(getRecoverySharedPath(), recoverySharedCount);
}
@Override
public void close() throws IOException {
try {
if (taskSharedCount != null) {
taskSharedCount.close();View on GitHub (pinned to 86d9c8fc54)