apache/iceberg · error · IllegalStateException
Unable to acquire lock
Error message
Unable to acquire lock
What it means
LockManagers' base implementation throws this IllegalStateException when a lock acquisition attempt never succeeds within the retry/timeout budget but no acquire callback ever returned true. It signals the caller could not become the lock owner for the given entityId, so the protected operation must not proceed.
Solutions
- Check the lock manager configuration (type, heartbeat interval, timeouts) and correct it
- Verify no other process holds the lock for the same entityId; release stale locks
- Increase acquire timeout/heartbeat intervals in the table properties
- Use a different LockManager implementation suited to your catalog
Example fix
// before
TableMetadata locked = commitWithLock(table); // fails when lock not acquired
// after
long timeoutMs = 300_000;
try (LockManager lock = LockManagers.from(properties)) {
lock.acquire("table-id", "owner-1"); // with larger acquire-timeout-ms property
} Defensive patterns
Strategy: retry
Validate before calling
// before acquiring
if (!lockManagerProperties.containsKey("acquire.timeout-ms") ||
Integer.parseInt(lockManagerProperties.get("acquire.timeout-ms")) < 1000) {
throw new IllegalArgumentException("acquire.timeout-ms must be configured with a sane value");
} Try / catch
try { lock.acquire(entityId, ownerId); } catch (IllegalStateException e) { /* check lock holder, backoff and retry */ } Prevention
- Configure lock.acquire.timeout-ms and heartbeat intervals explicitly
- Monitor for stale locks in the metastore
- Ensure unique ownerId per client instance
When it happens
Trigger: Calling acquire(entityId, ownerId) (via acquireOnce) on a lock manager whose underlying acquire call keeps failing or timing out — e.g. Hive/InMemory lock contention, wrong lock table configuration, or a stuck heartbeat.
Common situations: Two concurrent Spark/Flink jobs committing to the same table with a misconfigured Hive metastore lock; expired heartbeat threads; lock manager pointing at an unreachable metastore.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- Fail to acquire lock
- Cannot commit: Base metadata location
- Cannot commit changes based on stale table metadata
- Cannot commit because base metadata location ' ' is not…
- Cannot commit because Glue detected concurrent update
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/4a8c1e38de87a377.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/util/LockManagers.java:234
scheduler()
.scheduleAtFixedRate(
() -> {
InMemoryLockContent lastContent = LOCKS.get(entityId);
try {
long newExpiration = System.currentTimeMillis() + heartbeatTimeoutMs();
LOCKS.replace(
entityId, lastContent, new InMemoryLockContent(ownerId, newExpiration));
} catch (NullPointerException e) {
throw new RuntimeException(
"Cannot heartbeat to a deleted lock " + entityId, e);
}
},
0,
heartbeatIntervalMs(),
TimeUnit.MILLISECONDS));
} else {
throw new IllegalStateException("Unable to acquire lock " + entityId);
}
}
@Override
public boolean acquire(String entityId, String ownerId) {
try {
Tasks.foreach(entityId)
.retry(Integer.MAX_VALUE - 1)
.onlyRetryOn(IllegalStateException.class)
.throwFailureWhenFinished()
.exponentialBackoff(acquireIntervalMs(), acquireIntervalMs(), acquireTimeoutMs(), 1)
.run(id -> acquireOnce(id, ownerId));
return true;
} catch (IllegalStateException e) {
return false;
}
}
View on GitHub (pinned to 86d9c8fc54)