{"record":{"id":"49371d621ddc2f55","repo":"apache/iceberg","slug":"failed-to-heartbeat-for-hive-lock-while-committing","errorCode":null,"errorMessage":"Failed to heartbeat for hive lock while committing changes. This can lead to a concurrent commit attempt be able to overwrite this commit. Please check the commit history. If you are running into this issue, try reducing iceberg.hive.lock-heartbeat-interval-ms.","messagePattern":"Failed to heartbeat for hive lock while committing changes\\. This can lead to a concurrent commit attempt be able to overwrite this commit\\. Please check the commit history\\. If you are running into this issue, try reducing iceberg\\.hive\\.lock-heartbeat-interval-ms\\.","errorType":"exception","errorClass":"CommitStateUnknownException","httpStatus":null,"severity":"critical","filePath":"hive-metastore/src/main/java/org/apache/iceberg/hive/HiveTableOperations.java","lineNumber":355,"sourceCode":"          maxHiveTablePropertySize,\n          currentMetadataLocation());\n\n      if (!keepHiveStats) {\n        tbl.getParameters().remove(StatsSetupConst.COLUMN_STATS_ACCURATE);\n        tbl.getParameters().put(StatsSetupConst.DO_NOT_UPDATE_STATS, StatsSetupConst.TRUE);\n      }\n\n      lock.ensureActive();\n\n      try {\n        persistTable(\n            tbl, updateHiveTable, hiveLockEnabled(base, conf) ? null : baseMetadataLocation);\n        lock.ensureActive();\n\n        commitStatus = BaseMetastoreOperations.CommitStatus.SUCCESS;\n      } catch (LockException le) {\n        commitStatus = BaseMetastoreOperations.CommitStatus.UNKNOWN;\n        throw new CommitStateUnknownException(\n            \"Failed to heartbeat for hive lock while \"\n                + \"committing changes. This can lead to a concurrent commit attempt be able to overwrite this commit. \"\n                + \"Please check the commit history. If you are running into this issue, try reducing \"\n                + \"iceberg.hive.lock-heartbeat-interval-ms.\",\n            le);\n      } catch (org.apache.hadoop.hive.metastore.api.AlreadyExistsException e) {\n        throw new AlreadyExistsException(e, \"Table already exists: %s.%s\", database, tableName);\n\n      } catch (InvalidObjectException e) {\n        throw new ValidationException(e, \"Invalid Hive object for %s.%s\", database, tableName);\n\n      } catch (CommitFailedException | CommitStateUnknownException e) {\n        throw e;\n\n      } catch (Throwable e) {\n        if (e.getMessage() != null\n            && e.getMessage().contains(\"Table/View 'HIVE_LOCKS' does not exist\")) {\n          throw new RuntimeException(","sourceCodeStart":337,"sourceCodeEnd":373,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/hive-metastore/src/main/java/org/apache/iceberg/hive/HiveTableOperations.java#L337-L373","documentation":"While committing, the periodic heartbeat that keeps the acquired Hive lock alive failed (LockException), so the lock may have expired and another committer could overwrite this commit. The commit's outcome is genuinely unknown, so a CommitStateUnknownException is thrown instead of a retriable CommitFailedException.","triggerScenarios":"doCommit acquires a Hive lock (hive lock enabled) and the heartbeat (interval iceberg.hive.lock-heartbeat-interval-ms) fails inside persistTable due to HMS connection loss, HMS restart, GC pause longer than lock expiry, or heartbeat thread interruption.","commonSituations":"Long commits with a short heartbeat interval; flaky network to HMS; embedded/underprovisioned metastore dropping the lock table connection; long pauses (VM/swap) exceeding the lock timeout.","solutions":["Manually inspect the table's metadata_location / commit history to determine whether the commit landed before retrying","Increase iceberg.hive.lock-heartbeat-interval-ms headroom relative to commit duration and fix HMS connectivity","Ensure the lock's heartbeat thread is not starved (adequate client threads, no long GC pauses)","Use a metastore with reliable transactional locking (not embedded Derby)"],"exampleFix":"// before\nconf.set(\"iceberg.hive.lock-heartbeat-interval-ms\", \"3000\"); // too aggressive\n// after\nconf.set(\"iceberg.hive.lock-heartbeat-interval-ms\", \"60000\"); // tolerate long commits","handlingStrategy":"try-catch","validationCode":"// verify lock health before commit\nHiveLock lock = ...; if (lock == null || !lock.isActive()) { reAcquireLock(); }","typeGuard":null,"tryCatchPattern":"try {\n  table.newAppend().appendFile(f).commit();\n} catch (CommitStateUnknownException e) {\n  // do NOT blindly retry; check whether commit landed\n  Table refreshed = catalog.loadTable(ident);\n  if (!alreadyContains(refreshed, f)) { refreshed.newAppend().appendFile(f).commit(); }\n}","preventionTips":["Set a generous iceberg.hive.lock-heartbeat-interval-ms","Use a reliable remote HMS, not an embedded one","Monitor GC pauses and network stability toward HMS during writes","After CommitStateUnknownException, always inspect commit history before re-committing"],"tags":["hive-metastore","locking","heartbeat","commit-state-unknown"],"backgroundTag":"network-request-failed","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}