apache/seatunnel · error · SeaTunnelEngineException
Can not sync pipeline owned slot profiles with IMap
Error message
Can not sync pipeline owned slot profiles with IMap
What it means
JobMaster.syncPipelineOwnedSlotProfiles retries the IMap update of a pipeline's owned slot profiles using RetryUtils, tolerating only transient NullPointerExceptions while the job is running. If all retries (OPERATION_RETRY_TIME) are exhausted, it wraps the last exception in SeaTunnelEngineException 'Can not sync pipeline owned slot profiles with IMap'. The engine cannot persist which slots each pipeline owns, breaking later slot lookups.
Solutions
- Check the wrapped cause ('e') in the exception logs for the real IMap/serialization failure and fix it.
- Retry the job; transient IMap/partition issues during cluster churn often resolve after the cluster stabilizes.
- Verify master node health and Hazelcast connectivity; restart the master if the IMap service is stuck.
- If it recurs during job stop, avoid cancelling jobs while they are still initializing slot profiles.
Defensive patterns
Strategy: try-catch
Try / catch
try {
submitJob(conf);
} catch (SeaTunnelEngineException e) {
if (e.getMessage().contains("Can not sync pipeline owned slot profiles")) {
log.error("slot sync failed; cause:", e.getCause());
// check cluster health before retrying
}
throw e;
} Prevention
- Keep the master node healthy; monitor Hazelcast IMap health and partition migrations.
- Avoid cancelling jobs during initialization while slot profiles are being synced.
- Investigate the wrapped cause immediately - it names the real IMap failure.
When it happens
Trigger: Writing owned slot profiles to the IMap repeatedly fails (NPE while isRunning, or another exception) until the retry budget is exhausted - e.g. IMap unavailability, serialization failures, or job being shut down concurrently during initialization.
Common situations: Cluster under heavy load or during master failover causing IMap operations to fail; job cancelled/stopped mid-initialization making isRunning false so NPE retries stop helping; Hazelcast instance issues (partition migration, connectivity).
Understand the failure class
Background: "This is a bug, please report it": internal invariant violations, unreachable panics, and SNH errors explained — this error's family across 47 libraries.
Related errors
- can't find task group address from taskGroupLocation
- COMMON_SQL_OPERATION_FAILED
- CommonErrorCodeDeprecated.FLUSH_DATA_FAILED
- FLUSH_DATA_FAILED
- Get connection failed after retry times
AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10).
Data as JSON: /api/errors/39df2dc94d0a4204.
Report an issue: GitHub.
Appendix: source
Thrown at seatunnel-engine/seatunnel-engine-server/src/main/java/org/apache/seatunnel/engine/server/master/JobMaster.java:1556
}
}
public void setOwnedSlotProfiles(
@NonNull PipelineLocation pipelineLocation,
@NonNull Map<TaskGroupLocation, SlotProfile> pipelineOwnedSlotProfiles) {
ownedSlotProfilesIMap.put(pipelineLocation, pipelineOwnedSlotProfiles);
try {
RetryUtils.retryWithException(
() ->
pipelineOwnedSlotProfiles.equals(
ownedSlotProfilesIMap.get(pipelineLocation)),
new RetryUtils.RetryMaterial(
Constant.OPERATION_RETRY_TIME,
true,
exception -> exception instanceof NullPointerException && isRunning,
Constant.OPERATION_RETRY_SLEEP));
} catch (Exception e) {
throw new SeaTunnelEngineException(
"Can not sync pipeline owned slot profiles with IMap", e);
}
}
public SlotProfile getOwnedSlotProfiles(@NonNull TaskGroupLocation taskGroupLocation) {
Map<TaskGroupLocation, SlotProfile> taskGroupLocationSlotProfileMap =
ownedSlotProfilesIMap.get(
new PipelineLocation(
taskGroupLocation.getJobId(), taskGroupLocation.getPipelineId()));
if (taskGroupLocationSlotProfileMap == null) {
return null;
}
return taskGroupLocationSlotProfileMap.get(taskGroupLocation);
}
public ExecutorService getExecutorService() {
return executorService;View on GitHub (pinned to cf67b549a7)