apache/seatunnel · warning
Collect worker resource snapshot failed
Error message
Collect worker resource snapshot failed: {} What it means
PendingDiagnosticsCollector.collectWorkerResourceSnapshot builds a per-worker resource diagnostic projection from the master's ResourceManager. On exception, it logs 'Collect worker resource snapshot failed', marks the snapshot unavailable, and returns it so the collector can still emit a partial cluster snapshot.
Solutions
- Check the exception message to identify which worker failed to report.
- Confirm the worker process is running and reachable; restart it if dead.
- Re-run diagnostics after the cluster stabilizes to capture the full worker snapshot.
- If only some workers appear in snapshots, check master-worker RPC timeouts and network stability.
Example fix
// before: diagnose with missing worker data // after: restore worker then re-collect bin/seatunnel.sh --shutdown-worker <address> ; bin/seatunnel.sh --start-worker // then rerun pending diagnostics
Defensive patterns
Strategy: try-catch
Validate before calling
// Filter null/incomplete worker profiles before building the snapshot
List<WorkerProfile> profiles = workers.stream().filter(Objects::nonNull).collect(Collectors.toList());
if (profiles.isEmpty()) { log.warn("no worker profiles available; snapshot will be marked unavailable"); } Try / catch
try {
snapshot = collector.collectWorkerResourceSnapshot(resourceManager);
if (!snapshot.isAvailable()) {
// re-run later or fall back to last good snapshot
}
} catch (Exception e) {
log.warn("diagnostic collection degraded: {}", e.getMessage());
} Prevention
- Monitor worker liveness so resource queries don't hit dead nodes.
- Avoid running diagnostics during worker scale events.
- Treat snapshot.available=false as a signal to re-collect later.
- Increase master-worker RPC timeout if profiles are large or networks are slow.
When it happens
Trigger: Called from collectClusterSnapshot; fails when iterating worker profiles throws — a worker disconnected mid-iteration, its WorkerProfile is null/incomplete, or an RPC to fetch worker resource info fails.
Common situations: Flapping workers during scale-up/down, network partitions isolating a worker from the master, or running diagnostics on an already-degraded cluster where some workers cannot answer resource queries.
Related errors
- Collect worker count failed
- Close scanner from failed.
- CloseScanner result from
- Error scanning data from region.
- Failed to get checkpoint data from master node…
AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10).
Data as JSON: /api/errors/463f875be648757a.
Report an issue: GitHub.
Appendix: source
Thrown at seatunnel-engine/seatunnel-engine-server/src/main/java/org/apache/seatunnel/engine/server/diagnostic/PendingDiagnosticsCollector.java:296
snapshot.setCollectedAt(System.currentTimeMillis());
if (resourceManager == null) {
return snapshot;
}
try {
Map<Address, WorkerProfile> registerWorker = resourceManager.getRegisterWorker();
if (registerWorker == null) {
return snapshot;
}
List<WorkerResourceDiagnostic> workers =
registerWorker.entrySet().stream()
.map(entry -> convertWorker(entry.getKey(), entry.getValue()))
.filter(worker -> worker != null)
.sorted(Comparator.comparing(WorkerResourceDiagnostic::getAddress))
.collect(Collectors.toList());
snapshot.setAvailable(true);
snapshot.setWorkers(workers);
} catch (Exception e) {
log.warn("Collect worker resource snapshot failed: {}", ExceptionUtils.getMessage(e));
}
return snapshot;
}
private static WorkerResourceDiagnostic convertWorker(
Address registeredAddress, WorkerProfile workerProfile) {
if (workerProfile == null) {
return null;
}
WorkerResourceDiagnostic diagnostic = new WorkerResourceDiagnostic();
Address address = workerProfile.getAddress();
if (address == null) {
address = registeredAddress;
}
diagnostic.setAddress(address == null ? "UNKNOWN" : address.toString());
if (workerProfile.getAttributes() != null) {
diagnostic.setTags(new HashMap<>(workerProfile.getAttributes()));
} else {View on GitHub (pinned to cf67b549a7)