apache/seatunnel · warning
running-jobs summary slow: total=
Error message
running-jobs summary slow: total={}ms jobs={} decode={}ms What it means
getRunningJobsSummaryJson measures how long building the running-jobs summary takes. When total time exceeds 500ms it emits this warn with total duration, job count, and cumulative JSON decode time, so operators can see the summary endpoint is degrading.
Solutions
- Reduce running job count or poll frequency; cache summary responses on the client.
- Scale the master/REST node with more CPU/heap and tune GC to avoid long pauses.
- Profile the decode path; ensure runningJobBasicInfoCache is being used and not repeatedly decoding the same payloads.
- Monitor via the accompanying diagnostics log (errorIndex 3714) to distinguish decode cost from total cost.
Example fix
// before: polling every second with hundreds of jobs // curl http://host:8080/running-jobs/summary (every 1s) // after: back off and cache // every 10s -> fetch summary, cache locally, reuse between polls
Defensive patterns
Strategy: fallback
Validate before calling
// client: back off if previous summary took > 500ms
// if (lastSummaryMs > 500) { delay(nextPoll, backoffMs); } Prevention
- Poll the summary endpoint at a sane interval; cache results
- Watch job count; scale the cluster or limit concurrent running jobs
- Keep REST node heap/GC healthy
When it happens
Trigger: REST GET of the running-jobs summary endpoint while there are many running jobs, large per-job basic-info JSON payloads needing decode, or the JVM is CPU/GC-starved causing slow map iteration and JSON parsing.
Common situations: Hundreds of concurrently running jobs on a cluster; undersized REST node heap or heavy GC pauses; slow disk/network backing the Hazelcast maps; burst of REST dashboard polling amplifying load.
Understand the failure class
Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.
Related errors
- Failed to get HTTP port from member
- Failed to load finished job DAG for job
- Failed to load finished job metrics for job
- GET /overview dispatch delayed: dispatchDelayMs=
- GET /running-jobs dispatch delayed: dispatchDelayMs=
AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10).
Data as JSON: /api/errors/fdaff4fa907f6dc8.
Report an issue: GitHub.
Appendix: source
Thrown at seatunnel-engine/seatunnel-engine-server/src/main/java/org/apache/seatunnel/engine/server/rest/service/JobInfoService.java:249
} catch (Throwable ignored) {
// ignore
}
out.add(
new JsonObject()
.add(RestConstant.JOB_ID, String.valueOf(jobId))
.add(RestConstant.JOB_NAME, jobName)
.add(RestConstant.JOB_STATUS, jobStatus)
.add(RestConstant.CREATE_TIME, createTime));
}
if (!runningJobBasicInfoCache.isEmpty() && !jobIds.isEmpty()) {
runningJobBasicInfoCache.keySet().removeIf(id -> !jobIds.contains(id));
}
long totalMs = TimeUnit.NANOSECONDS.toMillis(System.nanoTime() - startNs);
if (totalMs > 500) {
log.warn(
"running-jobs summary slow: total={}ms jobs={} decode={}ms",
totalMs,
values.size(),
TimeUnit.NANOSECONDS.toMillis(decodeTotalNs));
Runtime rt = Runtime.getRuntime();
long usedBytes = rt.totalMemory() - rt.freeMemory();
log.warn(
"running-jobs summary slow diagnostics: totalMs={} jobs={} decodeMs={} "
+ "heapUsedMB={} heapTotalMB={} heapMaxMB={}",
totalMs,
values.size(),
TimeUnit.NANOSECONDS.toMillis(decodeTotalNs),
usedBytes / 1024 / 1024,
rt.totalMemory() / 1024 / 1024,
rt.maxMemory() / 1024 / 1024);
} else {
log.debug(
"running-jobs summary: total={}ms jobs={} decode={}ms",View on GitHub (pinned to cf67b549a7)