apache/seatunnel · warning

running-jobs summary slow: total=

Error message

running-jobs summary slow: total={}ms jobs={} decode={}ms

What it means

getRunningJobsSummaryJson measures how long building the running-jobs summary takes. When total time exceeds 500ms it emits this warn with total duration, job count, and cumulative JSON decode time, so operators can see the summary endpoint is degrading.

Solutions

  1. Reduce running job count or poll frequency; cache summary responses on the client.
  2. Scale the master/REST node with more CPU/heap and tune GC to avoid long pauses.
  3. Profile the decode path; ensure runningJobBasicInfoCache is being used and not repeatedly decoding the same payloads.
  4. Monitor via the accompanying diagnostics log (errorIndex 3714) to distinguish decode cost from total cost.

Example fix

// before: polling every second with hundreds of jobs
// curl http://host:8080/running-jobs/summary  (every 1s)
// after: back off and cache
// every 10s -> fetch summary, cache locally, reuse between polls
Defensive patterns

Strategy: fallback

Validate before calling

// client: back off if previous summary took > 500ms
// if (lastSummaryMs > 500) { delay(nextPoll, backoffMs); }

Prevention

When it happens

Trigger: REST GET of the running-jobs summary endpoint while there are many running jobs, large per-job basic-info JSON payloads needing decode, or the JVM is CPU/GC-starved causing slow map iteration and JSON parsing.

Common situations: Hundreds of concurrently running jobs on a cluster; undersized REST node heap or heavy GC pauses; slow disk/network backing the Hazelcast maps; burst of REST dashboard polling amplifying load.

Understand the failure class

Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/fdaff4fa907f6dc8. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-engine/seatunnel-engine-server/src/main/java/org/apache/seatunnel/engine/server/rest/service/JobInfoService.java:249

            } catch (Throwable ignored) {
                // ignore
            }

            out.add(
                    new JsonObject()
                            .add(RestConstant.JOB_ID, String.valueOf(jobId))
                            .add(RestConstant.JOB_NAME, jobName)
                            .add(RestConstant.JOB_STATUS, jobStatus)
                            .add(RestConstant.CREATE_TIME, createTime));
        }

        if (!runningJobBasicInfoCache.isEmpty() && !jobIds.isEmpty()) {
            runningJobBasicInfoCache.keySet().removeIf(id -> !jobIds.contains(id));
        }

        long totalMs = TimeUnit.NANOSECONDS.toMillis(System.nanoTime() - startNs);
        if (totalMs > 500) {
            log.warn(
                    "running-jobs summary slow: total={}ms jobs={} decode={}ms",
                    totalMs,
                    values.size(),
                    TimeUnit.NANOSECONDS.toMillis(decodeTotalNs));
            Runtime rt = Runtime.getRuntime();
            long usedBytes = rt.totalMemory() - rt.freeMemory();
            log.warn(
                    "running-jobs summary slow diagnostics: totalMs={} jobs={} decodeMs={} "
                            + "heapUsedMB={} heapTotalMB={} heapMaxMB={}",
                    totalMs,
                    values.size(),
                    TimeUnit.NANOSECONDS.toMillis(decodeTotalNs),
                    usedBytes / 1024 / 1024,
                    rt.totalMemory() / 1024 / 1024,
                    rt.maxMemory() / 1024 / 1024);
        } else {
            log.debug(
                    "running-jobs summary: total={}ms jobs={} decode={}ms",

View on GitHub (pinned to cf67b549a7)