apache/seatunnel · warning

running-jobs summary slow diagnostics: totalMs=

Error message

running-jobs summary slow diagnostics: totalMs={} jobs={} decodeMs={} heapUsedMB={} heapTotalMB={} heapMaxMB={}

What it means

Follow-up diagnostics emitted right after the slow-summary warning (errorIndex 3713) when the running-jobs summary exceeds 500ms. It adds heap diagnostics: heap used/total/max in MB, to correlate slowness with memory pressure (GC thrash, near-capacity heap).

Solutions

  1. Increase the REST/master node -Xmx based on heapUsedMB vs heapMaxMB in the log.
  2. Tune GC (e.g., G1 with shorter pause goals) to reduce decode-time inflation.
  3. Reduce per-job payload size / number of jobs returned per summary call.
  4. Correlate with GC logs; if heapUsedMB oscillates at max, the slowness is GC-driven, not logic-driven.

Example fix

// before
// JAVA_OPTS="-Xms2g -Xmx2g"
// after (heap headroom for REST summary endpoint)
// JAVA_OPTS="-Xms4g -Xmx4g -XX:+UseG1GC -XX:MaxGCPauseMillis=200"
Defensive patterns

Strategy: retry

Validate before calling

// before bulk REST polling, check headroom
// Runtime rt = Runtime.getRuntime();
// double usedRatio = (double)(rt.totalMemory()-rt.freeMemory())/rt.maxMemory();
// if (usedRatio > 0.85) { reducePolling(); }

Prevention

When it happens

Trigger: Same trigger as the slow summary (total > 500ms); logged whenever heapUsedMB is near heapMaxMB or decode time dominates, indicating GC pressure or payload-heavy decoding.

Common situations: REST node running close to -Xmx with heavy dashboard polling; frequent full GCs slowing map iteration and JSON decode; jobs with very large basic-info payloads.

Understand the failure class

Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/52d81d04080c731a. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-engine/seatunnel-engine-server/src/main/java/org/apache/seatunnel/engine/server/rest/service/JobInfoService.java:256

                            .add(RestConstant.JOB_NAME, jobName)
                            .add(RestConstant.JOB_STATUS, jobStatus)
                            .add(RestConstant.CREATE_TIME, createTime));
        }

        if (!runningJobBasicInfoCache.isEmpty() && !jobIds.isEmpty()) {
            runningJobBasicInfoCache.keySet().removeIf(id -> !jobIds.contains(id));
        }

        long totalMs = TimeUnit.NANOSECONDS.toMillis(System.nanoTime() - startNs);
        if (totalMs > 500) {
            log.warn(
                    "running-jobs summary slow: total={}ms jobs={} decode={}ms",
                    totalMs,
                    values.size(),
                    TimeUnit.NANOSECONDS.toMillis(decodeTotalNs));
            Runtime rt = Runtime.getRuntime();
            long usedBytes = rt.totalMemory() - rt.freeMemory();
            log.warn(
                    "running-jobs summary slow diagnostics: totalMs={} jobs={} decodeMs={} "
                            + "heapUsedMB={} heapTotalMB={} heapMaxMB={}",
                    totalMs,
                    values.size(),
                    TimeUnit.NANOSECONDS.toMillis(decodeTotalNs),
                    usedBytes / 1024 / 1024,
                    rt.totalMemory() / 1024 / 1024,
                    rt.maxMemory() / 1024 / 1024);
        } else {
            log.debug(
                    "running-jobs summary: total={}ms jobs={} decode={}ms",
                    totalMs,
                    values.size(),
                    TimeUnit.NANOSECONDS.toMillis(decodeTotalNs));
        }
        return out;
    }

View on GitHub (pinned to cf67b549a7)