{"record":{"id":"fdaff4fa907f6dc8","repo":"apache/seatunnel","slug":"running-jobs-summary-slow-total-ms-jobs-deco","errorCode":null,"errorMessage":"running-jobs summary slow: total={}ms jobs={} decode={}ms","messagePattern":"running-jobs summary slow: total=(.+?)ms jobs=(.+?) decode=(.+?)ms","errorType":"console","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"seatunnel-engine/seatunnel-engine-server/src/main/java/org/apache/seatunnel/engine/server/rest/service/JobInfoService.java","lineNumber":249,"sourceCode":"            } catch (Throwable ignored) {\n                // ignore\n            }\n\n            out.add(\n                    new JsonObject()\n                            .add(RestConstant.JOB_ID, String.valueOf(jobId))\n                            .add(RestConstant.JOB_NAME, jobName)\n                            .add(RestConstant.JOB_STATUS, jobStatus)\n                            .add(RestConstant.CREATE_TIME, createTime));\n        }\n\n        if (!runningJobBasicInfoCache.isEmpty() && !jobIds.isEmpty()) {\n            runningJobBasicInfoCache.keySet().removeIf(id -> !jobIds.contains(id));\n        }\n\n        long totalMs = TimeUnit.NANOSECONDS.toMillis(System.nanoTime() - startNs);\n        if (totalMs > 500) {\n            log.warn(\n                    \"running-jobs summary slow: total={}ms jobs={} decode={}ms\",\n                    totalMs,\n                    values.size(),\n                    TimeUnit.NANOSECONDS.toMillis(decodeTotalNs));\n            Runtime rt = Runtime.getRuntime();\n            long usedBytes = rt.totalMemory() - rt.freeMemory();\n            log.warn(\n                    \"running-jobs summary slow diagnostics: totalMs={} jobs={} decodeMs={} \"\n                            + \"heapUsedMB={} heapTotalMB={} heapMaxMB={}\",\n                    totalMs,\n                    values.size(),\n                    TimeUnit.NANOSECONDS.toMillis(decodeTotalNs),\n                    usedBytes / 1024 / 1024,\n                    rt.totalMemory() / 1024 / 1024,\n                    rt.maxMemory() / 1024 / 1024);\n        } else {\n            log.debug(\n                    \"running-jobs summary: total={}ms jobs={} decode={}ms\",","sourceCodeStart":231,"sourceCodeEnd":267,"githubUrl":"https://github.com/apache/seatunnel/blob/cf67b549a7a6c35fa0beb12d83c62892427ea919/seatunnel-engine/seatunnel-engine-server/src/main/java/org/apache/seatunnel/engine/server/rest/service/JobInfoService.java#L231-L267","documentation":"getRunningJobsSummaryJson measures how long building the running-jobs summary takes. When total time exceeds 500ms it emits this warn with total duration, job count, and cumulative JSON decode time, so operators can see the summary endpoint is degrading.","triggerScenarios":"REST GET of the running-jobs summary endpoint while there are many running jobs, large per-job basic-info JSON payloads needing decode, or the JVM is CPU/GC-starved causing slow map iteration and JSON parsing.","commonSituations":"Hundreds of concurrently running jobs on a cluster; undersized REST node heap or heavy GC pauses; slow disk/network backing the Hazelcast maps; burst of REST dashboard polling amplifying load.","solutions":["Reduce running job count or poll frequency; cache summary responses on the client.","Scale the master/REST node with more CPU/heap and tune GC to avoid long pauses.","Profile the decode path; ensure runningJobBasicInfoCache is being used and not repeatedly decoding the same payloads.","Monitor via the accompanying diagnostics log (errorIndex 3714) to distinguish decode cost from total cost."],"exampleFix":"// before: polling every second with hundreds of jobs\n// curl http://host:8080/running-jobs/summary  (every 1s)\n// after: back off and cache\n// every 10s -> fetch summary, cache locally, reuse between polls","handlingStrategy":"fallback","validationCode":"// client: back off if previous summary took > 500ms\n// if (lastSummaryMs > 500) { delay(nextPoll, backoffMs); }","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Poll the summary endpoint at a sane interval; cache results","Watch job count; scale the cluster or limit concurrent running jobs","Keep REST node heap/GC healthy"],"tags":["performance","rest-api","slow-endpoint","hazelcast"],"backgroundTag":"request-timeout","analyzedSha":"cf67b549a7a6c35fa0beb12d83c62892427ea919","analyzedAt":"2026-09-10T21:44:55.265Z","contentChangedAt":"2026-09-10T21:44:55.265Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}