{"record":{"id":"ee44b46e4b49c41a","repo":"risingwavelabs/risingwave","slug":"failed-to-query-prometheus","errorCode":null,"errorMessage":"failed to query Prometheus","messagePattern":"failed to query Prometheus","errorType":"exception","errorClass":"anyhow::Error","httpStatus":null,"severity":"error","filePath":"src/frontend/src/metrics_reader.rs","lineNumber":96,"sourceCode":"            channel_backpressure_result,\n        ) = {\n            let mut input_query = prometheus_client.query(channel_input_throughput_query);\n            let mut output_query = prometheus_client.query(channel_output_throughput_query);\n            let mut backpressure_query = prometheus_client.query(channel_backpressure_query);\n\n            // Set the evaluation time if provided\n            if let Some(at_time) = at_time {\n                input_query = input_query.at(at_time);\n                output_query = output_query.at(at_time);\n                backpressure_query = backpressure_query.at(at_time);\n            }\n\n            tokio::try_join!(\n                input_query.get(),\n                output_query.get(),\n                backpressure_query.get(),\n            )\n            .map_err(|e| anyhow!(e).context(\"failed to query Prometheus\"))?\n        };\n\n        // Process channel delta stats\n        let mut channel_data: HashMap<ChannelKey, ChannelDeltaStats> = HashMap::new();\n\n        // Collect input throughput\n        if let Some(channel_input_throughput_data) =\n            channel_input_throughput_result.data().as_vector()\n        {\n            for sample in channel_input_throughput_data {\n                if let Some(fragment_id_str) = sample.metric().get(\"fragment_id\")\n                    && let Some(upstream_fragment_id_str) =\n                        sample.metric().get(\"upstream_fragment_id\")\n                    && let (Ok(fragment_id), Ok(upstream_fragment_id)) = (\n                        fragment_id_str.parse::<u32>(),\n                        upstream_fragment_id_str.parse::<u32>(),\n                    )\n                {","sourceCodeStart":78,"sourceCodeEnd":114,"githubUrl":"https://github.com/risingwavelabs/risingwave/blob/6469eb736d691e8e9b8a419a57edd6429ca77417/src/frontend/src/metrics_reader.rs#L78-L114","documentation":"get_channel_delta_stats fetches input, output, and backpressure Prometheus queries in parallel via tokio::try_join!. If any of the three PromQL queries fails, the error is wrapped with context \"failed to query Prometheus\", indicating a metrics-collection failure rather than a data problem.","triggerScenarios":"Calling get_channel_delta_stats when the Prometheus endpoint is unreachable, returns a non-200 response, times out, or one of the input/output/backpressure instant queries fails to evaluate.","commonSituations":"Prometheus not running or misconfigured in the cluster, wrong metrics endpoint URL in config, network/firewall issues between frontend and Prometheus, or PromQL errors after Prometheus version upgrades.","solutions":["Verify the Prometheus endpoint is reachable: curl the configured URL's /api/v1/query endpoint.","Check the configured Prometheus address in RisingWave config and fix it if wrong.","Inspect the underlying error (chained via anyhow context) for timeout vs HTTP errors and address accordingly.","Ensure the metric series used by the queries actually exist (check scrape targets for compute nodes)."],"exampleFix":null,"handlingStrategy":"retry","validationCode":"// before relying on stats\ncurl -fsS \"$PROMETHEUS_URL/api/v1/query?query=up\" >/dev/null && echo ok","typeGuard":null,"tryCatchPattern":"match reader.get_channel_delta_stats(..).await {\n    Ok(stats) => stats,\n    Err(e) if e.to_string().contains(\"failed to query Prometheus\") => {\n        tracing::warn!(\"prometheus unavailable: {e:#}\");\n        retry_with_backoff(3).await\n    }\n    Err(e) => return Err(e),\n}","preventionTips":["Health-check the Prometheus endpoint at startup and alert on failure.","Keep the Prometheus address in one validated config value.","Set generous but bounded HTTP timeouts for metrics queries.","Verify scrape targets cover compute-node metrics before using auto-scaled stats."],"tags":["prometheus","metrics","network"],"backgroundTag":"http-request-failed","analyzedSha":"6469eb736d691e8e9b8a419a57edd6429ca77417","analyzedAt":"2026-09-11T21:06:21.487Z","contentChangedAt":"2026-09-11T21:06:21.487Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}