risingwavelabs/risingwave · error · anyhow::Error
failed to query Prometheus
Error message
failed to query Prometheus
What it means
get_channel_delta_stats fetches input, output, and backpressure Prometheus queries in parallel via tokio::try_join!. If any of the three PromQL queries fails, the error is wrapped with context "failed to query Prometheus", indicating a metrics-collection failure rather than a data problem.
Solutions
- Verify the Prometheus endpoint is reachable: curl the configured URL's /api/v1/query endpoint.
- Check the configured Prometheus address in RisingWave config and fix it if wrong.
- Inspect the underlying error (chained via anyhow context) for timeout vs HTTP errors and address accordingly.
- Ensure the metric series used by the queries actually exist (check scrape targets for compute nodes).
Defensive patterns
Strategy: retry
Validate before calling
// before relying on stats curl -fsS "$PROMETHEUS_URL/api/v1/query?query=up" >/dev/null && echo ok
Try / catch
match reader.get_channel_delta_stats(..).await {
Ok(stats) => stats,
Err(e) if e.to_string().contains("failed to query Prometheus") => {
tracing::warn!("prometheus unavailable: {e:#}");
retry_with_backoff(3).await
}
Err(e) => return Err(e),
} Prevention
- Health-check the Prometheus endpoint at startup and alert on failure.
- Keep the Prometheus address in one validated config value.
- Set generous but bounded HTTP timeouts for metrics queries.
- Verify scrape targets cover compute-node metrics before using auto-scaled stats.
When it happens
Trigger: Calling get_channel_delta_stats when the Prometheus endpoint is unreachable, returns a non-200 response, times out, or one of the input/output/backpressure instant queries fails to evaluate.
Common situations: Prometheus not running or misconfigured in the cluster, wrong metrics endpoint URL in config, network/firewall issues between frontend and Prometheus, or PromQL errors after Prometheus version upgrades.
Understand the failure class
Background: 'Something went wrong' / 'Request failed (500)' / 'HTTP error! status: 404' — what failed HTTP requests actually mean and how to find the real cause — this error's family across 28 libraries.
Related errors
- all request confluent registry all timeout
- {batch write failure, with context()}
- bigquery insert error
- bigquery insert error: end of resp stream
- cannot connect to kafka broker
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/ee44b46e4b49c41a.
Report an issue: GitHub.
Appendix: source
Thrown at src/frontend/src/metrics_reader.rs:96
channel_backpressure_result,
) = {
let mut input_query = prometheus_client.query(channel_input_throughput_query);
let mut output_query = prometheus_client.query(channel_output_throughput_query);
let mut backpressure_query = prometheus_client.query(channel_backpressure_query);
// Set the evaluation time if provided
if let Some(at_time) = at_time {
input_query = input_query.at(at_time);
output_query = output_query.at(at_time);
backpressure_query = backpressure_query.at(at_time);
}
tokio::try_join!(
input_query.get(),
output_query.get(),
backpressure_query.get(),
)
.map_err(|e| anyhow!(e).context("failed to query Prometheus"))?
};
// Process channel delta stats
let mut channel_data: HashMap<ChannelKey, ChannelDeltaStats> = HashMap::new();
// Collect input throughput
if let Some(channel_input_throughput_data) =
channel_input_throughput_result.data().as_vector()
{
for sample in channel_input_throughput_data {
if let Some(fragment_id_str) = sample.metric().get("fragment_id")
&& let Some(upstream_fragment_id_str) =
sample.metric().get("upstream_fragment_id")
&& let (Ok(fragment_id), Ok(upstream_fragment_id)) = (
fragment_id_str.parse::<u32>(),
upstream_fragment_id_str.parse::<u32>(),
)
{View on GitHub (pinned to 6469eb736d)