risingwavelabs/risingwave · error · DashboardError

Failed to get back pressure

Error message

Failed to get back pressure

What it means

When the dashboard aggregates streaming stats, it fans out GetStreamingStats gRPC requests to all workers and joins the futures. If any individual worker's request fails (gRPC error, worker down, timeout), the whole aggregation is discarded and replaced with this generic 'Failed to get back pressure' error — the original cause is dropped by the `map_err(|_| ...)`.

Solutions

  1. Check `SHOW NODES` / the dashboard worker list and confirm all workers are RUNNING; restart failed workers.
  2. Inspect the meta node logs around the request for the underlying gRPC error that was swallowed.
  3. Retry the dashboard request once the cluster is healthy; transient worker unavailability resolves this.
  4. Improve the code: replace `map_err(|_| ...)` with a closure that preserves the source error for diagnosis.

Example fix

// before
result.map_err(|_| anyhow!("Failed to get back pressure")).map_err(err)?
// after
result.map_err(|e| anyhow!("Failed to get back pressure: {e}")).map_err(err)?
Defensive patterns

Strategy: retry

Validate before calling

const health = await fetch('http://localhost:5691/api/v1/workers').then(r => r.json());
if (health.some(w => w.state !== 'RUNNING')) throw new Error('Not all workers are RUNNING; fix cluster health before querying streaming stats');

Try / catch

try { const stats = await getStreamingStats(); } catch (e) {
  if (String(e).includes('Failed to get back pressure')) {
    // check meta logs for the swallowed per-worker gRPC error, then retry with backoff
  }
}

Prevention

When it happens

Trigger: GET /api/v1/streaming_stats when at least one worker (compute node or meta) fails to respond to the GetStreamingStats RPC — e.g. the worker crashed, is restarting, or the gRPC channel is broken.

Common situations: Cluster during rolling upgrade; a compute node that was killed or OOMed; network partition between meta node and worker; worker still bootstrapping while the dashboard is polled.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/68358de09c7a152e. Report an issue: GitHub.

Appendix: source

Thrown at src/meta/src/dashboard/mod.rs:725

        let mut futures = Vec::new();

        for worker_node in worker_nodes {
            let client = srv.monitor_clients.get(&worker_node).await.map_err(err)?;
            let client = Arc::new(client);
            let fut = async move {
                let result = client.get_streaming_stats().await.map_err(err)?;
                Ok::<_, DashboardError>(result)
            };
            futures.push(fut);
        }
        let results = join_all(futures).await;

        let mut all = GetStreamingStatsResponse::default();

        for result in results {
            let result = result
                .map_err(|_| anyhow!("Failed to get back pressure"))
                .map_err(err)?;

            // Aggregate fragment_stats
            for (fragment_id, fragment_stats) in result.fragment_stats {
                if let Some(s) = all.fragment_stats.get_mut(&fragment_id) {
                    s.actor_count += fragment_stats.actor_count;
                    s.current_epoch = min(s.current_epoch, fragment_stats.current_epoch);
                } else {
                    all.fragment_stats.insert(fragment_id, fragment_stats);
                }
            }

            // Aggregate relation_stats
            for (relation_id, relation_stats) in result.relation_stats {
                if let Some(s) = all.relation_stats.get_mut(&relation_id) {
                    s.actor_count += relation_stats.actor_count;
                    s.current_epoch = min(s.current_epoch, relation_stats.current_epoch);
                } else {

View on GitHub (pinned to 6469eb736d)