risingwavelabs/risingwave · error · DashboardError
Failed to get back pressure
Error message
Failed to get back pressure
What it means
When the dashboard aggregates streaming stats, it fans out GetStreamingStats gRPC requests to all workers and joins the futures. If any individual worker's request fails (gRPC error, worker down, timeout), the whole aggregation is discarded and replaced with this generic 'Failed to get back pressure' error — the original cause is dropped by the `map_err(|_| ...)`.
Solutions
- Check `SHOW NODES` / the dashboard worker list and confirm all workers are RUNNING; restart failed workers.
- Inspect the meta node logs around the request for the underlying gRPC error that was swallowed.
- Retry the dashboard request once the cluster is healthy; transient worker unavailability resolves this.
- Improve the code: replace `map_err(|_| ...)` with a closure that preserves the source error for diagnosis.
Example fix
// before
result.map_err(|_| anyhow!("Failed to get back pressure")).map_err(err)?
// after
result.map_err(|e| anyhow!("Failed to get back pressure: {e}")).map_err(err)? Defensive patterns
Strategy: retry
Validate before calling
const health = await fetch('http://localhost:5691/api/v1/workers').then(r => r.json());
if (health.some(w => w.state !== 'RUNNING')) throw new Error('Not all workers are RUNNING; fix cluster health before querying streaming stats'); Try / catch
try { const stats = await getStreamingStats(); } catch (e) {
if (String(e).includes('Failed to get back pressure')) {
// check meta logs for the swallowed per-worker gRPC error, then retry with backoff
}
} Prevention
- Monitor worker liveness before polling streaming stats endpoints.
- Check meta node logs for the underlying gRPC failure hidden by the generic message.
- Avoid scraping this endpoint during rolling upgrades or restarts.
When it happens
Trigger: GET /api/v1/streaming_stats when at least one worker (compute node or meta) fails to respond to the GetStreamingStats RPC — e.g. the worker crashed, is restarting, or the gRPC channel is broken.
Common situations: Cluster during rolling upgrade; a compute node that was killed or OOMed; network partition between meta node and worker; worker still bootstrapping while the dashboard is polled.
Related errors
- Failed to get streaming stats from worker
- get none metadata in commit response for coordinated sink…
- Prometheus endpoint is not set
- should get start response but get
- {0}
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/68358de09c7a152e.
Report an issue: GitHub.
Appendix: source
Thrown at src/meta/src/dashboard/mod.rs:725
let mut futures = Vec::new();
for worker_node in worker_nodes {
let client = srv.monitor_clients.get(&worker_node).await.map_err(err)?;
let client = Arc::new(client);
let fut = async move {
let result = client.get_streaming_stats().await.map_err(err)?;
Ok::<_, DashboardError>(result)
};
futures.push(fut);
}
let results = join_all(futures).await;
let mut all = GetStreamingStatsResponse::default();
for result in results {
let result = result
.map_err(|_| anyhow!("Failed to get back pressure"))
.map_err(err)?;
// Aggregate fragment_stats
for (fragment_id, fragment_stats) in result.fragment_stats {
if let Some(s) = all.fragment_stats.get_mut(&fragment_id) {
s.actor_count += fragment_stats.actor_count;
s.current_epoch = min(s.current_epoch, fragment_stats.current_epoch);
} else {
all.fragment_stats.insert(fragment_id, fragment_stats);
}
}
// Aggregate relation_stats
for (relation_id, relation_stats) in result.relation_stats {
if let Some(s) = all.relation_stats.get_mut(&relation_id) {
s.actor_count += relation_stats.actor_count;
s.current_epoch = min(s.current_epoch, relation_stats.current_epoch);
} else {View on GitHub (pinned to 6469eb736d)