risingwavelabs/risingwave · error · anyhow
partial graph resetting
Error message
partial graph resetting
What it means
A TakeReceiver (exchange channel creation) request arrived while the target partial graph is being reset. Because the old graph is being torn down, the receiver cannot be taken and this error is returned to the remote requester. It is a transient, recovery-related rejection.
Solutions
- Retry the TakeReceiver request once the partial graph finishes resetting and is Running with a new term id.
- Confirm the meta node issued ResetPartialGraphs and that the new term id is used in follow-up requests.
- If this loops, check whether the reset itself keeps failing (look for ack_reset_partial_graph with a root error).
Example fix
// before: immediate take during reset worker.take_receiver(partial_graph_id, term_id, ids, request); // after: await ResetPartialGraph ack, then take with new term await_reset_ack(partial_graph_id).await; worker.take_receiver(partial_graph_id, new_term_id, ids, request);
Defensive patterns
Strategy: retry
Validate before calling
// avoid issuing take-receiver during reset window
while matches!(graph_status(partial_graph_id), Resetting) { tokio::time::sleep(poll_interval).await; } Type guard
fn is_resetting(err: &anyhow::Error) -> bool { err.to_string() == "partial graph resetting" } Try / catch
// wait-and-retry on reset rejection
match take_receiver().await {
Err(e) if is_resetting(&e) => { wait_reset_ack(partial_graph_id).await; take_receiver().await?; }
other => other,
} Prevention
- Serialize graph creation/reset with exchange channel setup via the control stream ordering.
- Treat Resetting as a normal transient state in retry logic, not a fatal error.
- Monitor reset duration; long resets indicate recurring root failures.
When it happens
Trigger: handle_actor_op matches PartialGraphStatus::Resetting while processing LocalActorOperation::TakeReceiver for that partial graph id.
Common situations: Concurrent exchange channel setup during fault recovery; upstream retries racing with a ResetPartialGraphs command from the meta node.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- partial graph suspended
- take receiver on unmatched partial graph term to current…
- actor exited unexpectedly
- anyhow!(message.to_owned())
- barrier reader closed unexpectedly
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/1af2d61cae269a34.
Report an issue: GitHub.
Appendix: source
Thrown at src/stream/src/task/barrier_worker/mod.rs:552
);
anyhow!(
"take receiver {:?} on unmatched partial graph term {} to current term {}",
ids,
term_id,
graph.local_barrier_manager.term_id
)
} else {
let (upstream_actor_id, actor_id) = ids;
graph.new_actor_output_request(
actor_id,
upstream_actor_id,
request,
);
return;
}
}
PartialGraphStatus::Suspended(_) => anyhow!("partial graph suspended"),
PartialGraphStatus::Resetting => anyhow!("partial graph resetting"),
PartialGraphStatus::Unspecified => unreachable!(),
},
Entry::Vacant(entry) => {
entry.insert(PartialGraphStatus::ReceivedExchangeRequest(vec![(
term_id, ids, request,
)]));
return;
}
};
if let TakeReceiverRequest::Remote { result_sender, .. } = request {
let _ = result_sender.send(Err(err.into()));
}
}
#[cfg(test)]
LocalActorOperation::GetCurrentLocalBarrierManager(sender) => {
let partial_graph_status = self
.state
.partial_graphsView on GitHub (pinned to 6469eb736d)