risingwavelabs/risingwave · warning · anyhow

partial graph reset

Error message

partial graph reset

What it means

During a partial graph reset in the stream barrier worker (`start_reset`), any pending remote exchange-take requests are answered with an error containing 'partial graph reset'. This is not a failure of user code but the controlled signal sent to remote readers whose upstream actors are being torn down as part of graph reconfiguration.

Solutions

  1. Generally no user action needed — consumers should propagate the error and let the new graph re-establish exchange channels.
  2. If it recurs during every barrier, check cluster stability and meta/stream logs for reset loops.
  3. Retry the affected query/stream consumption after the barrier completes.
  4. If the reset never completes (root_err missing), inspect actor deployment state and report with logs.
Defensive patterns

Strategy: try-catch

Try / catch

match res {
    Err(e) if e.to_string().contains("partial graph reset") => { /* transient during graph reconfig: retry / reconnect */ }
    other => other?,
}

Prevention

When it happens

Trigger: A barrier-driven partial graph reset occurs (e.g. actor reschedule, source drop, plan change) while remote exchange requests are pending in `ReceivedExchangeRequest` state; their result senders receive this error.

Common situations: Actors being migrated/rescheduled by the meta node; pausing/resuming or altering MVs causing upstream graph rebuilds; consumers of remote channels seeing a channel error during these operations.

Understand the failure class

Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/f56c0706223b4495. Report an issue: GitHub.

Appendix: source

Thrown at src/stream/src/task/barrier_worker/managed_state.rs:530

        let state = must_match!(replace(self, PartialGraphStatus::Unspecified), PartialGraphStatus::Running(state) => state);
        *self = PartialGraphStatus::Suspended(SuspendedPartialGraphState::new(
            state,
            Some((failed_actor, err)),
            completing_futures,
        ));
    }

    pub(super) fn start_reset(
        &mut self,
        partial_graph_id: PartialGraphId,
        completing_futures: Option<FuturesOrdered<AwaitEpochCompletedFuture>>,
        table_ids_to_clear: &mut HashSet<TableId>,
    ) -> BoxFuture<'static, ResetPartialGraphOutput> {
        match replace(self, PartialGraphStatus::Resetting) {
            PartialGraphStatus::ReceivedExchangeRequest(pending_requests) => {
                for (_, _, request) in pending_requests {
                    if let TakeReceiverRequest::Remote { result_sender, .. } = request {
                        let _ = result_sender.send(Err(anyhow!("partial graph reset").into()));
                    }
                }
                async move { ResetPartialGraphOutput { root_err: None } }.boxed()
            }
            PartialGraphStatus::Running(state) => {
                assert_eq!(partial_graph_id, state.partial_graph_id);
                info!(
                    %partial_graph_id,
                    "start partial graph reset from Running"
                );
                table_ids_to_clear.extend(state.table_ids.iter().copied());
                SuspendedPartialGraphState::new(state, None, completing_futures)
                    .reset()
                    .boxed()
            }
            PartialGraphStatus::Suspended(state) => {
                assert!(
                    completing_futures.is_none(),

View on GitHub (pinned to 6469eb736d)