risingwavelabs/risingwave · error · anyhow::Error
receiving commit request from non-running handle
Error message
receiving commit request from non-running handle {}, running handles: {:?} What it means
The steady-state loop tracks which handles are currently running; a commit request arriving from a handle not in `running_handles` is rejected. This means a stopped/aborted writer is still sending commits, which would write data from a handle that was logically torn down or replaced (e.g. during a parallelism change).
Solutions
- Ensure the writer stops sending commits immediately upon receiving StopCoordination and drains in-flight sends before exiting.
- Verify the coordinator's Stop path marks handles non-running only after the writer confirmed shutdown, and that the writer confirms before finishing.
- Check for stale HandleId reuse on reconnect — issue a fresh HandleId per incarnation.
- Restart the affected sink to clear the mismatched state, then investigate the writer's shutdown ordering.
Defensive patterns
Strategy: validation
Validate before calling
// writer-side precheck before sending a commit
if self.stop_requested { return Err(anyhow!("handle stopped; refusing to commit")); } Type guard
fn is_running(running: &HashSet<HandleId>, id: &HandleId) -> bool { running.contains(id) } Try / catch
// restart the affected sink to clear state, then audit shutdown ordering
if err.contains("non-running handle") { restart_sink(); inspect_writer_shutdown(); } Prevention
- Writers must stop emitting commits immediately on StopCoordination and drain in-flight sends first.
- Issue a fresh HandleId per writer incarnation to avoid stale-id reuse.
- Only mark handles non-running after the writer confirms shutdown.
When it happens
Trigger: A writer whose Stop was processed by the coordinator sends a buffered commit; a new writer reuses the HandleId before being registered as running; a writer continues emitting commits after being stopped in `alter_parallelisms`.
Common situations: Parallelism decrease where the removed writer has in-flight buffered data; races between Stop acknowledgement and the writer's commit send; duplicated HandleIds after a writer reconnect.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- Actor exited unexpectedly
- empty sink metadata
- end of writer request stream
- expect AlignInitialEpoch but got
- expect new handle during init, but got
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/b683bb8180e5348a.
Report an issue: GitHub.
Appendix: source
Thrown at src/meta/src/manager/sink_coordination/coordinator_worker.rs:799
"committing"
);
})
.await;
match commit_res {
Ok(_) => {
two_phase_handler.ack_committed(epoch).await?;
}
Err(e) => {
two_phase_handler.failed_committed(epoch, e);
}
}
continue;
}
};
if !running_handles.contains(&handle_id) {
bail!(
"receiving commit request from non-running handle {}, running handles: {:?}",
handle_id,
running_handles
);
}
pending_epochs.entry(epoch).or_default().add_new_request(
handle_id,
commit_request,
self.handle_manager.vnode_bitmap(handle_id),
)?;
if pending_epochs
.first_key_value()
.expect("non-empty")
.1
.aligned()
{
let (epoch, commit_requests) = pending_epochs.pop_first().expect("non-empty");
let mut metadatas = Vec::with_capacity(commit_requests.requests.len());View on GitHub (pinned to 6469eb736d)