risingwavelabs/risingwave · error · anyhow::Error
expect AlignInitialEpoch but got
Error message
expect AlignInitialEpoch but got {} What it means
`handle_init_requests_impl` collects initial epoch alignment requests from newly registered handles and expects each to be `AlignInitialEpoch`. Any other event type breaks the init protocol and returns this error. It indicates a handle is sending commits or stops during the alignment collection window.
Solutions
- Read the event name in the message to identify the offending writer behavior and fix that writer to align first.
- Ensure all sink writers are upgraded to a version supporting the AlignInitialEpoch handshake.
- Handle Stop events during init batch collection gracefully (drop the handle) rather than failing the entire batch.
- Add trace logs around `try_handle_init_requests` to capture which handle sent the unexpected event.
Defensive patterns
Strategy: try-catch
Validate before calling
// writer-side: ensure AlignInitialEpoch is the first event after StartResponse when alignment is requested
if self.requires_alignment && !self.aligned { align_before_commit()?; } Type guard
fn is_align_initial_epoch(ev: &CoordinationHandleManagerEvent) -> bool { matches!(ev, CoordinationHandleManagerEvent::AlignInitialEpoch(_)) } Try / catch
// coordinator: catch and identify the straggler handle from logs, then restart the batch
Err(e) if e.to_string().contains("expect AlignInitialEpoch") => restart_init_batch(), Prevention
- Upgrade all writers to versions that implement the alignment handshake.
- Treat writer Stop during init batch collection as a normal drop, not a batch failure.
- Trace init batches to spot handles that skip alignment.
When it happens
Trigger: While draining init requests for a batch of new handles, one handle emits `CommitRequest`, `Stop`, or `NewHandle` instead of `AlignInitialEpoch` — e.g. a writer that skips alignment, aborts, or double-registers during `try_handle_init_requests`.
Common situations: Mixed writer versions where some do not participate in initial epoch alignment; a writer failing right as the coordinator collects init requests; a race where the same HandleId re-registers mid-batch.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- expect new handle during init, but got
- empty sink metadata
- end of writer request stream
- failed to ack aligned initial epoch
- failed to acknowledge the commit for epoch
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/6f27c5a4e17c6f90.
Report an issue: GitHub.
Appendix: source
Thrown at src/meta/src/manager/sink_coordination/coordinator_worker.rs:630
pending_handle_ids: impl IntoIterator<Item = HandleId>,
) -> anyhow::Result<()> {
let log_store_rewind_start_epoch = self.last_writer_acked_epoch;
self.handle_manager
.start(log_store_rewind_start_epoch, pending_handle_ids)?;
if log_store_rewind_start_epoch.is_none() {
let mut align_requests = AligningRequests::default();
while !align_requests.aligned() {
let (handle_id, event) = self.handle_manager.next_event().await?;
match event {
CoordinationHandleManagerEvent::AlignInitialEpoch(initial_epoch) => {
align_requests.add_new_request(
handle_id,
initial_epoch,
self.handle_manager.vnode_bitmap(handle_id),
)?;
}
other => {
return Err(anyhow!("expect AlignInitialEpoch but got {}", other.name()));
}
}
}
let aligned_initial_epoch = align_requests
.requests
.into_iter()
.max()
.expect("non-empty");
self.handle_manager
.ack_aligned_initial_epoch(aligned_initial_epoch)?;
}
Ok(())
}
async fn next_event(
&mut self,
two_phase_handler: &mut TwoPhaseCommitHandler,
) -> anyhow::Result<CoordinatorWorkerEvent> {View on GitHub (pinned to 6469eb736d)