risingwavelabs/risingwave · info · MetaError
adhoc recovery triggered
Error message
adhoc recovery triggered
What it means
MetaError::AdhocRecovery signals that recovery of stream fragments/jobs was triggered manually rather than by automatic failure detection. It is thrown by the meta recovery machinery when an operator explicitly requests recovery (e.g. via a catalog manager or admin operation). It is a control-flow signal, not a fault: its message has no fields and it exists so callers can distinguish manual recovery requests from real errors.
Solutions
- Treat this variant as an expected control signal: match it separately and return success once recovery completes
- Ensure only one adhoc recovery runs at a time; concurrent triggers may log noise
- Check recovery status afterwards (streaming job states) rather than retrying
- If this fires unintentionally, audit the code path that calls manual recovery
Example fix
// before
meta_recovery_err.map_err(to_grpc)?;
// after
match meta_recovery_err {
MetaError::AdhocRecovery => Ok(recovery_completed_response()),
other => Err(to_grpc(other)),
} Defensive patterns
Strategy: try-catch
Type guard
fn is_adhoc_recovery(e: &MetaError) -> bool {
matches!(e, MetaError::AdhocRecovery)
} Try / catch
match recovery_result {
Err(MetaError::AdhocRecovery) => info!("manual recovery requested; not an error"),
Err(e) => return Err(e.into()),
Ok(v) => return Ok(v),
} Prevention
- Treat this variant as a control-flow signal, never retry it as a failure
- Serialize manual recovery triggers behind a mutex/leader to avoid concurrent runs
- Document the manual-recovery entry point for operators
When it happens
Trigger: An explicit ad-hoc recovery call: e.g. `recover_driver.adhoc_recovery()` or manual recovery triggered from cluster management operations (such as the catalog handler `metadata_manager`/stream manager requesting recovery when something looks stale).
Common situations: Administrators manually trigger recovery after noticing stuck stream jobs or after applying hot-reload/config changes; operator tooling calls recovery to reschedule fragments after a transient glitch.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- has been deprecated, please use instead.
- a stream has reached the end but some other stream has not…
- Barrier read is unavailable for now. Likely the cluster is…
- Cannot alter the job
- Cannot alter the job
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/602604b9d190675d.
Report an issue: GitHub.
Appendix: source
Thrown at src/meta/src/error.rs:132
ConnectorError,
),
#[error("Sink error: {0}")]
Sink(
#[from]
#[backtrace]
SinkError,
),
#[error(transparent)]
Internal(
#[from]
#[backtrace]
anyhow::Error,
),
// Indicates that recovery was triggered manually.
#[error("adhoc recovery triggered")]
AdhocRecovery,
#[error("Integrity check failed")]
IntegrityCheckFailed,
#[error("{0} has been deprecated, please use {1} instead.")]
Deprecated(String, String),
#[error(transparent)]
NotImplemented(#[from] NotImplemented),
#[error("Secret error: {0}")]
SecretError(
#[from]
#[backtrace]
SecretError,
),
}View on GitHub (pinned to 6469eb736d)