risingwavelabs/risingwave · error

trigger_manual_compaction wait report failed…

Error message

trigger_manual_compaction wait report failed. compaction_group {}

What it means

After dispatching a manual compaction task, trigger_manual_compaction awaits a one-shot report channel (report_rx). If the sender is dropped without sending (the compactor never reported), the await returns Err and the waiter is removed, yielding this error.

Solutions

  1. Retry the manual compaction after confirming a compactor is healthy.
  2. Check compactor logs/metrics for the task_id and why it exited without reporting.
  3. Verify compactor-meta connectivity and that heartbeat/worker timeouts are not too aggressive.

Example fix

// before
let result = report_rx.await.map_err(|_| anyhow!("trigger_manual_compaction wait report failed. compaction_group {}", group))?;
// after
// add a timeout so a silent compactor is detected explicitly and the task is retried:
let result = tokio::time::timeout(Duration::from_secs(600), report_rx).await
    .map_err(|_| anyhow!("compaction report timed out for group {}", group))?
    .map_err(|_| anyhow!("compactor dropped report channel for group {}", group))?
Defensive patterns

Strategy: retry

Validate before calling

// ensure at least one healthy compactor is registered before triggering
assert!(compactor_heartbeats.iter().any(|h| h.recent()), "no healthy compactor registered");

Try / catch

match trigger_manual_compaction(opt).await {
    Err(e) if e.to_string().contains("wait report failed") => {
        tokio::time::sleep(backoff).await; // compactor died/crashed mid-task; retry after backoff
        retry_with_backoff();
    },
    other => other?,
}

Prevention

When it happens

Trigger: The compactor that received the task crashes, disconnects, or is cancelled before reporting; the report-waiter is dropped due to manager state cleanup; the oneshot channel closes without a send.

Common situations: Compactor node killed or OOMed mid-task; network partition between meta and compactor; meta failover during a long-running manual compaction; shutting down the cluster right after triggering compaction.

Understand the failure class

Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/b6dff520712ccba4. Report an issue: GitHub.

Appendix: source

Thrown at src/meta/src/hummock/manager/compaction/mod.rs:1158

        let report_rx = self.register_compaction_task_report_waiter(task_id);
        if let Err(err) = compactor
            .send_event(ResponseEvent::CompactTask(compact_task.into()))
            .with_context(|| {
                format!(
                    "Failed to trigger compaction task for compaction_group {}",
                    compaction_group,
                )
            })
        {
            self.remove_compaction_task_report_waiter(task_id);
            return Err(err.into());
        }

        let report_result = match report_rx.await {
            Ok(result) => result,
            Err(_) => {
                self.remove_compaction_task_report_waiter(task_id);
                return Err(anyhow::anyhow!(
                    "trigger_manual_compaction wait report failed. compaction_group {}",
                    compaction_group
                )
                .into());
            }
        };
        if !report_result.reported {
            return Err(anyhow::anyhow!(
                "trigger_manual_compaction report not accepted. task_id {}",
                report_result.task_id
            )
            .into());
        }

        if report_result.task_status == TaskStatus::NoAvailCpuResourceCanceled
            || report_result.task_status == TaskStatus::NoAvailMemoryResourceCanceled
        {
            return Ok(ManualCompactionTriggerResult::Retry);

View on GitHub (pinned to 6469eb736d)