risingwavelabs/risingwave · error
Iceberg compaction task report timed out for sink
Error message
Iceberg compaction task report timed out for sink {} What it means
A manual Iceberg compaction waiter received an Err because the dispatched compaction task did not report within its report_deadline. The scheduler's finish_timed_out_compaction_tasks marks the task failed and resolves any pending manual-compaction oneshot waiter with this timeout error.
Solutions
- Check worker/meta logs around the timeout for the sink's task to find why no report arrived
- Retry the manual compaction after the track resets to Idle (the timed-out task is already marked failed)
- Increase the report timeout if tasks legitimately run longer than the deadline
- Fix worker health: connectivity to meta, resource saturation, or crashes
Example fix
// before
let result = rx.await??; // Err: report timed out
// after
match rx.await {
Ok(Ok(task_id)) => Ok(task_id),
Ok(Err(e)) if e.to_string().contains("report timed out") => {
warn!("compaction report timed out; will retry after track resets");
retry_manual_compaction(sink_id).await
}
r => r,
} Defensive patterns
Strategy: try-catch
Try / catch
let result = rx.await;
match result {
Ok(Ok(task_id)) => Ok(task_id),
Ok(Err(e)) if e.to_string().contains("report timed out") => {
warn!(sink_id, "manual compaction timed out; scheduler will reset track");
retry_after_reset(sink_id).await
}
Ok(Err(e)) => Err(e),
Err(_recv_err) => Err(anyhow!("waiter dropped by scheduler")),
} Prevention
- Keep the oneshot receiver alive until resolution; dropping it cancels the waiter
- Size report timeouts to accommodate worst-case compaction durations
- Monitor for repeated report timeouts, which indicate worker health or network issues
- Retry after the timeout because the scheduler automatically marks the task failed
When it happens
Trigger: A manual compaction was triggered (caller holds the oneshot::Receiver from start_manual_compaction); the in-flight task misses its report deadline, so finish_timed_out_compaction_tasks sends this error through the waiter channel.
Common situations: Worker executing the compaction task is overloaded, partitioned from meta, or crashed without reporting; task legitimately takes longer than the configured report timeout; long GC pause on the worker.
Understand the failure class
Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.
- Timeouts: ETIMEDOUT, deadlines, and hung requests — what actually expires when a request times out.
Related errors
- bounded compaction is not supported for copy-on-write tasks
- compact_with_plan returned no result for a non-empty…
- `compaction.delete_files_count_threshold` must be greater…
- `compaction_interval_sec` must be greater than 0
- `compaction-interval-sec` must be greater than 0 when…
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/d6ba4ae3fb13d910.
Report an issue: GitHub.
Appendix: source
Thrown at src/meta/src/manager/iceberg_compaction/schedule.rs:1048
.take(n)
.filter_map(|(sink_id, _)| {
let track = guard.sink_schedules.get_mut(&sink_id)?;
let attempt = track.start_attempt();
Some(IcebergCompactionHandle::new(
sink_id,
attempt,
self.inner.clone(),
self.metadata_manager.clone(),
))
})
.collect();
(handles, timed_out_waiters)
};
for (sink_id, waiter) in timed_out_waiters {
let _ = waiter.send(Err(anyhow!(
"Iceberg compaction task report timed out for sink {}",
sink_id
)
.into()));
}
handles
}
pub fn clear_iceberg_maintenance_by_sink_id(&self, sink_id: SinkId) {
let (task_to_cancel, waiter) = {
let mut guard = self.inner.write();
let task_to_cancel = Self::remove_sink_schedule(&mut guard, sink_id);
guard.snapshot_expiration_sink_ids.remove(&sink_id);
guard.manifest_rewrite_sink_ids.remove(&sink_id);
let waiter = guard.manual_compaction_waiters.remove(&sink_id);
(task_to_cancel, waiter)
};View on GitHub (pinned to 6469eb736d)