Hmbown/CodeWhale · error

Fleet run is already terminal ( )

Error message

Fleet run {} is already terminal ({lifecycle:?})

What it means

FleetManager::activate_run transitions a durable queued run to Running for managed clients, but terminal runs are never reactivated through this method (crates/tui/src/fleet/manager.rs:484). If the run's effective lifecycle (status override or recorded status) is already Completed, Failed, or Cancelled, activation is refused — a terminal run is final, and worker restart is a separate explicit control.

Solutions

  1. Check the run's status first (run_status / ledger state) and skip activation when it is already terminal.
  2. Create a new run instead of reactivating the completed one — terminal is final by design.
  3. If individual workers need rework, use the explicit worker restart control rather than run activation.

Example fix

// before
manager.activate_run(&run_id)?;
// after
if !matches!(manager.run_status(&run_id)?.status, FleetRunStatus::Completed | FleetRunStatus::Failed | FleetRunStatus::Cancelled) {
    manager.activate_run(&run_id)?;
} else {
    // start a new run or restart individual workers instead
}
Defensive patterns

Strategy: validation

Validate before calling

let status = manager.run_status(&run_id)?;
if matches!(status.status, FleetRunStatus::Completed | FleetRunStatus::Failed | FleetRunStatus::Cancelled) {
    // skip activation; run is final
}

Type guard

fn is_terminal_run_status(s: &FleetRunStatus) -> bool { matches!(s, FleetRunStatus::Completed | FleetRunStatus::Failed | FleetRunStatus::Cancelled) }

Try / catch

if let Err(e) = manager.activate_run(&run_id) {
    if e.to_string().contains("already terminal") { /* treat as success/no-op or create a new run */ }
    else { return Err(e.into()); }
}

Prevention

When it happens

Trigger: Calling activate_run (or start_run, which delegates to it) on a run whose ledger status or status override is Completed, Failed, or Cancelled — e.g. re-invoking start after the run finished, or after an operator cancelled it via a separate interrupt process.

Common situations: A managed client retries activation after a crash without checking run state; an operator re-runs a CLI start command for an already-finished run; a resumption script races with a cancellation.

Understand the failure class

Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.

Related errors


AI-assisted analysis of Hmbown/CodeWhale@73e0f67d83 (2026-09-22). Data as JSON: /api/errors/5fe9eba075589d4e. Report an issue: GitHub.

Appendix: source

Thrown at crates/tui/src/fleet/manager.rs:484

    pub fn activate_run(&self, run_id: &FleetRunId) -> Result<FleetRunReport> {
        let state = self.ledger.rebuild_state()?;
        let run =
            state
                .runs
                .get(&run_id.0)
                .cloned()
                .ok_or_else(|| FleetControlError::UnknownRun {
                    run_id: run_id.0.clone(),
                })?;
        let lifecycle = state
            .run_status_overrides
            .get(&run_id.0)
            .unwrap_or(&run.status);
        if matches!(
            lifecycle,
            FleetRunStatus::Completed | FleetRunStatus::Failed | FleetRunStatus::Cancelled
        ) {
            bail!("Fleet run {} is already terminal ({lifecycle:?})", run_id.0);
        }
        if !matches!(lifecycle, FleetRunStatus::Running) {
            self.ledger
                .update_run_status(run_id, FleetRunStatus::Running, &timestamp())?;
        }
        let state = self.ledger.rebuild_state()?;
        let snapshot = self.status_from_state(Some(run_id), &state);
        Ok(FleetRunReport {
            run_id: run.id,
            task_count: run.task_specs.len(),
            leased: 0,
            queued: snapshot.queued,
            worker_ids: run
                .worker_specs
                .iter()
                .map(|worker| worker.id.clone())
                .collect(),
            warnings: Vec::new(),

View on GitHub (pinned to 73e0f67d83)