Hmbown/CodeWhale · error

Fleet run does not exist

Error message

Fleet run {} does not exist

What it means

Thrown by launch_queued_work in the Fleet scheduler when the ledger state has no run matching the requested run_id. The scheduler rebuilds authoritative state from the ledger on each tick and looks up the run by id; an absent id means the run record was never created, was completed and pruned, or the id is wrong. It is a lookup failure before any worker launch decision.

Solutions

  1. Re-check the run id against the ledger (`fleet run list` or equivalent) and use a valid, existing run id.
  2. Re-create the Fleet run if its ledger entry was deleted and work must still execute.
  3. Discard stale queued work referencing the nonexistent run instead of retrying the tick.
  4. If the run should exist, inspect ledger storage for corruption or an incorrect data directory (e.g. wrong profile/home pointing at another ledger).

Example fix

// before
scheduler.tick_run(RunId("run-1042"))?; // deleted
// after
if ledger.rebuild_state()?.runs.contains_key("run-1042") {
    scheduler.tick_run(RunId("run-1042"))?;
}
Defensive patterns

Strategy: try-catch

Validate before calling

// Rust
let state = ledger.rebuild_state()?;
if !state.runs.contains_key(&run_id.0) {
    eprintln!("run {} not in ledger; skipping", run_id.0);
    return;
}

Try / catch

// Rust
if let Err(e) = scheduler.tick_run(run_id) {
    if e.to_string().contains("does not exist") {
        log::warn!("stale run {}: dropping queued work", run_id.0);
    } else { return Err(e); }
}

Prevention

When it happens

Trigger: tick_run calls launch_queued_work with a run id whose entry is missing from ledger.rebuild_state(): launching work for a run that was cancelled/cleaned between queueing and the tick, or passing a stale/mistyped run id.

Common situations: A queued run was deleted from the ledger while tasks were still pending; scheduler restarted against a fresh/rotated ledger holding different run ids; an off-by-one or truncated id copied from a UI.

Understand the failure class

Background: "Not found" and "does not exist" errors: why "Task not found", "No such folder", and "Can't find" fire when a lookup comes back empty — this error's family across 14 libraries.

Related errors


AI-assisted analysis of Hmbown/CodeWhale@73e0f67d83 (2026-09-22). Data as JSON: /api/errors/c0c2c00ae9da2d6d. Report an issue: GitHub.

Appendix: source

Thrown at crates/tui/src/fleet/scheduler.rs:407

            worker_id,
            task.entry.attempts,
            task_spec,
            FleetAlertEventClass::RestartExhausted,
        )?;
        Ok(())
    }

    fn launch_queued_work(
        &self,
        run_id: &FleetRunId,
        report: &mut FleetSchedulerReport,
    ) -> Result<()> {
        loop {
            let state = self.ledger.rebuild_state()?;
            let run = state
                .runs
                .get(&run_id.0)
                .ok_or_else(|| anyhow!("Fleet run {} does not exist", run_id.0))?;
            let active = active_tasks_for_run(&state, run_id);
            if active.len() >= self.policy.max_workers_per_run {
                return Ok(());
            }
            let counts = active_counts(&state, run);
            let Some((worker_id, task)) = self.next_launch(run, &state, &counts) else {
                return Ok(());
            };
            let lease_expires_at = self.lease_expires_at();
            if !self.ledger.start_task_if_enqueued(
                &task.entry.run_id,
                &task.entry.task_id,
                &worker_id,
                &self.timestamp(),
                Some(&lease_expires_at),
                Some(self.policy.max_workers_per_run),
                vec![
                    FleetWorkerEventPayload::Leased {

View on GitHub (pinned to 73e0f67d83)