Hmbown/CodeWhale · error

rate-limit interruption should preserve checkpoint

Error message

rate-limit interruption should preserve checkpoint

What it means

The interrupted sub-agent's snapshot has `checkpoint: None`, so `.expect("rate-limit interruption should preserve checkpoint")` panics. The library contract is that a rate-limited interruption records a resumable checkpoint (reason "api_rate_limited", continuable) plus a needs_input marker. A missing checkpoint means the interruption path skipped persistence.

Solutions

  1. Locate the interruption handler that sets SubAgentStatus::Interrupted and confirm it also builds and stores the checkpoint
  2. Verify the checkpoint reason is set to "api_rate_limited" and continuable=true on the rate-limit path
  3. Check whether checkpoint serialization/persistence errors are swallowed, leaving checkpoint None
  4. Add an assertion or log immediately after interruption to confirm the checkpoint was written
Defensive patterns

Strategy: validation

Validate before calling

assert!(snapshot.checkpoint.is_some(), "interrupted agent must carry a checkpoint");

Type guard

fn resumable(s: &SubAgentSnapshot) -> bool { matches!(&s.status, SubAgentStatus::Interrupted(_)) && s.checkpoint.as_ref().map(|c| c.continuable).unwrap_or(false) }

Try / catch

let cp = snapshot.checkpoint.as_ref().ok_or_else(|| anyhow!("rate-limit interruption lost checkpoint for {:?}", snapshot.status))?;

Prevention

When it happens

Trigger: The sub-agent is interrupted with reason containing "rate-limited provider response" but the code path that builds the checkpoint (reason=api_rate_limited, continuable=true) was not executed — e.g. the interruption aborted before snapshotting conversation state.

Common situations: Refactoring the retry loop to bail out instead of checkpointing; a new early-return on provider 429 responses; checkpoint persistence failing silently and leaving the field None.

Understand the failure class

Background: "This is a bug, please report it": internal invariant violations, unreachable panics, and SNH errors explained — this error's family across 47 libraries.

Related errors


AI-assisted analysis of Hmbown/CodeWhale@433685b202 (2026-09-15). Data as JSON: /api/errors/214ea653acff4c7a. Report an issue: GitHub.

Appendix: source

Thrown at crates/tui/src/tools/subagent/tests.rs:9400

        "rate-limit retries should be owned by the sub-agent retry loop"
    );
    let snapshot = {
        let manager = manager.read().await;
        manager
            .get_result(&agent_id)
            .expect("agent should stay registered")
    };
    let SubAgentStatus::Interrupted(reason) = &snapshot.status else {
        panic!("expected interrupted sub-agent, got {:?}", snapshot.status);
    };
    assert!(
        reason.contains("rate-limited provider response"),
        "reason should name the provider rate limit: {reason}"
    );
    let checkpoint = snapshot
        .checkpoint
        .as_ref()
        .expect("rate-limit interruption should preserve checkpoint");
    assert_eq!(checkpoint.reason, "api_rate_limited");
    assert!(checkpoint.continuable);
    assert!(snapshot.needs_input.is_some());
}

#[tokio::test]
async fn spawn_duplicate_session_name_error_names_conflicting_agent() {
    // #2656: the duplicate-name error must identify the conflicting agent so a
    // model can recover deterministically (reuse the id, or pick a new name).
    let manager = Arc::new(RwLock::new(SubAgentManager::new(PathBuf::from("."), 5)));
    let boot_id = manager.read().await.session_boot_id().to_string();
    let (input_tx, _input_rx) = mpsc::unbounded_channel();
    let mut existing = SubAgent::new(
        "test_agent_existing".to_string(),
        FleetRole::Scout,
        "scan".to_string(),
        make_assignment(),
        "deepseek-v4-flash".to_string(),

View on GitHub (pinned to 433685b202)