JuliusBrussee/caveman · error · CavemanRunError

cave_output_budget_exceeded

cave_output_budget_exceeded

Error message

output ${usage.outputTokens} exceeded budget ${definition.output.maxTokens}

What it means

After a successful, evidenced run, if `definition.output.maxTokens` was declared and the SDK's aggregate `usage.outputTokens` exceeds it, the executor throws post-hoc. The SDK had already run and spent — the error carries the receipt of what was actually spent rather than discarding it. This is a hard terminal ceiling on provider-reported output, not a soft clamp.

Solutions

  1. Raise `definition.output.maxTokens` to a realistic aggregate ceiling for the whole SDK loop.
  2. Reduce turns: narrow the task prompt or cap tool fan-out so total generated output fits.
  3. If the model is emitting large reasoning/tool payloads, switch to a model with tighter output or disable verbose tool schemas.
  4. Treat the carried receipt to see what actually exceeded the budget and size `maxTokens` from measured usage.

Example fix

// before
output: { maxTokens: 1024, schema }
// after (aggregate across the whole SDK loop)
output: { maxTokens: 8192, schema }
Defensive patterns

Strategy: validation

Validate before calling

// Size the aggregate budget from a measured receipt of a comparable run.
const measured = priorRun.receipt.calls.reduce((n, c) => n + c.outputTokens, 0);
const maxTokens = Math.ceil(measured * 1.5); // 50% headroom over observed aggregate
if (maxTokens < 1024) throw new Error("maxTokens too small for an SDK loop");

Type guard

function aggregateFits(calls: { outputTokens: number }[], maxTokens: number): boolean {
  return calls.reduce((n, c) => n + c.outputTokens, 0) <= maxTokens;
}

Try / catch

try {
  const r = await runClaudeAgent(opts);
} catch (e) {
  if (e instanceof CavemanRunError && e.code === "cave_output_budget_exceeded") {
    // e carries the receipt: read real outputTokens and raise maxTokens accordingly
  }
  throw e;
}

Prevention

When it happens

Trigger: Calling the Claude executor with `output: { schema?, maxTokens }` where the model's total generated output across the whole multi-turn SDK loop (all assistant messages, including tool-call arguments) exceeds `maxTokens`. Setting `maxTokens` near typical single-call limits while the SDK runs many turns is the classic trigger.

Common situations: Underestimating that `maxTokens` applies to the aggregate of the SDK's whole agent loop, not one completion; verbose tool calls with large arguments; reasoning-heavy models emitting long thinking/output; tasks that legitimately need long reports.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of JuliusBrussee/caveman@3ee70a1026 (2026-09-20). Data as JSON: /api/errors/ab84b7dbf48108fa. Report an issue: GitHub.

Appendix: source

Thrown at packages/agent/src/claude-runtime.ts:360

      }
      if (!Value.Check(definition.output.schema, parsed)) {
        throw carry("cave_output_schema_mismatch", "the SDK output did not match the declared schema");
      }
    }
    // Past the success gate, usage must be accountable. A success whose usage
    // validateProviderUsage rejected is a real evidence failure — surface it
    // carrying the receipt, never mask it.
    if (usage === undefined) {
      throw carry(
        usageError?.message ?? "cave_claude_terminal_missing",
        "the SDK reported no accountable usage on a successful result",
      );
    }
    if (definition.output !== undefined && usage.outputTokens > definition.output.maxTokens) {
      // The SDK already ran and spent by the time its aggregate usage is known,
      // so this post-hoc budget breach carries the receipt of what was spent
      // rather than throwing it away.
      throw carry(
        "cave_output_budget_exceeded",
        `output ${usage.outputTokens} exceeded budget ${definition.output.maxTokens}`,
      );
    }
    return {
      runId: runID,
      agentId: definition.id,
      text,
      contextIR: lowered.ir,
      contextBill: bill,
      cachePrefixSHA256: prefixSHA256,
      cacheBoundaryKnown: false,
      cacheBust: false,
      usageBasis: "provider_reported",
      inputTokens: usage.inputTokens,
      outputTokens: usage.outputTokens,
      cacheReadTokens: usage.cacheReadTokens,
      cacheWriteTokens: usage.cacheWriteTokens,

View on GitHub (pinned to 3ee70a1026)