{"record":{"id":"ab84b7dbf48108fa","repo":"JuliusBrussee/caveman","slug":"cave-output-budget-exceeded","errorCode":"cave_output_budget_exceeded","errorMessage":"output ${usage.outputTokens} exceeded budget ${definition.output.maxTokens}","messagePattern":"output (.+?) exceeded budget (.+?)","errorType":"error_code","errorClass":"CavemanRunError","httpStatus":null,"severity":"error","filePath":"packages/agent/src/claude-runtime.ts","lineNumber":360,"sourceCode":"      }\n      if (!Value.Check(definition.output.schema, parsed)) {\n        throw carry(\"cave_output_schema_mismatch\", \"the SDK output did not match the declared schema\");\n      }\n    }\n    // Past the success gate, usage must be accountable. A success whose usage\n    // validateProviderUsage rejected is a real evidence failure — surface it\n    // carrying the receipt, never mask it.\n    if (usage === undefined) {\n      throw carry(\n        usageError?.message ?? \"cave_claude_terminal_missing\",\n        \"the SDK reported no accountable usage on a successful result\",\n      );\n    }\n    if (definition.output !== undefined && usage.outputTokens > definition.output.maxTokens) {\n      // The SDK already ran and spent by the time its aggregate usage is known,\n      // so this post-hoc budget breach carries the receipt of what was spent\n      // rather than throwing it away.\n      throw carry(\n        \"cave_output_budget_exceeded\",\n        `output ${usage.outputTokens} exceeded budget ${definition.output.maxTokens}`,\n      );\n    }\n    return {\n      runId: runID,\n      agentId: definition.id,\n      text,\n      contextIR: lowered.ir,\n      contextBill: bill,\n      cachePrefixSHA256: prefixSHA256,\n      cacheBoundaryKnown: false,\n      cacheBust: false,\n      usageBasis: \"provider_reported\",\n      inputTokens: usage.inputTokens,\n      outputTokens: usage.outputTokens,\n      cacheReadTokens: usage.cacheReadTokens,\n      cacheWriteTokens: usage.cacheWriteTokens,","sourceCodeStart":342,"sourceCodeEnd":378,"githubUrl":"https://github.com/JuliusBrussee/caveman/blob/3ee70a102609e550bd2e68004bf5990a9341c851/packages/agent/src/claude-runtime.ts#L342-L378","documentation":"After a successful, evidenced run, if `definition.output.maxTokens` was declared and the SDK's aggregate `usage.outputTokens` exceeds it, the executor throws post-hoc. The SDK had already run and spent — the error carries the receipt of what was actually spent rather than discarding it. This is a hard terminal ceiling on provider-reported output, not a soft clamp.","triggerScenarios":"Calling the Claude executor with `output: { schema?, maxTokens }` where the model's total generated output across the whole multi-turn SDK loop (all assistant messages, including tool-call arguments) exceeds `maxTokens`. Setting `maxTokens` near typical single-call limits while the SDK runs many turns is the classic trigger.","commonSituations":"Underestimating that `maxTokens` applies to the aggregate of the SDK's whole agent loop, not one completion; verbose tool calls with large arguments; reasoning-heavy models emitting long thinking/output; tasks that legitimately need long reports.","solutions":["Raise `definition.output.maxTokens` to a realistic aggregate ceiling for the whole SDK loop.","Reduce turns: narrow the task prompt or cap tool fan-out so total generated output fits.","If the model is emitting large reasoning/tool payloads, switch to a model with tighter output or disable verbose tool schemas.","Treat the carried receipt to see what actually exceeded the budget and size `maxTokens` from measured usage."],"exampleFix":"// before\noutput: { maxTokens: 1024, schema }\n// after (aggregate across the whole SDK loop)\noutput: { maxTokens: 8192, schema }","handlingStrategy":"validation","validationCode":"// Size the aggregate budget from a measured receipt of a comparable run.\nconst measured = priorRun.receipt.calls.reduce((n, c) => n + c.outputTokens, 0);\nconst maxTokens = Math.ceil(measured * 1.5); // 50% headroom over observed aggregate\nif (maxTokens < 1024) throw new Error(\"maxTokens too small for an SDK loop\");","typeGuard":"function aggregateFits(calls: { outputTokens: number }[], maxTokens: number): boolean {\n  return calls.reduce((n, c) => n + c.outputTokens, 0) <= maxTokens;\n}","tryCatchPattern":"try {\n  const r = await runClaudeAgent(opts);\n} catch (e) {\n  if (e instanceof CavemanRunError && e.code === \"cave_output_budget_exceeded\") {\n    // e carries the receipt: read real outputTokens and raise maxTokens accordingly\n  }\n  throw e;\n}","preventionTips":["Remember maxTokens bounds the whole SDK agent loop, not one completion — budget aggregate.","Derive maxTokens from measured receipts of prior runs plus headroom.","Reduce loop verbosity: tighter prompts, fewer fan-out tool calls.","Monitor receipt outputTokens trends to catch creep before a breach."],"tags":["budget","tokens","output-limit","claude-agent-sdk"],"backgroundTag":"value-out-of-range","analyzedSha":"3ee70a102609e550bd2e68004bf5990a9341c851","analyzedAt":"2026-09-20T15:53:39.229Z","contentChangedAt":"2026-09-20T15:53:39.229Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}