JuliusBrussee/caveman · error · CavemanRunError

cave_output_schema_mismatch

cave_output_schema_mismatch

Error message

the SDK output did not match the declared schema

What it means

After the SDK output parses as JSON, it is validated with TypeBox's `Value.Check` against the agent's declared `output.schema`. A parse that produces JSON of the wrong shape fails this gate, throwing a CavemanRunError that carries the spend receipt. The library throws because the caller's declared output contract is part of the run's evidence; silently returning malformed data would break downstream consumers.

Solutions

  1. Log/inspect the parsed value and relax or correct the TypeBox schema to match what the model actually returns (e.g. allow `nullable`, fix field types).
  2. Strengthen the prompt with the exact JSON shape, field names, and types, ideally embedding the schema itself.
  3. Add a retry pass: re-invoke the agent asking it to fix its JSON to the schema when validation fails.
  4. Use TypeBox options deliberately (`additionalProperties`, `minimum`, etc.) so borderline outputs aren't rejected.

Example fix

// before
schema: Type.Object({ count: Type.Number() }) // model returned {"count": "12"}
// after
prompt: '... Return {"count": <number>} — count must be a JSON number, not a string.'
Defensive patterns

Strategy: validation

Validate before calling

import { Value } from "typebox/value";
// Validate the schema itself is what you intend before runs.
const schema = Type.Object({ count: Type.Number() });
if (!Value.Check(schema, { count: 0 })) throw new Error("declared schema rejects its own example");

Type guard

function matchesOutput(v: unknown, schema: TSchema): boolean {
  return Value.Check(schema, v);
}

Try / catch

try {
  const r = await runClaudeAgent({ ...opts });
} catch (e) {
  if (e instanceof CavemanRunError && e.code === "cave_output_schema_mismatch") {
    // retry once with a corrective message embedding the schema
  }
  throw e;
}

Prevention

When it happens

Trigger: Declaring `definition.output.schema` and receiving JSON that parses but violates it: missing required fields, wrong types (string vs number), extra/null fields under strict schemas, or the model returning an array where an object was declared.

Common situations: Schemas made stricter after agents were already prompted; models omitting optional-but-required-in-schema fields; numeric IDs returned as strings; LLM hallucinating extra wrapper keys like `{"result": {...}}` around the expected object.

Understand the failure class

Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.

Related errors


AI-assisted analysis of JuliusBrussee/caveman@3ee70a1026 (2026-09-20). Data as JSON: /api/errors/ec02b6cc7cdd54b0. Report an issue: GitHub.

Appendix: source

Thrown at packages/agent/src/claude-runtime.ts:344

        `the Claude Agent SDK ended with ${result?.subtype ?? "no result"}`,
      );
    }
    if (assistantModel !== undefined && normalizeClaudeModel(assistantModel) !== selected) {
      throw carry(
        "cave_provider_model_identity_mismatch",
        `expected ${selected}, the SDK answered as ${assistantModel}`,
      );
    }
    const text = result.result;
    if (definition.output?.schema !== undefined) {
      let parsed: unknown;
      try {
        parsed = result.structured_output ?? JSON.parse(text);
      } catch {
        throw carry("cave_output_schema_invalid_json", "the SDK output was not valid JSON");
      }
      if (!Value.Check(definition.output.schema, parsed)) {
        throw carry("cave_output_schema_mismatch", "the SDK output did not match the declared schema");
      }
    }
    // Past the success gate, usage must be accountable. A success whose usage
    // validateProviderUsage rejected is a real evidence failure — surface it
    // carrying the receipt, never mask it.
    if (usage === undefined) {
      throw carry(
        usageError?.message ?? "cave_claude_terminal_missing",
        "the SDK reported no accountable usage on a successful result",
      );
    }
    if (definition.output !== undefined && usage.outputTokens > definition.output.maxTokens) {
      // The SDK already ran and spent by the time its aggregate usage is known,
      // so this post-hoc budget breach carries the receipt of what was spent
      // rather than throwing it away.
      throw carry(
        "cave_output_budget_exceeded",
        `output ${usage.outputTokens} exceeded budget ${definition.output.maxTokens}`,

View on GitHub (pinned to 3ee70a1026)