Hmbown/CodeWhale · error · RuntimeError
Codewhale terminal receipt model did not match rollout
Error message
Codewhale terminal receipt model did not match rollout
What it means
Post-run receipt validation requires terminal.model to equal ctx.model — the model name the rollout requested via OPENAI_MODEL/CODEWHALE_MODEL and the --model argv flag. A mismatch means the binary ran a different model than the evaluation rollout expects, invalidating the trace.
Solutions
- Compare terminal.model in the receipt with ctx.model and align them exactly (string-equal)
- Remove any CODEWHALE_MODEL/OPENAI_MODEL overrides in config.resolved_env that differ from ctx.model
- Use the pinned binary (drop custom binary_path) so model forwarding matches the harness contract
Defensive patterns
Strategy: try-catch
Try / catch
try:
result = await harness.launch(ctx, trace, runtime, endpoint, secret, mcp_urls)
except RuntimeError as e:
if "model did not match rollout" in str(e):
# compare receipt terminal.model with ctx.model and fix overrides
...
else:
raise Prevention
- Avoid CODEWHALE_MODEL/OPENAI_MODEL overrides that diverge from ctx.model
- Use exact model names without aliases or vendor prefixes
- Confirm any custom binary forwards --model verbatim
When it happens
Trigger: The binary ignores --model, normalizes the model name differently (e.g. aliases or adds a prefix), or a stale env value overrides the model so the receipt reports a name differing from ctx.model during launch().
Common situations: Model aliasing in the local binary's config; ctx.model changed between request construction and launch; a custom binary_path whose config hardcodes a model; case/format differences like 'gpt-4o-mini' vs 'openai/gpt-4o-mini'.
Understand the failure class
Background: "invalid response format", "malformed payload", "missing data field": when an API returns 200 but the response shape is wrong — this error's family across 23 libraries.
Related errors
- Codewhale successful run contained an error event
- Codewhale terminal receipt did not confirm auto tools
- Codewhale terminal receipt did not use provider openai
- Codewhale terminal receipt sandbox did not match launch
- Choose both a provider and a model.
AI-assisted analysis of Hmbown/CodeWhale@433685b202 (2026-09-15).
Data as JSON: /api/errors/945289a603679330.
Report an issue: GitHub.
Appendix: source
Thrown at integrations/verifiers-codewhale/codewhale_harness/harness.py:230
"--output-format",
"stream-json",
]
if self.config.max_turns is not None:
argv.extend(["--max-turns", str(self.config.max_turns)])
if self.config.disabled_tools:
argv.extend(["--disallowed-tools", ",".join(self.config.disabled_tools)])
if system:
argv.extend(["--append-system-prompt", system])
argv.extend(["--", str(prompt or "")])
result = await runtime.run_program(argv, env)
if result.exit_code == 0:
receipt = _parse_stream_receipt(result.stdout)
terminal = receipt["terminal"]
if terminal.get("provider") != "openai":
raise RuntimeError("Codewhale terminal receipt did not use provider openai")
if terminal.get("model") != ctx.model:
raise RuntimeError("Codewhale terminal receipt model did not match rollout")
if terminal.get("approval_posture") != "auto_tools":
raise RuntimeError("Codewhale terminal receipt did not confirm auto tools")
if terminal.get("sandbox_posture") != sandbox:
raise RuntimeError("Codewhale terminal receipt sandbox did not match launch")
if receipt["events"].get("error", 0) != 0:
raise RuntimeError("Codewhale successful run contained an error event")
if terminal.get("status") != "completed":
raise RuntimeError("Codewhale terminal receipt did not report completion")
if terminal.get("termination_reason") != "resolved":
raise RuntimeError("Codewhale terminal receipt was not resolved")
trace.info["codewhale"] = receipt
return result
def _has_version(output: str, version: str) -> bool:
return (
re.search(
rf"(?<![0-9A-Za-z.+-]){re.escape(version)}(?![0-9A-Za-z.+-])",View on GitHub (pinned to 433685b202)