JuliusBrussee/caveman · error · Error
${where} dataset carries ${roleCounts.target_failure} target
Error message
${where} dataset carries ${roleCounts.target_failure} target-failure cases, want ${expectedTargets} What it means
Thrown when the number of dataset cases with role "target_failure" does not equal min(4, dataset.target_failure_cases.length). Spec 18.2 caps recorded target-failure replays at four and requires every listed target-failure case to actually appear (with that role) in dataset.cases.
Source
Thrown at packages/shared/contracts/scripts/validate-continuous-improvement.mjs:303
//
// The boundary and stale-input generators perturb the localization evidence
// a confidence guard reads. A ChangeSet whose applicability never reads that
// evidence has no threshold to sit beside, so those two cases must NOT be
// generated for it — a "boundary" case against a guard that does not exist
// tests nothing while counting as adversarial coverage.
const confidenceGuard = item.change_set.applicability.all.some((condition) =>
condition === "failure_location_confidence >= 0.90" ||
condition === "symbol_resolution == unique" ||
condition === "targeted_test_reproduces == true");
const perturbations = new Set(dataset.cases.map((entry) => entry.perturbation));
for (const [perturbation, label] of [["guard_threshold_boundary", "boundary"], ["stale_input", "stale-input"]]) {
if (perturbations.has(perturbation) !== confidenceGuard) {
throw new Error(`${where} ${confidenceGuard ? "omits" : "generated"} a ${label} case ${confidenceGuard ? "for" : "against"} a change set that ${confidenceGuard ? "declares" : "declares no"} confidence guard`);
}
}
const expectedTargets = Math.min(4, dataset.target_failure_cases.length);
const expectedPriors = Math.min(2, dataset.prior_success_cases.length);
if (roleCounts.target_failure !== expectedTargets) throw new Error(`${where} dataset carries ${roleCounts.target_failure} target-failure cases, want ${expectedTargets}`);
if (roleCounts.prior_success !== expectedPriors) throw new Error(`${where} dataset carries ${roleCounts.prior_success} prior-success cases, want ${expectedPriors}`);
if (roleCounts.boundary > 1) throw new Error(`${where} dataset carries ${roleCounts.boundary} boundary cases, want at most one`);
if (roleCounts.adversarial > (confidenceGuard ? 2 : 1)) throw new Error(`${where} dataset carries ${roleCounts.adversarial} adversarial cases`);
if (roleCounts.boundary !== dataset.boundary_cases.length) throw new Error(`${where} boundary case list disagrees with the composed dataset`);
if (roleCounts.adversarial !== dataset.generated_cases.length) throw new Error(`${where} generated case list disagrees with the composed dataset`);
const composed = roleCounts.target_failure + roleCounts.prior_success + roleCounts.boundary + roleCounts.adversarial;
if (composed !== dataset.cases.length) throw new Error(`${where} dataset roles (${composed}) do not account for its ${dataset.cases.length} cases`);
if (dataset.replay_case_ids.length !== datasetCaseIDs.size) throw new Error(`${where} replay manifest carries ${dataset.replay_case_ids.length} ids for ${datasetCaseIDs.size} dataset cases`);
for (const datasetCaseID of dataset.replay_case_ids) {
if (!datasetCaseIDs.has(datasetCaseID)) throw new Error(`${where} replay manifest case ${datasetCaseID} is not a composed dataset case`);
}
// The guard grader exists exactly when there is a perturbed case to catch a
// candidate on; a required grader with no case behind it is decoration.
const perturbed = dataset.cases.some((entry) => entry.perturbation !== "none");
if (perturbed !== item.eval_pack.graders.includes("guard_respected")) {
throw new Error(`${where} guard_respected grader ${perturbed ? "missing for" : "declared without"} perturbed dataset cases`);
}View on GitHub (pinned to 27d5a3981a)
Solutions
- Make target-failure entries in dataset.cases exactly match dataset.target_failure_cases one-for-one, capped at 4.
- If more than 4 failures exist, select the 4 to replay and list only those in both places.
- Verify each entry's role field is exactly "target_failure" (typos create unknown roles and count mismatches).
Example fix
// before: 5 cases with role "target_failure", list has 5 dataset.target_failure_cases.length = 5 // expectedTargets = 4 // after: keep 4 in both places dataset.cases role "target_failure" count = 4 dataset.target_failure_cases.length = 4
Defensive patterns
Strategy: validation
Validate before calling
const targetCount = dataset.cases.filter((c) => c.role === "target_failure").length; const expected = Math.min(4, dataset.target_failure_cases.length); const ok = targetCount === expected;
Prevention
- Compose dataset.cases and the role lists (target_failure_cases, prior_success_cases, boundary_cases, generated_cases) in one pass so they cannot disagree.
- Cap target-failure selection at 4 in the composer, not after.
When it happens
Trigger: dataset.target_failure_cases lists more or fewer entries than appear with role target_failure in dataset.cases; or more than 4 are replayed, so the composed count (capped at 4) no longer equals the actual count.
Common situations: Appending target-failure cases to dataset.cases without updating the target_failure_cases list; listing 5+ failures when the cap is 4.
Related errors
- ${where} dataset carries ${roleCounts.prior_success} prior-s
- ${where} dataset carries ${roleCounts.boundary} boundary cas
- ${where} dataset carries ${roleCounts.adversarial} adversari
- ${at} appears twice in the dataset
- ${at} names source unit ${entry.source_unit_id} that is not
AI-assisted analysis of JuliusBrussee/caveman@27d5a3981a (2026-08-15).
Data as JSON: /api/errors/ceec37b430c58036.
Report an issue: GitHub.