{"record":{"id":"7100f2fd64bb9873","repo":"JuliusBrussee/caveman","slug":"where-dataset-carries-rolecounts-adversarial","errorCode":null,"errorMessage":"${where} dataset carries ${roleCounts.adversarial} adversarial cases","messagePattern":"(.+?) dataset carries (.+?) adversarial cases","errorType":"exception","errorClass":"Error","httpStatus":null,"severity":"error","filePath":"packages/shared/contracts/scripts/validate-continuous-improvement.mjs","lineNumber":306,"sourceCode":"    // evidence has no threshold to sit beside, so those two cases must NOT be\n    // generated for it — a \"boundary\" case against a guard that does not exist\n    // tests nothing while counting as adversarial coverage.\n    const confidenceGuard = item.change_set.applicability.all.some((condition) =>\n      condition === \"failure_location_confidence >= 0.90\" ||\n      condition === \"symbol_resolution == unique\" ||\n      condition === \"targeted_test_reproduces == true\");\n    const perturbations = new Set(dataset.cases.map((entry) => entry.perturbation));\n    for (const [perturbation, label] of [[\"guard_threshold_boundary\", \"boundary\"], [\"stale_input\", \"stale-input\"]]) {\n      if (perturbations.has(perturbation) !== confidenceGuard) {\n        throw new Error(`${where} ${confidenceGuard ? \"omits\" : \"generated\"} a ${label} case ${confidenceGuard ? \"for\" : \"against\"} a change set that ${confidenceGuard ? \"declares\" : \"declares no\"} confidence guard`);\n      }\n    }\n    const expectedTargets = Math.min(4, dataset.target_failure_cases.length);\n    const expectedPriors = Math.min(2, dataset.prior_success_cases.length);\n    if (roleCounts.target_failure !== expectedTargets) throw new Error(`${where} dataset carries ${roleCounts.target_failure} target-failure cases, want ${expectedTargets}`);\n    if (roleCounts.prior_success !== expectedPriors) throw new Error(`${where} dataset carries ${roleCounts.prior_success} prior-success cases, want ${expectedPriors}`);\n    if (roleCounts.boundary > 1) throw new Error(`${where} dataset carries ${roleCounts.boundary} boundary cases, want at most one`);\n    if (roleCounts.adversarial > (confidenceGuard ? 2 : 1)) throw new Error(`${where} dataset carries ${roleCounts.adversarial} adversarial cases`);\n    if (roleCounts.boundary !== dataset.boundary_cases.length) throw new Error(`${where} boundary case list disagrees with the composed dataset`);\n    if (roleCounts.adversarial !== dataset.generated_cases.length) throw new Error(`${where} generated case list disagrees with the composed dataset`);\n    const composed = roleCounts.target_failure + roleCounts.prior_success + roleCounts.boundary + roleCounts.adversarial;\n    if (composed !== dataset.cases.length) throw new Error(`${where} dataset roles (${composed}) do not account for its ${dataset.cases.length} cases`);\n    if (dataset.replay_case_ids.length !== datasetCaseIDs.size) throw new Error(`${where} replay manifest carries ${dataset.replay_case_ids.length} ids for ${datasetCaseIDs.size} dataset cases`);\n    for (const datasetCaseID of dataset.replay_case_ids) {\n      if (!datasetCaseIDs.has(datasetCaseID)) throw new Error(`${where} replay manifest case ${datasetCaseID} is not a composed dataset case`);\n    }\n\n    // The guard grader exists exactly when there is a perturbed case to catch a\n    // candidate on; a required grader with no case behind it is decoration.\n    const perturbed = dataset.cases.some((entry) => entry.perturbation !== \"none\");\n    if (perturbed !== item.eval_pack.graders.includes(\"guard_respected\")) {\n      throw new Error(`${where} guard_respected grader ${perturbed ? \"missing for\" : \"declared without\"} perturbed dataset cases`);\n    }\n    for (const proof of item.replay.trial_proofs) {\n      if (!datasetCaseIDs.has(proof.dataset_case_id)) throw new Error(`${where} replay trial ${proof.id} replays a case outside the composed dataset`);\n      for (const arm of [proof.baseline, proof.candidate]) {","sourceCodeStart":288,"sourceCodeEnd":324,"githubUrl":"https://github.com/JuliusBrussee/caveman/blob/27d5a3981a347890211bb1bf2439e5c821a63bc9/packages/shared/contracts/scripts/validate-continuous-improvement.mjs#L288-L324","documentation":"Thrown when the count of role \"adversarial\" cases exceeds the cap, which is 2 when the change set declares a confidence guard and 1 when it does not. The guard-aware cap exists because a guarded change set legitimately earns one extra adversarial slot (the stale-input case); an unguarded one must not pad its adversarial coverage.","triggerScenarios":"Three adversarial cases on a guarded change set; two adversarial cases on an unguarded one (e.g., both misleading_tool_output and stale_input kept after removing the guard from applicability).","commonSituations":"Removing a confidence-guard condition from applicability without recomposing the dataset; adding extra adversarial cases by hand.","solutions":["Count role \"adversarial\" entries and trim to 2 with a guard, 1 without.","If a guard was removed, drop the stale-input case (and boundary case — see error 495) along with it.","Regenerate the dataset from the current applicability rather than accumulating cases."],"exampleFix":"// before: no guard in applicability.all\nroles: [\"adversarial\",\"adversarial\"]\n\n// after: unguarded cap is 1\nroles: [\"adversarial\"]","handlingStrategy":"validation","validationCode":"const GUARD_CONDITIONS = new Set([\n  \"failure_location_confidence >= 0.90\",\n  \"symbol_resolution == unique\",\n  \"targeted_test_reproduces == true\",\n]);\nconst hasGuard = item.change_set.applicability.all.some((c) => GUARD_CONDITIONS.has(c));\nconst adversarialCount = dataset.cases.filter((c) => c.role === \"adversarial\").length;\nconst ok = adversarialCount <= (hasGuard ? 2 : 1);","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Compute the adversarial cap from the current applicability at composition time; never accumulate cases across edits.","When removing a guard condition, remove the stale-input and boundary cases it justified in the same change."],"tags":["validation","dataset","composition"],"backgroundTag":null,"analyzedSha":"27d5a3981a347890211bb1bf2439e5c821a63bc9","analyzedAt":"2026-08-15T09:26:11.751Z","schemaVersion":2},"datasetVersion":"2026-08-15T22:17:37.221Z"}