{"record":{"id":"abfa20aa6000bd4c","repo":"affaan-m/ECC","slug":"atomic-rollback-compare-and-swap-failed","errorCode":null,"errorMessage":"atomic rollback compare-and-swap failed","messagePattern":"atomic rollback compare-and-swap failed","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"ecc2/src/session/store.rs","lineNumber":5529,"sourceCode":"                anyhow::bail!(\"health check result does not match persisted assertion\");\n            }\n            Ok(healthy)\n        });\n        let healthy = matches!(health_result, Ok(true));\n        let event_type = match &health_result {\n            Ok(true) => \"promoted\",\n            Ok(false) => \"promotion_rolled_back\",\n            Err(_) => \"health_check_error_rolled_back\",\n        };\n        let health_check_status = match &health_result {\n            Ok(true) => \"healthy\",\n            Ok(false) => \"unhealthy\",\n            Err(_) => \"error\",\n        };\n        if !healthy {\n            let restored = tx.execute(\"UPDATE active_harness_config SET candidate_id = ?1, updated_at = ?2 WHERE slot = 'default' AND candidate_id = ?3\", rusqlite::params![stored_baseline_id, now, stored_candidate_id])?;\n            if restored != 1 {\n                anyhow::bail!(\"atomic rollback compare-and-swap failed\");\n            }\n        }\n        let health_json = health_evidence.canonical_json()?;\n        let health_digest = health_evidence.digest()?;\n        tx.execute(\"INSERT INTO harness_evaluations (candidate_id, baseline_id, evaluator, samples_json, policy_json, comparison_json, evidence_ref, health_evidence_json, health_evidence_sha256, asserted_health, health_check_status, legacy_unverifiable, created_at) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, 0, ?12)\", rusqlite::params![stored_candidate_id, stored_baseline_id, evaluator, serde_json::to_string(samples)?, serde_json::to_string(&policy)?, serde_json::to_string(&comparison)?, evidence_ref, health_json, health_digest, health_evidence.asserted_healthy, health_check_status, now])?;\n        let evaluation_id = tx.last_insert_rowid();\n        tx.execute(\"INSERT INTO harness_eval_audit (event_type, candidate_id, prior_candidate_id, evaluation_id, evidence_ref, health_evidence_json, health_evidence_sha256, asserted_health, health_check_status, legacy_unverifiable, created_at) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, 0, ?10)\", rusqlite::params![event_type, stored_candidate_id, stored_baseline_id, evaluation_id, evidence_ref, health_json, health_digest, health_evidence.asserted_healthy, health_check_status, now])?;\n        tx.commit()?;\n        let failures = match health_result {\n            Ok(true) => Vec::new(),\n            Ok(false) => vec![\"post-promotion health check returned false\".to_string()],\n            Err(error) => vec![format!(\"health check error: {error:#}\")],\n        };\n        Ok(HarnessPromotionOutcome {\n            evaluation_id: Some(evaluation_id),\n            promoted: healthy,\n            rolled_back: !healthy,\n            failures,","sourceCodeStart":5511,"sourceCodeEnd":5547,"githubUrl":"https://github.com/affaan-m/ECC/blob/8321021c54d670126ce3b2969d5deb880b4b0c2a/ecc2/src/session/store.rs#L5511-L5547","documentation":"When the post-promotion health check fails or errors, the store rolls the active slot back to the baseline with its own compare-and-swap UPDATE (WHERE candidate_id = the just-promoted candidate). If that update affects zero rows, the promoted configuration was changed concurrently and the rollback cannot proceed safely, so the transaction aborts with this error rather than leaving ambiguous state.","triggerScenarios":"A concurrent writer (another promotion, rollback, or manual UPDATE) changing active_harness_config between the promotion UPDATE and the rollback UPDATE; an external tool rewriting the slot inside that window; nested/parallel evaluate_promote_and_health_check calls on the same store.","commonSituations":"Two operators acting on the same deployment simultaneously — one promoting, one investigating/rolling back; automation that forcibly overwrites active_harness_config without the compare-and-swap protocol; a watchdog process repairing the slot mid-transaction.","solutions":["Re-read the current active slot to determine its actual state, then decide explicitly whether to re-attempt promotion or restore the baseline, and retry.","Eliminate concurrent writers to active_harness_config: use a single coordinator or advisory lock around promotion/rollback flows.","Audit harness_eval_audit rows to reconstruct what happened between the promotion and rollback before acting.","If an external tool bypassed the CAS protocol, fix it to use evaluate_promote_and_health_check or the same WHERE-guarded updates."],"exampleFix":"// before: unguarded external overwrite causing CAS failure\nUPDATE active_harness_config SET candidate_id = 'x' WHERE slot = 'default'; -- bypasses CAS\n\n// after: only mutate via guarded updates / promotion API\nUPDATE active_harness_config SET candidate_id = ?1, updated_at = ?2\nWHERE slot = 'default' AND candidate_id = ?3; -- CAS, matches store semantics\n-- or simply: store.evaluate_promote_and_health_check(...)","handlingStrategy":"try-catch","validationCode":"// Ensure no other writers: acquire your deployment lock before promoting/rolling back\nlet _guard = deployment_lock.lock().await;\n// confirm expected active id before the call\nlet active: Option<String> = /* query active_harness_config WHERE slot = 'default' */;\nanyhow::ensure!(active.is_some(), \"no active config to promote against\");","typeGuard":null,"tryCatchPattern":"match store.evaluate_promote_and_health_check(...) {\n    Err(e) if e.to_string().contains(\"atomic rollback compare-and-swap failed\") => {\n        // another writer changed the slot mid-flight: re-read the slot,\n        // reconcile (promote again or restore baseline), and audit harness_eval_audit\n    }\n    other => other?,\n}","preventionTips":["Never write active_harness_config outside the library's CAS-guarded updates","Use a single coordinator/lock for all promotion and rollback operations","After any CAS failure, read the audit log to reconstruct state before retrying"],"tags":["concurrency","optimistic-locking","database","rust"],"backgroundTag":"invalid-state-transition","analyzedSha":"8321021c54d670126ce3b2969d5deb880b4b0c2a","analyzedAt":"2026-09-16T10:08:13.343Z","contentChangedAt":"2026-09-16T10:08:13.343Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}