affaan-m/ECC · error
atomic rollback compare-and-swap failed
Error message
atomic rollback compare-and-swap failed
What it means
When the post-promotion health check fails or errors, the store rolls the active slot back to the baseline with its own compare-and-swap UPDATE (WHERE candidate_id = the just-promoted candidate). If that update affects zero rows, the promoted configuration was changed concurrently and the rollback cannot proceed safely, so the transaction aborts with this error rather than leaving ambiguous state.
Solutions
- Re-read the current active slot to determine its actual state, then decide explicitly whether to re-attempt promotion or restore the baseline, and retry.
- Eliminate concurrent writers to active_harness_config: use a single coordinator or advisory lock around promotion/rollback flows.
- Audit harness_eval_audit rows to reconstruct what happened between the promotion and rollback before acting.
- If an external tool bypassed the CAS protocol, fix it to use evaluate_promote_and_health_check or the same WHERE-guarded updates.
Example fix
// before: unguarded external overwrite causing CAS failure UPDATE active_harness_config SET candidate_id = 'x' WHERE slot = 'default'; -- bypasses CAS // after: only mutate via guarded updates / promotion API UPDATE active_harness_config SET candidate_id = ?1, updated_at = ?2 WHERE slot = 'default' AND candidate_id = ?3; -- CAS, matches store semantics -- or simply: store.evaluate_promote_and_health_check(...)
Defensive patterns
Strategy: try-catch
Validate before calling
// Ensure no other writers: acquire your deployment lock before promoting/rolling back let _guard = deployment_lock.lock().await; // confirm expected active id before the call let active: Option<String> = /* query active_harness_config WHERE slot = 'default' */; anyhow::ensure!(active.is_some(), "no active config to promote against");
Try / catch
match store.evaluate_promote_and_health_check(...) {
Err(e) if e.to_string().contains("atomic rollback compare-and-swap failed") => {
// another writer changed the slot mid-flight: re-read the slot,
// reconcile (promote again or restore baseline), and audit harness_eval_audit
}
other => other?,
} Prevention
- Never write active_harness_config outside the library's CAS-guarded updates
- Use a single coordinator/lock for all promotion and rollback operations
- After any CAS failure, read the audit log to reconstruct state before retrying
When it happens
Trigger: A concurrent writer (another promotion, rollback, or manual UPDATE) changing active_harness_config between the promotion UPDATE and the rollback UPDATE; an external tool rewriting the slot inside that window; nested/parallel evaluate_promote_and_health_check calls on the same store.
Common situations: Two operators acting on the same deployment simultaneously — one promoting, one investigating/rolling back; automation that forcibly overwrites active_harness_config without the compare-and-swap protocol; a watchdog process repairing the slot mid-transaction.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- atomic promotion compare-and-swap failed
- baseline is not the active harness configuration
- an active harness configuration already exists
- candidate id collision with different immutable content
- Context graph observation #
AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16).
Data as JSON: /api/errors/abfa20aa6000bd4c.
Report an issue: GitHub.
Appendix: source
Thrown at ecc2/src/session/store.rs:5529
anyhow::bail!("health check result does not match persisted assertion");
}
Ok(healthy)
});
let healthy = matches!(health_result, Ok(true));
let event_type = match &health_result {
Ok(true) => "promoted",
Ok(false) => "promotion_rolled_back",
Err(_) => "health_check_error_rolled_back",
};
let health_check_status = match &health_result {
Ok(true) => "healthy",
Ok(false) => "unhealthy",
Err(_) => "error",
};
if !healthy {
let restored = tx.execute("UPDATE active_harness_config SET candidate_id = ?1, updated_at = ?2 WHERE slot = 'default' AND candidate_id = ?3", rusqlite::params![stored_baseline_id, now, stored_candidate_id])?;
if restored != 1 {
anyhow::bail!("atomic rollback compare-and-swap failed");
}
}
let health_json = health_evidence.canonical_json()?;
let health_digest = health_evidence.digest()?;
tx.execute("INSERT INTO harness_evaluations (candidate_id, baseline_id, evaluator, samples_json, policy_json, comparison_json, evidence_ref, health_evidence_json, health_evidence_sha256, asserted_health, health_check_status, legacy_unverifiable, created_at) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, 0, ?12)", rusqlite::params![stored_candidate_id, stored_baseline_id, evaluator, serde_json::to_string(samples)?, serde_json::to_string(&policy)?, serde_json::to_string(&comparison)?, evidence_ref, health_json, health_digest, health_evidence.asserted_healthy, health_check_status, now])?;
let evaluation_id = tx.last_insert_rowid();
tx.execute("INSERT INTO harness_eval_audit (event_type, candidate_id, prior_candidate_id, evaluation_id, evidence_ref, health_evidence_json, health_evidence_sha256, asserted_health, health_check_status, legacy_unverifiable, created_at) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, 0, ?10)", rusqlite::params![event_type, stored_candidate_id, stored_baseline_id, evaluation_id, evidence_ref, health_json, health_digest, health_evidence.asserted_healthy, health_check_status, now])?;
tx.commit()?;
let failures = match health_result {
Ok(true) => Vec::new(),
Ok(false) => vec!["post-promotion health check returned false".to_string()],
Err(error) => vec![format!("health check error: {error:#}")],
};
Ok(HarnessPromotionOutcome {
evaluation_id: Some(evaluation_id),
promoted: healthy,
rolled_back: !healthy,
failures,View on GitHub (pinned to 8321021c54)