affaan-m/ECC · error
atomic promotion compare-and-swap failed
Error message
atomic promotion compare-and-swap failed
What it means
Promotion uses an optimistic compare-and-swap: a single UPDATE of active_harness_config requires the slot to still hold the baseline candidate (WHERE candidate_id = baseline) and must affect exactly one row. If the row count is not 1, the active configuration changed between the earlier read and the update (a concurrent promotion/rollback), so the promotion aborts instead of overwriting an unknown state.
Solutions
- Simply retry the whole evaluate_promote_and_health_check call: re-read the current active id and use it as the new baseline.
- Coordinate promotions (single writer, lock, or queue) so only one process mutates active_harness_config at a time.
- Inspect harness_eval_audit to see which event changed the active slot and reconcile before retrying.
- Keep transactions short and avoid long-running work (e.g. slow health checks) outside the compare-and-swap window where possible.
Example fix
// before: fire-and-forget promotion that can race
store.evaluate_promote_and_health_check(&cand, &base, ...)?;
// after: retry on CAS failure
match store.evaluate_promote_and_health_check(&cand, &base, ...) {
Err(e) if e.to_string().contains("compare-and-swap failed") => {
let current = store.active_harness_id()?.unwrap();
store.evaluate_promote_and_health_check(&cand, ¤t, ...)?;
}
r => r?,
} Defensive patterns
Strategy: retry
Validate before calling
// Preflight: confirm the slot still holds your expected baseline let active: Option<String> = /* query active_harness_config WHERE slot = 'default' */; anyhow::ensure!(active.as_deref() == Some(stored_baseline_id), "active config changed; re-base before promoting");
Try / catch
const MAX_RETRIES: usize = 3;
for attempt in 0..MAX_RETRIES {
match store.evaluate_promote_and_health_check(...) {
Err(e) if e.to_string().contains("atomic promotion compare-and-swap failed") && attempt + 1 < MAX_RETRIES => continue,
other => { other?; break }
}
} Prevention
- Single-writer discipline for active_harness_config (lock or queue promotions)
- Keep promotion transactions short; do slow work outside the CAS window
- On CAS failure, always re-read state — never assume the update landed
When it happens
Trigger: Two concurrent evaluate_promote_and_health_check calls (or a concurrent rollback) mutating active_harness_config between the transaction's baseline read and the UPDATE; a baseline that was replaced after the row was read; any external writer updating the slot within the transaction window.
Common situations: Multiple CI jobs or operators promoting candidates in parallel against the same store; a health-check-triggered rollback landing between your read and update; scripted re-runs racing each other.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- atomic rollback compare-and-swap failed
- baseline is not the active harness configuration
- an active harness configuration already exists
- candidate id collision with different immutable content
- Context graph observation #
AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16).
Data as JSON: /api/errors/d3b65828725ada7c.
Report an issue: GitHub.
Appendix: source
Thrown at ecc2/src/session/store.rs:5507
if active != stored_baseline_id {
anyhow::bail!("baseline is not the active harness configuration");
}
let now = chrono::Utc::now().to_rfc3339();
if !comparison.passed {
tx.execute("INSERT INTO harness_evaluations (candidate_id, baseline_id, evaluator, samples_json, policy_json, comparison_json, evidence_ref, legacy_unverifiable, created_at) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, 0, ?8)", rusqlite::params![stored_candidate_id, stored_baseline_id, evaluator, serde_json::to_string(samples)?, serde_json::to_string(&policy)?, serde_json::to_string(&comparison)?, evidence_ref, now])?;
let evaluation_id = tx.last_insert_rowid();
tx.execute("INSERT INTO harness_eval_audit (event_type, candidate_id, prior_candidate_id, evaluation_id, evidence_ref, legacy_unverifiable, created_at) VALUES ('promotion_rejected', ?1, ?2, ?3, ?4, 0, ?5)", rusqlite::params![stored_candidate_id, stored_baseline_id, evaluation_id, evidence_ref, now])?;
tx.commit()?;
return Ok(HarnessPromotionOutcome {
evaluation_id: Some(evaluation_id),
promoted: false,
rolled_back: false,
failures: comparison.failures,
});
}
let changed = tx.execute("UPDATE active_harness_config SET candidate_id = ?1, updated_at = ?2 WHERE slot = 'default' AND candidate_id = ?3", rusqlite::params![stored_candidate_id, now, stored_baseline_id])?;
if changed != 1 {
anyhow::bail!("atomic promotion compare-and-swap failed");
}
let health_result = health_check(candidate_id).and_then(|healthy| {
if healthy != health_evidence.asserted_healthy {
anyhow::bail!("health check result does not match persisted assertion");
}
Ok(healthy)
});
let healthy = matches!(health_result, Ok(true));
let event_type = match &health_result {
Ok(true) => "promoted",
Ok(false) => "promotion_rolled_back",
Err(_) => "health_check_error_rolled_back",
};
let health_check_status = match &health_result {
Ok(true) => "healthy",
Ok(false) => "unhealthy",
Err(_) => "error",
};View on GitHub (pinned to 8321021c54)