JuliusBrussee/caveman · critical
: database or journal removed or replaced; connection…
Error message
%w: database or journal removed or replaced; connection quarantined until process exit
What it means
checkGeneration detected that the database or one of its journal files was removed or replaced since the Store opened. The connection is terminally quarantined: the library deliberately keeps (never closes) the stale SQLite connection until process exit, because closing could checkpoint an obsolete WAL into the new main file and corrupt the replacement database. Every subsequent Store operation and Close returns this error.
Solutions
- Stop replacing the database file under a running process: perform restores/replacements while the engine is stopped, then start it fresh.
- If the replacement was intentional, restart the process so a new Store opens the new database — automatic adoption of a replacement is deliberately unsafe.
- Check for concurrent processes creating/renaming ccr.db and serialize them or give each process its own store path.
- Restore the original file (same inode) if the swap was accidental — note quarantine is latched regardless; a process restart is still required.
- Persist recovery data outside the swapped window or snapshot the whole directory atomically (e.g. via rename of the directory, with the process restarted).
Example fix
// before mv ccr.db ccr.db.bak && cp new/ccr.db ccr.db # while process running // after systemctl stop caveman-agent mv ccr.db ccr.db.bak && cp new/ccr.db ccr.db systemctl start caveman-agent
Defensive patterns
Strategy: try-catch
Validate before calling
// compare inode before/after external operations:
before := inodeOf(dbPath) // os.Stat -> syscall.Stat_t.Ino
// ...external restore...
if inodeOf(dbPath) != before { restartEngine() } Try / catch
if errors.Is(err, ccr.ErrStorageChanged) {
// terminal quarantine: fail the request and restart the process/worker
// so a fresh Store opens the current database
os.Exit(1) // or supervisor restart
} Prevention
- Stop the engine before any db restore, replace, or rollback
- Restore the database atomically only while no process holds it open
- Give concurrent processes separate store paths
- Use supervisor restarts on ErrStorageChanged instead of retrying the same Store
When it happens
Trigger: Any Store method (Put, Get, PutObject, Summary, ...) or Close runs after the on-disk file identity (dev/inode of the main DB or existing journal files) changed — the db file was deleted and recreated, swapped by rename, or restored from a different file while the process kept its old descriptor.
Common situations: A deployment/restore job replaces ccr.db while the long-running agent process is alive; a container restart of a sidecar wipes the volume; the database is recreated by a parallel process; snapshot/restore tooling rolls the file back to a different inode.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- : inspect database
- : refusing non-regular database
- budget is below CCR storage minimum
- cave_ccr_budget_exceeded
- ccr storage budget is below minimum
AI-assisted analysis of JuliusBrussee/caveman@3ee70a1026 (2026-09-20).
Data as JSON: /api/errors/1ef25dd2bc0b70db.
Report an issue: GitHub.
Appendix: source
Thrown at engine/ccr/store_generation.go:162
if errors.Is(err, errStorageUnverifiable) {
// "Could not look" is not "was replaced". A momentary stat/permission
// failure — an antivirus lock on Windows, EIO on a network home, a
// parent directory whose mode was loose for one instant — fails THIS
// operation closed, but must not latch the terminal state below and
// make already-stored recoveries unreadable for the rest of the process.
return err
}
if err != nil {
s.quarantined = err
return s.quarantined
}
if s.db != nil && s.files.same(files) {
return nil
}
// No SQLite calls, including Close, follow invalidation. The public driver
// exposes no descriptor identity or NO_CKPT_ON_CLOSE control. Keeping this
// one connection until exit avoids replaying its obsolete journal.
s.quarantined = fmt.Errorf("%w: database or journal removed or replaced; connection quarantined until process exit", ErrStorageChanged)
return s.quarantined
}
func withStore[T any](s *Store, fn func() (T, error)) (T, error) {
s.mu.Lock()
defer s.mu.Unlock()
var zero T
if err := s.checkGeneration(); err != nil {
return zero, err
}
value, err := fn()
if changed := s.checkGeneration(); changed != nil {
return zero, changed
}
return value, err
}
// Close releases a valid connection. A quarantined connection is deliberatelyView on GitHub (pinned to 3ee70a1026)