influxdata/influxdb · critical
there should be a node to update
Error message
there should be a node to update
What it means
A panic from `.expect()` when applying NodeCatalogOp::StopNode. The match arm only runs after `get_by_id` returned Some, but `update()` then failed to find the node, so the repository state changed between the get and the update. This indicates an inconsistency that the catalog treats as unrecoverable for that operation.
Solutions
- Ensure catalog mutations are serialized (single writer) so get-then-update pairs are atomic.
- Check whether another code path deletes nodes while StopNode batches are being applied.
- Restore the node entry or replay the catalog from a consistent snapshot.
- If a benign race, convert to the existing None-arm behavior (warn and skip) instead of panicking.
Example fix
// before
self.nodes
.update(node_batch.node_catalog_id, new_node)
.expect("there should be a node to update");
// after
match self.nodes.update(node_batch.node_catalog_id, new_node) {
Some(_) => true,
None => {
warn!(node_id = %node_batch.node_catalog_id, "node vanished before stop; skipping");
false
}
} Defensive patterns
Strategy: validation
Validate before calling
if nodes.get_by_id(&node_batch.node_catalog_id).is_none() {
// skip stop op or warn, mirroring the existing None arm
} Type guard
fn can_stop(nodes: &Nodes, id: NodeId) -> bool { nodes.get_by_id(&id).is_some() } Prevention
- Serialize get-then-update pairs on the node repository
- Do not delete nodes concurrently with StopNode batches
- Fall back to the warn-and-skip path instead of panicking on races
- Watch multi-node deployments for simultaneous stop/delete ops
When it happens
Trigger: StopNode applied for a node_catalog_id present at the get_by_id check but absent at update time — concurrent removal of the node or a race between replay threads.
Common situations: Concurrent StopNode and node deletion in an Enterprise multi-node deployment, or manual tampering with the catalog repository during replay.
Understand the failure class
Background: "This is a bug, please report it": internal invariant violations, unreachable panics, and SNH errors explained — this error's family across 47 libraries.
Related errors
- auto field family exists
- column id in series key should be valid
- database should exist by id
- existing database should be updated
- existing node should update
AI-assisted analysis of influxdata/influxdb@06200ef96b (2026-09-19).
Data as JSON: /api/errors/4b7aaa3509557dce.
Report an issue: GitHub.
Appendix: source
Thrown at influxdb3_catalog/src/catalog/versions/v2.rs:1746
row_delete_predicate_version: *row_delete_predicate_version,
});
self.nodes
.insert(node_batch.node_catalog_id, new_node)
.expect("there should not already be a node");
}
true
}
NodeCatalogOp::StopNode(StopNodeLog {
stopped_time_ns, ..
}) => match self.nodes.get_by_id(&node_batch.node_catalog_id) {
Some(mut new_node) => {
let n = Arc::make_mut(&mut new_node);
n.state = NodeState::Stopped {
stopped_time_ns: *stopped_time_ns,
};
self.nodes
.update(node_batch.node_catalog_id, new_node)
.expect("there should be a node to update");
true
}
None => {
warn!(
node_id = %&node_batch.node_catalog_id,
"cannot find node id, skipping stop operation"
);
false
}
},
};
}
Ok(updated)
}
fn apply_restore_batch(
&mut self,
restore_batch: &log::RestoreBatch,View on GitHub (pinned to 06200ef96b)