risingwavelabs/risingwave · warning · MetaError
user was concurrently dropped
Error message
user {} was concurrently dropped What it means
`ensure_user_id` verifies that a user referenced by an operation still exists by counting rows in the user table. If the count is zero the operation aborts with "user {} was concurrently dropped", because another transaction deleted the user between the caller's earlier lookup and this check. It is a deliberate race-detection guard, not a plain not-found error.
Solutions
- Re-run the operation; it will now fail fast with a normal 'user not found' error that can be handled cleanly.
- Before executing ownership/creation DDL, verify the target user exists and avoid concurrently dropping it (serialize DDL on the same user in your tooling).
- If this occurs in tests or CI, add synchronization between the DROP USER and the dependent DDL statements.
- Use a single transaction that both checks and uses the user reference so the meta layer's locking prevents the race.
Example fix
// before let user = get_user_by_name(name).await?; spawn_drop_user(user.id); // races create_database(owner_id: user.id).await?; // after let user = get_user_by_name(name).await?; create_database(owner_id: user.id).await?; // drop only after dependents are removed
Defensive patterns
Strategy: retry
Validate before calling
// Pre-check user existence before dependent DDL
if User::find_by_id(user_id).count(db).await? == 0 {
return Err("user no longer exists; abort DDL");
} Try / catch
match ensure-dependent-ddl().await {
Err(e) if e.to_string().contains("was concurrently dropped") => retry_or_skip(),
other => other,
} Prevention
- Serialize user DDL and dependent object DDL in your tooling
- Avoid dropping users while ownership/creation operations are in flight
- In CI, synchronize DROP USER with any dependent statements
When it happens
Trigger: Calling `ensure_user_id` (from alter_owner, alter_secret, create_database, create_schema, create_subscription_catalog, create_source) with a UserId whose row was deleted by a concurrent DROP USER while the first operation was in flight.
Common situations: Two sessions racing: one runs ALTER ... OWNER TO <user> while another drops that user; automation scripts issuing parallel DDL against the same user; retry logic replaying an operation after the user was removed.
Understand the failure class
Background: "Not found" and "does not exist" errors: why "Task not found", "No such folder", and "Can't find" fire when a lookup comes back empty — this error's family across 14 libraries.
Related errors
- concurrent backup job is not supported: existent job
- id not found
- named already exists
- actor count ( ) exceeds vnode count ( )
- anyhow!(message.to_owned())
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/f6589410ea1d282c.
Report an issue: GitHub.
Appendix: source
Thrown at src/meta/src/controller/utils.rs:590
{
let count = Object::find_by_id(job_id).count(db).await?;
if count == 0 {
return Err(MetaError::cancelled(format!(
"job {} might be cancelled manually or by recovery",
job_id
)));
}
Ok(())
}
/// `ensure_user_id` ensures the existence of target user in the cluster.
pub async fn ensure_user_id<C>(user_id: UserId, db: &C) -> MetaResult<()>
where
C: ConnectionTrait,
{
let count = User::find_by_id(user_id).count(db).await?;
if count == 0 {
return Err(anyhow!("user {} was concurrently dropped", user_id).into());
}
Ok(())
}
/// `check_database_name_duplicate` checks whether the database name is already used in the cluster.
pub async fn check_database_name_duplicate<C>(name: &str, db: &C) -> MetaResult<()>
where
C: ConnectionTrait,
{
let count = Database::find()
.filter(database::Column::Name.eq(name))
.count(db)
.await?;
if count > 0 {
assert_eq!(count, 1);
return Err(MetaError::catalog_duplicated("database", name));
}
Ok(())View on GitHub (pinned to 6469eb736d)