risingwavelabs/risingwave · warning · MetaError

user was concurrently dropped

Error message

user {} was concurrently dropped

What it means

`ensure_user_id` verifies that a user referenced by an operation still exists by counting rows in the user table. If the count is zero the operation aborts with "user {} was concurrently dropped", because another transaction deleted the user between the caller's earlier lookup and this check. It is a deliberate race-detection guard, not a plain not-found error.

Solutions

  1. Re-run the operation; it will now fail fast with a normal 'user not found' error that can be handled cleanly.
  2. Before executing ownership/creation DDL, verify the target user exists and avoid concurrently dropping it (serialize DDL on the same user in your tooling).
  3. If this occurs in tests or CI, add synchronization between the DROP USER and the dependent DDL statements.
  4. Use a single transaction that both checks and uses the user reference so the meta layer's locking prevents the race.

Example fix

// before
let user = get_user_by_name(name).await?;
spawn_drop_user(user.id); // races
create_database(owner_id: user.id).await?;
// after
let user = get_user_by_name(name).await?;
create_database(owner_id: user.id).await?; // drop only after dependents are removed
Defensive patterns

Strategy: retry

Validate before calling

// Pre-check user existence before dependent DDL
if User::find_by_id(user_id).count(db).await? == 0 {
    return Err("user no longer exists; abort DDL");
}

Try / catch

match ensure-dependent-ddl().await {
    Err(e) if e.to_string().contains("was concurrently dropped") => retry_or_skip(),
    other => other,
}

Prevention

When it happens

Trigger: Calling `ensure_user_id` (from alter_owner, alter_secret, create_database, create_schema, create_subscription_catalog, create_source) with a UserId whose row was deleted by a concurrent DROP USER while the first operation was in flight.

Common situations: Two sessions racing: one runs ALTER ... OWNER TO <user> while another drops that user; automation scripts issuing parallel DDL against the same user; retry logic replaying an operation after the user was removed.

Understand the failure class

Background: "Not found" and "does not exist" errors: why "Task not found", "No such folder", and "Can't find" fire when a lookup comes back empty — this error's family across 14 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/f6589410ea1d282c. Report an issue: GitHub.

Appendix: source

Thrown at src/meta/src/controller/utils.rs:590

{
    let count = Object::find_by_id(job_id).count(db).await?;
    if count == 0 {
        return Err(MetaError::cancelled(format!(
            "job {} might be cancelled manually or by recovery",
            job_id
        )));
    }
    Ok(())
}

/// `ensure_user_id` ensures the existence of target user in the cluster.
pub async fn ensure_user_id<C>(user_id: UserId, db: &C) -> MetaResult<()>
where
    C: ConnectionTrait,
{
    let count = User::find_by_id(user_id).count(db).await?;
    if count == 0 {
        return Err(anyhow!("user {} was concurrently dropped", user_id).into());
    }
    Ok(())
}

/// `check_database_name_duplicate` checks whether the database name is already used in the cluster.
pub async fn check_database_name_duplicate<C>(name: &str, db: &C) -> MetaResult<()>
where
    C: ConnectionTrait,
{
    let count = Database::find()
        .filter(database::Column::Name.eq(name))
        .count(db)
        .await?;
    if count > 0 {
        assert_eq!(count, 1);
        return Err(MetaError::catalog_duplicated("database", name));
    }
    Ok(())

View on GitHub (pinned to 6469eb736d)