toeverything/AFFiNE · critical · FailedToUpsertSnapshot

failed_to_upsert_snapshot

failed_to_upsert_snapshot

Error message

Failed to store doc snapshot.

What it means

The snapshot upsert path (createOrUpdate of the doc snapshot row, followed by emitting doc.snapshot.updated) is wrapped in a catch that increments the snapshot_upsert_failed metric, logs 'Failed to upsert snapshot', and rethrows FailedToUpsertSnapshot (internal_server_error / failed_to_upsert_snapshot). It means the database rejected the snapshot write after the attempt.

Solutions

  1. Read the paired server log line 'Failed to upsert snapshot' for the true database error
  2. Verify DB health, disk space, and connection-pool configuration
  3. Ensure migrations are applied and the prisma model matches the deployed schema
  4. Retry is safe: the merge job re-runs and upserts are idempotent once the underlying issue is fixed

Example fix

// before
await docAdapter.upsert(docId, snapshot);

// after
try {
  await docAdapter.upsert(docId, snapshot);
} catch (e) {
  if (e.extensions?.code === 'failed_to_upsert_snapshot') {
    metrics.increment('snapshot_retry_scheduled');
    return requeueMergeJob(spaceId, docId); // idempotent retry
  }
  throw e;
}
Defensive patterns

Strategy: retry

Validate before calling

// health-check the DB before kicking off snapshot-heavy jobs
await db.$queryRaw`SELECT 1`;
await upsertSnapshot(spaceId, docId, snapshot);

Type guard

function isFailedToUpsertSnapshot(e: unknown): boolean {
  return (e as { extensions?: { code?: string } }).extensions?.code === 'failed_to_upsert_snapshot';
}

Try / catch

try {
  await upsertSnapshot(spaceId, docId, snapshot);
} catch (e) {
  if (isFailedToUpsertSnapshot(e)) {
    return requeueMergeJob(spaceId, docId); // upsert is idempotent; safe to retry
  }
  throw e;
}

Prevention

When it happens

Trigger: Postgres outage or connection-pool exhaustion during snapshot merge; snapshot blob exceeding column/size limits; unique-constraint conflicts when two writers race despite the doc:update mutex; schema drift where the snapshot table columns no longer match the model.

Common situations: Long-running self-hosted servers hitting max_connections during bulk imports; failed migrations leaving the snapshot table in an old shape; disk-full or transient managed-DB restarts while the merge job runs.

Related errors


AI-assisted analysis of toeverything/AFFiNE@2af30773ae (2026-08-18). Data as JSON: /api/errors/d62b93b7922d2812. Report an issue: GitHub.

Appendix: source

Thrown at packages/backend/server/src/core/doc/adapters/workspace.ts:397

    if (this.isEmptyBin(snapshot.bin)) {
      return false;
    }

    try {
      const blob = Buffer.from(snapshot.bin);
      const updatedSnapshot = await this.models.doc.upsert({
        spaceId: snapshot.spaceId,
        docId: snapshot.docId,
        blob,
        timestamp: snapshot.timestamp,
        editorId: snapshot.editor,
      });

      return !!updatedSnapshot;
    } catch (e) {
      metrics.doc.counter('snapshot_upsert_failed').add(1);
      this.logger.error('Failed to upsert snapshot', e);
      throw new FailedToUpsertSnapshot();
    }
  }

  protected override async lockDocForUpdate(
    workspaceId: string,
    docId: string
  ) {
    const lock = await this.mutex.acquire(`doc:update:${workspaceId}:${docId}`);

    if (!lock) {
      throw new Error('Too many concurrent writings');
    }

    return lock;
  }

  protected async lastDocHistory(workspaceId: string, id: string) {
    return this.models.history.getLatest(workspaceId, id);

View on GitHub (pinned to 2af30773ae)