Mintplex-Labs/anything-llm · error · Error

Gemini Failed to embed

Error message

Gemini Failed to embed: ${error}

What it means

Thrown from GeminiEmbedder.embedChunks after the per-chunk embedding requests complete. The class fans chunks out (maxConcurrentChunks = 4) against Google's OpenAI-compatible endpoint, collects every distinct failure string into a Set, and throws with all of them joined by commas. The text after the colon is the real diagnosis: it is the raw error body returned by the Gemini endpoint.

Solutions

  1. Read the joined message after 'Gemini Failed to embed:' — each comma-separated entry maps to one failed chunk request
  2. For 401/403 errors, verify GEMINI_EMBEDDING_API_KEY is active in Google AI Studio and regenerate it if revoked
  3. For 429 RESOURCE_EXHAUSTED, wait for quota reset or request a quota increase, then re-embed the document
  4. Verify EMBEDDING_MODEL_PREF is a real model id (default 'gemini-embedding-001')
  5. If inputs are very large, lower EMBEDDING_MODEL_MAX_CHUNK_LENGTH so each request stays within model limits
Defensive patterns

Strategy: try-catch

Validate before calling

// Pre-flight: confirm the key can list models before embedding a whole workspace
async function geminiEmbedderReachable(openai) {
  try {
    await openai.models.list();
    return true;
  } catch (e) {
    console.error("Gemini pre-flight failed:", e.status, e.message);
    return false;
  }
}

Try / catch

try {
  const vectors = await embedder.embedTextInput(text);
} catch (e) {
  if (e.message.startsWith("Gemini Failed to embed:")) {
    const detail = e.message.slice("Gemini Failed to embed:".length);
    if (detail.includes("429") || detail.includes("RESOURCE_EXHAUSTED")) {
      // back off and retry this document later — quota is time-windowed
    } else if (detail.includes("401") || detail.includes("API_KEY")) {
      // stop the batch: key is invalid, retrying cannot succeed
    }
  } else throw e;
}

Prevention

When it happens

Trigger: A revoked or wrong API key (401/403 API_KEY_INVALID); hitting project quota or rate limits (429 RESOURCE_EXHAUSTED) because 4 concurrent batch requests are in flight; EMBEDDING_MODEL_PREF set to a model that does not exist on generativelanguage.googleapis.com; a batch or single input exceeding the model's token/request limits; network egress blocked to googleapis.com.

Common situations: Embedding a large workspace for the first time and tripping Gemini free-tier quota; rotating an AI Studio key but not updating GEMINI_EMBEDDING_API_KEY; setting a preview/GA model name that is not available to the account; corporate proxies that block or MITM the Google API.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@20f6d3546c (2026-08-18). Data as JSON: /api/errors/17a3c83851b3a267. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/EmbeddingEngines/gemini/index.js:127

        .flat();
      if (errors.length > 0) {
        let uniqueErrors = new Set();
        errors.map((error) =>
          uniqueErrors.add(`[${error.type}]: ${error.message}`)
        );

        return {
          data: [],
          error: Array.from(uniqueErrors).join(", "),
        };
      }
      return {
        data: results.map((res) => res?.data || []).flat(),
        error: null,
      };
    });

    if (!!error) throw new Error(`Gemini Failed to embed: ${error}`);
    return data.length > 0 &&
      data.every((embd) => embd.hasOwnProperty("embedding"))
      ? data.map((embd) => embd.embedding)
      : null;
  }
}

module.exports = {
  GeminiEmbedder,
};

View on GitHub (pinned to 20f6d3546c)