Mintplex-Labs/anything-llm · error · Error
Gemini Failed to embed
Error message
Gemini Failed to embed: ${error} What it means
Thrown from GeminiEmbedder.embedChunks after the per-chunk embedding requests complete. The class fans chunks out (maxConcurrentChunks = 4) against Google's OpenAI-compatible endpoint, collects every distinct failure string into a Set, and throws with all of them joined by commas. The text after the colon is the real diagnosis: it is the raw error body returned by the Gemini endpoint.
Solutions
- Read the joined message after 'Gemini Failed to embed:' — each comma-separated entry maps to one failed chunk request
- For 401/403 errors, verify GEMINI_EMBEDDING_API_KEY is active in Google AI Studio and regenerate it if revoked
- For 429 RESOURCE_EXHAUSTED, wait for quota reset or request a quota increase, then re-embed the document
- Verify EMBEDDING_MODEL_PREF is a real model id (default 'gemini-embedding-001')
- If inputs are very large, lower EMBEDDING_MODEL_MAX_CHUNK_LENGTH so each request stays within model limits
Defensive patterns
Strategy: try-catch
Validate before calling
// Pre-flight: confirm the key can list models before embedding a whole workspace
async function geminiEmbedderReachable(openai) {
try {
await openai.models.list();
return true;
} catch (e) {
console.error("Gemini pre-flight failed:", e.status, e.message);
return false;
}
} Try / catch
try {
const vectors = await embedder.embedTextInput(text);
} catch (e) {
if (e.message.startsWith("Gemini Failed to embed:")) {
const detail = e.message.slice("Gemini Failed to embed:".length);
if (detail.includes("429") || detail.includes("RESOURCE_EXHAUSTED")) {
// back off and retry this document later — quota is time-windowed
} else if (detail.includes("401") || detail.includes("API_KEY")) {
// stop the batch: key is invalid, retrying cannot succeed
}
} else throw e;
} Prevention
- Wrap per-document embedding so one failed document doesn't kill a workspace-wide embed job
- Log the full joined message — it contains every distinct upstream error, not just the first
- For large workspaces, embed during off-peak quota windows or request quota increases up front
When it happens
Trigger: A revoked or wrong API key (401/403 API_KEY_INVALID); hitting project quota or rate limits (429 RESOURCE_EXHAUSTED) because 4 concurrent batch requests are in flight; EMBEDDING_MODEL_PREF set to a model that does not exist on generativelanguage.googleapis.com; a batch or single input exceeding the model's token/request limits; network egress blocked to googleapis.com.
Common situations: Embedding a large workspace for the first time and tripping Gemini free-tier quota; rotating an AI Studio key but not updating GEMINI_EMBEDDING_API_KEY; setting a preview/GA model name that is not available to the account; corporate proxies that block or MITM the Google API.
Related errors
- LiteLLM Failed to embed
- e.message
- GenericOpenAI Failed to embed
- Mistral Failed to embed
- Lemonade Failed to embed
AI-assisted analysis of Mintplex-Labs/anything-llm@20f6d3546c (2026-08-18).
Data as JSON: /api/errors/17a3c83851b3a267.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/EmbeddingEngines/gemini/index.js:127
.flat();
if (errors.length > 0) {
let uniqueErrors = new Set();
errors.map((error) =>
uniqueErrors.add(`[${error.type}]: ${error.message}`)
);
return {
data: [],
error: Array.from(uniqueErrors).join(", "),
};
}
return {
data: results.map((res) => res?.data || []).flat(),
error: null,
};
});
if (!!error) throw new Error(`Gemini Failed to embed: ${error}`);
return data.length > 0 &&
data.every((embd) => embd.hasOwnProperty("embedding"))
? data.map((embd) => embd.embedding)
: null;
}
}
module.exports = {
GeminiEmbedder,
};
View on GitHub (pinned to 20f6d3546c)