Mintplex-Labs/anything-llm · error · Error
GenericOpenAI Failed to embed
Error message
GenericOpenAI Failed to embed: ${error.message} What it means
Thrown from GenericOpenAiEmbedder.embedChunks when any single chunk request fails; the loop aborts the whole sequence because partial embeddings would be incomplete. The error object is normalized first (type from error.code/status, message from the response body), so the thrown text after the colon is the upstream endpoint's own message.
Solutions
- Inspect error.message — it is the literal response body from your server
- For 401s, set GENERIC_OPEN_AI_EMBEDDING_API_KEY (or fix it) to a valid key for that endpoint
- For 404s, correct EMBEDDING_BASE_PATH so it is the OpenAI-compatible root (commonly ends in /v1)
- For 429s, set GENERIC_OPEN_AI_EMBEDDING_API_DELAY_MS (minimum 500) to throttle between batches
- For context-length errors, lower EMBEDDING_MODEL_MAX_CHUNK_LENGTH
- Verify EMBEDDING_MODEL_PREF exactly matches a model id the endpoint serves
Example fix
# before EMBEDDING_BASE_PATH=http://localhost:8080 # 404: SDK posts to /embeddings at server root # after EMBEDDING_BASE_PATH=http://localhost:8080/v1 GENERIC_OPEN_AI_EMBEDDING_API_DELAY_MS=500
Defensive patterns
Strategy: try-catch
Validate before calling
// Smoke-test the endpoint before a bulk run: one tiny embedding must round-trip
async function genericOpenAiEmbedderHealthy(openai, model) {
try {
const res = await openai.embeddings.create({ model, input: ["ping"] });
return Array.isArray(res?.data?.[0]?.embedding);
} catch (e) {
console.error("Embedding pre-flight failed:", e.status, e.message);
return false;
}
} Try / catch
try {
const vectors = await embedder.embedTextInput(text);
} catch (e) {
if (e.message.startsWith("GenericOpenAI Failed to embed:")) {
const msg = e.message;
if (/401|unauthorized/i.test(msg)) { /* fix GENERIC_OPEN_AI_EMBEDDING_API_KEY; do not retry */ }
else if (/429|rate/i.test(msg)) { /* wait, then retry; consider GENERIC_OPEN_AI_EMBEDDING_API_DELAY_MS */ }
else if (/404|not found/i.test(msg)) { /* fix EMBEDDING_BASE_PATH or EMBEDDING_MODEL_PREF; do not retry */ }
else throw e;
} else throw e;
} Prevention
- Set GENERIC_OPEN_AI_EMBEDDING_API_DELAY_MS (>=500) when the backend is rate-limit-prone
- Keep EMBEDDING_MODEL_MAX_CHUNK_LENGTH within the target model's context so oversized chunks never reach the API
- Remember the loop aborts on first error — after a fix, re-embed the whole document, not the remainder
When it happens
Trigger: 401 from a server that requires a key when GENERIC_OPEN_AI_EMBEDDING_API_KEY is null or wrong; POST {EMBEDDING_BASE_PATH}/embeddings returning 404 because the base path is wrong (missing /v1 or points at a non-OpenAI route); model name in EMBEDDING_MODEL_PREF unknown to the server; 429 rate limiting on small self-hosted or shared endpoints; input text longer than the server's context window.
Common situations: Self-hosted single-threaded backends that 429 under AnythingLLM's batch load; pointing at Ollama's /v1 with a model id that is not pulled; using an OpenAI-compatible facade that does not implement /embeddings; very large documents blowing the max chunk length.
Related errors
- Gemini Failed to embed
- LiteLLM Failed to embed
- Mistral Failed to embed
- GenericOpenAI must have a valid base path to use for the…
- Lemonade Failed to embed
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/360ff021173fe66d.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/EmbeddingEngines/genericOpenAi/index.js:156
.create({
model: this.model,
input: chunk,
})
.then((result) => resolve({ data: result?.data, error: null }))
.catch((e) => {
e.type =
e?.response?.data?.error?.code ||
e?.response?.status ||
"failed_to_embed";
e.message = e?.response?.data?.error?.message || e.message;
resolve({ data: [], error: e });
});
});
// If any errors were returned from OpenAI abort the entire sequence because the embeddings
// will be incomplete.
if (error)
throw new Error(`GenericOpenAI Failed to embed: ${error.message}`);
allResults.push(...(data || []));
reportEmbeddingProgress(allResults.length, textChunks.length);
if (this.apiRequestDelay) await this.runDelay();
}
return allResults.length > 0 &&
allResults.every((embd) => embd.hasOwnProperty("embedding"))
? allResults.map((embd) => embd.embedding)
: null;
}
}
module.exports = {
GenericOpenAiEmbedder,
};
View on GitHub (pinned to 3aec848f28)