Mintplex-Labs/anything-llm · error
Ollama Failed to embed
Error message
Ollama Failed to embed: ${error} What it means
This is the aggregate wrapper thrown at the end of OllamaEmbedder.embedChunks() when any batch in the loop failed. The inner error's message is captured and logged, accumulated embeddings are discarded (data = []), and the loop breaks — so partial results are never returned. The text after the colon is the real cause: an Ollama 500, a model load failure or OOM, a mid-run connection reset, or the empty-embeddings error above.
Solutions
- Read the suffix of the message — it is the underlying batch error and dictates the real fix.
- Set OLLAMA_EMBEDDING_BATCH_SIZE=1 to serialize batches and lower peak memory use.
- Reduce the embedding max chunk length so num_ctx (and therefore VRAM footprint) shrinks.
- On low-memory hosts, avoid running the LLM and the embedder on the same Ollama simultaneously, or raise keep_alive so the model is not unloaded between batches.
Example fix
# .env — before # OLLAMA_EMBEDDING_BATCH_SIZE unset (defaults to 1, but often raised) OLLAMA_EMBEDDING_BATCH_SIZE=10 # .env — after OLLAMA_EMBEDDING_BATCH_SIZE=1
Defensive patterns
Strategy: try-catch
Try / catch
try {
const vectors = await embedder.embedChunks(chunks);
} catch (err) {
if (err.message.startsWith("Ollama Failed to embed:")) {
const cause = err.message.split(":").slice(1).join(":").trim();
// classify `cause`: OOM/model-load → reduce batch size & num_ctx; network → retry later
} else {
throw err;
}
} Prevention
- Set OLLAMA_EMBEDDING_BATCH_SIZE=1 on memory-constrained hosts.
- Keep embedding max chunk length modest so num_ctx does not exhaust VRAM.
- Give Ollama a dedicated keep_alive for the embedder on small machines so it is not reloaded per batch.
When it happens
Trigger: Ollama reloads or evicts the model between batches on low-memory machines (keep-alive expiry), causing load failures; num_ctx set from a very large max chunk length exhausting VRAM; the connection drops or the server crashes mid-run; the model was deleted from Ollama while embedding was in progress.
Common situations: Embedding large documents on machines with just enough RAM for either the LLM or the embedder (the code comments call this out explicitly); running the chat model and embedder on the same small GPU; OLLAMA_EMBEDDING_BATCH_SIZE greater than 1 increasing memory pressure.
Related errors
- No embedding model was set.
- Ollama returned empty embeddings for batch!
- Ollama service could not be reached. Is Ollama running?
- OpenAI Failed to embed
- OpenRouter Failed to embed
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/3c53ae46507b1823.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/EmbeddingEngines/ollama/index.js:133
throw new Error("Ollama returned empty embeddings for batch!");
// Using prompt param in embed() would return a single embedding (number[])
// but input param returns an array of embeddings (number[][]) for batch processing.
// This is why we spread the embeddings array into the data array.
data.push(...embeddings);
reportEmbeddingProgress(data.length, textChunks.length);
this.log(
`Batch ${currentBatch}/${totalBatches}: Embedded ${embeddings.length} chunks. Total: ${data.length}/${textChunks.length}`
);
} catch (err) {
this.log(err.message);
error = err.message;
data = [];
break;
}
}
if (!!error) throw new Error(`Ollama Failed to embed: ${error}`);
return data.length > 0 ? data : null;
}
}
module.exports = {
OllamaEmbedder,
};
View on GitHub (pinned to 3aec848f28)