Mintplex-Labs/anything-llm · error
Ollama returned empty embeddings for batch!
Error message
Ollama returned empty embeddings for batch!
What it means
Ollama answered the /api/embed call with HTTP 200, but the response's `embeddings` field was missing or an empty array. Because the code sends the batch via the `input` parameter, it expects exactly one embedding per input string; zero results means the model ran but produced nothing. The most common cause is that the configured model is a generative/LLM model with no embedding capability, or the Ollama version is too old to return the batch `embeddings` shape.
Solutions
- Point EMBEDDING_MODEL_PREF at a real embedding model (nomic-embed-text, mxbai-embed-large, all-minilm, snowflake-arctic-embed) and `ollama pull` it.
- Match the model name from `ollama list` exactly, including the tag.
- Upgrade Ollama to a recent release so /api/embed with `input` arrays returns `embeddings`.
- Reproduce outside AnythingLLM: `curl http://localhost:11434/api/embed -d '{"model":"nomic-embed-text","input":["hi"]}'` and confirm embeddings is non-empty.
Example fix
# .env — before (llama3 is a chat model and cannot embed) EMBEDDING_MODEL_PREF=llama3 # .env — after EMBEDDING_MODEL_PREF=nomic-embed-text
Defensive patterns
Strategy: try-catch
Validate before calling
async function modelProducesEmbeddings(basePath, model) {
const res = await fetch(`${basePath}/api/embed`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ model, input: ["ping"] }),
});
const body = await res.json();
return Array.isArray(body.embeddings) && body.embeddings.length > 0;
}
if (!(await modelProducesEmbeddings(basePath, model))) {
throw new Error(`${model} is not an embedding model — pick one from \`ollama list\` that embeds.`);
} Type guard
function hasEmbeddings(payload) {
return (
payload !== null &&
typeof payload === "object" &&
Array.isArray(payload.embeddings) &&
payload.embeddings.length > 0 &&
payload.embeddings.every((e) => Array.isArray(e))
);
} Try / catch
try {
const vectors = await embedder.embedChunks(chunks);
} catch (err) {
if (err.message.includes("empty embeddings")) {
// model is not an embedding model: fix EMBEDDING_MODEL_PREF, do not retry
} else {
throw err;
}
} Prevention
- Validate the embedder model with a one-string smoke test when saving settings.
- Pull embedding models explicitly (ollama pull nomic-embed-text) rather than relying on auto-pull of chat tags.
- Keep model ids in config exactly as shown by ollama list.
When it happens
Trigger: EMBEDDING_MODEL_PREF points at a chat model (llama3, mistral, qwen) instead of an embedding model; a model-name/tag mismatch so Ollama resolves a different model; an Ollama version that predates the /api/embed batch endpoint's response shape; a server returning an error-shaped body with status 200.
Common situations: Reusing the same model name for the LLM and the embedder in settings; forgetting `ollama pull nomic-embed-text` on a new machine; very old Ollama installs on NAS devices or routers.
Understand the failure class
Background: "empty response", "returned no data", "empty embeddings": what HTTP 200-with-empty-body errors mean across libraries — this error's family across 36 libraries.
Related errors
- Mistral returned empty embeddings for batch
- No embedding model was set.
- Ollama Failed to embed
- Ollama service could not be reached. Is Ollama running?
- ChromaCloud::Embedding dimension too large
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/c4fefaea93306e88.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/EmbeddingEngines/ollama/index.js:115
for (let i = 0; i < textChunks.length; i += this.maxConcurrentChunks) {
const batch = textChunks.slice(i, i + this.maxConcurrentChunks);
currentBatch++;
try {
// Use input param instead of prompt param to support batch processing
const res = await this.client.embed({
model: this.model,
input: batch,
options: {
// Always set the num_ctx to the max chunk length defined by the user in the settings
// so that the maximum context window is used and content is not truncated.
num_ctx: this.embeddingMaxChunkLength,
},
});
const { embeddings } = res;
if (!Array.isArray(embeddings) || embeddings.length === 0)
throw new Error("Ollama returned empty embeddings for batch!");
// Using prompt param in embed() would return a single embedding (number[])
// but input param returns an array of embeddings (number[][]) for batch processing.
// This is why we spread the embeddings array into the data array.
data.push(...embeddings);
reportEmbeddingProgress(data.length, textChunks.length);
this.log(
`Batch ${currentBatch}/${totalBatches}: Embedded ${embeddings.length} chunks. Total: ${data.length}/${textChunks.length}`
);
} catch (err) {
this.log(err.message);
error = err.message;
data = [];
break;
}
}
if (!!error) throw new Error(`Ollama Failed to embed: ${error}`);View on GitHub (pinned to 3aec848f28)