Mintplex-Labs/anything-llm · error

Ollama Failed to embed

Error message

Ollama Failed to embed: ${error}

What it means

This is the aggregate wrapper thrown at the end of OllamaEmbedder.embedChunks() when any batch in the loop failed. The inner error's message is captured and logged, accumulated embeddings are discarded (data = []), and the loop breaks — so partial results are never returned. The text after the colon is the real cause: an Ollama 500, a model load failure or OOM, a mid-run connection reset, or the empty-embeddings error above.

Solutions

  1. Read the suffix of the message — it is the underlying batch error and dictates the real fix.
  2. Set OLLAMA_EMBEDDING_BATCH_SIZE=1 to serialize batches and lower peak memory use.
  3. Reduce the embedding max chunk length so num_ctx (and therefore VRAM footprint) shrinks.
  4. On low-memory hosts, avoid running the LLM and the embedder on the same Ollama simultaneously, or raise keep_alive so the model is not unloaded between batches.

Example fix

# .env — before
# OLLAMA_EMBEDDING_BATCH_SIZE unset (defaults to 1, but often raised)
OLLAMA_EMBEDDING_BATCH_SIZE=10

# .env — after
OLLAMA_EMBEDDING_BATCH_SIZE=1
Defensive patterns

Strategy: try-catch

Try / catch

try {
  const vectors = await embedder.embedChunks(chunks);
} catch (err) {
  if (err.message.startsWith("Ollama Failed to embed:")) {
    const cause = err.message.split(":").slice(1).join(":").trim();
    // classify `cause`: OOM/model-load → reduce batch size & num_ctx; network → retry later
  } else {
    throw err;
  }
}

Prevention

When it happens

Trigger: Ollama reloads or evicts the model between batches on low-memory machines (keep-alive expiry), causing load failures; num_ctx set from a very large max chunk length exhausting VRAM; the connection drops or the server crashes mid-run; the model was deleted from Ollama while embedding was in progress.

Common situations: Embedding large documents on machines with just enough RAM for either the LLM or the embedder (the code comments call this out explicitly); running the chat model and embedder on the same small GPU; OLLAMA_EMBEDDING_BATCH_SIZE greater than 1 increasing memory pressure.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/3c53ae46507b1823. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/EmbeddingEngines/ollama/index.js:133

          throw new Error("Ollama returned empty embeddings for batch!");

        // Using prompt param in embed() would return a single embedding (number[])
        // but input param returns an array of embeddings (number[][]) for batch processing.
        // This is why we spread the embeddings array into the data array.
        data.push(...embeddings);
        reportEmbeddingProgress(data.length, textChunks.length);
        this.log(
          `Batch ${currentBatch}/${totalBatches}: Embedded ${embeddings.length} chunks. Total: ${data.length}/${textChunks.length}`
        );
      } catch (err) {
        this.log(err.message);
        error = err.message;
        data = [];
        break;
      }
    }

    if (!!error) throw new Error(`Ollama Failed to embed: ${error}`);
    return data.length > 0 ? data : null;
  }
}

module.exports = {
  OllamaEmbedder,
};

View on GitHub (pinned to 3aec848f28)