Mintplex-Labs/anything-llm · error

Ollama service could not be reached. Is Ollama running?

Error message

Ollama service could not be reached. Is Ollama running?

What it means

Before embedding, embedChunks() calls the private #isAlive(), which does a plain fetch() against EMBEDDING_BASE_PATH and requires an ok (2xx) response. The error is thrown when that fetch rejects (connection refused, DNS failure, timeout) or returns a non-2xx status (e.g. 401 from an auth gateway). This is a pre-flight health-check failure, not a failure of the embed API itself.

Solutions

  1. Start Ollama (`ollama serve` or the desktop app) and verify with `curl -i http://localhost:11434` that it returns 200.
  2. Fix EMBEDDING_BASE_PATH — inside docker use http://host.docker.internal:11434 (or the host IP), not localhost.
  3. If Ollama sits behind an auth proxy, set OLLAMA_AUTH_TOKEN to the bearer token the proxy expects.
  4. When Ollama runs on another host, start it with OLLAMA_HOST=0.0.0.0 so it accepts remote connections.

Example fix

# .env — before (localhost inside a container points at the container itself)
EMBEDDING_BASE_PATH=http://localhost:11434

# .env — after
EMBEDDING_BASE_PATH=http://host.docker.internal:11434
Defensive patterns

Strategy: validation

Validate before calling

async function ollamaAlive(basePath, token) {
  try {
    const res = await fetch(basePath, {
      headers: token ? { Authorization: `Bearer ${token}` } : {},
    });
    return res.ok;
  } catch {
    return false;
  }
}
if (
  !(await ollamaAlive(
    process.env.EMBEDDING_BASE_PATH,
    process.env.OLLAMA_AUTH_TOKEN
  ))
) {
  throw new Error("Ollama is not reachable — start it or fix EMBEDDING_BASE_PATH before embedding.");
}

Try / catch

try {
  const vectors = await embedder.embedChunks(chunks);
} catch (err) {
  if (err.message.includes("could not be reached")) {
    // connectivity/config problem, not a data problem:
    // surface 'start Ollama / check EMBEDDING_BASE_PATH' guidance to the user
  } else {
    throw err;
  }
}

Prevention

When it happens

Trigger: Ollama is not running (serve crashed or was never started); EMBEDDING_BASE_PATH is malformed or wrong (missing http://, wrong port, typo); AnythingLLM runs in docker and localhost resolves to the container, not the host; OLLAMA_AUTH_TOKEN is missing/wrong and a proxy in front of Ollama returns 401/403; a reverse proxy answering the root path with 502/504.

Common situations: Docker AnythingLLM with EMBEDDING_BASE_PATH=http://localhost:11434 instead of http://host.docker.internal:11434; Ollama bound to 127.0.0.1 while AnythingLLM runs on another machine; macOS Ollama desktop app not launched; systemd Ollama unit stopped after a reboot.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/e01dbd01db259f6d. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/EmbeddingEngines/ollama/index.js:81

  }

  /**
   * This function takes an array of text chunks and embeds them using the Ollama API.
   * Chunks are processed in batches based on the maxConcurrentChunks setting to balance
   * resource usage on the Ollama endpoint.
   *
   * We will use the num_ctx option to set the maximum context window to the max chunk length defined by the user in the settings
   * so that the maximum context window is used and content is not truncated.
   *
   * We also assume the default keep alive option. This could cause issues with models being unloaded and reloaded
   * on low memory machines, but that is simply a user-end issue we cannot control. If the LLM and embedder are
   * constantly being loaded and unloaded, the user should use another LLM or Embedder to avoid this issue.
   * @param {string[]} textChunks - An array of text chunks to embed.
   * @returns {Promise<Array<number[]>>} - A promise that resolves to an array of embeddings.
   */
  async embedChunks(textChunks = []) {
    if (!(await this.#isAlive()))
      throw new Error(
        `Ollama service could not be reached. Is Ollama running?`
      );
    this.log(
      `Embedding ${textChunks.length} chunks of text with ${this.model} in batches of ${this.maxConcurrentChunks}.`
    );

    let data = [];
    let error = null;

    // Process chunks in batches based on maxConcurrentChunks
    const totalBatches = Math.ceil(
      textChunks.length / this.maxConcurrentChunks
    );
    let currentBatch = 0;

    for (let i = 0; i < textChunks.length; i += this.maxConcurrentChunks) {
      const batch = textChunks.slice(i, i + this.maxConcurrentChunks);
      currentBatch++;

View on GitHub (pinned to 3aec848f28)