Mintplex-Labs/anything-llm · error

Mistral returned empty embeddings for batch

Error message

Mistral returned empty embeddings for batch

What it means

Thrown inside MistralEmbedder.embedChunks when the API call succeeded (HTTP 200) but response.data contains no usable embeddings, i.e. the mapped array has length 0. It is a data-integrity guard: a batch of N inputs must yield N embeddings. Note it is thrown inside the try block, so it is re-caught and re-wrapped as 'Mistral Failed to embed: Mistral returned empty embeddings for batch'.

Solutions

  1. Set EMBEDDING_MODEL_PREF to a real Mistral embedding model (mistral-embed) — chat models return unusable payloads
  2. Filter empty/whitespace-only strings out of the input before embedding
  3. Split oversized batches (fewer chunks per call, e.g. via EMBEDDING_MODEL_MAX_CHUNK_LENGTH) and retry
  4. If transient (rare 200-with-no-data), retry the batch — this is a fresh call, not a cached failure

Example fix

// before
const embeddings = await embedder.embedChunks(textChunks); // batch contains ""

// after
const clean = textChunks.filter((t) => typeof t === 'string' && t.trim().length > 0);
const embeddings = await embedder.embedChunks(clean);
Defensive patterns

Strategy: validation

Validate before calling

// Guard the exact condition: every batch input must be a non-empty string
function validEmbeddingBatch(chunks) {
  return (
    Array.isArray(chunks) &&
    chunks.length > 0 &&
    chunks.every((c) => typeof c === "string" && c.trim().length > 0)
  );
}
if (!validEmbeddingBatch(textChunks)) {
  textChunks = textChunks.filter((c) => typeof c === "string" && c.trim().length > 0);
}

Type guard

function isEmbeddingResponseUsable(response) {
  return (
    !!response &&
    Array.isArray(response.data) &&
    response.data.length > 0 &&
    response.data.every((d) => Array.isArray(d.embedding) && d.embedding.length > 0)
  );
}

Try / catch

try {
  const vectors = await embedder.embedChunks(batch);
} catch (e) {
  if (e.message.includes("Mistral returned empty embeddings for batch")) {
    // filter/resize the batch (drop empty strings, split large batches) and retry once
  } else throw e;
}

Prevention

When it happens

Trigger: A batch whose inputs are all empty strings or get filtered server-side; batch size or total token count tripping Mistral-side limits such that it returns a degenerate body; EMBEDDING_MODEL_PREF pointing to a non-embedding (chat) model, which can return data without embedding fields; upstream API quirk returning data: null on overload.

Common situations: Embedding pre-chunked text where some batches contain only whitespace; switching the model preference from mistral-embed to a chat model like mistral-large; very large document batches that occasionally get truncated responses.

Understand the failure class

Background: "empty response", "returned no data", "empty embeddings": what HTTP 200-with-empty-body errors mean across libraries — this error's family across 36 libraries.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@f92433b4ea (2026-09-22). Data as JSON: /api/errors/432efb9586be3c27. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/EmbeddingEngines/mistral/index.js:122

      if (errors.length > 0) {
        let uniqueErrors = new Set();
        errors.map((error) =>
          uniqueErrors.add(`[${error.type}]: ${error.message}`)
        );
        return { data: [], error: Array.from(uniqueErrors).join(", ") };
      }
      return {
        data: results.map((res) => res?.data || []).flat(),
        error: null,
      };
    });

    if (!!error) throw new Error(`Mistral Failed to embed: ${error}`);

    // Throw rather than return null so a document is never silently embedded with empty vectors (#5513).
    const embeddings = data.map((emb) => emb.embedding);
    if (embeddings.length === 0)
      throw new Error("Mistral returned empty embeddings for batch");
    return embeddings;
  }
}

module.exports = {
  MistralEmbedder,
};

View on GitHub (pinned to f92433b4ea)