{"record":{"id":"432efb9586be3c27","repo":"Mintplex-Labs/anything-llm","slug":"mistral-returned-empty-embeddings-for-batch","errorCode":null,"errorMessage":"Mistral returned empty embeddings for batch","messagePattern":"Mistral returned empty embeddings for batch","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"server/utils/EmbeddingEngines/mistral/index.js","lineNumber":122,"sourceCode":"      if (errors.length > 0) {\n        let uniqueErrors = new Set();\n        errors.map((error) =>\n          uniqueErrors.add(`[${error.type}]: ${error.message}`)\n        );\n        return { data: [], error: Array.from(uniqueErrors).join(\", \") };\n      }\n      return {\n        data: results.map((res) => res?.data || []).flat(),\n        error: null,\n      };\n    });\n\n    if (!!error) throw new Error(`Mistral Failed to embed: ${error}`);\n\n    // Throw rather than return null so a document is never silently embedded with empty vectors (#5513).\n    const embeddings = data.map((emb) => emb.embedding);\n    if (embeddings.length === 0)\n      throw new Error(\"Mistral returned empty embeddings for batch\");\n    return embeddings;\n  }\n}\n\nmodule.exports = {\n  MistralEmbedder,\n};\n","sourceCodeStart":104,"sourceCodeEnd":130,"githubUrl":"https://github.com/Mintplex-Labs/anything-llm/blob/f92433b4ea0598492a1e6645ea22addbb4dd1287/server/utils/EmbeddingEngines/mistral/index.js#L104-L130","documentation":"Thrown inside MistralEmbedder.embedChunks when the API call succeeded (HTTP 200) but response.data contains no usable embeddings, i.e. the mapped array has length 0. It is a data-integrity guard: a batch of N inputs must yield N embeddings. Note it is thrown inside the try block, so it is re-caught and re-wrapped as 'Mistral Failed to embed: Mistral returned empty embeddings for batch'.","triggerScenarios":"A batch whose inputs are all empty strings or get filtered server-side; batch size or total token count tripping Mistral-side limits such that it returns a degenerate body; EMBEDDING_MODEL_PREF pointing to a non-embedding (chat) model, which can return data without embedding fields; upstream API quirk returning data: null on overload.","commonSituations":"Embedding pre-chunked text where some batches contain only whitespace; switching the model preference from mistral-embed to a chat model like mistral-large; very large document batches that occasionally get truncated responses.","solutions":["Set EMBEDDING_MODEL_PREF to a real Mistral embedding model (mistral-embed) — chat models return unusable payloads","Filter empty/whitespace-only strings out of the input before embedding","Split oversized batches (fewer chunks per call, e.g. via EMBEDDING_MODEL_MAX_CHUNK_LENGTH) and retry","If transient (rare 200-with-no-data), retry the batch — this is a fresh call, not a cached failure"],"exampleFix":"// before\nconst embeddings = await embedder.embedChunks(textChunks); // batch contains \"\"\n\n// after\nconst clean = textChunks.filter((t) => typeof t === 'string' && t.trim().length > 0);\nconst embeddings = await embedder.embedChunks(clean);","handlingStrategy":"validation","validationCode":"// Guard the exact condition: every batch input must be a non-empty string\nfunction validEmbeddingBatch(chunks) {\n  return (\n    Array.isArray(chunks) &&\n    chunks.length > 0 &&\n    chunks.every((c) => typeof c === \"string\" && c.trim().length > 0)\n  );\n}\nif (!validEmbeddingBatch(textChunks)) {\n  textChunks = textChunks.filter((c) => typeof c === \"string\" && c.trim().length > 0);\n}","typeGuard":"function isEmbeddingResponseUsable(response) {\n  return (\n    !!response &&\n    Array.isArray(response.data) &&\n    response.data.length > 0 &&\n    response.data.every((d) => Array.isArray(d.embedding) && d.embedding.length > 0)\n  );\n}","tryCatchPattern":"try {\n  const vectors = await embedder.embedChunks(batch);\n} catch (e) {\n  if (e.message.includes(\"Mistral returned empty embeddings for batch\")) {\n    // filter/resize the batch (drop empty strings, split large batches) and retry once\n  } else throw e;\n}","preventionTips":["Filter whitespace-only strings from chunks before embedding — they can shrink a batch to nothing server-side","Keep EMBEDDING_MODEL_PREF on a real embedding model (mistral-embed), never a chat model","Cap batch sizes for very large documents so responses stay well-formed"],"tags":["mistral","embeddings","empty-response","data-validation"],"backgroundTag":"empty-api-response","analyzedSha":"f92433b4ea0598492a1e6645ea22addbb4dd1287","analyzedAt":"2026-09-22T02:13:38.154Z","contentChangedAt":"2026-09-22T02:13:38.154Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}