Mintplex-Labs/anything-llm · error · Error

LMStudio Failed to embed

Error message

LMStudio Failed to embed: ${Array.from(uniqueErrors).join(", ")}

What it means

Thrown from LMStudioEmbedder.embedChunks after the sequential per-chunk loop. Because LM Studio drops queued requests, chunks are embedded one at a time; any failure sets hasError and stops the loop. Each error is normalized to '[type]: message' (type from the HTTP error code/status, default 'failed_to_embed', or the sentinel 'EMPTY_ARR'), de-duplicated, and joined with commas.

Solutions

  1. Match on the [type] segment: EMPTY_ARR means the server answered but produced no usable embedding; an HTTP code means the server rejected the request
  2. Load a real embedding model (e.g. nomic-embed-text-v1.5) and set EMBEDDING_MODEL_PREF to exactly that id
  3. If chunks overflow context, raise the model's context in LM Studio or lower EMBEDDING_MODEL_MAX_CHUNK_LENGTH
  4. Confirm the model stays loaded for the whole embedding (disable JIT unload / keep model in memory)
  5. Re-run the document embed after fixing — the loop aborts on first error so later chunks were skipped

Example fix

# before: chat model selected, fails with EMPTY_ARR or 4xx
EMBEDDING_MODEL_PREF=mistral-7b-instruct

# after: dedicated embedding model loaded in LM Studio
EMBEDDING_MODEL_PREF=text-embedding-nomic-embed-text-v1.5
Defensive patterns

Strategy: try-catch

Validate before calling

// Pre-flight: the configured model must be an embedding model that round-trips one input
async function lmStudioCanEmbed(openai, model) {
  try {
    const res = await openai.embeddings.create({ model, input: "ping", encoding_format: "base64" });
    const emb = res.data?.[0]?.embedding;
    return Array.isArray(emb) && emb.length > 0;
  } catch {
    return false;
  }
}

Try / catch

try {
  const vectors = await embedder.embedTextInput(text);
} catch (e) {
  if (e.message.startsWith("LMStudio Failed to embed:")) {
    if (e.message.includes("[EMPTY_ARR]")) { /* model returned no embedding — usually a chat model selected; switch EMBEDDING_MODEL_PREF */ }
    else if (/\[(4\d\d|5\d\d)\]/.test(e.message)) { /* server rejected the request — check model id / context length */ }
    else throw e;
  } else throw e;
}

Prevention

When it happens

Trigger: EMBEDDING_MODEL_PREF points at a loaded chat/LLM model rather than an embedding model (server rejects or returns unusable output); model id not loaded (404 from the server); a chunk exceeding the loaded model's context length; embedding returns an empty array, producing "[EMPTY_ARR]: The embedding was empty from LMStudio"; server stopped between the #isAlive check and the embedding call.

Common situations: Selecting a conversational model (e.g. a Mistral or Llama chat model) as the embedder; nomic-embed-text loaded with a tiny context so large chunks fail; LM Studio's JIT server unloading the model mid-run.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/88ecd3f8bf9e205b. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/EmbeddingEngines/lmstudio/index.js:110

      );
    }

    // Accumulate errors from embedding.
    // If any are present throw an abort error.
    const errors = results
      .filter((res) => !!res.error)
      .map((res) => res.error)
      .flat();

    if (errors.length > 0) {
      let uniqueErrors = new Set();
      console.log(errors);
      errors.map((error) =>
        uniqueErrors.add(`[${error.type}]: ${error.message}`)
      );

      if (errors.length > 0)
        throw new Error(
          `LMStudio Failed to embed: ${Array.from(uniqueErrors).join(", ")}`
        );
    }

    const data = results.map((res) => res?.data || []);
    return data.length > 0 ? data : null;
  }
}

module.exports = {
  LMStudioEmbedder,
};

View on GitHub (pinned to 3aec848f28)