Mintplex-Labs/anything-llm · error · Error

e.message

Error message

e.message

What it means

FoundryLLM.getChatCompletion wraps the OpenAI-compatible chat.completions.create call in a .catch that rethrows `new Error(e.message)`, flattening whatever the SDK produced. This is a passthrough: the real failure is a Foundry Local HTTP/transport error (model crashed mid-request, malformed request, context overflow of max_completion_tokens, connection dropped). The rethrow exists to normalize SDK APIError objects into plain Errors for AnythingLLM's handler, but it discards status codes and the original stack.

Solutions

  1. Read the inner message — it is the literal Foundry Local error (e.g. 'max_completion_tokens too large', 'model not found') and dictates the fix
  2. Reduce the prompt size / lower GENERIC token limits or clear chat history if the context overflows the small default 16K window
  3. Reload or pick a smaller model if inference crashed (FoundryLLM forgets evicted models, so a retry may reload it)
  4. Update Foundry Local and restart the service if the error indicates a schema/protocol problem
Defensive patterns

Strategy: try-catch

Try / catch

try {
  return await llm.getChatCompletion(messages);
} catch (err) {
  // err.message is the raw Foundry Local error text; branch on its content
  if (/context|token/i.test(err.message)) return respondTrimmedAndRetry(messages);
  if (/Premature close|socket/i.test(err.message)) return retryOnce(); // model eviction
  return respond(`Foundry inference failed: ${err.message}`);
}

Prevention

When it happens

Trigger: The HTTP POST to Foundry Local's /chat/completions returning 4xx/5xx: requesting more max_completion_tokens than the loaded model supports, a request payload the local service rejects, the model process dying mid-inference, or the socket closing after an eviction ('Premature close' — see #handleStreamFailure which treats this case).

Common situations: Context larger than the local model's window; model evicted between the load check and the request; Foundry Local version mismatch producing a schema the SDK rejects; laptop under memory pressure so inference crashes.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/3742ab0d21bab956. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/AiProviders/foundry/index.js:324

    if (!this.model)
      throw new Error(
        `Foundry chat: ${this.model} is not valid or defined model for chat completion!`
      );

    // max_completion_tokens is required by Foundry (it caps output at 1024
    // otherwise), so the window has to be resolved before the request is built.
    await this.assertModelContextLimits();
    await this.assertModelLoaded();
    const result = await LLMPerformanceMonitor.measureAsyncFunction(
      this.openai.chat.completions
        .create({
          model: this.model,
          messages,
          temperature,
          max_completion_tokens: this.promptWindowLimit(),
        })
        .catch((e) => {
          throw new Error(e.message);
        })
    );

    if (
      !result.output.hasOwnProperty("choices") ||
      result.output.choices.length === 0
    )
      return null;

    return {
      textResponse: result.output.choices[0].message.content,
      metrics: {
        prompt_tokens: result.output.usage.prompt_tokens || 0,
        completion_tokens: result.output.usage.completion_tokens || 0,
        total_tokens: result.output.usage.total_tokens || 0,
        outputTps: result.output.usage.completion_tokens / result.duration,
        duration: result.duration,
        model: this.model,

View on GitHub (pinned to 3aec848f28)