Mintplex-Labs/anything-llm · error · Error

e.message

Error message

e.message

What it means

GenericOpenAiLLM.getChatCompletion rethrows e.message from the OpenAI SDK call against the arbitrary user-configured endpoint. The generic provider has isValidChatCompletionModel hard-coded to true ('short circuit since we have no idea if the model is valid'), so there is no pre-flight validation — any endpoint error (404 wrong path/model, 401 missing key, connection refused, non-OpenAI schema responses) arrives here raw. The max_tokens sent is GENERIC_OPEN_AI_MAX_TOKENS (default 1024).

Solutions

  1. Read the inner message: '404' → fix base path/model id; '401' → set GENERIC_OPEN_AI_API_KEY; 'ECONNREFUSED' → start the local service
  2. Confirm the endpoint answers: curl <GENERIC_OPEN_AI_BASE_PATH>/models and match the model id exactly
  3. Ensure the base path includes the version segment (usually /v1)
  4. If the service has its own auth (e.g. bearer token), put it in GENERIC_OPEN_AI_API_KEY

Example fix

# before (.env)
GENERIC_OPEN_AI_BASE_PATH=http://localhost:11434  # missing /v1 -> 404 at chat

# after (.env)
GENERIC_OPEN_AI_BASE_PATH=http://localhost:11434/v1
Defensive patterns

Strategy: try-catch

Validate before calling

// generic provider can't validate models (isValidChatCompletionModel is always true),
// so probe the endpoint yourself before chatting
const res = await fetch(`${process.env.GENERIC_OPEN_AI_BASE_PATH}/models`);
if (!res.ok) throw new Error(`Endpoint unreachable or wrong base path (HTTP ${res.status}).`);
const { data } = await res.json();
if (!data.some((m) => m.id === model)) throw new Error(`Model '${model}' not served by endpoint.`);

Try / catch

try {
  return await llm.getChatCompletion(messages);
} catch (err) {
  if (/ECONNREFUSED|fetch failed/i.test(err.message)) return respond("Local inference server is down — start it and retry.");
  if (/404/.test(err.message)) return respond("Wrong base path or model id — check /v1 and the model name.");
  if (/401|api key/i.test(err.message)) return respond("Endpoint requires GENERIC_OPEN_AI_API_KEY.");
  throw err;
}

Prevention

When it happens

Trigger: chat.completions.create failing against GENERIC_OPEN_AI_BASE_PATH: base path missing /v1 (404), model id not served by the endpoint, GENERIC_OPEN_AI_API_KEY required but unset/wrong, local service down (ECONNREFUSED), or the endpoint not actually OpenAI-compatible (SDK fails parsing).

Common situations: LM Studio/Ollama/vLLM not running or listening on a different port; path set to http://localhost:11434 without /v1; endpoint needs a key the user assumed optional; model renamed on the local server after the workspace pref was saved.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/41ccbb313fb5987f. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/AiProviders/genericOpenAi/index.js:234

    if (process.env.GENERIC_OPEN_AI_REPORT_USAGE !== "true") return {};
    return {
      stream_options: {
        include_usage: true,
      },
    };
  }

  async getChatCompletion(messages = null, { temperature = 0.7 }) {
    const result = await LLMPerformanceMonitor.measureAsyncFunction(
      this.openai.chat.completions
        .create({
          model: this.model,
          messages,
          temperature,
          max_tokens: this.maxTokens,
        })
        .catch((e) => {
          throw new Error(e.message);
        })
    );

    if (
      !result.output.hasOwnProperty("choices") ||
      result.output.choices.length === 0
    )
      return null;

    const usage = {
      prompt_tokens: result.output?.usage?.prompt_tokens || 0,
      completion_tokens: result.output?.usage?.completion_tokens || 0,
      total_tokens: result.output?.usage?.total_tokens || 0,
      duration: result.duration,
    };
    this.#extractLlamaCppTimings(result.output, usage);

    return {

View on GitHub (pinned to 3aec848f28)