Mintplex-Labs/anything-llm · error

${e.message}

Error message

${e.message}

What it means

KoboldCPP's chat completion wrapper re-throws any error from the underlying HTTP client (OpenAI SDK pointed at the KoboldCPP endpoint) while calling the /v1/chat/completions route. The message is the raw client error, so it can be a network failure, a non-JSON response, or an API error like an invalid model. After the call, the code also expects result.output.choices to exist; a malformed provider response surfaces here first as this thrown error.

Solutions

  1. Verify KoboldCPP is running and its base URL (e.g. http://127.0.0.1:5001/v1) is correct in the AI provider settings.
  2. Check that the configured model matches a model actually loaded in KoboldCPP.
  3. Curl the endpoint manually (curl http://host:5001/v1/chat/completions) to see the raw error/response.
  4. Upgrade KoboldCPP if its OpenAI-compatible API is missing or outdated.
  5. Inspect the wrapped e.message in the logs for the root cause (ECONNREFUSED vs 404 vs parse error).

Example fix

// before
.catch((e) => {
  throw new Error(e.message);
})
// after
.catch((e) => {
  throw new Error(`KoboldCPP::getChatCompletion failed: ${e.message}`);
})
Defensive patterns

Strategy: try-catch

Validate before calling

// health check before calling
const res = await fetch(`${koboldBaseURL}/v1/models`);
if (!res.ok) throw new Error(`KoboldCPP unreachable: ${res.status}`);

Type guard

function hasChoices(output) {
  return output && typeof output === 'object' &&
    Array.isArray(output.choices) && output.choices.length > 0;
}

Try / catch

try {
  const text = await provider.getChatCompletion(messages);
} catch (e) {
  if (/ECONNREFUSED|fetch failed/i.test(e.message)) {
    // provider offline: check KoboldCPP server/base URL
  } else if (/40[13]|model/i.test(e.message)) {
    // wrong model or auth: verify loaded model
  } else {
    throw e; // unknown: surface to user
  }
}

Prevention

When it happens

Trigger: Calling getChatCompletion when the KoboldCPP server is unreachable, returns a non-200 or non-JSON response, rejects the model name, or the request payload (temperature/max_tokens) is rejected by the backend.

Common situations: KoboldCPP not running or wrong KoboldCPP base URL configured; model loaded in KoboldCPP does not match the configured model; KoboldCPP version whose OpenAI-compatible API differs (older builds lacking /v1 routes); reverse proxy returning HTML error pages the SDK cannot parse.

Understand the failure class

Background: "API request failed": what wrapped HTTP errors from external APIs mean and how to find the real cause — this error's family across 29 libraries.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@f92433b4ea (2026-09-22). Data as JSON: /api/errors/ff336ddc416c4a87. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/AiProviders/koboldCPP/index.js:180

    if (
      !!message?.reasoning_content &&
      message.reasoning_content.trim().length > 0
    )
      textResponse = `<think>${message.reasoning_content}</think>${textResponse}`;
    return textResponse;
  }

  async getChatCompletion(messages = null, { temperature = 0.7 }) {
    const result = await LLMPerformanceMonitor.measureAsyncFunction(
      this.openai.chat.completions
        .create({
          model: this.model,
          messages,
          temperature,
          ...(this.maxTokens ? { max_tokens: this.maxTokens } : {}),
        })
        .catch((e) => {
          throw new Error(e.message);
        })
    );

    if (
      !result.output.hasOwnProperty("choices") ||
      result.output.choices.length === 0
    )
      return null;

    return {
      textResponse: this.#parseReasoningFromResponse(result.output.choices[0]),
      metrics: {
        prompt_tokens: result.output.usage?.prompt_tokens || 0,
        completion_tokens: result.output.usage?.completion_tokens || 0,
        total_tokens: result.output.usage?.total_tokens || 0,
        outputTps:
          (result.output.usage?.completion_tokens || 0) / result.duration,
        duration: result.duration,

View on GitHub (pinned to f92433b4ea)