Mintplex-Labs/anything-llm · error
${e.message}
Error message
${e.message} What it means
KoboldCPP's chat completion wrapper re-throws any error from the underlying HTTP client (OpenAI SDK pointed at the KoboldCPP endpoint) while calling the /v1/chat/completions route. The message is the raw client error, so it can be a network failure, a non-JSON response, or an API error like an invalid model. After the call, the code also expects result.output.choices to exist; a malformed provider response surfaces here first as this thrown error.
Solutions
- Verify KoboldCPP is running and its base URL (e.g. http://127.0.0.1:5001/v1) is correct in the AI provider settings.
- Check that the configured model matches a model actually loaded in KoboldCPP.
- Curl the endpoint manually (curl http://host:5001/v1/chat/completions) to see the raw error/response.
- Upgrade KoboldCPP if its OpenAI-compatible API is missing or outdated.
- Inspect the wrapped e.message in the logs for the root cause (ECONNREFUSED vs 404 vs parse error).
Example fix
// before
.catch((e) => {
throw new Error(e.message);
})
// after
.catch((e) => {
throw new Error(`KoboldCPP::getChatCompletion failed: ${e.message}`);
}) Defensive patterns
Strategy: try-catch
Validate before calling
// health check before calling
const res = await fetch(`${koboldBaseURL}/v1/models`);
if (!res.ok) throw new Error(`KoboldCPP unreachable: ${res.status}`); Type guard
function hasChoices(output) {
return output && typeof output === 'object' &&
Array.isArray(output.choices) && output.choices.length > 0;
} Try / catch
try {
const text = await provider.getChatCompletion(messages);
} catch (e) {
if (/ECONNREFUSED|fetch failed/i.test(e.message)) {
// provider offline: check KoboldCPP server/base URL
} else if (/40[13]|model/i.test(e.message)) {
// wrong model or auth: verify loaded model
} else {
throw e; // unknown: surface to user
}
} Prevention
- Health-check the KoboldCPP endpoint at startup and before long jobs.
- Keep the configured model name in sync with the model loaded in KoboldCPP.
- Pin and test your KoboldCPP version against its OpenAI-compatible API.
- Log the underlying client error, not just the re-thrown message.
When it happens
Trigger: Calling getChatCompletion when the KoboldCPP server is unreachable, returns a non-200 or non-JSON response, rejects the model name, or the request payload (temperature/max_tokens) is rejected by the backend.
Common situations: KoboldCPP not running or wrong KoboldCPP base URL configured; model loaded in KoboldCPP does not match the configured model; KoboldCPP version whose OpenAI-compatible API differs (older builds lacking /v1 routes); reverse proxy returning HTML error pages the SDK cannot parse.
Understand the failure class
Background: "API request failed": what wrapped HTTP errors from external APIs mean and how to find the real cause — this error's family across 29 libraries.
Related errors
- LMStudio:getModelInfo
- An error occurred while deleting the model
- Failed to fetch generated image
- HTTP
- LLM processing failed
AI-assisted analysis of Mintplex-Labs/anything-llm@f92433b4ea (2026-09-22).
Data as JSON: /api/errors/ff336ddc416c4a87.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/AiProviders/koboldCPP/index.js:180
if (
!!message?.reasoning_content &&
message.reasoning_content.trim().length > 0
)
textResponse = `<think>${message.reasoning_content}</think>${textResponse}`;
return textResponse;
}
async getChatCompletion(messages = null, { temperature = 0.7 }) {
const result = await LLMPerformanceMonitor.measureAsyncFunction(
this.openai.chat.completions
.create({
model: this.model,
messages,
temperature,
...(this.maxTokens ? { max_tokens: this.maxTokens } : {}),
})
.catch((e) => {
throw new Error(e.message);
})
);
if (
!result.output.hasOwnProperty("choices") ||
result.output.choices.length === 0
)
return null;
return {
textResponse: this.#parseReasoningFromResponse(result.output.choices[0]),
metrics: {
prompt_tokens: result.output.usage?.prompt_tokens || 0,
completion_tokens: result.output.usage?.completion_tokens || 0,
total_tokens: result.output.usage?.total_tokens || 0,
outputTps:
(result.output.usage?.completion_tokens || 0) / result.duration,
duration: result.duration,View on GitHub (pinned to f92433b4ea)