Mintplex-Labs/anything-llm · error · Error
e.message
Error message
e.message
What it means
GenericOpenAiLLM.getChatCompletion rethrows e.message from the OpenAI SDK call against the arbitrary user-configured endpoint. The generic provider has isValidChatCompletionModel hard-coded to true ('short circuit since we have no idea if the model is valid'), so there is no pre-flight validation — any endpoint error (404 wrong path/model, 401 missing key, connection refused, non-OpenAI schema responses) arrives here raw. The max_tokens sent is GENERIC_OPEN_AI_MAX_TOKENS (default 1024).
Solutions
- Read the inner message: '404' → fix base path/model id; '401' → set GENERIC_OPEN_AI_API_KEY; 'ECONNREFUSED' → start the local service
- Confirm the endpoint answers: curl <GENERIC_OPEN_AI_BASE_PATH>/models and match the model id exactly
- Ensure the base path includes the version segment (usually /v1)
- If the service has its own auth (e.g. bearer token), put it in GENERIC_OPEN_AI_API_KEY
Example fix
# before (.env) GENERIC_OPEN_AI_BASE_PATH=http://localhost:11434 # missing /v1 -> 404 at chat # after (.env) GENERIC_OPEN_AI_BASE_PATH=http://localhost:11434/v1
Defensive patterns
Strategy: try-catch
Validate before calling
// generic provider can't validate models (isValidChatCompletionModel is always true),
// so probe the endpoint yourself before chatting
const res = await fetch(`${process.env.GENERIC_OPEN_AI_BASE_PATH}/models`);
if (!res.ok) throw new Error(`Endpoint unreachable or wrong base path (HTTP ${res.status}).`);
const { data } = await res.json();
if (!data.some((m) => m.id === model)) throw new Error(`Model '${model}' not served by endpoint.`); Try / catch
try {
return await llm.getChatCompletion(messages);
} catch (err) {
if (/ECONNREFUSED|fetch failed/i.test(err.message)) return respond("Local inference server is down — start it and retry.");
if (/404/.test(err.message)) return respond("Wrong base path or model id — check /v1 and the model name.");
if (/401|api key/i.test(err.message)) return respond("Endpoint requires GENERIC_OPEN_AI_API_KEY.");
throw err;
} Prevention
- Verify base path + model id with a /models probe when settings are saved
- Keep local inference servers under a process supervisor so they restart automatically
- Include the /v1 suffix in the configured base path
When it happens
Trigger: chat.completions.create failing against GENERIC_OPEN_AI_BASE_PATH: base path missing /v1 (404), model id not served by the endpoint, GENERIC_OPEN_AI_API_KEY required but unset/wrong, local service down (ECONNREFUSED), or the endpoint not actually OpenAI-compatible (SDK fails parsing).
Common situations: LM Studio/Ollama/vLLM not running or listening on a different port; path set to http://localhost:11434 without /v1; endpoint needs a key the user assumed optional; model renamed on the local server after the workspace pref was saved.
Related errors
- e.message
- e.message
- e.message
- GenericOpenAI must have a valid base path to use for the…
- Could not load into Foundry Local
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/41ccbb313fb5987f.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/AiProviders/genericOpenAi/index.js:234
if (process.env.GENERIC_OPEN_AI_REPORT_USAGE !== "true") return {};
return {
stream_options: {
include_usage: true,
},
};
}
async getChatCompletion(messages = null, { temperature = 0.7 }) {
const result = await LLMPerformanceMonitor.measureAsyncFunction(
this.openai.chat.completions
.create({
model: this.model,
messages,
temperature,
max_tokens: this.maxTokens,
})
.catch((e) => {
throw new Error(e.message);
})
);
if (
!result.output.hasOwnProperty("choices") ||
result.output.choices.length === 0
)
return null;
const usage = {
prompt_tokens: result.output?.usage?.prompt_tokens || 0,
completion_tokens: result.output?.usage?.completion_tokens || 0,
total_tokens: result.output?.usage?.total_tokens || 0,
duration: result.duration,
};
this.#extractLlamaCppTimings(result.output, usage);
return {View on GitHub (pinned to 3aec848f28)