Mintplex-Labs/anything-llm · error · Error
e.message
Error message
e.message
What it means
FoundryLLM.getChatCompletion wraps the OpenAI-compatible chat.completions.create call in a .catch that rethrows `new Error(e.message)`, flattening whatever the SDK produced. This is a passthrough: the real failure is a Foundry Local HTTP/transport error (model crashed mid-request, malformed request, context overflow of max_completion_tokens, connection dropped). The rethrow exists to normalize SDK APIError objects into plain Errors for AnythingLLM's handler, but it discards status codes and the original stack.
Solutions
- Read the inner message — it is the literal Foundry Local error (e.g. 'max_completion_tokens too large', 'model not found') and dictates the fix
- Reduce the prompt size / lower GENERIC token limits or clear chat history if the context overflows the small default 16K window
- Reload or pick a smaller model if inference crashed (FoundryLLM forgets evicted models, so a retry may reload it)
- Update Foundry Local and restart the service if the error indicates a schema/protocol problem
Defensive patterns
Strategy: try-catch
Try / catch
try {
return await llm.getChatCompletion(messages);
} catch (err) {
// err.message is the raw Foundry Local error text; branch on its content
if (/context|token/i.test(err.message)) return respondTrimmedAndRetry(messages);
if (/Premature close|socket/i.test(err.message)) return retryOnce(); // model eviction
return respond(`Foundry inference failed: ${err.message}`);
} Prevention
- Cap prompt size to the model's real window (FoundryLLM defaults to 16K) before sending
- Retry once on socket/premature-close errors — the class already forgets evicted models so the retry reloads
- Watch Foundry Local logs alongside your own to correlate the rethrown message with a status code
When it happens
Trigger: The HTTP POST to Foundry Local's /chat/completions returning 4xx/5xx: requesting more max_completion_tokens than the loaded model supports, a request payload the local service rejects, the model process dying mid-inference, or the socket closing after an eviction ('Premature close' — see #handleStreamFailure which treats this case).
Common situations: Context larger than the local model's window; model evicted between the load check and the request; Foundry Local version mismatch producing a schema the SDK rejects; laptop under memory pressure so inference crashes.
Related errors
- e.message
- e.message
- e.message
- Could not load into Foundry Local
- Foundry chat: is not valid or defined model for chat…
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/3742ab0d21bab956.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/AiProviders/foundry/index.js:324
if (!this.model)
throw new Error(
`Foundry chat: ${this.model} is not valid or defined model for chat completion!`
);
// max_completion_tokens is required by Foundry (it caps output at 1024
// otherwise), so the window has to be resolved before the request is built.
await this.assertModelContextLimits();
await this.assertModelLoaded();
const result = await LLMPerformanceMonitor.measureAsyncFunction(
this.openai.chat.completions
.create({
model: this.model,
messages,
temperature,
max_completion_tokens: this.promptWindowLimit(),
})
.catch((e) => {
throw new Error(e.message);
})
);
if (
!result.output.hasOwnProperty("choices") ||
result.output.choices.length === 0
)
return null;
return {
textResponse: result.output.choices[0].message.content,
metrics: {
prompt_tokens: result.output.usage.prompt_tokens || 0,
completion_tokens: result.output.usage.completion_tokens || 0,
total_tokens: result.output.usage.total_tokens || 0,
outputTps: result.output.usage.completion_tokens / result.duration,
duration: result.duration,
model: this.model,View on GitHub (pinned to 3aec848f28)